ML PHD without A* Publications [D]
I know top ML PhD admissions are insanely competitive, so I’m trying to figure out if it’s even worth applying or if I should just focus seriously on jobs instead.For context, I’m doing my MS at a...
View ArticleTransformers vs RNNs vs SSMs: Where Does Memory Actually Live? [D]
Someone who has always loved looking at the space between different AI techniques, this time I went a little deeper into the memory trade-offs between RNNs, Transformers and SSMs. I found it...
View ArticleAFP-GIC: Controllable Generative Image Compression [R]
Hi ML Community,I am excited to share our latest framework, AFP-GIC, officially published in IEEE Access (2026). We have released the deployment codebase and hosted an interactive visual...
View ArticleLearning to Learn a Language: in-context learning of natural language from a...
Learning from data as we observe it is easy for humans, but most machine learning models have limited ability to learn from new data that they have not seen during training. Prior-fitted networks (the...
View ArticleNeurIPS workshop registration for registered author [D]
For NeurIPS, I already registered a few months ago (for Atlanta) when there were news that Sydney is sold out. However, I forgot to register for the workshop. Now, when I go to register just for the...
View ArticleSWE-Race: a coding-agent benchmark of 188 real concurrency bugs, with results...
We've been building a benchmark out of real concurrency bugs (race conditions, deadlocks, cancellation issues) taken from merged PRs in about 100 Python projects. Each task gets graded by the project's...
View ArticleNeurIPS 2026 Financial Assistance [D]
Is anyone else unable to open the form for financial assistance despite the deadline still being in the future? submitted by /u/lcj29 [link] [comments]
View ArticleI have trained a model to predict my blood sugar (Part 2) [P]
This is related to my previous post where I shared an encoder-only transformer model trained on ohiot1dm + shanghait1dm + azt1d datasets. This time I trained the model on the outputs of my T1DM patient...
View ArticleA chunking lib in Rust that is ~20x faster [P]
Hey, I wanted a faster chunking library for my system without affecting the overall accuracy. Did not find many options. So I've build https://github.com/d1pankarmedhi/chunkrIt has most of the chunking...
View ArticleDistilling Stockfish on a Billion Positions, Full 3.9B Dataset Available [P]
In this project, I distilled the Stockfish value function into a ResNet/ViT model using 1 billion positions from the Gigafish dataset. The 3.9 billion position dataset is available on huggingface:...
View ArticleEmbedding Every Font with Neural Networks makes some Nice Structures...
I've been working on a font searching tool for about a year now, and my investigations have centered around pre-training neural networks to produce embeddings of each font. I usually then need to...
View ArticleSona: one transformer replaced our 15+ candidate generators, pre-ranker and...
Our production recommender at Yandex Music has 15+ candidate generators feeding pre-ranking and ranking models with hundreds of features. LLMs showed that one end-to-end model can take over work that...
View ArticleWithdrawing an accepted paper before camera-ready due to zero funding? (ACML...
Hi everyone, I recently had a paper accepted at ACML 2026, but I just found out that I have absolutely no funding to cover the registration fee or travel expenses. Because of this, I need to withdraw...
View ArticleLanguage barrier, shadier terms and jargon fog [D]
Hey all,I don't know if you guys are experiencing the same thing, but there is this behaviour that i have been noticing on the latest models on openAI (since sol 5.6) and Anthropic since Fable 5.1 and...
View ArticleTop ARC-ΑGI-3 scores on Kaggle just went from 7% to 56% [N]
https://preview.redd.it/gkjgii48gfth1.png?width=575&format=png&auto=webp&s=4e1c35bdacf41d18a6409baedcddc9432ec79814This happened over the past 30 days. So smallish local models (Kagglers...
View Articlethe official ICLR template .bib has had Bengio listed twice since 2019 [D]
i work on reference checking stuff so i was reading through the ICLR 2027 author guidelines and style files this week the sample .bib that ships with the template has the Deep Learning book as...
View ArticleWorking with an AI Company That Does Things You Disagree With [D]
I'm a PhD student in machine learning in the EU and was looking for internships at exciting companies.I shortlisted few and applied by reaching out to people and now reading project descriptions sent...
View ArticleA Minimal Interpretable Architecture for Zero-Shot Reconstruction of...
In our #NeurIPS2026 paper “A Minimal Interpretable Architecture for Zero-Shot Reconstruction of Dynamical Systems (DS)” (preprint: https://arxiv.org/abs/2607.14937) we reduce a DS foundation model to...
View ArticleASRN Adaptive Sparse Recurrence Network [N]
A copy layer for language models that finds earlier occurrences of the current context with learned hash tables and copies what came next — with memory linear in sequence length. submitted by...
View ArticleNeurIPS Workshops [D]
Anyone going to WMHS, Physworld AI or Vercodegen? Would love to get to know some people beforehand! submitted by /u/navalsaras [link] [comments]
View Article