Curated Papers
Foundational and breakthrough machine learning papers every engineer should understand, alongside live daily trending research.
Must-Read Foundational Papers
The paper that started it all. Introduced the Transformer architecture, replacing RNNs entirely. Everything in modern deep learning traces back to this one.
Empirically proved that model performance follows predictable power laws with compute, data, and parameters. This gave AI labs a roadmap to GPT-4 and beyond.
The paper behind InstructGPT and ChatGPT. Introduced RLHF as the standard recipe for making LLMs useful and safe — arguably the most impactful alignment paper to date.
Made fine-tuning massive LLMs accessible on consumer hardware. LoRA is now the default PEFT method — injecting trainable low-rank matrices while freezing base weights.
Shows that pure RL without supervised fine-tuning can produce strong chain-of-thought reasoning. The open-source release changed the landscape for reasoning models.
First serious Transformer alternative that scales well. Selective state spaces let the model choose what to remember — a fundamental rethink of sequence modeling.
Showed that vision-language models can directly output robot actions, transferring internet-scale knowledge to physical manipulation. A landmark in generalist robot policies.
DreamerV3 learns a world model from raw pixels and plans within it — achieving human-level performance across 150+ tasks including robotics, without task-specific tuning.
Decoded imagined handwriting from motor cortex at 90 characters/min — a breakthrough in intracortical BCI speed. Demonstrates that neural population geometry can encode fine motor sequences.
Compact depthwise CNN that generalises across BCI paradigms (P300, SSVEP, ERN, MI) with very few parameters. The go-to baseline architecture for EEG decoding research.
Trending Daily Research
Live from HuggingFace Daily Papers
Natural-language service requests can require a language-model decision before execution starts, consuming part of the request's latency budget. We integrate Jev's decision-oriented application programming interface (API) into edge service…
Video generation models are increasingly being explored as world simulators for embodied planning and learning. To do so effectively, these models must not only generate visually appealing frames, but also predict how environments dynamical…
An assistant that serves the same user over a long horizon has to answer from what that user has revealed: which preferences still hold, which were revised, and which constraints apply now. Retaining that information is not the same as acti…
Video generation has rapidly progressed from short, low-quality clips to high-resolution, long-duration sequences with complex spatiotemporal dynamics. Despite strong generative priors learned through large-scale pretraining, pretrained vid…
An activation-steering coefficient sets intervention strength, but requesting a particular degree of persona expression requires a behavioral scale. We study persona dosing: controlling a language model through a trait description and a req…
Pretrained transformers use little of their depth to follow references in context. Thirteen base models reliably follow only 1.4-3.6 lines, and extra pretrained loops add little. A task-trained rank-8 LoRA at one early layer extends this co…
Video world models require persistent scene memory to maintain consistency during long-horizon video generation. Existing spatial memories accumulate RGB observations or latent features, increasing storage requirements as generation proceed…
Multi-step agents are trained on flat action streams: SFT and RLVR weight every token uniformly and ignore the sub-procedures that recur across tasks, the hierarchy that lets humans plan top-down from reusable routines. This structure sits…
Multidisciplinary tumor boards integrate multimodal clinical observations and longitudinal patient histories through specialist discussions, yet benchmarks rarely capture these real-world trajectories. We introduce OpenTumorBoard, a benchma…
Keyword-matching benchmarks can credit small models for tool use they never perform. We document such a false positive in a matched-architecture pair of Spanish security language models and propose a ladder of strict, cheap diagnostics. A 6…
What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahead of them at sensing which prior idea, buried in an ever-growing archive of research, a new problem needs. To study thi…
On-policy learning has been argued to reduce catastrophic forgetting, produce sparser parameter updates, and improve generalisation. However, existing comparisons between supervised fine-tuning and reinforcement learning vary many factors s…
Coding agents are beginning to move beyond purely digital tasks to tackle physical-world challenges, particularly in robotics. Existing robotics benchmarks, however, primarily focus on the performance of individual artifacts, such as polici…
Recent video-based world models pair the scalability of autoregressive (AR) prediction with the visual quality of diffusion models. The choice of scene tokenizer is paramount for the optimal performance of each of these, both in terms of fi…
On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a frozen or EMA version of themselves that receives privileged information. Its application to mul…