TechGita
Research Compendium

Curated Papers

Foundational and breakthrough machine learning papers every engineer should understand, alongside live daily trending research.

Must-Read Foundational Papers

Attention Is All You Need

Ashish Vaswani, Noam Shazeer, et al.·NeurIPS 2017·2017
The paper that started it all. Introduced the Transformer architecture, replacing RNNs entirely. Everything in modern deep learning traces back to this one.
TransformersNLPArchitecturearXiv:1706.03762

Scaling Laws for Neural Language Models

Jared Kaplan, Sam McCandlish, et al.·arXiv·2020
Empirically proved that model performance follows predictable power laws with compute, data, and parameters. This gave AI labs a roadmap to GPT-4 and beyond.
ScalingLLMsTheoryarXiv:2001.08361

Training Language Models to Follow Instructions with Human Feedback

Long Ouyang, Jeff Wu, et al.·NeurIPS 2022·2022
The paper behind InstructGPT and ChatGPT. Introduced RLHF as the standard recipe for making LLMs useful and safe — arguably the most impactful alignment paper to date.
RLHFAlignmentLLMsarXiv:2203.02155

LoRA: Low-Rank Adaptation of Large Language Models

Edward Hu, Yelong Shen, et al.·ICLR 2022·2021
Made fine-tuning massive LLMs accessible on consumer hardware. LoRA is now the default PEFT method — injecting trainable low-rank matrices while freezing base weights.
Fine-tuningPEFTLLMsarXiv:2106.09685

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

DeepSeek-AI·arXiv·2025
Shows that pure RL without supervised fine-tuning can produce strong chain-of-thought reasoning. The open-source release changed the landscape for reasoning models.
LLMsReasoningRLarXiv:2501.12948

Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Albert Gu, Tri Dao·arXiv·2023
First serious Transformer alternative that scales well. Selective state spaces let the model choose what to remember — a fundamental rethink of sequence modeling.
ArchitectureSSMEfficiencyarXiv:2312.00752

RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Anthony Brohan, et al. (Google DeepMind)·CoRL 2023·2023
Showed that vision-language models can directly output robot actions, transferring internet-scale knowledge to physical manipulation. A landmark in generalist robot policies.
RoboticsVLMRobot LearningarXiv:2307.15818

Mastering Diverse Domains through World Models

Danijar Hafner, Jurgis Pasukonis, et al.·arXiv·2023
DreamerV3 learns a world model from raw pixels and plans within it — achieving human-level performance across 150+ tasks including robotics, without task-specific tuning.
RoboticsWorld ModelsRLarXiv:2301.04104

High-performance brain-to-text communication via handwriting

Frank R. Willett, Donald T. Avansino, et al.·Nature·2021
Decoded imagined handwriting from motor cortex at 90 characters/min — a breakthrough in intracortical BCI speed. Demonstrates that neural population geometry can encode fine motor sequences.
BCINeural DecodingNeuroscience

EEGNet: A Compact Convolutional Neural Network for EEG-based Brain-Computer Interfaces

Vernon J. Lawhern, Amelia J. Solon, et al.·Journal of Neural Engineering·2018
Compact depthwise CNN that generalises across BCI paradigms (P300, SSVEP, ERN, MI) with very few parameters. The go-to baseline architecture for EEG decoding research.
BCIEEGDeep LearningarXiv:1611.08024

Trending Daily Research

Live from HuggingFace Daily Papers

Replacing Large Language Models with Jev Decision Models for Low-Latency Edge Service Orchestration

Delong Li, Xu Wang, Haochen Gong +2· 2026

Natural-language service requests can require a language-model decision before execution starts, consuming part of the request's latency budget. We integrate Jev's decision-oriented application programming interface (API) into edge service…

Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation

Patrick Amadeus Irawan, Iskandar Muda Rizky Parlambang, Rava Maulana +10· 2026

Video generation models are increasingly being explored as world simulators for embodied planning and learning. To do so effectively, these models must not only generate visually appealing frames, but also predict how environments dynamical…

MemFold: Learning Compact Soft Memory for Long-Context Personalization via On-Policy Optimization

Jingxuan Wu, Yuzhe Yang, Yiqiao Huang +6· 2026

An assistant that serves the same user over a long horizon has to answer from what that user has revealed: which preferences still hold, which were revised, and which constraints apply now. Retaining that information is not the same as acti…

Video Generation Models: A Survey of Post-Training and Alignment

Chaoyu Li, Xiaoyi Gu, Yogesh Kulkarni +10· 2026

Video generation has rapidly progressed from short, low-quality clips to high-resolution, long-duration sequences with complex spatiotemporal dynamics. Despite strong generative priors learned through large-scale pretraining, pretrained vid…

Persona Dosing: Calibrated Activation Steering for Graded Trait Control

Zehao Jin, Junran Wang, Ruixuan Deng +4· 2026

An activation-steering coefficient sets intervention strength, but requesting a particular degree of persona expression requires a behavioral scale. We study persona dosing: controlling a language model through a trait description and a req…

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Zehao Jin, Ruixuan Deng, Junran Wang· 2026

Pretrained transformers use little of their depth to follow references in context. Thirteen base models reliably follow only 1.4-3.6 lines, and extra pretrained loops add little. A task-trained rank-8 LoRA at one early layer extends this co…

Honeycomb: Constant-Size Scene Memory Representation for Video World Models

Jack Wei Lun Shi, Kaichen Zhou, Haoyu Chen +6· 2026

Video world models require persistent scene memory to maintain consistency during long-horizon video generation. Existing spatial memories accumulate RGB observations or latent features, increasing storage requirements as generation proceed…

X-Tree: Tokenizing Reusable Experience for Efficient Agent Generalization

Sitao Cheng, Xunjian Yin, Zhiyuan Sun +4· 2026

Multi-step agents are trained on flat action streams: SFT and RLVR weight every token uniformly and ignore the sub-procedures that recur across tasks, the hierarchy that lets humans plan top-down from reusable routines. This structure sits…

OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories

Anqi Li, Zhixuan Ge, Yixuan Duan +9· 2026

Multidisciplinary tumor boards integrate multimodal clinical observations and longitudinal patient histories through specialist discussions, yet benchmarks rarely capture these real-world trajectories. We introduce OpenTumorBoard, a benchma…

Keyword Harnesses Fail Open: A Cheap Diagnostic Ladder for Tool-Use Claims in Small Language Models

Juan S. Santillana· 2026

Keyword-matching benchmarks can credit small models for tool use they never perform. We document such a false positive in a matched-architecture pair of Spanish security language models and propose a ladder of strict, cheap diagnostics. A 6…

ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research

Sohyeon Kim, Yoonho Lee, Bo Liu +11· 2026

What makes great scientists great? Even as AI systems start to make progress on open problems, scientists remain far ahead of them at sensing which prior idea, buried in an ever-growing archive of research, a new problem needs. To study thi…

On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics

Julianna Piskorz, Antonin Berthon, Mihaela van der Schaar· 2026

On-policy learning has been argued to reduce catastrophic forgetting, produce sparser parameter updates, and improve generalisation. However, existing comparisons between supervised fine-tuning and reinforcement learning vary many factors s…

RLE-Bench: A Qualifying Exam for Coding Agents as Robot Learning Engineers

Haitong Ma, Chenxiao Gao, Rushi Qiang +2· 2026

Coding agents are beginning to move beyond purely digital tasks to tackle physical-world challenges, particularly in robotics. Existing robotics benchmarks, however, primarily focus on the performance of individual artifacts, such as polici…

SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation

Mikhail Dereviannykh, Vikram Voleti, Simon Donne +3· 2026

Recent video-based world models pair the scalability of autoregressive (AR) prediction with the visual quality of diffusion models. The choice of scene tokenizer is paramount for the optimal performance of each of these, both in terms of fi…

Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes

Sophia Sirko-Galouchenko, Monika Wysoczanska, Andrei Bursuc +2· 2026

On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a frozen or EMA version of themselves that receives privileged information. Its application to mul…