[TechGita]

Must-Read Papers

Papers every AI engineer should know

Attention Is All You Need

Ashish Vaswani, Noam Shazeer, et al.·NeurIPS 2017·2017
The paper that started it all. Introduced the Transformer architecture, replacing RNNs entirely. Everything in modern deep learning traces back to this one.
TransformersNLPArchitecturearXiv:1706.03762

Scaling Laws for Neural Language Models

Jared Kaplan, Sam McCandlish, et al.·arXiv·2020
Empirically proved that model performance follows predictable power laws with compute, data, and parameters. This gave AI labs a roadmap to GPT-4 and beyond.
ScalingLLMsTheoryarXiv:2001.08361

Training Language Models to Follow Instructions with Human Feedback

Long Ouyang, Jeff Wu, et al.·NeurIPS 2022·2022
The paper behind InstructGPT and ChatGPT. Introduced RLHF as the standard recipe for making LLMs useful and safe — arguably the most impactful alignment paper to date.
RLHFAlignmentLLMsarXiv:2203.02155

LoRA: Low-Rank Adaptation of Large Language Models

Edward Hu, Yelong Shen, et al.·ICLR 2022·2021
Made fine-tuning massive LLMs accessible on consumer hardware. LoRA is now the default PEFT method — injecting trainable low-rank matrices while freezing base weights.
Fine-tuningPEFTLLMsarXiv:2106.09685

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

DeepSeek-AI·arXiv·2025
Shows that pure RL without supervised fine-tuning can produce strong chain-of-thought reasoning. The open-source release changed the landscape for reasoning models.
LLMsReasoningRLarXiv:2501.12948

Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Albert Gu, Tri Dao·arXiv·2023
First serious Transformer alternative that scales well. Selective state spaces let the model choose what to remember — a fundamental rethink of sequence modeling.
ArchitectureSSMEfficiencyarXiv:2312.00752

RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Anthony Brohan, et al. (Google DeepMind)·CoRL 2023·2023
Showed that vision-language models can directly output robot actions, transferring internet-scale knowledge to physical manipulation. A landmark in generalist robot policies.
RoboticsVLMRobot LearningarXiv:2307.15818

Mastering Diverse Domains through World Models

Danijar Hafner, Jurgis Pasukonis, et al.·arXiv·2023
DreamerV3 learns a world model from raw pixels and plans within it — achieving human-level performance across 150+ tasks including robotics, without task-specific tuning.
RoboticsWorld ModelsRLarXiv:2301.04104

High-performance brain-to-text communication via handwriting

Frank R. Willett, Donald T. Avansino, et al.·Nature·2021
Decoded imagined handwriting from motor cortex at 90 characters/min — a breakthrough in intracortical BCI speed. Demonstrates that neural population geometry can encode fine motor sequences.
BCINeural DecodingNeuroscience

EEGNet: A Compact Convolutional Neural Network for EEG-based Brain-Computer Interfaces

Vernon J. Lawhern, Amelia J. Solon, et al.·Journal of Neural Engineering·2018
Compact depthwise CNN that generalises across BCI paradigms (P300, SSVEP, ERN, MI) with very few parameters. The go-to baseline architecture for EEG decoding research.
BCIEEGDeep LearningarXiv:1611.08024

Trending on HuggingFace

Updated on each site deploy · 50 papers

GaussianSelector: Lightweight Human-Guided Object Selection in 3D Gaussian Splatting with Graph Optimization

Baihan Yang, Tiexin Li, Yuheng Liu +4· 2026

Selecting a complete 3D object from a reconstructed scene with minimal user effort is essential for practical scene editing and embodied interaction. Existing 3DGS-based methods either retrain the Gaussian representation to embed per-object

Weights or Skills? A Survey of Robot-Learning Techniques: from Action-Predicting Weights to Robots that Write their Own Skills

Gaytri Jena, Kapil Wanaskar, Vinija Jain +3· 2026

Robot learning is splitting into two bets: policies that bake competence into frozen weights (vision-language-action, or VLA, models), and agents that write and refine their own executable skills as code. This survey organises the field aro

FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban Worlds

Kapil Wanaskar, Gaytri Jena, Aman Chadha +3· 2026

World models have attracted significant attention for their ability to capture and predict the structure and dynamics of the physical world. In this emerging landscape, Joint Embedding Predictive Architectures (JEPA) offer a particularly co

KVAE: Family of Tokenizers for Multimodal Generative Models

Andrey Shutkin, Denis Parkhomenko, Ivan Kirillov +11· 2026

Latent diffusion modeling (LDM), a prominent paradigm, utilizes tokenizers to map input signal to compressed representation. This dependency positions tokenizer as an integral part of generation process itself, since it affects learning spe

Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay

Nossa Iyamu· 2026

Computer-use agents pay full frontier inference to re-derive routines their user has already performed, because an agent's memory today records what the user said, not what the user did. We compile passively captured screen activity into ag

DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces

Boyan Li, Zhuowen Liang, Yupeng Xie +11· 2026

Data agents enable natural-language analytics over organizational workspaces, where relevant evidence may be scattered across databases, structured files, long documents, and multimedia. Existing benchmarks largely isolate structured queryi

Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval

Ilia Semenkov, Daria Kleeva, Ivan Dakhtin +2· 2026

Short segments of perceived speech can be retrieved from non-invasive magnetoencephalographic (MEG) recordings by deep networks trained with a CLIP-style objective against wav2vec 2.0 audio embeddings. Yet their weights do not map onto elec

MameLoshnLM: Yiddish Language Model and Evaluation Benchmark

Uri Katz, Omer Goldman, Tomasz Limisiewicz +2· 2026

We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish. Despite Yiddish's rich textual tradition, its limited digital presence and the scarcity of reliable evaluation resources have constrai

Continual Learning in Transition

Zhiyan Hou, Dan Zhang, Tao Feng +10· 2026

Classical continual learning (CL) has primarily focused on enabling models to update and retain knowledge through parameter-centric mechanisms, e.g., training strategies, architectural designs, and weight adaptation. However, emerging parad

Task-Conditional Flow Matching for Balanced Multilingual Text Embedding Adaptation

Tirth Bhatt, Naren Kumar S, Mayank Singh· 2026

Multilingual text embedding models are commonly adapted using a single training objective across diverse tasks, despite different tasks requiring fundamentally different optimization strategies. We introduce Task-Conditional Flow Matching (

Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains

Ayoub Kirouane, Christos Petrocheilos· 2026

Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major multilingual retrieval benchmarks, despite being important for retrieval-augmented generation (RAG) in legal, energy, financial, and medical applications. We pres

From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models

Jiale Han, Xiang Li, Jing Qian +7· 2026

Economic World Models (EWMs) are generative economic models that simulate how economies evolve from within by modeling heterogeneous agents, their beliefs and actions, and the market and institutional mechanisms through which their interact

MASS: Multiplayer World Models with Authoritative Shared State

Ziqi Cai, Siqi Yang, Yimu Wang +6· 2026

Current video world models struggle in multiplayer environments because they entangle world state with view-dependent visual latents, leading to redundant compute, view inconsistencies, and poor scalability. We propose MAS (Multiplayer worl

ContextMaster: Interactive Multi-Shot Video Creation via Fixed-Budget Sparse Context Routing

Xu Guo, Zhengxuan Wei, Xinghui Li +11· 2026

Recent video models increasingly support generation, reference conditioning, and editing within a single model, yet typically expose them as separate operations over fixed inputs. Practical creation unfolds across multiple shots, requiring

EffectLearner: World-Aware Object-Effect Reasoning for Real-World Video Object Removal

Feier Wu, Wanke Xia, Xu He +8· 2026

Video object removal must eliminate not only the target object but also its induced effects while maintaining high-fidelity and spatiotemporally coherent restoration. Existing methods mainly learn object-effect correspondences implicitly fr

Invisible Shortcuts: Why Vision Encoders Know Your Camera

Vladan Stojnić, Ryan Ramos, Giorgos Kordopatis-Zilos +2· 2026

Deep vision models exploit shortcuts, relying on cues that correlate with supervision signals. Prior work has focused on visible biases, such as object-background or texture correlations. We identify a different source of shortcut learning:

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

Qiushi Sun, Kanzhi Cheng, Yian Wang +20· 2026

Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is central to CUA evaluation, data curatio

PaDoc: Layout-Grounded Parallel Decoding for Document Parsing

Hao Yu, Jiabo Zhan, Kang Liu +8· 2026

End-to-end document parsers provide a unified interface, but serialize page layouts and regional contents into one autoregressive sequence. This formulation forces independent regions onto a decoding path whose length grows with the total c

World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation

Yuhao Pan, Haosong Peng, Zhengshen Zhang +8· 2026

Vision-language-action (VLA) models often treat main-view and wrist-view observations as parallel visual inputs, overlooking their distinct roles in robot manipulation. Fine-grained manipulation, however, benefits from anticipating how wris

EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning

Zishan Xu, Zhiyuan Yao, Yuxin Chen +9· 2026

Training large language model agents for long-horizon tool use typically relies on interactions with real or synthesized executable environments, whose construction and verification are costly, or on external simulators that are difficult t

SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding

Yue Zhang, Yingzhao Jian, Yunqiu Xu +2· 2026

Understanding 3D scenes is fundamental to embodied intelligence, requiring joint reasoning over heterogeneous information from multiple modalities, including visual and geometric cues. However, the relevance of these modalities often varies

CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks

Fanzhe Meng, Guoxin Chen, Jiale Zhao +6· 2026

Training terminal agents requires executable and verifiable tasks that are not merely solvable, but appropriately challenging for learning. Executable validation establishes feasibility, yet does not reveal how a task behaves relative to a

AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning

Zi-Han Wang, Zhengxi Lu, Zhiyuan Yao +10· 2026

Reinforcement learning (RL) with verifiable rewards constructs trajectory-level advantage estimates, yet it often fails to credit the few pivotal decisions that determine outcomes in long-horizon, multi-turn agentic tasks. Recent work intro

WorldClaw: Agentic 3D Open-World Generation at Scale

Chunchao Guo, Jinpeng Li, Yang Li +1· 2026

Generating large-scale, freely explorable 3D worlds from open-ended text remains challenging because a system must jointly maintain global spatial coherence, rich local content, and explicit assets suitable for downstream editing and reuse.

GST-Bench: Can VLMs Develop Global Spatial Awareness from Video?

Qifeng Zhang, Kaixiang Huang, Heng Dong +6· 2026

Spatial intelligence is fundamental to embodied agents, yet existing benchmarks focus on local spatial perception from single or few viewpoints, overlooking global spatial awareness over continuous, long-horizon visual streams. To address t

On-Policy Delta Distillation for Multilingual Math Reasoning

Byeongho Heo, Jaehui Hwang, Sangdoo Yun +1· 2026

On-Policy Distillation (OPD) is emerging as a promising alternative to reinforcement learning for LLM post-training, yet its effectiveness in multilingual settings remains underexplored. We study OPD and its advanced variant, On-Policy Delt

Learning from Failures: Retrieval-Centric CoT via Hard Negatives for Unified Multimodal Retrieval

Zelong Sun, Jun Wang, Kaicheng Yang +3· 2026

Unified multimodal retrieval aims to identify candidates that satisfy complex user intent expressed through heterogeneous inputs. Although Large Vision-Language Model (LVLM)-based retrievers are efficient and scalable, directly encoding raw

ChronoVision: Temporal Reasoning via Latent State Reconstruction

Yifan Shen, Jian Xu, Boyi Li +6· 2026

Multimodal large language models excel at passive perception but struggle with complex visual cognitive tasks requiring multi-step temporal reasoning. This degradation largely stems from the inherent ambiguity of language-based reasoning, w

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization

Varun Ursekar, Apaar Shanker, Yash Maurya +4· 2026

As LLMs are increasingly deployed within agentic systems, their capabilities depend not only on the model weights but also on the harness: the prompts, tools, control flow, memory, and orchestration code surrounding them. This makes automat

DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation

Junfeng Li, Junjie He, Zhide Zhong +12· 2026

Vision-Language-Action (VLA) models have become a powerful paradigm for robot manipulation, but training a single generalist policy for heterogeneous robot embodiments remains an open problem. Existing methods have two main limitations. Fir

Recursive Synthesis for Long-Horizon Terminal Tasks

Zhongzhi Li, Yucheng Shi, Zongxia Li +8· 2026

High-quality long-horizon training data for terminal agents is expensive to produce, often costing hundreds to thousands of dollars per task, because each task must keep the instruction, environment, reference solution, and verifier mutuall

Helping Music Co-Creation Agents 'Listen' Well: Hierarchical Self-Supervised World Models for Understanding and Generation

Scott H. Hawley· 2026

Collaborative music agents need internal representations rich enough to support both understanding and generation, yet flexible enough for a workflow where the human retains agency. We present a hierarchical self-supervised ``world model''

FinanceHarness: Autonomous Financial Deep Research Framework

Yijia Xiao, Rujun Han, Yanfei Chen +8· 2026

Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most deep research systems write general-purpose reports, which are inadequate for financial deep research

What AI Red-Team Evaluations Can and Cannot Prove

Bandana Kaur· 2026

Red-team evaluations of AI models support some claims and not others, and the boundary between the two is calculable rather than merely a matter of judgment. We define the evidential ceiling of an evaluation as the largest factor by which o

SIGNPOST-Bench: Benchmarking Text-Vision Conflict Resolution in Multimodal Large Language Models

Sirun Li, Minghao Liu, Ling Dai +4· 2026

Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitrate between these evidence sources when they conflict. We int

Lossless Tensor Compression as Program Synthesis

Jieke Shi, Junda He, Wenjia Jiang +11· 2026

Model checkpoints are growing in both number and size, which makes archival, transfer, and deployment increasingly costly. General-purpose compressors can reduce storage requirements but ignore tensor structure, whereas existing tensor-spec

Resume Means Resume: A Machine-Checked Conformance Contract for Checkpoint, Interrupt, and Resume Semantics in Workflow Persistence Layers

Sajjad Khan· 2026

A framework that persists execution state so a run can be interrupted, survive a crash, and continue must decide what a resume means for effects that already fired. Five widely deployed agent workflow frameworks answer differently, none exp

DRIFT: Derailing Denoising Trajectories of Flow-Matching VLAs with Adversarial Patch Attack

Hoseong Tae, Jong-Seok Lee· 2026

Flow-matching vision-language-action (VLA) models such as pi0 generate robot actions by integrating a learned denoising velocity field, and have been reported to resist adversarial perturbations that readily fool autoregressive VLAs. We sho

Self-Evolving Coding Agents

Hao Zhou, Haichuan Hu, Ye Shang +1· 2026

Large language models are increasingly embedded in software engineering workflows as coding agents that can inspect repositories, invoke tools, execute tests, debug failures, and generate patches. Yet most existing agents remain largely sta

ABSeeker: Training Long-Horizon Search Agents via Answer-Backtracked Credit Assignment

Yijun Lu, Rui Ye, Jiajun Wang +4· 2026

Long-horizon search agents must make multiple sequential actions (steps) to search, retrieve, verify, and integrate evidence to reach a final answer. However, existing methods for training these agents typically treat all steps within a tra

FocusMem: Factorizing Content, Readout, and Trust in Latent GUI Memory

Zhuoran Zhang, Bowen Li, Jingcheng Ju +5· 2026

GUI agents must remember both useful experience from earlier tasks and unfinished progress in the current interaction. Latent memory offers a compact solution by compressing multimodal trajectories into a few continuous tokens. Existing met

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks

Leijun Zhou, Zhihao Liu, Xiang Qu +9· 2026

Agent self-evolution updates an agent's persistent state from prior experience and reuses it to solve related tasks more effectively. Evaluating self-evolution is difficult: existing benchmarks provide limited coverage of economically valua

AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities

Yuqing Wen, Yukai Huang, Qianqian Xie +6· 2026

While instruction-based video editing has advanced rapidly, real-world videos contain tightly coupled audio and visual signals, and editing one modality often requires coordinated changes in the other. Existing benchmarks primarily evaluate

Agent Against Agent: An Agentic System for Automatic Prompt Injection Red Teaming

Yanting Wang, Chenlong Yin, Runpeng Geng +1· 2026

Prompt injection poses significant security risks to LLM agents. Efficient and effective red-teaming is therefore critical, both for evaluating these risks and for collecting training data to improve defenses. Existing state-of-the-art prom

Consistency-Driven Co-Evolution for Self-Supervised Cross-Representation Learning

Xuehang Guo, Pengyuan Li, Tom Hope +3· 2026

As chart images, tabular data, and visualization code play increasingly important roles across diverse domains, cross-representation understanding across these modalities poses fundamental challenges for AI systems: the relationships across

Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data

Ye Wang, Pei Lin, Xiong-Hui Chen +12· 2026

Learning generalizable robot manipulation policies requires large-scale and diverse demonstration data. Egocentric human manipulation videos offer rich scene and task diversity, and prior work has shown that retargeting and rendering such v

K-EXAONE 2.0 Technical Report

Eunbi Choi, Kibong Choi, Sehyun Chun +74· 2026

This technical report presents K-EXAONE 2.0, an open-weight multilingual foundation model developed by LG AI Research as a step in our effort toward global frontier-scale foundation models. Rather than training from scratch, we upcycle K-EX

Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance

Zhuowen Han, Jinwei Xiao, Zhengxi Lu +9· 2026

Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training large language models (LLMs). While Group Relative Policy Optimization (GRPO) is widely adopted, it suffers from sparse reward signals an

BridgeVLA++: A Data-Efficient, Generalizable, and Memory-Augmented Vision-Language-Action Framework for 3D Manipulation

Peiyan Li, Yuze Zhu, Yixiang Chen +10· 2026

Leveraging pre-trained vision-language models (VLMs) to construct vision-language-action (VLA) models has emerged as a promising paradigm for 3D robot manipulation. However, existing 3D VLA methods remain data-hungry, exhibit limited genera

WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models

Bohai Gu, Yueyang Yuan, Taiyi Wu +9· 2026

Interactive video world models are essential for long-horizon planning and exploration, yet they suffer from compounding errors. Post-training methods such as reinforcement learning (RL) can improve these models, but they hit a verification