Back to AI Lab

ArXiv AI Papers

Latest artificial intelligence and machine learning research papers from ArXiv.

Showing 50 of 258 items

Causal Memory Policy: Making Memory Utility Identifiable by Intervening on Retrieval

Treats an agent’s memory as a causal system and probes it by intervening on what gets retrieved. If you are tuning retrieval or tool memory, this gives a more principled way to measure which memories actually matter.

Arman Behnam, Binghui Wang

TopK-Guided: Adaptive, Budget-Aware Activation Sparsity for Efficient LLM Inference

Introduces a method that activates only the most useful neurons in each layer, under a fixed compute budget. If you care about cutting the cost of running the AI without throwing away too much quality, this is directly actionable.

Mukund Agarwalla, Chih-Jen Lin

VideoEvolve: Evolving Agent Harnesses for Video Temporal Grounding

Optimizes the agent harness around video models that need to find events in long clips. If you are building video agents, this shows that harness design and curriculum matter as much as the base model.

Bingjun Luo, Yuhuan Fan

Not All Experience Belongs in the Weights: Component Routing for Self-Improving GUI Agents

Shows a way for GUI agents to improve by routing experience into modular components instead of always changing core weights. That matters if you want agents that get better on real apps without constant retraining runs.

Beining Wu, Zihao Ding

Continuous Process-Level Evaluation for Evolving Enterprise AI Agent Skills

Proposes always on evaluation of enterprise agents at the business process level, not just individual prompts. If you run pilots inside big companies, this gives a structure for tracking whether agents actually improve operations over time.

Ngoc Phuoc An Vo, Aarya Doshi

Code Owns the Simulation, Jev Owns the Evaluation

Argues that you should treat evaluation code as a first class object, separate from the environment being simulated. Useful for anyone designing complex agent benches, where leakiness between sim code and scoring quietly corrupts results.

Yaodong Yang, Hongyao Tang

Can AI Oversight Be Zero Knowledge?

Applies zero knowledge ideas so overseers can check model behavior without seeing sensitive inputs or outputs. If you work on safety for regulated sectors, this sketches how to watch powerful models without breaking privacy rules.

Alessandro Chiesa, Ziyi Guan

Mingbird: A Local-First Agent Harness Enabling Small Open Models to Complete Real Tasks

Introduces an agent harness tuned for small open models running on local hardware. If you want agents that respect data boundaries and still finish end to end tasks, this paper is worth copying from.

Hao Wang, Ting Huang

Mimir: Physics-Grounded LLM Agents for Long-Horizon Irrigation Control

Combines physics based simulators with language model agents to control farm irrigation over long periods. It is a concrete template for tying agents to physical models so they do not drift into unphysical decisions.

Yimeng Liu, Mi Zhang

External Observers May See More Clearly: Cross-Model Span-Level Hallucination Detection in Large Language Models via Hidden State Probing

Uses a separate probe model to watch another model’s hidden states and flag hallucinated spans. If you ship high stakes apps, this is a concrete recipe for wrapping your main model with an independent fact checker.

Kingshuk Gupta, Davide Buscaldi

VISTA: A Visual Harness for Reasoning in an Interactive World

Shows that a smart vision harness can unlock long horizon reasoning in existing text image models without retraining them. If you build embodied agents or robotics stacks, this is a blueprint for wrapping foundation models instead of training new ones from scratch.

Qiushi Han, Keya Hu

ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research

Defines a benchmark that scores search systems on whether they surface papers that actually spark new follow up work, not just high citation counts. Useful if you build research agents or citation tools and need a target beyond vanilla relevance and click metrics.

Sohyeon Kim, Yoonho Lee

Emergent Intelligence: Resonant Oscillators Produce Proactive Adaptive Behavior

The authors show that a tiny network of coupled oscillators, with no training, can explore, adapt, and change behavior based on feedback. It is a toy but suggests new foundations for "curious" agents.

Alex Fedosov, Maxim Yakimenko

The Weight Is Over - Interactive Diffusion on Consumer GPUs

This work refactors diffusion pipelines so mid-range GPUs can handle image generation interactively. They mix model surgery, compression, and scheduling to keep quality while shrinking memory and latency.

Frieder Ganz, Maximilian Müller

LLMs as Feature Engineers for Text-and-Tabular Prediction

They use one model to propose human-readable categorical features from text, and another to extract them as table columns. A standard tabular model then scores their value and drives an iterative search.

Merwan Barlier, Blaz Skrlj

Detecting Pretraining Data in Large Language Models from a Free-Energy Perspective

The authors argue you should judge membership by loss relative to uncertainty, not loss alone. Their "free-energy" style detector separates memorized data from truly predictable text much more cleanly.

Chenye Ke, Zirui Liu

ExpBoN: Exponential-Noise Best-of-n for Efficient Test-Time LLM Alignment

The paper introduces ExpBoN, a softer version of "best-of-n" sampling that adds exponential noise so you can dial reward gains against distribution shift. It greatly cuts the number of samples needed for strong test-time alignment, making long "thinking" runs more affordable.

Yanxiao Liu, Sicheng Wan

Can 4D Foundation Models Remember?

The authors probe video and 4D models to see how well they retain scene details across time and camera motion. They find serious memory gaps, especially for small or briefly seen objects.

Guangzhao He, Hadar Averbuch-Elor

Stress-testing Alignment Midtraining

Alignment midtraining mixes safety-relevant documents into pretraining instead of relying only on post-training. This paper builds tests showing where that approach helps, and where dangerous behaviors still sneak through.

Sid Baines, Jonathan Bostock

JEPA-Anything: Learning Predictive Models across Different Worlds

JEPA-Anything trains a single predictive model across very different domains by factorizing what changes and what stays stable. It shows that one learning recipe can cover physics, games, and more.

Taoyong Cui, Zhongyao Wang

Refuse, Decompose, Refresh: A Claim-Safe Protocol for Closed-Loop AI Evaluation

They propose a three-step evaluation protocol so closed-loop agent tests do not oversell claims. The method flags when data cannot support a claim, then decomposes and re-runs tests until it can.

Peiying Zhu, Sidi Chang

The Organization of Inference: Information, Resource Constraints, and AI Production

Using controlled coding workflows, the authors treat AI systems like factories with planning and execution stages. They quantify how shifting token budgets and information between stages changes throughput and reliability.

Yukun Zhang, Kemu Xu

How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents

This paper A/B tests real enterprise agents with and without meaningful plans and release controls. It shows structured guidance and staged rollout sharply cut silent failures at similar headline success rates.

Yukun Zhang, Kemu Xu

SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness

SoL-Pi automates experiments on agent harness designs across many environments and models, then feeds the findings back into the harness. It improves coding-agent performance while cutting token and dollar costs.

Haozhe Liu, Tian Ye

An Empirical Study of Harness Design for Coding Agents

They vary planning, tools, and context handling across 176 harness configurations for coding agents. Results show context strategy matters most under tight budgets, while planning mainly shifts cost, not accuracy.

Run-Ze Fan, Zihao Zhang

AgentPProf: Semantic Profiler for Long Horizon AI Agents

AgentPProf logs agent behavior like a performance profiler, but in semantic space instead of CPU cycles. It helps teams see which tasks burn tokens, fail often, or cause risky side effects.

Yusheng Zheng, Chaokun Chang

SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking

SAS learns which tokens matter most by feeding a learned gate directly into the attention scores. That lets gradients flow from the language loss into the selector. If you care about cheaper long context LLMs, this offers a practical recipe beyond hand tuned pruning tricks.

Zhiwei Li, Lei Zhu

Rethinking Heterogeneous System Disaggregation for Subquadratic Attention

They design a new way to split LLM workloads across mixed hardware by treating subquadratic attention separately from dense layers. Their SQD scheme boosts tokens per joule on future GPU plus memory systems and tightens latency under fixed power. If you run large models on weird new hardware, this gives concrete design and deployment ideas.

Arya Tschand, Yaosheng Fu

Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models

They build TAM, a benchmark where models must follow real manuals for medical coding and criminal sentencing. Even strong models barely solve any tasks end to end. If you trust “multi hop reasoning” scores, this is a sharp reminder that long rule books are still a brick wall.

Utkarsh Soni, Syed Shariyar Murtaza

K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments

They show that “unlearned” secrets still leak once the model runs as a tool using agent. K‑Bench checks every channel an agent exposes, not just the final answer. This is a must read if you plan to forget training data or user prompts while still keeping agents useful.

Guangsheng Yu, Yanna Jiang

EduFair-Bench: Evaluating Pedagogical Fairness of LLM Tutors Across Student Demographics

They simulate students with different gender, language, immigration background, and income, then measure how LLM tutors treat them. The benchmark scores both teaching quality and systematic gaps. If you ship AI tutors or classroom tools, this is a template for checking who your model quietly leaves behind.

Jiaxu Zhao, Bahar Radmehr

How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks

Physicists regrade several headline physics benchmarks and discover many supposed model failures were actually bad keys and ambiguous questions. After fixing tests, GPT‑5.6 Sol and peers nearly max out multiple physics suites. Anyone tracking “models still fail physics” needs to revisit those claims and pay more attention to benchmark design.

Ali Ansari, Haoran Sun

Autonomous Research for Open-Ended Problems: A Case Study on Telecom Ticket Retrieval

An LLM based agent iteratively designs, trains, and evaluates models to solve messy telecom support retrieval tasks. It operates over large search spaces where classic automated ML tools struggle. Use this as a concrete reference when arguing about how far “automated research interns” already go on real business problems.

Junghyun Min, Huseyin Uzunalioglu

Embodied-BenchForge: A Closed-Loop Agentic Workflow for Embodied Benchmark Construction

They build an agent pipeline that turns a high level evaluation idea into a full robot or embodied benchmark. The system both generates tasks and automatically checks and repairs flawed intermediate assets. If you care about realistic tests for robot agents, this is a blueprint for automating that grind.

Baoyang Jiang, Fengchun Zhang

Post-Training Language Models for Gold-Medal Performance in Coding Competitions

NVIDIA fine tunes Nemotron models on curated programming problems plus synthetic reasoning traces, then layers a test time solve refine loop. Their system beats the top human at IOI 2026 under contest rules. If you use models for serious coding, this shows how much is still on the table with better post training and runtime strategies.

Aleksander Ficek, Sean Narenthiran

CC-Mediation: Evaluating Large Language Models for Cross-Cultural Conflict Mediation

The authors build a benchmark where mediators must respect different cultural norms while resolving conflicts. Models often import biases from their training data and mishandle culturally sensitive moves. Teams using LLMs in HR, counseling, or moderation should treat this as an early warning about one size fits all advice.

Suhyun Lee, Wenxuan Zhang

Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs

This work dissects how specific layers and tokens support reasoning inside chain of thought traces. The authors link internal activations to logical operations rather than treating traces as opaque text. If you care about mechanistic interpretability, these tools help pinpoint where models actually reason versus just narrate.

Seogyeong Jeong, Jaehui Hwang

BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference

BeaconKV tracks a small set of "beacon" queries that predict which past tokens reasoning models will revisit later. Keeping only keys and values tied to these beacons shrinks memory use without breaking long chain reasoning. If context costs dominate your bills, this is a concrete alternative to brute force KV caching.

Janghyeon Kim, Minsoo Kim

Extremely Sparse Supervision Incentivizes Reasoning Ability

The paper finds that rewarding only one or two key tokens per reasoning trace can match or beat full token supervision. Reasoning gains hold across models, tasks, and even RL with verifiable rewards. If you run expensive training runs, this challenges the assumption that you must grade every step to get strong reasoning.

Zhishuai Liu, Xingzi Xu

SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations Center

SENTINEL-RL moves graph style attack path reasoning into a dedicated module rather than asking the language model to juggle all structure. The agent queries this module for reachability and risk, which cuts hallucinated attack paths. If you build security copilot systems, this is a blueprint for separating math from prose.

Uday Vallabhaneni, Cassie L. Cagwin

Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents

The authors create CONFLICTGUI, a benchmark where instructions clash with each other or with what is on screen. Many GUI agents keep clicking even when no safe action exists. They propose CONFLICTGUARD to teach agents to stop instead of barging ahead, which is critical if your agent can touch production systems.

Zhaoyuan Huang, Tianjie Ju

DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice Agents

This benchmark stresses voice agents on realistic tasks like when to interrupt, backchannel, or stay silent based only on a role description. Current systems either ignore persona cues or mishandle timing when instructions conflict. Anyone building real time speech agents should treat these test cases as minimum acceptance criteria.

Puneet Mathur, Dinesh Manocha

Dude: A Dual-Detection Multi-Agent System for Paper-Code Discrepancy Detection

Dude pairs two specialized agents to cross check research papers against their code and avoid both missed bugs and false alarms. Careful negotiation between text and code views improves F1 by large margins over single agent baselines. If you rely on public repos and papers, this hints at how automated code review will look.

Weijie Liu, Running Zhao

Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation

The authors show that conversational models often adopt a user's one sided story in moral dilemmas without questioning missing perspectives. Multi turn narration shifts model judgments far beyond single turn baselines. If you ship advisory agents, you need tests for this failure mode and guardrails that demand counter perspectives.

Yuhe Wu, Guangyu Wang

Fresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory

PlanFence forces an agent to check whether the specific facts that justified an action are still valid, not just whether the state is recent. It prevents agents from executing old plans after the world has changed. If you run multi agent systems touching real infrastructure, this is a pattern to copy.

Evan Chen, Shiqiang Wang

Speculative Macro Commit for Faster Tool-Using Agents

Two cooperating models reuse common action patterns so agents can pre-execute multi step tool calls safely. This cuts wall clock latency without hurting success rates. If you are building heavy tool-using agents, treat this as a concrete recipe for speeding them up before you reach for bigger models.

Zeyu Liu, Souvik Kundu

Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge

ElephantBench tests whether language models can hold multiple conflicting real world facts instead of collapsing to a single popular story. Results show even top models often recall only one side, so long tail knowledge and minority views quietly disappear.

Zhuoshi Pan, Junru Lu

ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL

ContextPilot lets long running agents edit their own working memory with planning, long term recall, and smarter compression tools. A custom reinforcement learning scheme focuses training on key context edits, keeping answers strong while shrinking context size on long reasoning tasks.

Zhuoshi Pan, Qizhi Pei

Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning

Physical scenes get encoded as executable code, so the model reasons over objects, forces, and dynamics instead of raw pixels. Those code worlds then supervise a vision language model on precise physics questions, beating strong closed models on quantitative benchmarks.

Hanyang Wang, Yimo Cai

TerraNova: A Foundation Model for the Anthropocene

Builds a single model that jointly learns from global physical data and country-level social indicators, keeping both geometries intact. It can fill in missing climate fields and reason about national metrics, which is a big deal for climate risk, policy modeling, and long-term planning tools.

Carlos Rodriguez-Pardo, Massimo Tavoni