ArXiv AI Papers
Latest artificial intelligence and machine learning research papers from ArXiv.
Showing 50 of 258 items
Causal Memory Policy: Making Memory Utility Identifiable by Intervening on Retrieval
Treats an agent’s memory as a causal system and probes it by intervening on what gets retrieved. If you are tuning retrieval or tool memory, this gives a more principled way to measure which memories actually matter.
TopK-Guided: Adaptive, Budget-Aware Activation Sparsity for Efficient LLM Inference
Introduces a method that activates only the most useful neurons in each layer, under a fixed compute budget. If you care about cutting the cost of running the AI without throwing away too much quality, this is directly actionable.
VideoEvolve: Evolving Agent Harnesses for Video Temporal Grounding
Optimizes the agent harness around video models that need to find events in long clips. If you are building video agents, this shows that harness design and curriculum matter as much as the base model.
Not All Experience Belongs in the Weights: Component Routing for Self-Improving GUI Agents
Shows a way for GUI agents to improve by routing experience into modular components instead of always changing core weights. That matters if you want agents that get better on real apps without constant retraining runs.
Continuous Process-Level Evaluation for Evolving Enterprise AI Agent Skills
Proposes always on evaluation of enterprise agents at the business process level, not just individual prompts. If you run pilots inside big companies, this gives a structure for tracking whether agents actually improve operations over time.
Code Owns the Simulation, Jev Owns the Evaluation
Argues that you should treat evaluation code as a first class object, separate from the environment being simulated. Useful for anyone designing complex agent benches, where leakiness between sim code and scoring quietly corrupts results.
Can AI Oversight Be Zero Knowledge?
Applies zero knowledge ideas so overseers can check model behavior without seeing sensitive inputs or outputs. If you work on safety for regulated sectors, this sketches how to watch powerful models without breaking privacy rules.
Mingbird: A Local-First Agent Harness Enabling Small Open Models to Complete Real Tasks
Introduces an agent harness tuned for small open models running on local hardware. If you want agents that respect data boundaries and still finish end to end tasks, this paper is worth copying from.
Mimir: Physics-Grounded LLM Agents for Long-Horizon Irrigation Control
Combines physics based simulators with language model agents to control farm irrigation over long periods. It is a concrete template for tying agents to physical models so they do not drift into unphysical decisions.
External Observers May See More Clearly: Cross-Model Span-Level Hallucination Detection in Large Language Models via Hidden State Probing
Uses a separate probe model to watch another model’s hidden states and flag hallucinated spans. If you ship high stakes apps, this is a concrete recipe for wrapping your main model with an independent fact checker.
VISTA: A Visual Harness for Reasoning in an Interactive World
Shows that a smart vision harness can unlock long horizon reasoning in existing text image models without retraining them. If you build embodied agents or robotics stacks, this is a blueprint for wrapping foundation models instead of training new ones from scratch.
ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research
Defines a benchmark that scores search systems on whether they surface papers that actually spark new follow up work, not just high citation counts. Useful if you build research agents or citation tools and need a target beyond vanilla relevance and click metrics.
Emergent Intelligence: Resonant Oscillators Produce Proactive Adaptive Behavior
The authors show that a tiny network of coupled oscillators, with no training, can explore, adapt, and change behavior based on feedback. It is a toy but suggests new foundations for "curious" agents.
The Weight Is Over - Interactive Diffusion on Consumer GPUs
This work refactors diffusion pipelines so mid-range GPUs can handle image generation interactively. They mix model surgery, compression, and scheduling to keep quality while shrinking memory and latency.
LLMs as Feature Engineers for Text-and-Tabular Prediction
They use one model to propose human-readable categorical features from text, and another to extract them as table columns. A standard tabular model then scores their value and drives an iterative search.
Detecting Pretraining Data in Large Language Models from a Free-Energy Perspective
The authors argue you should judge membership by loss relative to uncertainty, not loss alone. Their "free-energy" style detector separates memorized data from truly predictable text much more cleanly.
ExpBoN: Exponential-Noise Best-of-n for Efficient Test-Time LLM Alignment
The paper introduces ExpBoN, a softer version of "best-of-n" sampling that adds exponential noise so you can dial reward gains against distribution shift. It greatly cuts the number of samples needed for strong test-time alignment, making long "thinking" runs more affordable.
Can 4D Foundation Models Remember?
The authors probe video and 4D models to see how well they retain scene details across time and camera motion. They find serious memory gaps, especially for small or briefly seen objects.
Stress-testing Alignment Midtraining
Alignment midtraining mixes safety-relevant documents into pretraining instead of relying only on post-training. This paper builds tests showing where that approach helps, and where dangerous behaviors still sneak through.
JEPA-Anything: Learning Predictive Models across Different Worlds
JEPA-Anything trains a single predictive model across very different domains by factorizing what changes and what stays stable. It shows that one learning recipe can cover physics, games, and more.
Refuse, Decompose, Refresh: A Claim-Safe Protocol for Closed-Loop AI Evaluation
They propose a three-step evaluation protocol so closed-loop agent tests do not oversell claims. The method flags when data cannot support a claim, then decomposes and re-runs tests until it can.
The Organization of Inference: Information, Resource Constraints, and AI Production
Using controlled coding workflows, the authors treat AI systems like factories with planning and execution stages. They quantify how shifting token budgets and information between stages changes throughput and reliability.
How Do Agent Harnesses Create Value? Planning Information and Release Control in Stateful LLM Agents
This paper A/B tests real enterprise agents with and without meaningful plans and release controls. It shows structured guidance and staged rollout sharply cut silent failures at similar headline success rates.
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness
SoL-Pi automates experiments on agent harness designs across many environments and models, then feeds the findings back into the harness. It improves coding-agent performance while cutting token and dollar costs.
An Empirical Study of Harness Design for Coding Agents
They vary planning, tools, and context handling across 176 harness configurations for coding agents. Results show context strategy matters most under tight budgets, while planning mainly shifts cost, not accuracy.
AgentPProf: Semantic Profiler for Long Horizon AI Agents
AgentPProf logs agent behavior like a performance profiler, but in semantic space instead of CPU cycles. It helps teams see which tasks burn tokens, fail often, or cause risky side effects.
SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking
SAS learns which tokens matter most by feeding a learned gate directly into the attention scores. That lets gradients flow from the language loss into the selector. If you care about cheaper long context LLMs, this offers a practical recipe beyond hand tuned pruning tricks.
Rethinking Heterogeneous System Disaggregation for Subquadratic Attention
They design a new way to split LLM workloads across mixed hardware by treating subquadratic attention separately from dense layers. Their SQD scheme boosts tokens per joule on future GPU plus memory systems and tightens latency under fixed power. If you run large models on weird new hardware, this gives concrete design and deployment ideas.
Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models
They build TAM, a benchmark where models must follow real manuals for medical coding and criminal sentencing. Even strong models barely solve any tasks end to end. If you trust “multi hop reasoning” scores, this is a sharp reminder that long rule books are still a brick wall.
K-Bench: A Benchmark for LLM Unlearning in Agentic Deployments
They show that “unlearned” secrets still leak once the model runs as a tool using agent. K‑Bench checks every channel an agent exposes, not just the final answer. This is a must read if you plan to forget training data or user prompts while still keeping agents useful.
EduFair-Bench: Evaluating Pedagogical Fairness of LLM Tutors Across Student Demographics
They simulate students with different gender, language, immigration background, and income, then measure how LLM tutors treat them. The benchmark scores both teaching quality and systematic gaps. If you ship AI tutors or classroom tools, this is a template for checking who your model quietly leaves behind.
How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks
Physicists regrade several headline physics benchmarks and discover many supposed model failures were actually bad keys and ambiguous questions. After fixing tests, GPT‑5.6 Sol and peers nearly max out multiple physics suites. Anyone tracking “models still fail physics” needs to revisit those claims and pay more attention to benchmark design.
Autonomous Research for Open-Ended Problems: A Case Study on Telecom Ticket Retrieval
An LLM based agent iteratively designs, trains, and evaluates models to solve messy telecom support retrieval tasks. It operates over large search spaces where classic automated ML tools struggle. Use this as a concrete reference when arguing about how far “automated research interns” already go on real business problems.
Embodied-BenchForge: A Closed-Loop Agentic Workflow for Embodied Benchmark Construction
They build an agent pipeline that turns a high level evaluation idea into a full robot or embodied benchmark. The system both generates tasks and automatically checks and repairs flawed intermediate assets. If you care about realistic tests for robot agents, this is a blueprint for automating that grind.
Post-Training Language Models for Gold-Medal Performance in Coding Competitions
NVIDIA fine tunes Nemotron models on curated programming problems plus synthetic reasoning traces, then layers a test time solve refine loop. Their system beats the top human at IOI 2026 under contest rules. If you use models for serious coding, this shows how much is still on the table with better post training and runtime strategies.
CC-Mediation: Evaluating Large Language Models for Cross-Cultural Conflict Mediation
The authors build a benchmark where mediators must respect different cultural norms while resolving conflicts. Models often import biases from their training data and mishandle culturally sensitive moves. Teams using LLMs in HR, counseling, or moderation should treat this as an early warning about one size fits all advice.
Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs
This work dissects how specific layers and tokens support reasoning inside chain of thought traces. The authors link internal activations to logical operations rather than treating traces as opaque text. If you care about mechanistic interpretability, these tools help pinpoint where models actually reason versus just narrate.
BeaconKV: Key-Value Cache Compression Guided by Beacon Queries for Efficient Large Reasoning Model Inference
BeaconKV tracks a small set of "beacon" queries that predict which past tokens reasoning models will revisit later. Keeping only keys and values tied to these beacons shrinks memory use without breaking long chain reasoning. If context costs dominate your bills, this is a concrete alternative to brute force KV caching.
Extremely Sparse Supervision Incentivizes Reasoning Ability
The paper finds that rewarding only one or two key tokens per reasoning trace can match or beat full token supervision. Reasoning gains hold across models, tasks, and even RL with verifiable rewards. If you run expensive training runs, this challenges the assumption that you must grade every step to get strong reasoning.
SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations Center
SENTINEL-RL moves graph style attack path reasoning into a dedicated module rather than asking the language model to juggle all structure. The agent queries this module for reachability and risk, which cuts hallucinated attack paths. If you build security copilot systems, this is a blueprint for separating math from prose.
Do GUI Agents Know When Not to Act? Enabling Conflict-Aware Termination for Multimodal GUI Agents
The authors create CONFLICTGUI, a benchmark where instructions clash with each other or with what is on screen. Many GUI agents keep clicking even when no safe action exists. They propose CONFLICTGUARD to teach agents to stop instead of barging ahead, which is critical if your agent can touch production systems.
DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice Agents
This benchmark stresses voice agents on realistic tasks like when to interrupt, backchannel, or stay silent based only on a role description. Current systems either ignore persona cues or mishandle timing when instructions conflict. Anyone building real time speech agents should treat these test cases as minimum acceptance criteria.
Dude: A Dual-Detection Multi-Agent System for Paper-Code Discrepancy Detection
Dude pairs two specialized agents to cross check research papers against their code and avoid both missed bugs and false alarms. Careful negotiation between text and code views improves F1 by large margins over single agent baselines. If you rely on public repos and papers, this hints at how automated code review will look.
Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation
The authors show that conversational models often adopt a user's one sided story in moral dilemmas without questioning missing perspectives. Multi turn narration shifts model judgments far beyond single turn baselines. If you ship advisory agents, you need tests for this failure mode and guardrails that demand counter perspectives.
Fresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory
PlanFence forces an agent to check whether the specific facts that justified an action are still valid, not just whether the state is recent. It prevents agents from executing old plans after the world has changed. If you run multi agent systems touching real infrastructure, this is a pattern to copy.
Speculative Macro Commit for Faster Tool-Using Agents
Two cooperating models reuse common action patterns so agents can pre-execute multi step tool calls safely. This cuts wall clock latency without hurting success rates. If you are building heavy tool-using agents, treat this as a concrete recipe for speeding them up before you reach for bigger models.
Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge
ElephantBench tests whether language models can hold multiple conflicting real world facts instead of collapsing to a single popular story. Results show even top models often recall only one side, so long tail knowledge and minority views quietly disappear.
ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL
ContextPilot lets long running agents edit their own working memory with planning, long term recall, and smarter compression tools. A custom reinforcement learning scheme focuses training on key context edits, keeping answers strong while shrinking context size on long reasoning tasks.
Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning
Physical scenes get encoded as executable code, so the model reasons over objects, forces, and dynamics instead of raw pixels. Those code worlds then supervise a vision language model on precise physics questions, beating strong closed models on quantitative benchmarks.
TerraNova: A Foundation Model for the Anthropocene
Builds a single model that jointly learns from global physical data and country-level social indicators, keeping both geometries intact. It can fill in missing climate fields and reason about national metrics, which is a big deal for climate risk, policy modeling, and long-term planning tools.