Memory & Continual Learning
Long-context understanding, persistent memory, RAG systems, and lifelong learning. Giving AI the ability to remember and learn continuously.
Key Benchmarks
Recent Papers
Causal Memory Policy: Making Memory Utility Identifiable by Intervening on Retrieval
Arman Behnam, Binghui Wang
ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research
Sohyeon Kim, Yoonho Lee, Bo Liu +11 more
Not All Experience Belongs in the Weights: Component Routing for Self-Improving GUI Agents
Beining Wu, Zihao Ding, Jun Huang
OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction
Xiangyu Zeng, Yuandong Yang, Zhiqiu Zhang +13 more
World Observer: Joint Actor-Observer Generation for Persistent World Modeling
Hyunwook Choi, Dahyun Chung, Hyunsung Kim +4 more
Can 4D Foundation Models Remember?
Guangzhao He, Hadar Averbuch-Elor, Wei-Chiu Ma
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression
DeepSeek-AI, Anyi Xu, B. Li +2 more
SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking
Zhiwei Li, Lei Zhu, Hao Gu +6 more
Rethinking Heterogeneous System Disaggregation for Subquadratic Attention
Arya Tschand, Yaosheng Fu, Vikram Sharma Mailthody +6 more
Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models
Utkarsh Soni, Syed Shariyar Murtaza, Yifan Nie +2 more
Recent Milestones
Claude Haiku 5.5 slashes small‑model costs
Anthropic released Claude Haiku 5.5 on October 7, 2026, as its fastest and cheapest small model in the Claude 5.5 family. The model offers a 1M token context window, multimodal support and around 75 to 90 percent lower prices than Haiku 4.5 for most workloads, and is now live on Anthropic’s platform and major clouds.
Kolibri: EU sovereign 78B open model lands
On October 4, 2026 The Frontier detailed Aleph Alpha’s Kolibri-1, a 78.1 billion parameter German English mixture of experts model whose open weights were released on Hugging Face under an Apache 2.0 license. The analysis notes Kolibri’s 3.46 billion active parameters per token, up to 1 million token validated context window and positioning for sovereign European deployments.
Tencent opens 770B Hy4 with 1M context
Tencent’s Hunyuan team released and open sourced the Hy4 preview large language model on August 28, 2026 and global tech outlets published detailed breakdowns on August 30, 2026. Hy4 uses a 770 billion parameter mixture of experts architecture with 49 billion active parameters and a context window of over 1 million tokens, with weights released under an Apache 2.0 style license.
WikiSkill: small agents beat bigger models with memory
A Google Research and Virginia Tech team has introduced WikiSkill, a framework where AI agents log their successes and failures into a wiki‑like knowledge base and evolve reusable skills over time. In experiments, a Qwen‑3.5‑9B model equipped with WikiSkill‑learned skills outperformed a larger Qwen‑3.6‑27B model without skills across multiple benchmarks. ([xenospectrum.com](https://xenospectrum.com/en/google-wikiskill-agent-memory/))
EVAF Adds Durable Goals to Long‑Running Agents
On June 25, 2026, a preprint by Haoliang Han introduced EVAF, a gated LoRA-based consolidation mechanism that writes long-term goals into a small parametric store so agents retain behavior even after context is cleared. On June 28, 24 AI’s "Today in AI" digest spotlighted the work as a key advance in persistent memory for long-running language agents. ([arxiv.org](https://arxiv.org/abs/2606.26806?utm_source=openai))
GLM‑5.2: 1M‑Token Open Coding Model
On June 13, 2026, Zhipu AI’s international brand Z.ai rolled out its GLM‑5.2 model to all GLM Coding Plan users, featuring a 1‑million‑token context window and new ‘High’ and ‘Max’ reasoning modes. The company says an API and MIT‑licensed open‑weight release will follow next week, positioning GLM‑5.2 as its most capable open model for long‑horizon coding and agents.
ChatGPT gets scalable long‑term memory
On June 4, 2026, OpenAI detailed a new "dreaming"-based memory system for ChatGPT designed to synthesize and refresh long‑term user memories at scale. The rollout aims to improve how ChatGPT recalls user preferences, projects and context across multi‑year interactions.
NVIDIA Context Memory hits mainstream servers
On May 26, 2026, AIC announced via PRNewswire that it will showcase new AI storage and compute platforms and co-host a ‘Breaking the Memory Wall’ panel with NVIDIA and VAST Data at Computex 2026 in Taipei. The company will demonstrate systems built around NVIDIA’s Context Memory Platform, BlueField-4 and high-density GPU servers aimed at long-context LLMs, video analytics, and agentic AI workloads.
Gemini 1.5 Pro 2M Context
Google expands Gemini 1.5 Pro to 2 million token context window.
GPT-4 Turbo 128K Context
OpenAI expands GPT-4 Turbo to 128K tokens with improved retrieval over long documents.
Claude 3 200K Context
Anthropic releases Claude 3 with 200K context window and improved long-context performance.
Gemini 1.5 Pro 1M Context Window
Google releases Gemini 1.5 Pro with 1 million token context window, 10x previous limits.