Back to Frontiers

Memory & Continual Learning

Steady48%

Long-context understanding, persistent memory, RAG systems, and lifelong learning. Giving AI the ability to remember and learn continuously.

long-contextRAGepisodic-memorycontinual-learningretrievalmemory
120
Papers
12
Milestones
$0
Funding
2
Benchmarks

Key Benchmarks

RULER

Long-context benchmark testing retrieval and reasoning over 128K+ tokens

91.1%Human: 98%
Leader: NVIDIA Nemotron 3 Nano 4Bmedium saturation

InfiniteBench

Ultra-long context benchmark testing 100K+ token understanding

44%Human: 95%
Leader: Llama-3.1-8Blow saturation

Recent Papers

Recent Milestones

Claude Haiku 5.5 slashes small‑model costs

Anthropic released Claude Haiku 5.5 on October 7, 2026, as its fastest and cheapest small model in the Claude 5.5 family. The model offers a 1M token context window, multimodal support and around 75 to 90 percent lower prices than Haiku 4.5 for most workloads, and is now live on Anthropic’s platform and major clouds.

Oct 7, 2026releaseImpact: 70/100

Kolibri: EU sovereign 78B open model lands

On October 4, 2026 The Frontier detailed Aleph Alpha’s Kolibri-1, a 78.1 billion parameter German English mixture of experts model whose open weights were released on Hugging Face under an Apache 2.0 license. The analysis notes Kolibri’s 3.46 billion active parameters per token, up to 1 million token validated context window and positioning for sovereign European deployments.

Oct 4, 2026releaseImpact: 80/100

Tencent opens 770B Hy4 with 1M context

Tencent’s Hunyuan team released and open sourced the Hy4 preview large language model on August 28, 2026 and global tech outlets published detailed breakdowns on August 30, 2026. Hy4 uses a 770 billion parameter mixture of experts architecture with 49 billion active parameters and a context window of over 1 million tokens, with weights released under an Apache 2.0 style license.

Aug 30, 2026releaseImpact: 90/100

WikiSkill: small agents beat bigger models with memory

A Google Research and Virginia Tech team has introduced WikiSkill, a framework where AI agents log their successes and failures into a wiki‑like knowledge base and evolve reusable skills over time. In experiments, a Qwen‑3.5‑9B model equipped with WikiSkill‑learned skills outperformed a larger Qwen‑3.6‑27B model without skills across multiple benchmarks. ([xenospectrum.com](https://xenospectrum.com/en/google-wikiskill-agent-memory/))

Aug 30, 2026breakthroughImpact: 70/100

EVAF Adds Durable Goals to Long‑Running Agents

On June 25, 2026, a preprint by Haoliang Han introduced EVAF, a gated LoRA-based consolidation mechanism that writes long-term goals into a small parametric store so agents retain behavior even after context is cleared. On June 28, 24 AI’s "Today in AI" digest spotlighted the work as a key advance in persistent memory for long-running language agents. ([arxiv.org](https://arxiv.org/abs/2606.26806?utm_source=openai))

Jun 25, 2026paperImpact: 70/100

GLM‑5.2: 1M‑Token Open Coding Model

On June 13, 2026, Zhipu AI’s international brand Z.ai rolled out its GLM‑5.2 model to all GLM Coding Plan users, featuring a 1‑million‑token context window and new ‘High’ and ‘Max’ reasoning modes. The company says an API and MIT‑licensed open‑weight release will follow next week, positioning GLM‑5.2 as its most capable open model for long‑horizon coding and agents.

Jun 13, 2026releaseImpact: 80/100

ChatGPT gets scalable long‑term memory

On June 4, 2026, OpenAI detailed a new "dreaming"-based memory system for ChatGPT designed to synthesize and refresh long‑term user memories at scale. The rollout aims to improve how ChatGPT recalls user preferences, projects and context across multi‑year interactions.

Jun 4, 2026releaseImpact: 80/100

NVIDIA Context Memory hits mainstream servers

On May 26, 2026, AIC announced via PRNewswire that it will showcase new AI storage and compute platforms and co-host a ‘Breaking the Memory Wall’ panel with NVIDIA and VAST Data at Computex 2026 in Taipei. The company will demonstrate systems built around NVIDIA’s Context Memory Platform, BlueField-4 and high-density GPU servers aimed at long-context LLMs, video analytics, and agentic AI workloads.

May 26, 2026releaseImpact: 70/100

Gemini 1.5 Pro 2M Context

Google expands Gemini 1.5 Pro to 2 million token context window.

May 14, 2024benchmarkImpact: 88/100

GPT-4 Turbo 128K Context

OpenAI expands GPT-4 Turbo to 128K tokens with improved retrieval over long documents.

Apr 9, 2024releaseImpact: 78/100

Claude 3 200K Context

Anthropic releases Claude 3 with 200K context window and improved long-context performance.

Mar 4, 2024releaseImpact: 85/100

Gemini 1.5 Pro 1M Context Window

Google releases Gemini 1.5 Pro with 1 million token context window, 10x previous limits.

Feb 15, 2024releaseImpact: 92/100

Leading Organizations

OpenAI
Anthropic
Google
Meta
Cohere

ArXiv Categories

cs.LGcs.AIcs.CLcs.IR

Related Frontiers