Back to Frontiers

Foundation Models & Reasoning

Steady42%

Core model architectures, training methods, chain-of-thought reasoning, and test-time compute scaling. The backbone of modern AI capabilities.

transformersscaling lawschain-of-thoughto1reasoningtest-time computeworld models
250
Papers
183
Milestones
$97.2B
Funding
3
Benchmarks

Key Benchmarks

GPQA Diamond

Graduate-level science questions requiring PhD-level expertise

95.5%Human: 70%
Leader: Fuguhigh saturation

MMLU-Pro

Massive Multitask Language Understanding - Pro version with 10 answer choices and harder reasoning

91.2%Human: 89%
Leader: Gemini 3.1 Pro (tied with Llama-3.1-Nemotron-70B-Instruct-HF)high saturation

HLE (Humanity's Last Exam)

2,500 questions at the frontier of human knowledge across 100+ subjects

46.4%Human: 95%
Leader: Gemini 3.1 Pro Preview (thinking high)low saturation

Recent Papers

AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers

Tianyu Huai, Tingshuo Fan, Xinchi Chen +5 more

Jul 31, 2026ArXivPDF

DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat

Ismayil Ismayilov, Atakan Kara, Kaan Oktay

Jul 31, 2026ArXivPDF

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning

Luca Viano, Antoine Moulin, Audrey Huang +3 more

Jul 31, 2026ArXivPDF

The Parts Are Greater Than the Sum: Automated Task Sequencing for Efficient Training of Multi-Policy LLMs

Jiajia Tang, Sizhe Yuen, Francisco Gomez Medina +2 more

Jul 31, 2026ArXivPDF

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback

Manith Adikari, Bei Peng, Samuele Vinanzi +1 more

Jul 31, 2026ArXivPDF

A Human-Centered Validation of the Explainability-Performance Coefficient

Christian Oliva, Luis F. Lago-Fernández

Jul 31, 2026ArXivPDF

TerraNova: A Foundation Model for the Anthropocene

Carlos Rodriguez-Pardo, Massimo Tavoni

Jul 31, 2026ArXivPDF

MOT-SR: Multi-Objective Tool-Augmented Scientific Equation Discovery with Large Language Models

Boxiao Wang, Runxiang Wang, Kai Li +4 more

Jul 31, 2026ArXivPDF

AREX: Towards a Recursively Self-Improving Agent for Deep Research

Shuqi Lu, Chaofan Li, Kun Luo +21 more

Jul 24, 2026ArXivPDF

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation

Junsong Chen, Jincheng Yu, Yitong Li +11 more

Jul 24, 2026HuggingFacePDF

Recent Milestones

DeepMind Shakeup, New Discovery Loop Lab

On August 5, 2026, Google DeepMind cofounder Demis Hassabis said he is stepping down as CEO to become chair of DeepMind and chief scientist at Alphabet. Google DeepMind CTO Koray Kavukcuoglu will run the lab’s day-to-day AI work, while longtime Google chief scientist Jeff Dean and colleague Sanjay Ghemawat are leaving to launch a new AI startup called Discovery Loop.

Aug 5, 2026fundingImpact: 70/100

Anthropic inks $10B Volta deal for frontier compute

TechCrunch reported on August 4, 2026 that Anthropic has agreed a roughly 10 billion dollar, six‑year compute deal with AI cloud startup Volta. Volta, working with Bitdeer and Nvidia hardware in a 133 megawatt Norway data center, will provide cloud capacity to support Anthropic’s Claude models.

Aug 4, 2026fundingImpact: 90/100

EU backs 7 AI gigafactories with €10B

On July 30, 2026, the European Union announced a 10 billion euro public funding package to support seven AI gigafactories across the bloc. The European Commission says the sites will each house at least 100,000 advanced AI chips and more than double the EU’s current AI computing capacity. The program is designed to catalyze a further 20 billion euros in private investment.

Jul 30, 2026fundingImpact: 80/100

Claude Mythos hits expert level cryptanalysis

Anthropic reported on July 28, 2026 that its unreleased Claude Mythos Preview model found improved attacks on the HAWK post‑quantum signature scheme and a reduced‑round version of AES. The company says the AI‑assisted work halves HAWK’s effective key strength and speeds a known 7‑round AES attack by 200 to 800 times, though no production systems are affected.

Jul 29, 2026breakthroughImpact: 90/100

Recursive Superintelligence signs $410M AWS deal

On July 28, 2026, Recursive Superintelligence said it had signed a multi‑year compute agreement with Amazon Web Services worth about $410 million to power its self‑improving AI research. Techmeme and multiple briefings report that Amazon confirmed the deal, which represents the bulk of Recursive’s funding to date and secures priority access to AWS’s latest high‑end instances.

Jul 28, 2026fundingImpact: 80/100

France opens IPCEI to fund industrial AI stack

On July 28, 2026, the French government opened a call for projects under the new IPCEI AI program to support large industrial AI initiatives across Europe. The scheme, tied to the France 2030 strategy, targets foundation models, AI cloud, energy efficiency, secure data access and sectoral AI deployments.

Jul 28, 2026fundingImpact: 70/100

Moonshot opens 2.8T param Kimi K3 weights

Chinese lab Moonshot AI released open weights, a technical report and key infra components for its 2.8 trillion parameter Kimi K3 model on July 27, 2026. The company published the model on Hugging Face and GitHub under a Kimi K3 specific license, making what is likely the largest open weight model to date available for external researchers and platforms.([ithome.com](https://www.ithome.com/0/982/259.htm))

Jul 27, 2026releaseImpact: 90/100

Celeris-1 diffusion LLM targets real-time AGI

San Francisco startup Celeris announced Celeris-1 on July 27, 2026, a diffusion-based large language model that aims to match near GPT-5 performance while delivering up to 15 times faster responses. The company says Celeris-1 can generate around 1,664 tokens per second and is available immediately through an OpenAI-compatible API for latency sensitive workloads like agents and real time voice.([prnewswire.com](https://www.prnewswire.com/news-releases/celeris-unveils-celeris-1-unlocking-real-time-ai-through-diffusion-based-language-generation-302835273.html))

Jul 27, 2026releaseImpact: 80/100

Kimi K3: Open 2.8T model with 1M context

On July 27, 2026 multiple technical blogs reported that Moonshot AI has released free public download weights for its Kimi K3 model, a 2.8 trillion parameter open weight system with a 1 million token context window. The weights reportedly went live late on July 26 US time, fulfilling Moonshot’s promise to open the model by July 27 after launching it via API on July 16.

Jul 27, 2026releaseImpact: 90/100

Chinese frontier models crash global AI prices

Fortune reports that Chinese labs Moonshot AI, Z.AI and DeepSeek are now offering frontier‑level models like Kimi K3 and GLM‑5.2 at a fraction of U.S. competitors’ prices, with some tokens priced under 2 percent of Anthropic’s Fable. The story, published July 26, 2026 at 5:00 PM ET, details how these models are gaining traction with U.S. developers and enterprises despite export controls. ([fortune.com](https://fortune.com/2026/07/26/china-moonshot-deepseek-zai-kimi-challenging-us-ai-cost/))

Jul 26, 2026releaseImpact: 90/100

Kimi K3: giant Chinese open model goes global

Chinese startup Moonshot AI’s Kimi K3, a massive open‑weight model launched in mid‑July, is rapidly gaining global users and is being adopted by US developers drawn to its low cost and strong performance. An AP feature on July 26, 2026 highlights how Kimi K3 and other Chinese open models are winning users like Mozilla CTO Raffi Krikorian and pushing firms such as Alibaba to roll out increasingly capable systems, putting pressure on OpenAI and Anthropic in key workloads.

Jul 26, 2026releaseImpact: 80/100

Korea ties $950B tech deals to AI-era push

At 07:02 KST on July 26, 2026, Yonhap reported that President Lee Jae Myung told Korean expatriates in San Francisco that the AI era is a “truly new opportunity” for South Korea. He highlighted recently announced business agreements between Nvidia, OpenAI, Anthropic and Korean firms such as Samsung Electronics and SK hynix, which a presidential official valued at about 950 billion dollars.

Jul 25, 2026fundingImpact: 80/100

DeepSeek pauses $1.5B after viral founder remarks

Chinese AI startup DeepSeek has suspended its second fundraising round, in which it aimed to raise at least 10 billion yuan, after comments attributed to founder Liang Wenfeng about US–China AI competition went viral. Bloomberg, via Business Standard, reported on July 25 that DeepSeek informed prospective investors the deal is on hold, even as it considers an IPO and emphasizes a long term push toward AGI with open source models.([business-standard.com](https://www.business-standard.com/world-news/deepseek-puts-fundraising-on-hold-after-founder-s-remarks-go-viral-report-126072501034_1.html))

Jul 25, 2026fundingImpact: 70/100

Claude Opus 5 nears frontier at half the price

On July 25, 2026, Swedish outlet Omni reported that Anthropic has launched Claude Opus 5, its fourth new model in two months, describing it as at least as capable as Claude Fable 5 at roughly half the cost. Independent tracking sites note that Opus 5, released July 24, keeps Opus tier pricing at $5 per million input tokens and $25 per million output tokens while moving closer to frontier benchmarks.

Jul 25, 2026releaseImpact: 80/100

Claude Opus 5: near‑Fable power at half cost

Anthropic released its Claude Opus 5 model on July 24, 2026, positioning it as a near‑Fable‑level system for complex coding, research and knowledge work at roughly half the price of Fable 5. The model is now the default for Claude Max and the top option for Claude Pro, and it is available via Anthropic’s API as well as on Amazon Bedrock and Claude Platform on AWS.

Jul 24, 2026releaseImpact: 80/100

DeepSeek goes AGI-first with agents and lifelong learning

On July 23, 2026, TechNode reported that Chinese AI lab DeepSeek is prioritizing artificial general intelligence research over short‑term product or revenue goals, based on a leaked four‑hour investor meeting transcript. Founder Liang Wenfeng reportedly told investors that open‑source reasoning models, coding agents and continual learning are the company’s focus, with commercial APIs mainly funding long‑term AGI work.

Jul 23, 2026fundingImpact: 70/100

Chinese open‑weight models reach frontier tier

On July 23, 2026, Colombian outlet Semana, citing experts interviewed by Deutsche Welle, reported that Chinese AI firms like DeepSeek, Zhipu and Moonshot AI are closing the performance gap with US models through open weight systems such as GLM 5.2 and Kimi K3. The article ties these advances to messages from China’s leadership at the World Artificial Intelligence Conference in Shanghai and the launch of a China led World AI Cooperation Organisation.

Jul 23, 2026breakthroughImpact: 80/100

AMD, Anthropic ink multi‑GW GPU and $5B pact

On July 23, 2026, a Reuters report said AMD is set to showcase new AI hardware, including its first Helios server racks and Venice data center CPUs, at an event in San Francisco. The article also noted AMD’s plan, announced earlier in the week, to sell up to 2 GW of Instinct MI450 chips to Anthropic and invest up to 5 billion dollars in the Claude maker starting in 2027.

Jul 23, 2026fundingImpact: 90/100

India funds 20 sovereign AI models at scale

On July 23, 2026, India’s IT minister Ashwini Vaishnaw told Parliament that the government has selected 20 “sovereign” Indian AI model projects for support under the IndiaAI Mission. The portfolio includes large language models from Sarvam AI, speech models from Gnani.ai, multilingual models from BharatGen and video generation work from Avataar AI, alongside expanded compute and skilling programmes.

Jul 23, 2026fundingImpact: 70/100

HKGAI Plans Global Push for V3 City Models

On July 21, 2026, Hong Kong’s generative AI research center HKGAI announced a “large-model go-abroad strategy” at WAIC 2026 in Shanghai. The plan aims to export Hong Kong-developed V3 large models and AI “city solutions” globally, leveraging Hong Kong as a bridge between mainland AI research and overseas markets.

Jul 21, 2026fundingImpact: 70/100

Leading Organizations

OpenAI
DeepMind
Anthropic
Meta

ArXiv Categories

cs.LGcs.AIcs.CL

Related Frontiers