Clockwork.io announced a $31 million funding round on October 5, 2026, to expand its fault-tolerance software for large AI GPU clusters. LinkedIn, Together AI and WhiteFiber are already using its LinkPass and TorchPass tools to prevent tens of thousands of GPU-hours of downtime each month.
This article aggregates reporting from 5 news sources. The TL;DR is AI-generated from original reporting. Race to AGI's analysis provides editorial context on implications for AGI development.
This round is part of a broader shift from “more GPUs” to “better used GPUs.” Training and serving frontier‑scale models has exposed how fragile large clusters are: a single failed link or GPU can idle thousands of chips and waste hours of work. Clockwork sits squarely in that bottleneck, offering a software layer that treats failure as normal and keeps workloads running via live migration, network failover and now TorchSnap distributed checkpointing.([techfundingnews.com](https://techfundingnews.com/clockwork-io-lands-31m-to-eliminate-wasted-gpu-hours-with-ai-infrastructure-fault-tolerance/))
For the race to AGI, this kind of plumbing matters more than it looks. As clusters move from thousands to tens of thousands of accelerators, fault‑tolerance stops being a nice‑to‑have and becomes a gating factor on how fast and cheaply labs can push model scale and RL experiments. Every hour lost to restarts is an hour not spent on new training runs or safety evaluations. Tools that turn unreliable hardware into dependable “AI factories” effectively lower the real cost of compute, which strengthens the business case for ever larger runs.
The customer list is another signal. When LinkedIn, Together AI and neoclouds like WhiteFiber standardize on a resilience layer, it nudges the ecosystem toward a de facto expectation that serious AI clusters will bake in fault tolerance. That, in turn, will influence how hyperscalers and sovereign AI projects architect their next-generation infrastructure.


