Clockwork.io Raises $31 Million To Expand AI Infrastructure Fault-Tolerance Platform

Clockwork.io has raised $31 million in new funding to expand its fault-tolerance software for artificial intelligence infrastructure, bringing total funding to $73 million.

The round was co-led by Premji Invest, Wing Venture Capital and Seligman Ventures, with participation from existing investors NEA and e& Capital.

Clockwork.io plans to use the financing to accelerate deployment of its fault-tolerance technology across AI training, inference and reinforcement learning, expand enterprise adoption and scale distribution through cloud partners.

The company also announced production deployments at LinkedIn and Together AI and an expanded relationship with WhiteFiber.

Clockwork.io develops software designed to keep large distributed AI workloads running when GPUs, network links or servers fail.

Its LinkPass product reroutes network traffic around failed links, while TorchPass moves workloads from failing GPUs to healthy hardware rather than requiring a job to restart from an earlier checkpoint.

The company introduced two additional TorchPass capabilities alongside the financing.

Multi-node platform snapshots capture the state of a distributed AI job across all of its nodes without requiring changes to the training code, allowing the workload to be restored following a major failure.

Clockwork.io also introduced asynchronous application checkpoints that operate in the background while jobs continue running.

The feature is designed in part to accelerate reinforcement learning by delivering updated model weights to inference replicas more quickly.

LinkedIn has deployed LinkPass across its AI infrastructure fleet and says the technology prevents tens of thousands of GPU-hours of downtime each month.

Together AI plans to offer TorchPass as a service on its GPU Clusters, while WhiteFiber is expanding Clockwork.io across its global GPU-as-a-service infrastructure.

Clockwork.io said its approach addresses a growing infrastructure challenge as AI jobs expand across thousands of GPUs and individual hardware or networking problems can disrupt an entire workload.

KEY QUOTES:

“Failures are inevitable at AI scale. Losing hours of useful work to them should not be. Fault tolerance is a goodput multiplier: it keeps GPUs doing useful work instead of waiting for recovery or repeating work already done. We built our software alongside enterprises and cloud providers operating some of the largest GPU fleets, so it handles the failures they actually see. That protection belongs in the infrastructure enterprises and cloud providers rely on every day.”

Suresh Vasudevan, CEO of Clockwork.io

“At AI infrastructure scale, a single network issue should never sideline healthy GPUs or interrupt running workloads. Before Clockwork.io, one InfiniBand NIC flap could remove an eight-GPU server from service, while a switch port flap could drain a second server, doubling the impact to 16 GPUs. Clockwork.io helped transform that operating model. Its network fault-tolerance technology automatically reroutes traffic onto healthy paths, allowing jobs to continue uninterrupted while link, optic, cable, or NIC faults are repaired. In aggregate, Clockwork.io prevents tens of thousands of GPU-hours of downtime per month across our fleet. By turning what were once disruptive operational incidents into manageable maintenance events, Clockwork.io has helped improve infrastructure utilization and operational efficiency.”

Raghu Hiremagalur, SVP and CTO Infrastructure at LinkedIn

“Our customers grade us on goodput, the share of their GPU-hours that actually move the model forward. Node repair already detects faults and provisions replacement capacity automatically. Clockwork.io’s TorchPass and LinkPass build on that foundation and are designed to keep jobs moving through GPU faults and link failures, preserving progress. We are bringing them to market as the next layer of resilience in the platform.”

Pavneet Ahluwalia, Product Lead at Together AI

“Pressure-testing a cluster’s reliability before it reaches production is critical, because a customer who inherits a hidden fabric fault pays for it later in failed jobs and lost GPU-hours. Marginal optics, misconfigured NICs, and links that pass a basic test but degrade under load can slip through. Clockwork.io’s automated fleet audit validates every link and node at once, localizes faults in minutes, and lets us correct them before acceptance. We bring clusters up faster, and a customer’s first training run lands on a fabric validated end-to-end, not just powered on. With market demand growing as rapidly as it is, getting validated capacity to customers quickly is critical to our business, and it is why we are expanding Clockwork.io across our clusters.”

Tom Sanfilippo, Chief Technology Officer of WhiteFiber

“Cluster fault tolerance used to be a training problem. It is now an inference problem too. In our ClusterMAX, TorchPass cuts training goodput loss from 14% to under 3% for a gold-rated neocloud. Reinforcement Learning (RL) ties the two together: inference replicas generate rollouts, the trainer learns from them, and the updated weights go back to the replicas. Clockwork.io keeps replicas serving through link flaps and network failures. Its extremely fast checkpoints accelerate weight transfer back into the rollout fleet, so neither direction stalls the run. One fault-tolerance layer under training, inference, and RL is where this has to be solved.”

Dylan Patel, Founder, CEO and Chief Analyst at SemiAnalysis

“The one thing that scales perfectly is unreliability: put enough GPUs in one machine and something is always failing. The old playbook: stop the job, reload a checkpoint, makes no sense at today’s scale. Clockwork treats failure as the normal state: TorchPass migrates training off a failing GPU live, and now snapshots an entire running job with no code changes. It’s already saving tens of thousands of GPU-hours a month. We first backed Clockwork.io in 2021 and are thrilled to keep supporting them as they define the performance layer of the AI cluster.”

Greg Papadopoulos, Venture Partner at NEA