Spot GPU Interruption Rates Are Low Enough Now: How Heartbeat-Based Checkpoint and Resume Makes Preemptible Instances Viable for Agent Workloads
Spot GPU interruption rates are under 5%/day. Learn how checkpoint-and-resume lets preemptible instances cut costs 50–90% for agent workloads.
TL;DR: Spot GPU interruption rates have dropped low enough, often under 5% per day on major clouds, that preemptible instances are now genuinely viable for background agent workloads. Heartbeat-based checkpoint-and-resume via durable workflow engines eliminates the operational risk by automatically replaying interrupted tasks from their last saved state. AWS Spot can reduce compute costs by 50–90% compared to on-demand, a spread that now outweighs the engineering overhead for most long-running, interruption-tolerant AI jobs.
Your team is paying $9 per GPU-hour out of fear of an interruption that, if it happened, might cost you under two minutes of rework. Here is what the actual math says.
Key takeawaysH100 spot instances trade as low as $1.35/GPU-hour against on-demand rates of $7–$13, a substantial reduction for teams willing to build around interruptions.Spot GPU interruption rates in 2026 are low enough to make spot the sound operational choice for most non-user-facing workloads.Agent workloads need task-boundary checkpoints of 30–120 seconds, not multi-hour training intervals.Real-time and user-facing inference don't belong on spot; batch eval, data enrichment, and async agents do.Cross-cloud routing hedges against localized capacity crunches.
Introduction
Most ML platform teams running background agent workloads still pay on-demand GPU rates. Not because spot is unavailable, but because interruptions once felt catastrophic. That calculus has changed. H100 SXM5 spot capacity now trades at $1.35/GPU-hour while on-demand sits at $7–$13. The interruption problem has not disappeared; it has been re-engineered around.
What are actual spot GPU interruption rates on AWS, GCP, and Azure in 2026?
In 2026, spot GPU interruption rates run from under 5% per day on GCP's best regions to roughly 7–8% of requests per week on AWS and GCP at standard price points, low enough to make spot operationally sound for non-user-facing workloads.
GCP's a2-highgpu-1g hits under 5% per day in europe-west4. AWS terminates roughly 7.2% of spot GPU requests per week at the $3.52 price point; GCP preemptible VMs show roughly 8.4% over comparable weekly periods. Spot instances broadly run about 30% below on-demand with interruption rates under 5% for non-critical training workloads.
The table below reflects verified data as of Q2 2026. Footnote markers identify the source for each discount figure.
| Cloud provider | Instance type | Interruption rate | Spot discount vs. on-demand | Best region |
|---|---|---|---|---|
| GCP | a2-highgpu-1g | <5% / day [^1] | ~30% [^3] | europe-west4 |
| AWS | GPU spot | ~7.2% / week [^2] | 50–90% [^4] | Varies |
| GCP | Preemptible VMs | ~8.4% / week [^2] | ~30% [^3] | Varies |
| Neoclouds (H100 SXM5) | H100 SXM5 spot | Not published | Significant vs. on-demand | $1.35/hr cited [^5] |
How does heartbeat-based checkpoint and resume actually work for GPU agent tasks?
Heartbeat-based checkpoint and resume works by having a durable workflow engine emit a liveness signal every few seconds from a running GPU task; a missed heartbeat triggers automatic rescheduling with no human intervention.
The distinction worth understanding is task-level reschedule versus checkpoint restore. When tool calls and state writes are idempotent, a re-queued task produces the same result as the original. This is the task-boundary checkpoint pattern, and preemption rewinds only that task unit's work.

Which agent workloads should and should not run on spot GPUs, and what is the break-even math?
| Workload type | User-facing? | Task duration | Spot viable? | Reasoning |
|---|---|---|---|---|
| Batch model evaluation | No | 30–120s | Yes | Low blast radius; reschedulable |
| Data enrichment / entity extraction | No | 60–300s | Yes | Idempotent writes; pipeline tolerant |
| Offline research / RAG agents | No | 1–10 min | Yes | Async; latency tolerant |
| Async fine-tuning jobs | No | Hours | Conditional | Checkpoint overhead rises; use epoch-level saves |
| Synchronous user-facing inference | Yes | <5s | No | Interruption means visible failure; SLA breach |
| Real-time streaming pipelines | Yes | Continuous | No | No natural task boundary |
The price gap is significant. On-demand GPU rates run $7–$13/hr while spot capacity on H100 SXM5 is cited at $1.35/hr, with spot starting as low as $2.10/hr on some providers. AWS Spot can reduce compute costs by 50–90% when paired with graceful interruption handling. For teams running many GPU-hours of async, non-user-facing work, the savings compound quickly against any reasonable engineering investment in idempotency.

How does cross-cloud spot routing reduce interruption exposure further?
Even at 7–8% weekly rates, single-cloud deployment carries unnecessary concentration risk. A regional capacity event can spike local interruption rates while another provider remains stable. The cross-cloud arbitrage pattern is now practical tooling: maintain a priority-ordered spot market list, query availability and price at task submission, and route to the cheapest option with headroom. This is a documented 2026 operational pattern. Spheron spot runs 10–26% below on-demand with less contention than hyperscaler pools, making it a reliable overflow tier when AWS or GCP spot tightens.
Splitting a GPU fleet across AWS us-east-1 and GCP europe-west4, for example, means a provider-level capacity event affects only the portion routed there. Pipeline throughput takes a partial hit rather than a full one, and the healthy provider absorbs rescheduled tasks.
Frequently asked questions
What is a spot GPU interruption rate and why does it matter for AI workloads?
A spot GPU interruption rate is the percentage of preemptible instances a cloud provider reclaims before the workload finishes. Modern durable workflow engines detect a missed heartbeat and reschedule automatically, making interruptions an operational nuisance rather than a catastrophic failure for properly designed pipelines.
At what interruption rate do spot GPU savings stop being worth it?
As a practical rule of thumb used in this guide: for tasks under 2 minutes with heartbeat rescheduling active, weekly interruption rates below roughly 15% are generally acceptable because rescheduled work is measured in seconds. Above that threshold, the cumulative replay time begins to erode savings meaningfully. The break-even degrades further when tasks run for hours with high checkpoint overhead, which is why long training jobs have lower interruption tolerance than short agent tasks. Teams should treat this as a starting heuristic and validate against their own task durations and observed interruption rates.
How is checkpoint and resume different for agent workloads versus training jobs?
Training jobs serialize gigabytes of weights and optimizer state on long intervals with complex restore logic. Agent pipeline checkpointing is task idempotency: no weights, no restore logic. Design idempotent tool calls and task boundaries, and the workflow engine handles the rest. The engineering surface area is an order of magnitude smaller.
Conclusion
The obstacle to running agent workloads on spot GPUs was never the interruption rate; it was the mental model inherited from training jobs, where interruptions cost hours. For a 90-second agent task, the blast radius is seconds.
The 2026 numbers are clear: H100 SXM5 spot at $1.35/hr against $7–$13 on-demand represents a dramatic reduction available today. AWS Spot can cut compute costs by 50–90% when interruption handling is in place. The engineering ask is idempotent tasks and a heartbeat timeout, not a full checkpoint system. Every hour you stay on on-demand for async, non-user-facing, task-shaped workloads is a savings opportunity deferred.
Audit your background agent pipelines against the spot viability matrix above. Find the first async, non-user-facing, task-shaped workload. Run a two-week spot pilot. At the current spread between spot and on-demand GPU rates, payback comes fast.
Learn from me

Forward Deployed Engineering Bootcamp for Full-Stack Developers, my Maven cohort. Build and ship complete AI products end to end, from React and Node.js frontends to deployed models with caching and observability. Join the next cohort →
Hire us
Traversaal.ai. We're a team of forward deployed engineers solving the toughest AI problems for Fortune 100 companies: document intelligence, agentic data platforms, and real-time web intelligence, deployed in production. Work with our team to deploy your next agentic ecosystem. Talk to Traversaal.ai →
Join us
Want to solve these problems with us? We're always looking for forward deployed engineers who want to ship production AI. jobs@traversaal.ai