The H200 is not a faster H100. It is the same Hopper chip with a bigger fuel tank. This guide is about whether the upgrade is worth it for your workload; for today's numbers use the live H100 vs H200 cloud price & availability comparison, which is refreshed hourly across every provider we track.
Same 1,979 TFLOPS FP16. Same NVLink 4.0. Same 700W TDP. NVIDIA swapped the memory subsystem: 80 GB HBM3 to 141 GB HBM3e, with bandwidth up 43%. The question is whether that memory upgrade justifies the price premium - and the answer depends entirely on your workload.
What is actually different
The H200 uses the identical GH100 die as the H100 - same 132 streaming multiprocessors, same 4th-gen Tensor Cores, same compute capability. The only change is the memory:
| Spec | H100 SXM5 | H200 SXM |
|---|---|---|
| Architecture | Hopper (GH100) | Hopper (GH100) - same die |
| Memory | 80 GB HBM3 | 141 GB HBM3e |
| Memory bandwidth | 3.35 TB/s | 4.8 TB/s |
| FP16 Tensor Core | 1,979 TFLOPS | 1,979 TFLOPS - identical |
| NVLink | 4.0 (900 GB/s) | 4.0 (900 GB/s) - identical |
| TDP | 700W | 700W - identical |
| Release | 2023 | 2024 |
HBM3e raises per-pin transfer rate from 6.4 Gbps to 9.8 Gbps and adds a sixth memory stack. The result: 76% more capacity and 43% more bandwidth, with zero change to compute throughput.
Both use the SXM5 socket and are mechanically compatible with the same NVLink Switch systems. The H200 is a drop-in replacement in DGX H100 baseplates. Same CUDA compute capability (9.0), same driver stack - zero code changes required to migrate.
Cloud pricing comparison
Here is what H100 and H200 cost per hour across the providers we track, pulled live from our database:
| Provider | H100 /hr | H200 /hr | Difference |
|---|---|---|---|
| Vast | $1.25Out | $2.63Out | +110% |
| Hyperbolic | $1.29No data | - | - |
| Lium | $1.30In stock | $3.00Stale · 31h ago | +131% |
| Seeweb | $1.89No data | $2.60No data | +38% |
| PrimeIntellect | $1.90Out | $2.00Out | +5% |
| Runpod | $1.99Limited | $3.59In stock | +80% |
| Theta EdgeCloud | $2.29No data | $3.69No data | +61% |
| Hyperstack | $2.50Limited | - | - |
| Shadeform | $2.50In stock | - | - |
| Scaleway | $2.52In stock | - | - |
| Lambda | $3.29Out | - | - |
| Verda | $3.56Out | $4.83Out | +36% |
| Nebius | $4.50In stock | $5.40In stock | +20% |
| Digital Ocean | $6.74Stale · 46d ago | - | - |
| AWS | $6.88No data | - | - |
| Azure | $6.98No data | - | - |
| Google Cloud | $10.56No data | - | - |
| Yotta | - | $2.10No data | - |
The Difference column and median beneath the table are calculated from those same live rows. Premiums vary by provider, so use them instead of a fixed headline percentage. For spot decisions, compare the current H100 and H200 pages; spot supply and launch capacity can change independently of listed on-demand prices.
When the H200 is worth the premium
Large model inference (70B+ parameters)
This is where the H200 earns its price. Llama 3 70B in FP16 uses roughly 140 GB of VRAM - it fits on a single H200 but needs two H100s.
The following is a fixed worked example using prices captured on 22 August 2026, not a current quote:
- 1x H200 at median $2.80/hr
- 2x H100 at median $2.45/hr each = $4.90/hr total
Under those dated assumptions, the H200 is 43% cheaper and eliminates tensor parallelism overhead entirely. For production inference serving, this also means half the instances to manage, half the egress surface, and no cross-GPU communication latency. See our egress fees comparison for the hidden cost that stacks up here.
Long-context serving (128K+ tokens)
KV cache for long-context models is where the H200 pulls furthest ahead. At 128K context length, a single sequence with a 70B model can push 60-80 GB of KV cache on top of the model weights. The H100 at 80 GB total gets evicted constantly, while the H200 holds full context in memory.
The throughput difference is not 43% - it is often 3-5x because you eliminate KV cache offloading entirely. For document QA, RAG with large retrievals, or multi-turn agents running long sessions, this is the deciding factor.
Mixtral and large MoE models
Mixtral 8x7B in FP16 uses roughly 90 GB. One H200 handles it; one H100 cannot. Mistral Large (123B) requires two H100s or one H200 with quantization. If you are running mixture-of-experts architectures, the extra VRAM eliminates the need for model sharding.
When the H100 is the better choice
Models under 70B parameters
Llama 3 8B, Mistral 7B, CodeLlama 13B at INT8, Stable Diffusion XL - all run identically on H100 and H200. If your workload does not use the extra memory, any H200 premium shown in the live table buys capacity you will never touch.
For inference serving of sub-70B models, the H100 is the right call.
Compute-bound training
Same TFLOPS means same training throughput. Distributed training on 70B+ models is NVLink-bound and network-bound, not memory-bandwidth-bound per card. The 43% memory bandwidth uplift helps data loading and optimizer steps marginally - expect 5-10% end-to-end speedup at best for multi-node training. Use the live premium above to test whether that uplift justifies H200 for your cluster; otherwise spend the budget on more H100 nodes.
Spot pricing advantage
H100 spot supply is generally broader and more mature than H200 spot supply. For batch workloads and preemptible inference jobs where interruption is tolerable, compare the current spot rows and reliability signals on each GPU page before choosing on unit economics.
Cost breakdown by scenario
This table is an illustrative 22 August 2026 cost model, using the fixed hourly assumptions above. It is not live; recalculate with the current table before budgeting.
| Scenario | H100 cost | H200 cost | Winner |
|---|---|---|---|
| 7B model training, 100 hours | $129 (1x) | $214 (1x) | H100 saves 40% |
| 70B inference, 24/7, monthly | $3,528 (2x) | $2,016 (1x) | H200 saves 43% |
| Fine-tuning 70B, 10 hours | $49 (2x) | $28 (1x) | H200 saves 43% |
| 8B inference serving, monthly | $929 (1x) | $2,016 (1x) | H100 saves 54% |
Under the dated assumptions above, running ten 70B inference instances saves roughly $15,000 per month by choosing H200 over 2x H100 configurations.
Rule of thumb: if your model fits comfortably in 80 GB, default to H100. If it needs 120 GB or more, the H200 pays for itself within the first month of continuous use.
Availability and migration
The H100 has historically had broader provider and spot coverage, while H200 capacity is concentrated in newer deployments. Treat that as editorial context, not a live stock claim: the current status badges and timestamps above are the decision signal.
Migration from H100 to H200 requires zero software changes. Same CUDA 9.0, same driver stack, same NVLink topology. The only migration work is re-benchmarking memory-bound workloads and adjusting batch sizes upward to exploit the extra 61 GB.
Both are 700W SXM - same rack density, same cooling contracts, same PDU math. If you have qualified a facility for H100, the H200 drops in without renegotiating power.
Verdict
- Model fits in 80 GB? H100. Compare the live premium before paying for unused memory.
- Need 120 GB+ VRAM? H200 can avoid multi-GPU sharding; recalculate with the live rates.
- Budget-constrained? Compare current H100 spot rows and reliability.
- Production inference, 70B+? H200. Fewer instances, lower total cost.
- Distributed training? H100. Same compute throughput, better economics.
The H200 is not a generational leap - it is a targeted memory upgrade for workloads that were memory-constrained on the H100. If that describes your workload, the premium is justified. If not, the H100 remains the best value in cloud GPUs.
Compare live pricing for both: H100 cloud pricing | H200 cloud pricing | How we collect data