- Data-center GPU shortlisting
- Large-model memory planning
- Hopper-to-Blackwell evaluation
- Capacity briefs with explicit tradeoffs
NVIDIA data center comparison
H100 vs H200 vs B200: choose by bottleneck and availability.
H100 offers 80 GB HBM3 and an established Hopper path; H200 keeps Hopper but expands memory to 141 GB HBM3e; B200 moves to Blackwell with 180 GB HBM3e. All three have public GPURento planning rates and request paths. Benchmark the same workload before comparing cost per token, checkpoint or completed job.
Reviewed by the GPURento infrastructure team · September 1, 2026
| H100 SXM | 80 GB · $2.69/hr | $64.56 for a 24-hour compute scenario. |
|---|---|---|
| H200 SXM | 141 GB · $3.59/hr | $86.16 for a 24-hour compute scenario. |
| B200 | 180 GB · $5.98/hr | $143.52 for a 24-hour compute scenario. |
| Decision rule | Memory, then time/job | Use compatible software and identical inputs. |
- Treating peak specs as a benchmark
- Ignoring end-to-end system cost
- Comparing different model quality or software versions
Decision method
What to verify before you deploy capacity.
The content below separates published specifications, GPURento catalog references and decisions that still require a workload benchmark.
H100: established Hopper performance
Choose H100 when 80 GB is enough and the software stack benefits from Hopper transformer acceleration. It has the broadest GPURento request geography of the three and the lower public reference rate than H200.
H200: solve a memory-bound workload
Choose H200 when 141 GB and higher memory bandwidth reduce sharding, increase context or raise concurrency enough to offset the higher hourly reference. If the workload remains compute-bound and fits 80 GB, validate the gain.
- Measure GPUs required for the same model
- Keep context and concurrency fixed
- Compare cost per accepted token or checkpoint
B200: use Blackwell where it pays
B200 offers the largest memory of the three and a newer architecture at $5.98 per GPU-hour. Use it when Blackwell throughput, 180 GB memory or FP4 support offsets the higher hourly rate.
Benchmark the same system boundary
Pin the model and dataset revision, container digest, precision, context or batch, quality target and checkpoint policy. Record warm-up separately, then compare p50 and p95 end-to-end runtime, peak memory and total compute cost. Manufacturer peaks are specifications, not results from GPURento infrastructure.
- Use the exact GPU form factor shown in the request
- Include data loading, preprocessing and checkpoint I/O
- Report cost per accepted token, checkpoint or completed job
Catalog shortlist
Relevant GPU options.
NVIDIA H100 SXM
80 GB HBM3 · $2.69/hr
Intensive training and FP8 inference
NVIDIA H200 SXM
141 GB HBM3e · $3.59/hr
Large-model inference and HPC
NVIDIA B200
180 GB HBM3e · $5.98/hr
Frontier training and inference
Questions
Clear answers, including the limits.
Does H200 have more VRAM than H100?+
Yes. The listed H200 has 141 GB HBM3e compared with 80 GB HBM3 on H100 SXM.
Is B200 available to rent on GPURento?+
B200 is listed as a requestable configuration at $5.98 per GPU-hour for Paris and Frankfurt. The catalog does not guarantee immediate stock; capacity is checked when requested.
Which is cheapest per hour?+
H100 starts at $2.69/hour, H200 at $3.59/hour and B200 at $5.98/hour. Cost per job still requires a benchmark.
- NVIDIA H100 product pageArchitecture and memory specifications from the manufacturer.
- NVIDIA H200 product pageH200 memory capacity and bandwidth from the manufacturer.
- NVIDIA HGX reference architectureManufacturer reference for H100, H200 and B200 memory configurations.
Last reviewed September 1, 2026. GPURento rates are current catalog references; provisioning state remains visible in the workspace, and manufacturer specifications do not substitute for workload benchmarks.
Continue the research
Deployment plan
Choose the bottleneck you need to remove.
Use the catalog and pricing calculator to prepare an H100, H200 or B200 capacity request with every assumption visible.