One hour
$3.59
1 GPU × 1h · compute only
High-memory Hopper GPU
NVIDIA H200 pairs Hopper compute with 141 GB of HBM3e, making it useful when model weights, KV cache or scientific data exceed an 80 GB device. GPURento lists a $3.59 per GPU-hour on-demand rate with self-service creation in Paris and Frankfurt.
Reviewed by the GPURento infrastructure team · September 1, 2026
Cost scenarios
One hour
$3.59
1 GPU × 1h · compute only
One day
$86.16
1 GPU × 24h · compute only
Seven days
$603.12
1 GPU × 168h · compute only
Average month · 730h
$2620.70
1 GPU × 730h · compute only
| Memory | 141 GB HBM3e | More single-GPU memory than H100 SXM. |
|---|---|---|
| Memory bandwidth | 4.8 TB/s | Manufacturer specification for H200. |
| On-demand rate | $3.59 / GPU-hour | Compute only, before storage or optional services. |
| Available regions | Paris · Frankfurt | Both regions are selectable after funding. |
Decision method
The content below separates published specifications, GPURento catalog references and decisions that still require a workload benchmark.
H200 is not simply “a faster H100.” Its practical advantage is the larger, higher-bandwidth memory subsystem. That can let a model use fewer GPUs, hold a longer context or serve more concurrent sequences, depending on quantization and runtime.
For inference, record time to first token, output tokens per second, p95 latency and cost per million tokens at a fixed quality level. A high-memory GPU is valuable only if the runtime can use that memory and bandwidth efficiently.
The catalog exposes H200 in EU West Paris and EU Central Frankfurt. GPURento checks account funding, stores the configuration and starts the self-service provisioning workflow without a sales form.
Catalog shortlist
141 GB HBM3e · $3.59/hr
Large-model inference and HPC
80 GB HBM3 · $2.69/hr
Intensive training and FP8 inference
180 GB HBM3e · $5.98/hr
Frontier training and inference
Questions
The H200 configuration described here has 141 GB of HBM3e memory and a manufacturer-specified 4.8 TB/s memory bandwidth.
Choose H200 when 80 GB forces unwanted sharding, limits context or constrains concurrency. If your workload already fits and is compute-bound, benchmark both before paying the higher reference rate.
The current GPURento catalog exposes self-service H200 creation in EU West Paris and EU Central Frankfurt.
Last reviewed September 1, 2026. GPURento rates are current catalog references; provisioning state remains visible in the workspace, and manufacturer specifications do not substitute for workload benchmarks.
Continue the research
Deployment plan
Create a workspace and choose Paris or Frankfurt with your model size, runtime and target concurrency.