One work session
$2.72
1 GPU × 8h · compute only
Image generation workflows
Stable Diffusion and ComfyUI workloads are usually best matched by the exact graph, checkpoint, resolution, batch and auxiliary models. RTX 4090 offers a $0.34/hour GPURento reference with 24 GB; RTX 5090 provides 32 GB at $0.69/hour; L40S offers 48 GB at $0.79/hour when a workflow needs more headroom.
Reviewed by the GPURento infrastructure team · September 1, 2026
Transparent planning windows
These are compute rental windows, not predictions of job duration or performance. Replace the hours with a measured pilot and add storage in the calculator.
One work session
$2.72
1 GPU × 8h · compute only
Forty-hour window
$13.60
1 GPU × 40h · compute only
One-hundred-twenty hours
$40.80
1 GPU × 120h · compute only
| Value option | RTX 4090 · 24 GB | $0.34/hour reference. |
|---|---|---|
| More headroom | RTX 5090 · 32 GB | $0.69/hour reference. |
| Large graphs | L40S · 48 GB | $0.79/hour reference. |
| Compare by | $ / accepted output | Keep resolution, steps and model fixed. |
Decision method
The content below separates published specifications, GPURento catalog references and decisions that still require a workload benchmark.
Pin the container, ComfyUI version, custom-node commits, checkpoint hashes and input settings. A visual graph without versioned dependencies is not a reproducible workload.
The checkpoint is only one memory consumer. VAEs, ControlNet, upscalers, text encoders and video nodes can overlap. Run a representative high-resolution job and record peak allocation before choosing the cheapest card.
Record warm run time, rejected outputs and total metered GPU time. Model download and node installation make cold runs slower, which is why persistent storage may be valuable for repeat sessions.
Catalog shortlist
24 GB GDDR6X · $0.34/hr
Image, video and mid-size inference
32 GB GDDR7 · $0.69/hr
Fast single-GPU inference and media
48 GB GDDR6 ECC · $0.79/hr
Production inference, fine-tuning and video
Questions
RTX 4090 is a strong starting point for graphs that fit 24 GB. Choose RTX 5090 for 32 GB or L40S for 48 GB when the actual workflow requires more memory.
This page does not claim a one-click template. The GPU Cloud accepts container configuration; confirm the image and workflow requirements in the workspace.
Persist reusable models, stop compute after outputs are saved, and compare cost per accepted output rather than peak images per second.
Last reviewed September 1, 2026. GPURento rates are current catalog references; provisioning state remains visible in the workspace, and manufacturer specifications do not substitute for workload benchmarks.
Continue the research
Deployment plan
Create a workspace with the exact image, storage and region settings, then fund it before requesting capacity.