Can I run this model?

VRAM math for local models — language, image and video. What the weights cost, what everything around the weights costs, and the cheapest GPU that fits it.

The job

Model kind
Task

Verdict

—
— GB VRAM NEEDED
0 GB

Where the memory goes

ComponentSizeShare
Peak100%

Rentable GPUs that fit

ConfigurationProviderVRAM$ / hrYour job
The math, so you can check it

Everything here is arithmetic you can reproduce, not a benchmark. Language estimates land within roughly 10–15% of what nvidia-smi reports for llama.cpp at these settings.

vLLM works differently and the number above is not what it will show you. vLLM is budget-driven rather than demand-driven: it claims gpu_memory_utilization of the card up front (0.90 in Runpod's guide, 0.92 is vLLM's own current default), subtracts model weights, non-Torch memory and the profiled activation peak, and hands the entire remainder to the KV cache as paged blocks of 16 tokens. So it never asks whether your context fits — it asks how many tokens the leftover can hold. Two consequences worth knowing: your usable VRAM is about 10% less than the sticker figure, and raising Max Model Length does not reserve more cache, it just lets one request consume more of a fixed pool. The deploy panel works out that pool and the resulting concurrency for you.

LANGUAGE
weights   = params × bytes_per_weight
kv_cache  = 2 × layers × kv_heads × head_dim × context × kv_bytes × seqs
overhead  = 0.8 GB (CUDA context + runtime) + 5% of weights
QLoRA     = 4-bit base + adapters/optimizer + checkpointed activations
Full tune = 16 bytes/param (fp16 weights + grads + Adam states + fp32 master)

IMAGE & VIDEO
model     = params × bytes_per_weight
encoder   = text_encoder_params × encoder_bytes   ← the forgotten one
denoise   = base × (pixels ÷ 1024²) × batch × (frames ÷ 16)^exp
vae_decode = decode_peak × (pixels ÷ 1024²)   × 0.3 if tiled
work_set  = denoise + vae_decode
offloaded = max(encoder, model + vae + work_set) + 0.8 GB
resident  = encoder + model + vae + work_set + 0.8 GB

The text encoder is the trap. Flux and SD 3.5 ship with T5-XXL, which is 4.7B parameters — 9.1 GB in fp16, before the image model loads at all. Offloading it (what ComfyUI does by default) is usually the difference between fitting and not.

The working-set term is the roughest number on this page. It is calibrated against reported figures for SDXL and Flux at 1024 × 1024, and video scaling is superlinear because 3D attention sees every frame at once. Treat video numbers as a starting point; block-swapping and tiled VAE decode beat them substantially.

Video and audio together is its own case, and MiniMax-H3 is the model that forced the two extra terms above. exp is 1.3 for most video models because their memory grows faster than their frame count; for H3 it is 1.0, since a 3D DiT with FlashAttention holds O(tokens) rather than O(tokens²) and tokens are linear in frames. The vae_decode term is the other half: H3's VAE compresses 16× spatially and reconstructs through a 128-channel stack at full output resolution, one 17-frame clip at a time. That peak does not care how long your clip is — it is roughly 10 GB at 1024² whether you render 39 frames or 362 — so folding it into the frame-scaled term would have made short clips look free and long ones look impossible. Both are wrong. The audio side is genuinely cheap: 40 latent tokens per second of 32 kHz stereo, against roughly 5,700 video tokens per second at 768p. Note also that H3 ships CFG-distilled at cfg_scale 1.0, so unlike Flux or SDXL it does not run a second unconditional pass — the working set below is not doubled, and any guide telling you to budget for both branches is describing a different model.

Multi-GPU adds about 10% for tensor-parallel buffers, applied only when a job is actually sharded. MoE models hold every expert in memory but activate a few per token — memory tracks total params, speed tracks active params.

Prices were captured in August 2026 and drift constantly. Runpod figures are on-demand Secure Cloud, billed per second; Community Cloud and spot are cheaper. Lambda figures are published on-demand rates for single-GPU instances. Vast.ai is a marketplace where hosts set their own prices, so those rows are marked ~ and should be read as indicative — interruptible instances there run 30–50% below on-demand, at the cost of being interrupted.

One bias to correct for: Runpod has the most rows here because it publishes a full price list I could transcribe; Vast.ai has three because those are the only rates I could source with confidence. Vast.ai not appearing beside a given card does not mean it costs more there — it usually means I had no number. Check it yourself for anything above a couple of dollars an hour.