📊 Benchmarks & Hardware Record

actual measured runs — estimates live on the Model Catalog

Hardware inventory

HostGPUArchVRAMMem BWSoftware stackRole
vastnode002RTX PRO 6000Blackwell~96 GB~1.8 TB/s (vendor)CUDA · ComfyUI cu130MiniMax-H3 primary (:8000) — int8 t2v/i2v/r2v +audio
vastnode0032× RTX PRO 5000Blackwell48 GB ea~1.34 TB/sCUDA · vLLM + LightX2V · on VastGPU0 dense-27b NVFP4 (:8001) + GPU1 Wan T2V (:8000) — since 2026-08-06
gpuhost001RTX PRO 6000 WSBlackwell~96 GB~1.8 TB/s (vendor)CUDA · ComfyUI cu130MiniMax-H3 #2 (:8000) — int8; dense-27b moved off 2026-08-06
r730a-a50002× RTX A5000Ampere24.0 GB ea (measured 24564 MiB)768 GB/s (vendor)driver 580.159.03 · CUDA 13.0 (measured) · 16c/64GBOmni-30B-A3B AWQ across BOTH cards via PP-2 (65K ctx, :8901, max-num-seqs 24). ~16.6 tok/s single-stream, ~746 tok/s saturated aggregate. PP not TP — TP-2 = 1 tok/s (per-step all-reduce on PCIe-PHB). Image-edit retired. Cards pinned 190W.
3930aRTX 4070S TiAda16 GB~672 GB/s (vendor)CUDA · diffusers · vllmACE-Step 1.5 music gen (live; replaced VibeVoice)
3930bRTX 4070S TiAda16 GB~672 GB/s (vendor)CUDA · transformersQwen3-TTS clone+design (live)
vm9052× RX 7900 XTXRDNA324 GB ea960 GB/s (vendor)ROCm · llama.cpp · IPEXBOTH cards: 2× DiT-only edit workers (data-parallel ~18GB ea), proxy :8000, remote TE on the 4500 → 2× throughput (4500 embed fixed: bf16 ViT)
(candidate)Arc Pro B70Battlemage Xe232 GB608 GB/s (vendor)XPU · vLLM-XPU / IPEX-LLMEvaluating — not owned; no FP8/FP4

Benchmark log

DateModelQuantCard / HostFrameworkParamsSingle-streamBatched (n streams)VRAM peakNotes
2026-07-05Qwen3.6-27BNVFP4 (native FP4)1× PRO 6000 / vastnode002 (+I2V) @300WvLLM 0.24 + flashinfer 120f200K ctx, max-seqs 36, MTP n=2, util 0.54~78-80 tok/s warm (measured)~297 tok/s aggregate (measured)~19.5 GB wt (shares card w/ Wan I2V)sm120 FP4 crash-loop FIXED via flashinfer compute_120f (FLASHINFER_FORCE_SM=120f) — the AOT 120a cutlass FP4 GEMM was broken on Blackwell. n=3 MTP is 5× worse; autotune-on crashes. ~100 needs faster kernel (nightly staged).
2026-06-15Qwen3.6-27BFP81× PRO 6000 / vastnode002 (+I2V)vLLM 0.22.1200K ctx, max-seqs 32, MTP spec, util 0.5872.3 tok/s (measured)363 @8 · 1558 @32 streams (measured)55.6 GB (shares card w/ Wan I2V)Moved off 5000s→6000; ~1.8× single / ~1.5× batched vs the old 5000 (40/1044). Coexists w/ Wan I2V.
2026-06-14Qwen3-Omni-30B-A3B InstructAWQ-int4A5000 GPU1 / r730avLLM 0.22.120K ctx · image≤2 · video 1×128f · max-seqs 4133.8 tok/s (measured)426 tok/s @4 streams (107/strm, measured)19.2 GB wt / ~22.7 GB runMoE 3B-active → fast decode; image:3+ OOMs at video profiling
2026-06-14Qwen2.5-VL-7Bint87900 XTX GPU1 / vm905custom /vl-chat (ROCm)~200 tok gen~4 tok/s (measured)batch TBD~9 GB50.7s for ~200 tok — SLOW for 7B int8; investigate ROCm path / per-call overhead
2026-06-14LTX-2.3 22B (t2av)fp8 self-quantPRO 6000 / vastnode002LightX2V~10 s clip, 768p, +audio~33 s/clip (measured)— (1 stream)~27 GBNo longer the pipeline bottleneck
2026-06-14Qwen-Image-2512 (T2I)bnb int8 + remote 7B TEA5000 GPU0 / r730adiffusers1024², 40 steps247.7 s/img (measured)— (1 stream)20.9 GB peakRe-measured (was 190); remote-TE round-trip to vm905 adds latency. Pipeline bottleneck.
2026-06-23ACE-Step v1.5 turbo (music)2B DiT + 1.7B LM4070S Ti / 3930adiffusers + vllm30s track, audio_duration 30~12 s/track (measured)batch up to 4~10 GBTier-5 (12-16GB): max 8min/track w/ LM, 10min DiT-only. Apache, commercial-OK. Replaced VibeVoice.
2026-06-14VibeVoice-7B (retired)bnb int84070S Ti / 3930atransformers SDPA~70-word narration0.53× realtime (measured)batch TBD~13 GBRETIRED 2026-06-23 (replaced by ACE-Step; Qwen handles voice). 46.8s gen for 24.8s audio — slower than realtime.
2026-06-23Qwen-Image-Edit-2511 + 8-step LightningGGUF Q6_K + bf16 LoRA7900 XTX GPU0 / vm905diffusers GGUF (ROCm) + peft768², multi-image edit, 8-step cfg1.0~49 s/edit warm (measured)— serialized (lock)~17 GB8-step Lightning LoRA (lightx2v 2511) = 2.4× vs ~120s @20-step. peft req'd for LoRA-on-GGUF. Concurrent /edit serialized via lock (ROCm core-dumps on parallel pipe calls). Cold run +75s MIOpen compile.
2026-06-15Wan2.2 T2V-A14BNVFP4PRO 5000 ×2 / vastnode003LightX2V720p, 81 frames, 4-step distill~41 s/clip (measured)1 clip/worker; 2 workers ≈ 1.6× single-6000~40 GB peak (8 GB headroom)Moved 6000→5000s. 6000 was ~33s/clip same config. Server uses config res, ignores API params.
2026-07-09dense-27b (Qwen3.6-27B)NVFP4 + MTP n21× PRO 5000 (GPU1) / vastnode003vLLM nightly800 tok, temp0, max-model-len 150k, seqs 16, util 0.92~94.7 tok/s (measured)@8=510 · @16=970 · @24=828 tok/s agg (measured)44/49 GBOne PRO 5000 replica. Single-stream 94.7 = 82% of the 6000's ~115 (BW ratio). Agg still LINEAR 8→16 (510→970) = NOT saturated at seqs 16 — the ~1550 BW ceiling needs seqs~40, which 150k ctx won't allow on 48GB (KV 128KB/tok). @24 dips (seqs-16 admission cap). So 2×5000 @ this config ≈ ~1940 agg (parity w/ one 6000 @ high-conc), NOT the 1.5×; the 1.5× needs concurrency the big-context config forbids. 22W idle→check under load.

green = measured on our hardware  ·  TBD = not yet benchmarked  ·  (vendor) = spec sheet.
LLM rows record single-stream and batched aggregate with stream count. Append runs to BENCHMARKS in server.py.