Benchmarks

Measured on our own hardware, with the harness noted against every row. Rows we have not measured yet say pending — we would rather show you an empty column than an estimate.

Inference

ModelGPUStack ShapeThroughputTTFT $/M tokensMeasured
deepseek-ai/DeepSeek-R1-Distill-Qwen-7B
rev main
1× NVIDIA L40S vLLM 0.6.6
none
256 in / 128 out
concurrency 4
146.78 tok/s 511 ms $2.27 2026-07-01 run ↗
Live run on G293-1 (1×L40S), streaming TTFT, warmup excluded. 20 reqs @ concurrency 4.
meta-llama/Llama-3.1-8B-Instruct
rev TBD
1× NVIDIA L40S vLLM (version TBD)
none
512 in / 512 out
concurrency 1
pending pending pending pending
Median of N runs, warmup excluded. Pending real hardware run.

Launch latency

ScenarioGPUTimeMeasured
cold NVIDIA L40S pending pending
Account→running, cold node. Pending real run.
warm-template NVIDIA L40S pending pending
Pre-warmed template launch. Pending real run.

Last updated 2026-07-01. Figures we cannot source yet are marked rather than guessed.