Q3 observations.
Not Q4 speeds.
These IQ3-labelled experiments helped develop the recipe. They are single-repetition observations, not measurements of the Q4 artifacts now featured on the main page.
Return to the Q4 capacity evidence ↗Show the work.
Two model families. Two workloads.
Every number comes with its conditions.
Server-reported decode throughput. Excludes prefill; not an end-to-end speed measure.
One repetition per workload, no confidence intervals. Host CPU/RAM metadata is omitted from the shared record; these are not fully specified cross-machine comparisons. Different model families are not a like-for-like engine or GPU comparison. Fixed MTP4 is not the best depth for every task.
Exact settings, real drafted/accepted counts, artifact names, and the experiment binary digest are retained in the downloadable record.
Download the measurement recordInspect the test conditions
- Reference GPUs
- RX 6900 XT 16 GB · RTX 3080 10 GB
- Placement
- Layer split · Vulkan0, CUDA0 · 3:2
- Cache / execution
- q8_0 · ubatch32 · context4096 · 8 threads
- Speculation
- Embedded draft-MTP · fixed maximum depth4
- Generation
- Greedy · seed1234 · 512 tokens · one repetition
- Scope
- Pre-release experiment; optional placement policy disabled. Not a clean-checkout kit benchmark.
- Upstream pin
Loading…
| Model | Decode tok/s | End-to-end tok/s | Drafted / accepted |
|---|