Qwen3.8 · 27B
16.5 GB GGUF · exact artifact in the record
| Tokens reserved | AMD:NVIDIA | Short checks |
|---|---|---|
| 8,192 | 3:2 | 4 / 4 |
| 32,768 | 3:2 | 4 / 4 |
| 65,536 | 3:2 | 4 / 4 |
| 131,072 | 3:2 | Load failedChecks not reached |
| 131,072 | 2:1 | 4 / 4Separate retry |
A small Linux/Python companion kit for running one local GGUF across explicitly selected AMD Vulkan and NVIDIA CUDA devices, powered by a separately built llama.cpp.
Model capacity first, not a promise that mixed GPUs are faster. Device memory remains separate.
Experimental v0.1.0. Linux · Python 3.11+ · standard library.
16 GB device memory
10 GB device memory
Reservations and short-output checks.
Not full-length context validation.
One exact MoE artifact at 8K.
Short checks and owned teardown.
A separate, labelled experiment.
No borrowed Q4 speed numbers.
llama-server.Linux · Python 3.11+ · standard library. Requires a readable AMD junction sensor; the guard fails closed when unavailable. Infermeld does not build the engine, download weights or run as a persistent service.
Read the runtime and shutdown contract
Build the pinned engine
Troubleshoot the first launch
python3 infermeld.py devices \
--server ./llama.cpp/build/bin/llama-serverpython3 infermeld.py serve \
--server ./llama.cpp/build/bin/llama-server \
--model ./models/model.gguf \
--devices Vulkan0,CUDA0 --split 3,2 \
--confirm-devices --ctx-size 8192 \
--ubatch 32 --dry-runScroll horizontally on narrow screens. Replace the example paths and confirm physical identities before using these device IDs. Dry-run prints arguments only. After inspection, remove --dry-run to serve on loopback; Ctrl+C stops the owned server. MTP is off by default, unlike the allocation experiments below.
These are reserved-context allocation checks with short outputs, not full-length prompt ingestion or long-context recall. Neither model has a proven maximum usable context. Sustained Q4 throughput remains unqualified.
16.5 GB GGUF · exact artifact in the record
| Tokens reserved | AMD:NVIDIA | Short checks |
|---|---|---|
| 8,192 | 3:2 | 4 / 4 |
| 32,768 | 3:2 | 4 / 4 |
| 65,536 | 3:2 | 4 / 4 |
| 131,072 | 3:2 | Load failedChecks not reached |
| 131,072 | 2:1 | 4 / 4Separate retry |
22.7 GB GGUF · exact artifact in the record
| Tokens reserved | AMD:NVIDIA | Short checks |
|---|---|---|
| 8,192 | 3:2 | 4 / 4 |
| 32,768 | 3:2 | 4 / 4 |
| 65,536 | 3:2 | 4 / 4 |
| 131,072 | 3:2 | 4 / 4 |
Layer split (AMD:NVIDIA): relative layer-placement weights, in the selected device order. Not exact memory percentages, a speed ratio or unified memory. The example uses Vulkan0,CUDA0 only after physical-card confirmation.
4/4 short checks: exact output, JSON, short retrieval and nonempty completed prose sanity checks. These are not a broad quality score; short retrieval does not validate long-context recall.
Failure is retained: the dense 131,072-token attempt at 3:2 exited during loading. Checks were not reached. The passing 2:1 result is a separate retry.
Allocation conditions: MTP4 · q8_0 target/draft caches · batch 512 / microbatch 32.
Different from the launcher's default non-MTP recipe. Download the capacity record.
The unchanged launcher and recorded fresh engine cohort served the exact Qwen3.6-35B-A3B UD-Q4_K_M artifact. No dense-model acceptance or sustained performance claim is implied.
| Mode | Tokens reserved | Short checks | Teardown |
|---|---|---|---|
| Default · no MTP | 8,192 | 4 / 4 | Owned exit 143 |
| Explicit · MTP4 | 8,192 | 4 / 4 | Owned exit 143 |
Exact output, JSON, short retrieval and prose. Neither case qualifies full-length context or throughput. Inspect the cohort and checks.
IQ3-named artifacts, one repetition per workload, 512 output tokens and native MTP4. These rates are not Q4 speeds, fresh-build rates or a universal mixed-GPU speedup.
| Model | Workload | Decode | End-to-end | Prompt processing |
|---|---|---|---|---|
| 27B dense | code | 55.64 | 52.54 | 115.61 |
| 27B dense | prose | 31.48 | 30.45 | 133.21 |
| 35B-A3B MoE | code | 144.24 | 132.52 | 194.61 |
| 35B-A3B MoE | prose | 91.64 | 86.86 | 234.88 |
Code prompts: 56 uncached tokens; prose: 65. Prompt-processing rates are not sustained prefill. Output stopped at the token cap, not proof of completed-task quality. Host CPU/RAM metadata is omitted from the shared record, so these are not fully specified cross-machine comparisons. Explore the scoped historical chart. Download the scoped historical record.
Mixed-vendor execution comes from llama.cpp. Infermeld adds explicit placement, a guarded foreground wrapper, preflight checks and source-only packaging.
No model weights, native binaries, driver changes or private raw receipts are included. Original kit code is MIT; upstream and model licenses remain separate. Qwen / Unsloth attribution stays with the exact artifacts.
Read the MIT license · Source checks and contributing
Reference engine revision:
b92761a515ea31e852e7fbc1fad5f874b46f3718
Experimental source-only kit. No guaranteed speedup or maximum-context preset. Documentation links point to the source repository and require access while it remains private.