infermeld
Local inference / explicitly placed

Different silicon.
One model.

A small Linux/Python companion kit for running one local GGUF across explicitly selected AMD Vulkan and NVIDIA CUDA devices, powered by a separately built llama.cpp.

Model capacity first, not a promise that mixed GPUs are faster. Device memory remains separate.

Experimental v0.1.0. Linux · Python 3.11+ · standard library.

One local GGUF / tested reference pair
AMD / Vulkan

RX 6900 XT

16 GB device memory

NVIDIA / CUDA

RTX 3080

10 GB device memory

Explicit layer placement. Separate device memory. A conceptual grouping, not a concurrency or timing diagram.

Q4 allocation

Reservations and short-output checks.
Not full-length context validation.

Fresh CLI acceptance

One exact MoE artifact at 8K.
Short checks and owned teardown.

Historical IQ3 speeds

A separate, labelled experiment.
No borrowed Q4 speed numbers.

01 / Start with the actual hardware

Inspect. Select.
Preview locally.

  • Use a compatible, separately built CUDA + Vulkan llama-server.
  • Verify your local GGUF against its pinned SHA-256.
  • Confirm two different physical GPUs. Backend indices alone are not identities.

Linux · Python 3.11+ · standard library. Requires a readable AMD junction sensor; the guard fails closed when unavailable. Infermeld does not build the engine, download weights or run as a persistent service.

Read the runtime and shutdown contract
Build the pinned engine
Troubleshoot the first launch

1. Inspect the physical devices
python3 infermeld.py devices \
  --server ./llama.cpp/build/bin/llama-server
2. Preview, without starting a server
python3 infermeld.py serve \
  --server ./llama.cpp/build/bin/llama-server \
  --model ./models/model.gguf \
  --devices Vulkan0,CUDA0 --split 3,2 \
  --confirm-devices --ctx-size 8192 \
  --ubatch 32 --dry-run

Scroll horizontally on narrow screens. Replace the example paths and confirm physical identities before using these device IDs. Dry-run prints arguments only. After inspection, remove --dry-run to serve on loopback; Ctrl+C stops the owned server. MTP is off by default, unlike the allocation experiments below.

02 / Allocation and short serving

Q4 capacity, with the conditions.

These are reserved-context allocation checks with short outputs, not full-length prompt ingestion or long-context recall. Neither model has a proven maximum usable context. Sustained Q4 throughput remains unqualified.

DENSE / UD-Q4_K_M

Qwen3.8 · 27B

16.5 GB GGUF · exact artifact in the record

Reserved context, not proven useful context
Tokens reservedAMD:NVIDIAShort checks
8,1923:24 / 4
32,7683:24 / 4
65,5363:24 / 4
131,0723:2Load failedChecks not reached
131,0722:14 / 4Separate retry
MOE / UD-Q4_K_M

Qwen3.6 · 35B-A3B

22.7 GB GGUF · exact artifact in the record

Reserved context, not proven useful context
Tokens reservedAMD:NVIDIAShort checks
8,1923:24 / 4
32,7683:24 / 4
65,5363:24 / 4
131,0723:24 / 4

Layer split (AMD:NVIDIA): relative layer-placement weights, in the selected device order. Not exact memory percentages, a speed ratio or unified memory. The example uses Vulkan0,CUDA0 only after physical-card confirmation.

4/4 short checks: exact output, JSON, short retrieval and nonempty completed prose sanity checks. These are not a broad quality score; short retrieval does not validate long-context recall.

Failure is retained: the dense 131,072-token attempt at 3:2 exited during loading. Checks were not reached. The passing 2:1 result is a separate retry.

Allocation conditions: MTP4 · q8_0 target/draft caches · batch 512 / microbatch 32.
Different from the launcher's default non-MTP recipe. Download the capacity record.

Separate evidence / fresh engine + CLI

MoE Q4 at an 8K reservation.

The unchanged launcher and recorded fresh engine cohort served the exact Qwen3.6-35B-A3B UD-Q4_K_M artifact. No dense-model acceptance or sustained performance claim is implied.

3:2 layer split · microbatch 32 · owned SIGTERM shutdown
ModeTokens reservedShort checksTeardown
Default · no MTP8,1924 / 4Owned exit 143
Explicit · MTP48,1924 / 4Owned exit 143

Exact output, JSON, short retrieval and prose. Neither case qualifies full-length context or throughput. Inspect the cohort and checks.

03 / Distinct artifacts, distinct cohort

Historical speeds stay historical.

IQ3-named artifacts, one repetition per workload, 512 output tokens and native MTP4. These rates are not Q4 speeds, fresh-build rates or a universal mixed-GPU speedup.

View the four historical workload rows
Tokens per second · short uncached prompts · recorded pre-release experiment
ModelWorkloadDecodeEnd-to-endPrompt processing
27B densecode55.6452.54115.61
27B denseprose31.4830.45133.21
35B-A3B MoEcode144.24132.52194.61
35B-A3B MoEprose91.6486.86234.88

Code prompts: 56 uncached tokens; prose: 65. Prompt-processing rates are not sustained prefill. Output stopped at the token cap, not proof of completed-task quality. Host CPU/RAM metadata is omitted from the shared record, so these are not fully specified cross-machine comparisons. Explore the scoped historical chart. Download the scoped historical record.

04 / Inspect the source, not just the screenshot

Small kit. Checkable evidence.

Mixed-vendor execution comes from llama.cpp. Infermeld adds explicit placement, a guarded foreground wrapper, preflight checks and source-only packaging.

No model weights, native binaries, driver changes or private raw receipts are included. Original kit code is MIT; upstream and model licenses remain separate. Qwen / Unsloth attribution stays with the exact artifacts.

Read the MIT license · Source checks and contributing

Reference engine revision:
b92761a515ea31e852e7fbc1fad5f874b46f3718

Experimental source-only kit. No guaranteed speedup or maximum-context preset. Documentation links point to the source repository and require access while it remains private.