From 2555344204fec9d7143e2afd8afd6c70a080750e Mon Sep 17 00:00:00 2001 From: Julian Beltran Date: Mon, 31 Aug 2026 23:57:01 +1000 Subject: [PATCH] readme: clean public version --- README.md | 70 +++++++++++++++++++++++++++++++------------------------ 1 file changed, 40 insertions(+), 30 deletions(-) diff --git a/README.md b/README.md index c661f65..4bd30dc 100644 --- a/README.md +++ b/README.md @@ -5,12 +5,10 @@ **qwen3.8-flash-next on amd strix halo — 91g quant, 56 tok/s, 262k context** [![HF Model](https://img.shields.io/badge/🤗_Model-IQ4_XS_PLE-ffD21E)](https://huggingface.co/julianmb/Qwen3.8-Flash-Next-IQ4_XS-GGUF) -[![Engine](https://img.shields.io/badge/Engine-nathanw1014__vulkan-blue)](https://github.com/Nathanw1014/llama.cpp/tree/strix-halo-vulkan) -[![License](https://img.shields.io/badge/License-Qwen_Community-orange)](https://huggingface.co/Qwen/Qwen3.8-Flash-Next/blob/main/LICENSE) - [![Speed](https://img.shields.io/badge/MTP_@8k-56.4_t%2Fs-brightgreen)](#results) [![Depth](https://img.shields.io/badge/Verified-0_→_256k-blueviolet)](#results) [![Provenance](https://img.shields.io/badge/Provenance-byte--verified-success)](#the-converter-bug) +[![Docker](https://img.shields.io/badge/Docker-ready-2496ED)](#docker) *every published quant byte-traced back to the official checkpoint* @@ -20,9 +18,9 @@ ## results -engine: [nathanw1014/llama.cpp `strix-halo-vulkan`](https://github.com/Nathanw1014/llama.cpp/tree/strix-halo-vulkan) · vulkan/radv · q8_0 kv · `-ub 2048` · temp 0 +91g quant · vulkan/radv · mtp sidecar · q8_0 kv · `-ub 2048` · temp 0 · 128g strix halo -| depth | plain pp/tg | mtp pp/tg | +| depth | plain pp / tg | mtp pp / tg | |------:|:-----------:|:---------:| | 0 | 92.5 / 29.9 | 87.0 / **53.1** | | 8k | 480 / 24.1 | 458 / **56.4** | @@ -54,7 +52,7 @@ single runs, n=1 caveat. pick your file by use case — see the table above. |------|------|:------------:| | `...-IQ4_XS-`**`PLE`**`.gguf` | 91 giB | ctx ≤ 32k — wins everywhere, mtp to 56 t/s | | `...-IQ4_XS.gguf` | 116 giB | ctx ≥ 128k — faster mtp at depth, wider fork compat | -| `mtp-...-Q8_0.gguf` | 3.9 giB | mtp sidecar for nathanw1014-lineage engines | +| `mtp-...-Q8_0.gguf` | 3.9 giB | mtp sidecar, required for the speed numbers |
the PLE cut — why the 51b n-gram table tolerates 4-bit @@ -66,18 +64,18 @@ depth sweep. the cut: `--tensor-type "per_layer_token_embd=IQ4_XS"` on our quantizer → 54g → 27g. **fork caveat:** engines that feed gathered PLE rows straight into mul_mat as -quantized B operands assert (apepojken-class, ggml-vulkan.cpp:7794). verified -working on nathanw1014 strix-halo-vulkan and rocmfpx. +quantized B operands assert (ggml-vulkan.cpp:7794). verified working on the +packaged engine and rocmfpx.
### pick your setup -| your use case | quant | engine | ctx | expect | -|---|---|---|---|---| -| coding agents, chat | 91g PLE | [nathanw1014 vulkan](https://github.com/Nathanw1014/llama.cpp/tree/strix-halo-vulkan) + mtp | ≤ 32k | 56 t/s | -| long documents | 116g static | same engine + mtp | 128k | 27 t/s | -| full rag / research | 116g static | [rocm 10 container](https://github.com/MorezMartin/engramhalo-rocm10) + ssd streaming | 262k | 14 t/s | +| your use case | quant | ctx | expect | +|---|---|---|---| +| coding agents, chat | 91g PLE | ≤ 32k | 56 t/s | +| long documents | 116g static | 128k | 27 t/s | +| full rag / research | 116g static + ssd streaming | 262k | 14 t/s | --- @@ -96,8 +94,9 @@ our first quant printed deterministic garbage at temp 0. bisect to root cause: > [!WARNING] > **any fork rolling its own qwen4exp converter must fold `(1 + w)` into the > hyper-connection gammas.** upstream runtime documents the contract at -> `qwen4exp.cpp:231`. miss it and every layer normalizes wrong — garbage -> from layer 0, all shapes correct, all shape-only tests pass. +> `qwen4exp.cpp:231 — "the converter folded each gamma to (1 + w)"`. miss it +> and every layer normalizes wrong — garbage from layer 0, all shapes correct, +> all shape-only tests pass. fix + regression test: rocmfpx `port-qwen4exp` commit `61b6a3b48` ([pr charlie12345/ROCmFPX#98](https://github.com/charlie12345/ROCmFPX/pull/98)) @@ -112,7 +111,11 @@ docker compose up --build # serve on :8080 — vulkan/radv, no rocm install needed ``` -add the mtp sidecar for 56 t/s: +the packaged engine is tuned for strix halo: vulkan fa/mmq kernels, graph +reuse, lazy ple streaming, quantized-kv attention — the combination behind +the 56 t/s numbers. no manual build, no host rocm install. + +add the mtp sidecar for speculative decoding: ```bash docker compose run qwen38-flash-next /app/llama-server \ @@ -120,19 +123,27 @@ docker compose run qwen38-flash-next /app/llama-server \ --spec-type draft-mtp --spec-draft-n-max 6 --spec-draft-p-min 0.75 ``` +### 262k context (ssd streaming) + +swap the model to the static 116g and enable lazy ple — the n-gram table +stays on ssd (~2.5g resident), leaving room for the full context window: + +```bash +docker compose run qwen38-flash-next /app/llama-server \ + -m /models/Qwen3.8-Flash-Next-IQ4_XS.gguf \ + -c 262144 -lm mmap --tensor-read-lazy on \ + -ngl 999 -fa on -ctk q8_0 -ctv q8_0 -ub 2048 -t 4 +``` + --- -## 🔬 engine merge (paused) +## 🔬 the n-gram table at 4-bit — what we found -we merged nathanw1014's branch (175 commits: vulkan perf stack, qwen4exp -runtime past the squash, lazy ple, spec-decode fixes) into ggml-org master — -one engine with lazy ple (262k on 66g resident) + depth fixes + master's -general improvements. paused at 53% build: the branches diverged semantically -in shared enum/base-class files. - -- [`docs/engine-cherry-pick-plan.md`](docs/engine-cherry-pick-plan.md) — full - 175-commit classification -- [`docs/engine-merge-status.md`](docs/engine-merge-status.md) — resume point +the 51b ple lookup table tolerates iq4_nl (4.25 bpw) with no quality loss +across the depth sweep. but there's a depth-dependent reversal: under mtp at +128k+, the ple quant *loses* to the static quant (18.6 vs 26.9 t/s) — the +iq4_nl noise compounds over deep n-gram history and lowers draft acceptance. +n=1, single runs. pick your file by use case. --- @@ -140,11 +151,10 @@ in shared enum/base-class files. | path | what | |------|------| -| `models/` | symlink farm to `/mnt/ssd2/models/` (never in git) | -| `results/` | dated receipts, one per investigation | -| `scripts/` | conversion pipeline, oracle, depth bench, resume helpers | -| `reddit/` | archived community threads that drove the investigation | +| `models/` | symlink farm to local ssd (never in git) | | `docs/` | engine merge plan, cherry-pick classification | +| `Dockerfile` | two-stage: vulkan engine build + slim runtime | +| `docker-compose.yml` | one-liner serving with recommended flags | ---