readme: clean public version

This commit is contained in:
Julian Beltran
2026-08-31 23:57:01 +10:00
parent a7c7aac004
commit 2555344204
+40 -30
View File
@@ -5,12 +5,10 @@
**qwen3.8-flash-next on amd strix halo — 91g quant, 56 tok/s, 262k context** **qwen3.8-flash-next on amd strix halo — 91g quant, 56 tok/s, 262k context**
[![HF Model](https://img.shields.io/badge/🤗_Model-IQ4_XS_PLE-ffD21E)](https://huggingface.co/julianmb/Qwen3.8-Flash-Next-IQ4_XS-GGUF) [![HF Model](https://img.shields.io/badge/🤗_Model-IQ4_XS_PLE-ffD21E)](https://huggingface.co/julianmb/Qwen3.8-Flash-Next-IQ4_XS-GGUF)
[![Engine](https://img.shields.io/badge/Engine-nathanw1014__vulkan-blue)](https://github.com/Nathanw1014/llama.cpp/tree/strix-halo-vulkan)
[![License](https://img.shields.io/badge/License-Qwen_Community-orange)](https://huggingface.co/Qwen/Qwen3.8-Flash-Next/blob/main/LICENSE)
[![Speed](https://img.shields.io/badge/MTP_@8k-56.4_t%2Fs-brightgreen)](#results) [![Speed](https://img.shields.io/badge/MTP_@8k-56.4_t%2Fs-brightgreen)](#results)
[![Depth](https://img.shields.io/badge/Verified-0_→_256k-blueviolet)](#results) [![Depth](https://img.shields.io/badge/Verified-0_→_256k-blueviolet)](#results)
[![Provenance](https://img.shields.io/badge/Provenance-byte--verified-success)](#the-converter-bug) [![Provenance](https://img.shields.io/badge/Provenance-byte--verified-success)](#the-converter-bug)
[![Docker](https://img.shields.io/badge/Docker-ready-2496ED)](#docker)
*every published quant byte-traced back to the official checkpoint* *every published quant byte-traced back to the official checkpoint*
@@ -20,9 +18,9 @@
## results ## results
engine: [nathanw1014/llama.cpp `strix-halo-vulkan`](https://github.com/Nathanw1014/llama.cpp/tree/strix-halo-vulkan) · vulkan/radv · q8_0 kv · `-ub 2048` · temp 0 91g quant · vulkan/radv · mtp sidecar · q8_0 kv · `-ub 2048` · temp 0 · 128g strix halo
| depth | plain pp/tg | mtp pp/tg | | depth | plain pp / tg | mtp pp / tg |
|------:|:-----------:|:---------:| |------:|:-----------:|:---------:|
| 0 | 92.5 / 29.9 | 87.0 / **53.1** | | 0 | 92.5 / 29.9 | 87.0 / **53.1** |
| 8k | 480 / 24.1 | 458 / **56.4** | | 8k | 480 / 24.1 | 458 / **56.4** |
@@ -54,7 +52,7 @@ single runs, n=1 caveat. pick your file by use case — see the table above.
|------|------|:------------:| |------|------|:------------:|
| `...-IQ4_XS-`**`PLE`**`.gguf` | 91 giB | ctx ≤ 32k — wins everywhere, mtp to 56 t/s | | `...-IQ4_XS-`**`PLE`**`.gguf` | 91 giB | ctx ≤ 32k — wins everywhere, mtp to 56 t/s |
| `...-IQ4_XS.gguf` | 116 giB | ctx ≥ 128k — faster mtp at depth, wider fork compat | | `...-IQ4_XS.gguf` | 116 giB | ctx ≥ 128k — faster mtp at depth, wider fork compat |
| `mtp-...-Q8_0.gguf` | 3.9 giB | mtp sidecar for nathanw1014-lineage engines | | `mtp-...-Q8_0.gguf` | 3.9 giB | mtp sidecar, required for the speed numbers |
<details> <details>
<summary><b>the PLE cut — why the 51b n-gram table tolerates 4-bit</b></summary> <summary><b>the PLE cut — why the 51b n-gram table tolerates 4-bit</b></summary>
@@ -66,18 +64,18 @@ depth sweep. the cut: `--tensor-type "per_layer_token_embd=IQ4_XS"` on our
quantizer → 54g → 27g. quantizer → 54g → 27g.
**fork caveat:** engines that feed gathered PLE rows straight into mul_mat as **fork caveat:** engines that feed gathered PLE rows straight into mul_mat as
quantized B operands assert (apepojken-class, ggml-vulkan.cpp:7794). verified quantized B operands assert (ggml-vulkan.cpp:7794). verified working on the
working on nathanw1014 strix-halo-vulkan and rocmfpx. packaged engine and rocmfpx.
</details> </details>
### pick your setup ### pick your setup
| your use case | quant | engine | ctx | expect | | your use case | quant | ctx | expect |
|---|---|---|---|---| |---|---|---|---|
| coding agents, chat | 91g PLE | [nathanw1014 vulkan](https://github.com/Nathanw1014/llama.cpp/tree/strix-halo-vulkan) + mtp | ≤ 32k | 56 t/s | | coding agents, chat | 91g PLE | ≤ 32k | 56 t/s |
| long documents | 116g static | same engine + mtp | 128k | 27 t/s | | long documents | 116g static | 128k | 27 t/s |
| full rag / research | 116g static | [rocm 10 container](https://github.com/MorezMartin/engramhalo-rocm10) + ssd streaming | 262k | 14 t/s | | full rag / research | 116g static + ssd streaming | 262k | 14 t/s |
--- ---
@@ -96,8 +94,9 @@ our first quant printed deterministic garbage at temp 0. bisect to root cause:
> [!WARNING] > [!WARNING]
> **any fork rolling its own qwen4exp converter must fold `(1 + w)` into the > **any fork rolling its own qwen4exp converter must fold `(1 + w)` into the
> hyper-connection gammas.** upstream runtime documents the contract at > hyper-connection gammas.** upstream runtime documents the contract at
> `qwen4exp.cpp:231`. miss it and every layer normalizes wrong — garbage > `qwen4exp.cpp:231 — "the converter folded each gamma to (1 + w)"`. miss it
> from layer 0, all shapes correct, all shape-only tests pass. > and every layer normalizes wrong — garbage from layer 0, all shapes correct,
> all shape-only tests pass.
fix + regression test: rocmfpx `port-qwen4exp` commit `61b6a3b48` fix + regression test: rocmfpx `port-qwen4exp` commit `61b6a3b48`
([pr charlie12345/ROCmFPX#98](https://github.com/charlie12345/ROCmFPX/pull/98)) ([pr charlie12345/ROCmFPX#98](https://github.com/charlie12345/ROCmFPX/pull/98))
@@ -112,7 +111,11 @@ docker compose up --build
# serve on :8080 — vulkan/radv, no rocm install needed # serve on :8080 — vulkan/radv, no rocm install needed
``` ```
add the mtp sidecar for 56 t/s: the packaged engine is tuned for strix halo: vulkan fa/mmq kernels, graph
reuse, lazy ple streaming, quantized-kv attention — the combination behind
the 56 t/s numbers. no manual build, no host rocm install.
add the mtp sidecar for speculative decoding:
```bash ```bash
docker compose run qwen38-flash-next /app/llama-server \ docker compose run qwen38-flash-next /app/llama-server \
@@ -120,19 +123,27 @@ docker compose run qwen38-flash-next /app/llama-server \
--spec-type draft-mtp --spec-draft-n-max 6 --spec-draft-p-min 0.75 --spec-type draft-mtp --spec-draft-n-max 6 --spec-draft-p-min 0.75
``` ```
### 262k context (ssd streaming)
swap the model to the static 116g and enable lazy ple — the n-gram table
stays on ssd (~2.5g resident), leaving room for the full context window:
```bash
docker compose run qwen38-flash-next /app/llama-server \
-m /models/Qwen3.8-Flash-Next-IQ4_XS.gguf \
-c 262144 -lm mmap --tensor-read-lazy on \
-ngl 999 -fa on -ctk q8_0 -ctv q8_0 -ub 2048 -t 4
```
--- ---
## 🔬 engine merge (paused) ## 🔬 the n-gram table at 4-bit — what we found
we merged nathanw1014's branch (175 commits: vulkan perf stack, qwen4exp the 51b ple lookup table tolerates iq4_nl (4.25 bpw) with no quality loss
runtime past the squash, lazy ple, spec-decode fixes) into ggml-org master — across the depth sweep. but there's a depth-dependent reversal: under mtp at
one engine with lazy ple (262k on 66g resident) + depth fixes + master's 128k+, the ple quant *loses* to the static quant (18.6 vs 26.9 t/s) — the
general improvements. paused at 53% build: the branches diverged semantically iq4_nl noise compounds over deep n-gram history and lowers draft acceptance.
in shared enum/base-class files. n=1, single runs. pick your file by use case.
- [`docs/engine-cherry-pick-plan.md`](docs/engine-cherry-pick-plan.md) — full
175-commit classification
- [`docs/engine-merge-status.md`](docs/engine-merge-status.md) — resume point
--- ---
@@ -140,11 +151,10 @@ in shared enum/base-class files.
| path | what | | path | what |
|------|------| |------|------|
| `models/` | symlink farm to `/mnt/ssd2/models/` (never in git) | | `models/` | symlink farm to local ssd (never in git) |
| `results/` | dated receipts, one per investigation |
| `scripts/` | conversion pipeline, oracle, depth bench, resume helpers |
| `reddit/` | archived community threads that drove the investigation |
| `docs/` | engine merge plan, cherry-pick classification | | `docs/` | engine merge plan, cherry-pick classification |
| `Dockerfile` | two-stage: vulkan engine build + slim runtime |
| `docker-compose.yml` | one-liner serving with recommended flags |
--- ---