Files
haloq38flash/docker-compose.yml
T
hermes 7f3096e86e Dockerfile: fix build on ubuntu 24.04
- ppa:kisak/kisak does not exist (repo is kisak-mesa) and add-apt-repository
  failed outright; noble-updates already ships mesa 25.2.x with gfx1151 radv,
  so drop the PPA and use the distro mesa
- pin a modern shaderc/glslc (v2026.3) from conda-forge: noble's glslc
  (shaderc 2023.8) predates GL_NV_cooperative_matrix2 and the fork's
  coopmat2 vulkan kernels; the pinned stack is self-contained (rpath)
- add missing build deps this fork needs: pkg-config, spirv-headers
  (find_package(SPIRV-Headers) in ggml-vulkan), libssl-dev (httplib TLS),
  unzip/zstd for the glslc bundle
- copy all shared libs from bin/ (fork produces libmtmd, *-impl libs and
  backend .so's the old glob missed) + register /app via ld.so.conf
- drop LLAMA_CURL (deprecated no-op in this fork)
- CMD is now bare /app/llama-server so the image can be driven by an
  external runtime (llamaswap); host/port stay overridable via
  LLAMA_ARG_HOST/LLAMA_ARG_PORT env. compose passes the recommended flags

Verified by a clean-room build of the strix-halo-vulkan branch with
noble's toolchain + headers (gcc 13, vulkan-headers 1.3.275, openssl 3.0.13):
all three targets link and resolve against stock noble runtime libs,
llama-server/cli/bench boot successfully.
2026-09-01 14:44:28 +00:00

30 lines
910 B
YAML

services:
qwen38-flash-next:
build: .
image: haloq38flash:latest
container_name: qwen38-flash-next
devices:
- /dev/dri
group_add:
- video
- render
security_opt:
- seccomp=unconfined
volumes:
- /mnt/ssd2/models/qwen38-flash-next:/models:ro
ports:
- "8080:8080"
environment:
- LD_LIBRARY_PATH=/app
# the image CMD is bare (/app/llama-server) — pass the model + flags:
command: >-
-m /models/Qwen3.8-Flash-Next-IQ4_XS-PLE.gguf
-ngl 999 -fa on -ctk q8_0 -ctv q8_0
-c 32768 -ub 2048 -t 4 --jinja
# for mtp speculative decoding:
# command: >-
# -m /models/Qwen3.8-Flash-Next-IQ4_XS-PLE.gguf
# -md /models/mtp-Qwen3.8-Flash-Next-Q8_0.gguf
# --spec-type draft-mtp --spec-draft-n-max 6 --spec-draft-p-min 0.75
# -ngl 999 -fa on -ctk q8_0 -ctv q8_0 -c 32768 -ub 2048 -t 4 --jinja