# ComfyUI for gfx1151 (Ryzen AI MAX / Strix Halo) Dockerized ComfyUI with PyTorch & flash-attention for **gfx1151** (AMD Strix Halo, e.g. Ryzen AI Max+ 395 — Minisforum MS-S1 Max, Framework Desktop), based on AMD's pre-built and pre-configured `rocm/pytorch` image (no custom wheels needed). This is my working setup, adapted from [ignatberesnev/comfyui-gfx1151](https://hub.docker.com/r/ignatberesnev/comfyui-gfx1151) (see [Acknowledgements](#acknowledgements)). The difference here: this repo is meant to be **built locally** with `docker compose build` — no image pull required. Reported working on: * Minisforum MS-S1 Max (Ryzen AI Max+ 395, gfx1151) — daily driver * AMD Radeon RX 9700 (gfx1201), 32 GB — works out of the box (thanks!) Versions used: * ROCm: 7.2.4 * PyTorch: 2.9.1 * Python: 3.12 * ComfyUI: whatever you clone (see below) > [!NOTE] > I got this working after a day of digging, and the final solution turned out to be > much simpler than most of what's out there. I don't claim to understand every moving > part — if you know better, PRs and comments welcome. ## Get started ### 1. Clone ComfyUI yourself The container mounts ComfyUI from the host, so you can update it independently of the image. Clone it into this directory (or anywhere — just adjust the volume path in `docker-compose.yml`): ```bash git clone https://github.com/comfyanonymous/ComfyUI ``` If the mounted directory is empty when the container starts, a pre-cloned (baked-in) copy of ComfyUI is copied into it automatically — but bringing your own clone is the recommended way. ### 2. Build and run ```bash docker compose up -d --build ``` ComfyUI is available at http://localhost:8188. The starter templates should generate images without any issues. ### 3. Verify the GPU works While the container is running: ```bash # PyTorch sees the GPU docker exec -it comfyui /bin/bash /opt/comfyui-gfx1151-utils/test-pytorch.sh # flash-attention (Triton backend) works docker exec -it comfyui python3 /opt/comfyui-gfx1151-utils/test-pytorch-flashattention.py ``` Both should run without errors (sample output in the [upstream README](https://github.com/ignatberesnev/comfyui-gfx1151)). ## What's configured In [docker-compose.yml](docker-compose.yml): * `build: .` — image is built from the [Dockerfile](Dockerfile) in this repo * `shm_size: 8G` — shared memory for PyTorch/ComfyUI. Should NOT be larger than available RAM; this is **not** VRAM. Feel free to lower it. * `./ComfyUI:/opt/ComfyUI` — your ComfyUI clone (models, custom nodes, output all live in here) * Port `8188` exposed * `/dev/kfd`, `/dev/dri` + `video`/`render` groups — GPU access. Works as-is on Arch; may need extra steps on Ubuntu, untested. * `TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL=1` — enables the experimental AOTriton backend No reverse proxy / external network is configured — it binds `8188` directly. Put it behind whatever you like if you need auth/TLS. ## Updating * **ComfyUI**: `cd ComfyUI && git pull`, then `docker compose restart`. Custom nodes: the ComfyUI Manager is enabled (`--enable-manager`), or install from inside the container: ```bash docker exec -it comfyui /bin/bash cd /opt/ComfyUI && pip install -r requirements.txt ``` * **ROCm / PyTorch**: bump the `FROM rocm/pytorch:...` line in the [Dockerfile](Dockerfile) and the `image:` tag in `docker-compose.yml`, then `docker compose up -d --build`. ## What's inside / how it works The image is based on AMD's [rocm/pytorch](https://hub.docker.com/r/rocm/pytorch) image (Ubuntu 24.04, ROCm, Python 3.12, PyTorch) where everything is configured to work together — see [AMD's docs](https://rocm.docs.amd.com/projects/radeon-ryzen/en/latest/docs/install/installryz/native_linux/install-pytorch.html#use-docker-image-with-pre-installed-pytorch). Only two pieces are added: ComfyUI (nothing special — bring your own clone) and **flash-attention**. The latter "doesn't work" out of the box on AMD's image: it lacks the frontend APIs but ships the Triton backend. Setting `FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE` makes flash-attention install fast and route to Triton, which does the actual work. It's cloned from the `main_perf` branch of [ROCm/flash-attention](https://github.com/ROCm/flash-attention/) — that's what others (vLLM on Strix Halo, [kyuz0](https://github.com/kyuz0)'s setups) use, presumably for Triton support not yet in `main` (see [ROCm/flash-attention#27](https://github.com/ROCm/flash-attention/issues/27)). The `scripts/` directory holds the test scripts above; they're baked into the image at `/opt/comfyui-gfx1151-utils/`. ## What did NOT work (so you don't try it) * Custom-built wheel setups (e.g. [pccr10001/comfyui-gfx1151-fa](https://github.com/pccr10001/comfyui-gfx1151-fa)) — never worked for me, presumably non-locked dependencies pulling newer ROCm under old wheels. Huge thanks to them for the pointers though. * `ghcr.io/rocm/therock_pytorch_dev_ubuntu_24_04_gfx1151` — no longer published. * Speed "tuning" env vars (`PYTORCH_TUNABLEOP_ENABLED`, `MIOPEN_FIND_MODE`, `ROCBLAS_USE_HIPBLASLT`) — made things slower and (probably) crashed my X11 display server during SD runs. Don't add them unless you enjoy that. ## Acknowledgements * [ignatberesnev](https://github.com/ignatberesnev) — original repo this is based on * [pccr10001](https://github.com/pccr10001), [lhl](https://github.com/lhl), [kyuz0](https://github.com/kyuz0) — for setting everyone on the right path