diff --git a/.gitignore b/.gitignore new file mode 100644 index 0000000..b625809 --- /dev/null +++ b/.gitignore @@ -0,0 +1 @@ +/ComfyUI/ diff --git a/README.md b/README.md index 649344e..00ba8a0 100644 --- a/README.md +++ b/README.md @@ -1,196 +1,137 @@ -# ComfyUI for gfx1151 (Ryzen AI MAX) +# ComfyUI for gfx1151 (Ryzen AI MAX / Strix Halo) -Dockerized ComfyUI with PyTorch & flash-attention for gfx1151 (AMD Strix Halo, Ryzen AI Max+ 395), -relying on AMD's pre-built and pre-configured environment (no custom wheels). +Dockerized ComfyUI with PyTorch & flash-attention for **gfx1151** (AMD Strix Halo, +e.g. Ryzen AI Max+ 395 — Minisforum MS-S1 Max, Framework Desktop), based on AMD's +pre-built and pre-configured `rocm/pytorch` image (no custom wheels needed). + +This is my working setup, adapted from +[ignatberesnev/comfyui-gfx1151](https://hub.docker.com/r/ignatberesnev/comfyui-gfx1151) +(see [Acknowledgements](#acknowledgements)). The difference here: this repo is meant to +be **built locally** with `docker compose build` — no image pull required. + +Reported working on: + +* Minisforum MS-S1 Max (Ryzen AI Max+ 395, gfx1151) — daily driver +* AMD Radeon RX 9700 (gfx1201), 32 GB — works out of the box (thanks!) Versions used: -* ROCm: 7.2 + +* ROCm: 7.2.4 * PyTorch: 2.9.1 * Python: 3.12 -* ComfyUI (built-in): v0.15.0 +* ComfyUI: whatever you clone (see below) -**Last updated & tested**: Feb 25, 2026, on 6.18.9 (ArchLinux), AMD RYZEN AI MAX+ 395 (Framework Desktop), -with [opencl-amd](https://aur.archlinux.org/packages/opencl-amd) packages (7.2.0-1). +> [!NOTE] +> I got this working after a day of digging, and the final solution turned out to be +> much simpler than most of what's out there. I don't claim to understand every moving +> part — if you know better, PRs and comments welcome. -> [!CAUTION] -> I kinda understand what's going on here, but not fully. It took me most of the day to figure out how to run -> ComfyUI on my Framework Desktop without it crashing (which is absurd for a CPU with _AI_ in the name), and -> the final working solution turned out to be much simpler than what I was able to find initially. -> -> That being said, it works (as of Feb 25, 2026), but I'm not sure if there's an even better / more correct way of -> achieving the same thing. Same goes for the environment variables that are supposedly making ComfyUI -> faster / resource efficient. -> -> I just want to share this solution to save someone else a couple of hours ¯\_(ツ)_/¯ +## Get started -## Get started now +### 1. Clone ComfyUI yourself -The Docker image is published to [Docker Hub](https://hub.docker.com/r/ignatberesnev/comfyui-gfx1151), -so you can, but don't have to build it yourself. - -There are two options: - -* Copy [docker-compose.yml](docker-compose.yml) and run `docker compose up -d` -* Copy [docker-run.sh](docker-run.sh) and run `./docker-run.sh`. After the first run, use `docker start comfyui-gfx1151`. - -ComfyUI will be available at http://localhost:8188. - -The starter templates should generate images without any issues. - -Once you've verified that it works, feel free to use this repository as the foundation for your own setup or workflow. - -#### Parameters - -Both options have the same pre-configured parameters, which are: - -* Allocate 8GB of shared memory (`shm_size`) for internal PyTorch / ComfyUI shenanigans, this should be plenty, - feel free to lower it. This should NOT be > than available RAM. This is NOT allocating VRAM. -* Mount `./ComfyUI` for the root of [ComfyUI](https://github.com/comfyanonymous/ComfyUI). If the directory is empty - when the container starts, it will copy a pre-cloned (baked in) version of ComfyUI. If it's not empty, it will be - used to run ComfyUI located in it. You can update this directory manually to use newer version of ComfyUI without - having to re-download the image -* Expose port `8188` for ComfyUI -* Add video + rendering devices and groups. While this just works on Arch, it might require some pre-requisite steps - on Ubuntu, I haven't checked. - -There are a couple of scripts that can check that both PyTorch and flash-attention work, you can find them below. - -#### Updating ComfyUI / dependencies - -If you need to install custom nodes or refresh ComfyUI dependencies after a manual update, -you can do it from within the container (until #4 is resolved): +The container mounts ComfyUI from the host, so you can update it independently of the +image. Clone it into this directory (or anywhere — just adjust the volume path in +`docker-compose.yml`): ```bash -docker exec -it comfyui-gfx1151 /bin/bash - -cd /opt/ComfyUI - -pip install -r requirements.txt +git clone https://github.com/comfyanonymous/ComfyUI ``` -## What's inside / how to replicate +If the mounted directory is empty when the container starts, a pre-cloned (baked-in) +copy of ComfyUI is copied into it automatically — but bringing your own clone is the +recommended way. -This image is based on AMD's [rocm/pytorch](https://hub.docker.com/r/rocm/pytorch) image that has Ubuntu 24.04, -ROCm 7.2, Python 3.12 and PyTorch 2.9.1, in which everything is configured to work together and it just works. -You can find out more about this image in -[AMD's ROCm documentation](https://rocm.docs.amd.com/projects/radeon-ryzen/en/latest/docs/install/installryz/native_linux/install-pytorch.html#use-docker-image-with-pre-installed-pytorch). - -There are only two missing pieces which this image adds: [flash-attention](https://github.com/ROCm/flash-attention/) -and, well, ComfyUI. - -There's nothing specific about the ComfyUI installation, you can actually bring your own, it should work. - -**flash-attention**, however, "doesn't work" out of the box if you run AMD's image. I'm saying "doesn't work" because, -as far as I understand, it doesn't have the frontend for it (the APIs), but it does have the backend: **Triton**. -So flash-attention can be "installed" with a special env variable `FLASH_ATTENTION_TRITON_AMD_ENABLE`, which makes -ComfyUI and other tools using flash-attention think that flash-attention is installed and works (even though it's -triton under the hood, which is actually doing the job). You can see the lines that install it in -[Dockerfile](Dockerfile), and if you try to do it yourself, you'll notice that it executes very fast -(because flash-attention isn't actually built in full). - -It's worth noting that flash-attention is cloned from a specific branch `main_perf` -- I'm not sure why exactly, -I haven't checked, but I assume it's because it has (stable?) support for Triton which is not yet in the main branch, -see [this issue](https://github.com/ROCm/flash-attention/issues/27). I basically copy-pasted this part from other -installations ([vLLM](https://community.frame.work/t/compiling-vllm-from-source-on-strix-halo/77241) and repos by -[kyuz0](https://github.com/kyuz0)), so I hope they know what they're doing :D - -In [scripts](scripts) there are two scripts that can check if PyTorch and flash-attention work as expected and utilize -the iGPU. I used these when looking for a solution, they proved to be helpful, so I'm adding them to the image in case -something breaks or doesn't work as expected, maybe they'll help debug the problem or something. - -With that knowledge, you should be able to take [Dockerfile](Dockerfile) and build an image yourself. - -If any of this makes more sense to you than it does to me and you know how to improve something or can add a helpful -comment with additional context, please do! - -## What I tried that didn't work - -The majority of other solutions seem to rely on custom-built wheels, such as the image by -[pccr10001/comfyui-gfx1151-fa](https://github.com/pccr10001/comfyui-gfx1151-fa). - -I never managed to make these custom wheels work, presumably because of non-locked dependencies (pulling in newer -version of rocm/etc with old wheels). However, they made me begin to understand what was happening and how to move -forward, so huge thanks to everyone who left any comments on the topic. - -Some other solutions also relied on the image `ghcr.io/rocm/therock_pytorch_dev_ubuntu_24_04_gfx1151`, which is no -longer published, so I never got that working either. The image I'm referencing -([rocm/pytorch](https://hub.docker.com/r/rocm/pytorch)) seems like a replacement for it though? - -Initially, I [copied over](https://github.com/pccr10001/comfyui-gfx1151-fa/blob/e6e59be08ff439ab5f9799aa2161f70709fcd975/README.md?plain=1#L33) -some environment variables that were supposed to speed up ComfyUI / PyTorch and make it more resource efficient: -`PYTORCH_TUNABLEOP_ENABLED`, `MIOPEN_FIND_MODE` and `ROCBLAS_USE_HIPBLASLT` (not adding them as a codeblock to avoid -someone copy-pasting them). However, at least one of them not only made it worse when it comes to the speed, but I -believe it would crash my display server (X11) every now and then when running stable diffusion models. Apparently, -this is relatively common to see with AMD drivers in general, so I'm not entirely sure that those env variables were -100% responsible for the crashes (might've been something else), but removing all of them helped (at least for now), -so I've removed them from this repo's scripts too. If you also experience display server crashes, let me know. - -## Tests - -There are two scripts that you can use to test if everything works correctly - -#### Test PyTorch - -While the container is running, running +### 2. Build and run ```bash -docker exec -it comfyui-gfx1151 /bin/bash /opt/comfyui-gfx1151-utils/test-pytorch.sh +docker compose up -d --build ``` -should produce NO errors. The output should be something like: +ComfyUI is available at http://localhost:8188. The starter templates should generate +images without any issues. -```text -GPU: AMD Radeon Graphics | FlashAttn: True -Mean: -0.026233481243252754 -``` +### 3. Verify the GPU works -#### Test flash-attention - -While the container is working, running +While the container is running: ```bash -docker exec -it comfyui-gfx1151 python3 /opt/comfyui-gfx1151-utils/test-pytorch-flashattention.py +# PyTorch sees the GPU +docker exec -it comfyui /bin/bash /opt/comfyui-gfx1151-utils/test-pytorch.sh + +# flash-attention (Triton backend) works +docker exec -it comfyui python3 /opt/comfyui-gfx1151-utils/test-pytorch-flashattention.py ``` -should produce NO errors. The output should be something like: +Both should run without errors (sample output in the +[upstream README](https://github.com/ignatberesnev/comfyui-gfx1151)). -```text -=== PyTorch Installation Check === -PyTorch version: 2.9.1+rocm7.1.1.git351ff442 -PyTorch ROCm version: 7.1.52802-26aae437f6 -CUDA available: True -Device count: 1 -Device name: AMD Radeon Graphics +## What's configured -=== Flash Attention Support Check === -/usr/lib/python3.12/contextlib.py:105: FutureWarning: `torch.backends.cuda.sdp_kernel()` is deprecated. In the future, this context manager will be removed. Please see `torch.nn.attention.sdpa_kernel()` for the new context manager, with updated signature. - self.gen = func(*args, **kwds) -Available SDP backends: -Flash Attention backend enabled -Test tensors created on cuda -Flash Attention test successful! Output shape: torch.Size([2, 8, 128, 64]) +In [docker-compose.yml](docker-compose.yml): -=== AOTriton Check === -AOTriton not available: No module named 'pyaotriton' +* `build: .` — image is built from the [Dockerfile](Dockerfile) in this repo +* `shm_size: 8G` — shared memory for PyTorch/ComfyUI. Should NOT be larger than + available RAM; this is **not** VRAM. Feel free to lower it. +* `./ComfyUI:/opt/ComfyUI` — your ComfyUI clone (models, custom nodes, output all live + in here) +* Port `8188` exposed +* `/dev/kfd`, `/dev/dri` + `video`/`render` groups — GPU access. Works as-is on Arch; + may need extra steps on Ubuntu, untested. +* `TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL=1` — enables the experimental AOTriton + backend -=== Environment Variables === -ROCM_PATH: Not set -HIP_PATH: Not set -HIP_PLATFORM: Not set -HIP_ARCH: Not set -HSA_OVERRIDE_GFX_VERSION: Not set +No reverse proxy / external network is configured — it binds `8188` directly. Put it +behind whatever you like if you need auth/TLS. -=== Testing GFX Version Override === -Set HSA_OVERRIDE_GFX_VERSION=11.0.0 to test gfx110x mapping -✓ Flash Attention worked with GFX override! +## Updating -=== Flash Attention Backend Detection === -✓ flash backend works -✓ mem_efficient backend works -✓ math backend works -``` +* **ComfyUI**: `cd ComfyUI && git pull`, then `docker compose restart`. Custom nodes: + the ComfyUI Manager is enabled (`--enable-manager`), or install from inside the + container: + + ```bash + docker exec -it comfyui /bin/bash + cd /opt/ComfyUI && pip install -r requirements.txt + ``` + +* **ROCm / PyTorch**: bump the `FROM rocm/pytorch:...` line in the + [Dockerfile](Dockerfile) and the `image:` tag in `docker-compose.yml`, then + `docker compose up -d --build`. + +## What's inside / how it works + +The image is based on AMD's [rocm/pytorch](https://hub.docker.com/r/rocm/pytorch) +image (Ubuntu 24.04, ROCm, Python 3.12, PyTorch) where everything is configured to +work together — see +[AMD's docs](https://rocm.docs.amd.com/projects/radeon-ryzen/en/latest/docs/install/installryz/native_linux/install-pytorch.html#use-docker-image-with-pre-installed-pytorch). + +Only two pieces are added: ComfyUI (nothing special — bring your own clone) and +**flash-attention**. The latter "doesn't work" out of the box on AMD's image: it lacks +the frontend APIs but ships the Triton backend. Setting +`FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE` makes flash-attention install fast and route +to Triton, which does the actual work. It's cloned from the `main_perf` branch of +[ROCm/flash-attention](https://github.com/ROCm/flash-attention/) — that's what others +(vLLM on Strix Halo, [kyuz0](https://github.com/kyuz0)'s setups) use, presumably for +Triton support not yet in `main` (see +[ROCm/flash-attention#27](https://github.com/ROCm/flash-attention/issues/27)). + +The `scripts/` directory holds the test scripts above; they're baked into the image at +`/opt/comfyui-gfx1151-utils/`. + +## What did NOT work (so you don't try it) + +* Custom-built wheel setups (e.g. + [pccr10001/comfyui-gfx1151-fa](https://github.com/pccr10001/comfyui-gfx1151-fa)) — + never worked for me, presumably non-locked dependencies pulling newer ROCm under old + wheels. Huge thanks to them for the pointers though. +* `ghcr.io/rocm/therock_pytorch_dev_ubuntu_24_04_gfx1151` — no longer published. +* Speed "tuning" env vars (`PYTORCH_TUNABLEOP_ENABLED`, `MIOPEN_FIND_MODE`, + `ROCBLAS_USE_HIPBLASLT`) — made things slower and (probably) crashed my X11 display + server during SD runs. Don't add them unless you enjoy that. ## Acknowledgements -Big thanks to [pccr10001](https://github.com/pccr10001), [lhl](https://github.com/lhl) and [kyuz0](https://github.com/kyuz0) for setting me on the right path! - +* [ignatberesnev](https://github.com/ignatberesnev) — original repo this is based on +* [pccr10001](https://github.com/pccr10001), [lhl](https://github.com/lhl), + [kyuz0](https://github.com/kyuz0) — for setting everyone on the right path diff --git a/docker-compose.yml b/docker-compose.yml index 51a3608..f9763c9 100644 --- a/docker-compose.yml +++ b/docker-compose.yml @@ -1,11 +1,18 @@ services: comfyui: + # Built locally from the Dockerfile in this repo (docker compose build) image: comfyui-gfx1151:v0.8 container_name: comfyui - # Shared memory allocation, should not be > than available RAM, feel free to lower it - #shm_size: 8G + build: + context: . + # Bump this in the Dockerfile (FROM rocm/pytorch:...) and here when updating + # Shared memory for PyTorch / ComfyUI internals. + # Must NOT be larger than available RAM. This is NOT VRAM. + shm_size: 8G volumes: - - /data/@sd/ComfyUI:/opt/ComfyUI + # Clone ComfyUI into this directory first (see README): + # git clone https://github.com/comfyanonymous/ComfyUI + - ./ComfyUI:/opt/ComfyUI ports: - "8188:8188" environment: @@ -15,14 +22,9 @@ services: - /dev/dri:/dev/dri group_add: - video + - render cap_add: - SYS_PTRACE security_opt: - seccomp:unconfined - networks: - - proxy-network - -networks: - proxy-network: - name: proxy-network - external: true + ipc: host diff --git a/docker-run.sh b/docker-run.sh index a629a99..84ab513 100755 --- a/docker-run.sh +++ b/docker-run.sh @@ -1,6 +1,13 @@ #!/bin/bash +# Alternative to docker compose: build and run manually. +# Clone ComfyUI into ./ComfyUI first: +# git clone https://github.com/comfyanonymous/ComfyUI -docker run -it \ +set -e + +docker build -t comfyui-gfx1151:v0.8 . + +docker run -it --rm \ --cap-add=SYS_PTRACE \ --security-opt seccomp=unconfined \ --device=/dev/kfd \ @@ -12,5 +19,5 @@ docker run -it \ -p 8188:8188 \ -v $(pwd)/ComfyUI:/opt/ComfyUI \ --shm-size 8G \ - --name comfyui-gfx1151 \ - ignatberesnev/comfyui-gfx1151:v0.2 + --name comfyui \ + comfyui-gfx1151:v0.8