Rewrite README, build from compose, drop proxy network
- README: rewritten for this fork — build with docker compose, clone ComfyUI yourself as the mount, gfx1201 (RX 9700) confirmed working - docker-compose.yml: add build: context, local image tag, ./ComfyUI volume, restore shm_size/render group/ipc, remove external proxy-network - docker-run.sh: build locally, matching image tag and container name - add .gitignore for the ComfyUI mount
This commit is contained in:
@@ -0,0 +1 @@
|
|||||||
|
/ComfyUI/
|
||||||
@@ -1,196 +1,137 @@
|
|||||||
# ComfyUI for gfx1151 (Ryzen AI MAX)
|
# ComfyUI for gfx1151 (Ryzen AI MAX / Strix Halo)
|
||||||
|
|
||||||
Dockerized ComfyUI with PyTorch & flash-attention for gfx1151 (AMD Strix Halo, Ryzen AI Max+ 395),
|
Dockerized ComfyUI with PyTorch & flash-attention for **gfx1151** (AMD Strix Halo,
|
||||||
relying on AMD's pre-built and pre-configured environment (no custom wheels).
|
e.g. Ryzen AI Max+ 395 — Minisforum MS-S1 Max, Framework Desktop), based on AMD's
|
||||||
|
pre-built and pre-configured `rocm/pytorch` image (no custom wheels needed).
|
||||||
|
|
||||||
|
This is my working setup, adapted from
|
||||||
|
[ignatberesnev/comfyui-gfx1151](https://hub.docker.com/r/ignatberesnev/comfyui-gfx1151)
|
||||||
|
(see [Acknowledgements](#acknowledgements)). The difference here: this repo is meant to
|
||||||
|
be **built locally** with `docker compose build` — no image pull required.
|
||||||
|
|
||||||
|
Reported working on:
|
||||||
|
|
||||||
|
* Minisforum MS-S1 Max (Ryzen AI Max+ 395, gfx1151) — daily driver
|
||||||
|
* AMD Radeon RX 9700 (gfx1201), 32 GB — works out of the box (thanks!)
|
||||||
|
|
||||||
Versions used:
|
Versions used:
|
||||||
* ROCm: 7.2
|
|
||||||
|
* ROCm: 7.2.4
|
||||||
* PyTorch: 2.9.1
|
* PyTorch: 2.9.1
|
||||||
* Python: 3.12
|
* Python: 3.12
|
||||||
* ComfyUI (built-in): v0.15.0
|
* ComfyUI: whatever you clone (see below)
|
||||||
|
|
||||||
**Last updated & tested**: Feb 25, 2026, on 6.18.9 (ArchLinux), AMD RYZEN AI MAX+ 395 (Framework Desktop),
|
> [!NOTE]
|
||||||
with [opencl-amd](https://aur.archlinux.org/packages/opencl-amd) packages (7.2.0-1).
|
> I got this working after a day of digging, and the final solution turned out to be
|
||||||
|
> much simpler than most of what's out there. I don't claim to understand every moving
|
||||||
|
> part — if you know better, PRs and comments welcome.
|
||||||
|
|
||||||
> [!CAUTION]
|
## Get started
|
||||||
> I kinda understand what's going on here, but not fully. It took me most of the day to figure out how to run
|
|
||||||
> ComfyUI on my Framework Desktop without it crashing (which is absurd for a CPU with _AI_ in the name), and
|
|
||||||
> the final working solution turned out to be much simpler than what I was able to find initially.
|
|
||||||
>
|
|
||||||
> That being said, it works (as of Feb 25, 2026), but I'm not sure if there's an even better / more correct way of
|
|
||||||
> achieving the same thing. Same goes for the environment variables that are supposedly making ComfyUI
|
|
||||||
> faster / resource efficient.
|
|
||||||
>
|
|
||||||
> I just want to share this solution to save someone else a couple of hours ¯\_(ツ)_/¯
|
|
||||||
|
|
||||||
## Get started now
|
### 1. Clone ComfyUI yourself
|
||||||
|
|
||||||
The Docker image is published to [Docker Hub](https://hub.docker.com/r/ignatberesnev/comfyui-gfx1151),
|
The container mounts ComfyUI from the host, so you can update it independently of the
|
||||||
so you can, but don't have to build it yourself.
|
image. Clone it into this directory (or anywhere — just adjust the volume path in
|
||||||
|
`docker-compose.yml`):
|
||||||
There are two options:
|
|
||||||
|
|
||||||
* Copy [docker-compose.yml](docker-compose.yml) and run `docker compose up -d`
|
|
||||||
* Copy [docker-run.sh](docker-run.sh) and run `./docker-run.sh`. After the first run, use `docker start comfyui-gfx1151`.
|
|
||||||
|
|
||||||
ComfyUI will be available at http://localhost:8188.
|
|
||||||
|
|
||||||
The starter templates should generate images without any issues.
|
|
||||||
|
|
||||||
Once you've verified that it works, feel free to use this repository as the foundation for your own setup or workflow.
|
|
||||||
|
|
||||||
#### Parameters
|
|
||||||
|
|
||||||
Both options have the same pre-configured parameters, which are:
|
|
||||||
|
|
||||||
* Allocate 8GB of shared memory (`shm_size`) for internal PyTorch / ComfyUI shenanigans, this should be plenty,
|
|
||||||
feel free to lower it. This should NOT be > than available RAM. This is NOT allocating VRAM.
|
|
||||||
* Mount `./ComfyUI` for the root of [ComfyUI](https://github.com/comfyanonymous/ComfyUI). If the directory is empty
|
|
||||||
when the container starts, it will copy a pre-cloned (baked in) version of ComfyUI. If it's not empty, it will be
|
|
||||||
used to run ComfyUI located in it. You can update this directory manually to use newer version of ComfyUI without
|
|
||||||
having to re-download the image
|
|
||||||
* Expose port `8188` for ComfyUI
|
|
||||||
* Add video + rendering devices and groups. While this just works on Arch, it might require some pre-requisite steps
|
|
||||||
on Ubuntu, I haven't checked.
|
|
||||||
|
|
||||||
There are a couple of scripts that can check that both PyTorch and flash-attention work, you can find them below.
|
|
||||||
|
|
||||||
#### Updating ComfyUI / dependencies
|
|
||||||
|
|
||||||
If you need to install custom nodes or refresh ComfyUI dependencies after a manual update,
|
|
||||||
you can do it from within the container (until #4 is resolved):
|
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
docker exec -it comfyui-gfx1151 /bin/bash
|
git clone https://github.com/comfyanonymous/ComfyUI
|
||||||
|
|
||||||
cd /opt/ComfyUI
|
|
||||||
|
|
||||||
pip install -r requirements.txt
|
|
||||||
```
|
```
|
||||||
|
|
||||||
## What's inside / how to replicate
|
If the mounted directory is empty when the container starts, a pre-cloned (baked-in)
|
||||||
|
copy of ComfyUI is copied into it automatically — but bringing your own clone is the
|
||||||
|
recommended way.
|
||||||
|
|
||||||
This image is based on AMD's [rocm/pytorch](https://hub.docker.com/r/rocm/pytorch) image that has Ubuntu 24.04,
|
### 2. Build and run
|
||||||
ROCm 7.2, Python 3.12 and PyTorch 2.9.1, in which everything is configured to work together and it just works.
|
|
||||||
You can find out more about this image in
|
|
||||||
[AMD's ROCm documentation](https://rocm.docs.amd.com/projects/radeon-ryzen/en/latest/docs/install/installryz/native_linux/install-pytorch.html#use-docker-image-with-pre-installed-pytorch).
|
|
||||||
|
|
||||||
There are only two missing pieces which this image adds: [flash-attention](https://github.com/ROCm/flash-attention/)
|
|
||||||
and, well, ComfyUI.
|
|
||||||
|
|
||||||
There's nothing specific about the ComfyUI installation, you can actually bring your own, it should work.
|
|
||||||
|
|
||||||
**flash-attention**, however, "doesn't work" out of the box if you run AMD's image. I'm saying "doesn't work" because,
|
|
||||||
as far as I understand, it doesn't have the frontend for it (the APIs), but it does have the backend: **Triton**.
|
|
||||||
So flash-attention can be "installed" with a special env variable `FLASH_ATTENTION_TRITON_AMD_ENABLE`, which makes
|
|
||||||
ComfyUI and other tools using flash-attention think that flash-attention is installed and works (even though it's
|
|
||||||
triton under the hood, which is actually doing the job). You can see the lines that install it in
|
|
||||||
[Dockerfile](Dockerfile), and if you try to do it yourself, you'll notice that it executes very fast
|
|
||||||
(because flash-attention isn't actually built in full).
|
|
||||||
|
|
||||||
It's worth noting that flash-attention is cloned from a specific branch `main_perf` -- I'm not sure why exactly,
|
|
||||||
I haven't checked, but I assume it's because it has (stable?) support for Triton which is not yet in the main branch,
|
|
||||||
see [this issue](https://github.com/ROCm/flash-attention/issues/27). I basically copy-pasted this part from other
|
|
||||||
installations ([vLLM](https://community.frame.work/t/compiling-vllm-from-source-on-strix-halo/77241) and repos by
|
|
||||||
[kyuz0](https://github.com/kyuz0)), so I hope they know what they're doing :D
|
|
||||||
|
|
||||||
In [scripts](scripts) there are two scripts that can check if PyTorch and flash-attention work as expected and utilize
|
|
||||||
the iGPU. I used these when looking for a solution, they proved to be helpful, so I'm adding them to the image in case
|
|
||||||
something breaks or doesn't work as expected, maybe they'll help debug the problem or something.
|
|
||||||
|
|
||||||
With that knowledge, you should be able to take [Dockerfile](Dockerfile) and build an image yourself.
|
|
||||||
|
|
||||||
If any of this makes more sense to you than it does to me and you know how to improve something or can add a helpful
|
|
||||||
comment with additional context, please do!
|
|
||||||
|
|
||||||
## What I tried that didn't work
|
|
||||||
|
|
||||||
The majority of other solutions seem to rely on custom-built wheels, such as the image by
|
|
||||||
[pccr10001/comfyui-gfx1151-fa](https://github.com/pccr10001/comfyui-gfx1151-fa).
|
|
||||||
|
|
||||||
I never managed to make these custom wheels work, presumably because of non-locked dependencies (pulling in newer
|
|
||||||
version of rocm/etc with old wheels). However, they made me begin to understand what was happening and how to move
|
|
||||||
forward, so huge thanks to everyone who left any comments on the topic.
|
|
||||||
|
|
||||||
Some other solutions also relied on the image `ghcr.io/rocm/therock_pytorch_dev_ubuntu_24_04_gfx1151`, which is no
|
|
||||||
longer published, so I never got that working either. The image I'm referencing
|
|
||||||
([rocm/pytorch](https://hub.docker.com/r/rocm/pytorch)) seems like a replacement for it though?
|
|
||||||
|
|
||||||
Initially, I [copied over](https://github.com/pccr10001/comfyui-gfx1151-fa/blob/e6e59be08ff439ab5f9799aa2161f70709fcd975/README.md?plain=1#L33)
|
|
||||||
some environment variables that were supposed to speed up ComfyUI / PyTorch and make it more resource efficient:
|
|
||||||
`PYTORCH_TUNABLEOP_ENABLED`, `MIOPEN_FIND_MODE` and `ROCBLAS_USE_HIPBLASLT` (not adding them as a codeblock to avoid
|
|
||||||
someone copy-pasting them). However, at least one of them not only made it worse when it comes to the speed, but I
|
|
||||||
believe it would crash my display server (X11) every now and then when running stable diffusion models. Apparently,
|
|
||||||
this is relatively common to see with AMD drivers in general, so I'm not entirely sure that those env variables were
|
|
||||||
100% responsible for the crashes (might've been something else), but removing all of them helped (at least for now),
|
|
||||||
so I've removed them from this repo's scripts too. If you also experience display server crashes, let me know.
|
|
||||||
|
|
||||||
## Tests
|
|
||||||
|
|
||||||
There are two scripts that you can use to test if everything works correctly
|
|
||||||
|
|
||||||
#### Test PyTorch
|
|
||||||
|
|
||||||
While the container is running, running
|
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
docker exec -it comfyui-gfx1151 /bin/bash /opt/comfyui-gfx1151-utils/test-pytorch.sh
|
docker compose up -d --build
|
||||||
```
|
```
|
||||||
|
|
||||||
should produce NO errors. The output should be something like:
|
ComfyUI is available at http://localhost:8188. The starter templates should generate
|
||||||
|
images without any issues.
|
||||||
|
|
||||||
```text
|
### 3. Verify the GPU works
|
||||||
GPU: AMD Radeon Graphics | FlashAttn: True
|
|
||||||
Mean: -0.026233481243252754
|
|
||||||
```
|
|
||||||
|
|
||||||
#### Test flash-attention
|
While the container is running:
|
||||||
|
|
||||||
While the container is working, running
|
|
||||||
|
|
||||||
```bash
|
```bash
|
||||||
docker exec -it comfyui-gfx1151 python3 /opt/comfyui-gfx1151-utils/test-pytorch-flashattention.py
|
# PyTorch sees the GPU
|
||||||
|
docker exec -it comfyui /bin/bash /opt/comfyui-gfx1151-utils/test-pytorch.sh
|
||||||
|
|
||||||
|
# flash-attention (Triton backend) works
|
||||||
|
docker exec -it comfyui python3 /opt/comfyui-gfx1151-utils/test-pytorch-flashattention.py
|
||||||
```
|
```
|
||||||
|
|
||||||
should produce NO errors. The output should be something like:
|
Both should run without errors (sample output in the
|
||||||
|
[upstream README](https://github.com/ignatberesnev/comfyui-gfx1151)).
|
||||||
|
|
||||||
```text
|
## What's configured
|
||||||
=== PyTorch Installation Check ===
|
|
||||||
PyTorch version: 2.9.1+rocm7.1.1.git351ff442
|
|
||||||
PyTorch ROCm version: 7.1.52802-26aae437f6
|
|
||||||
CUDA available: True
|
|
||||||
Device count: 1
|
|
||||||
Device name: AMD Radeon Graphics
|
|
||||||
|
|
||||||
=== Flash Attention Support Check ===
|
In [docker-compose.yml](docker-compose.yml):
|
||||||
/usr/lib/python3.12/contextlib.py:105: FutureWarning: `torch.backends.cuda.sdp_kernel()` is deprecated. In the future, this context manager will be removed. Please see `torch.nn.attention.sdpa_kernel()` for the new context manager, with updated signature.
|
|
||||||
self.gen = func(*args, **kwds)
|
|
||||||
Available SDP backends: <contextlib._GeneratorContextManager object at 0x7f9f4568c1d0>
|
|
||||||
Flash Attention backend enabled
|
|
||||||
Test tensors created on cuda
|
|
||||||
Flash Attention test successful! Output shape: torch.Size([2, 8, 128, 64])
|
|
||||||
|
|
||||||
=== AOTriton Check ===
|
* `build: .` — image is built from the [Dockerfile](Dockerfile) in this repo
|
||||||
AOTriton not available: No module named 'pyaotriton'
|
* `shm_size: 8G` — shared memory for PyTorch/ComfyUI. Should NOT be larger than
|
||||||
|
available RAM; this is **not** VRAM. Feel free to lower it.
|
||||||
|
* `./ComfyUI:/opt/ComfyUI` — your ComfyUI clone (models, custom nodes, output all live
|
||||||
|
in here)
|
||||||
|
* Port `8188` exposed
|
||||||
|
* `/dev/kfd`, `/dev/dri` + `video`/`render` groups — GPU access. Works as-is on Arch;
|
||||||
|
may need extra steps on Ubuntu, untested.
|
||||||
|
* `TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL=1` — enables the experimental AOTriton
|
||||||
|
backend
|
||||||
|
|
||||||
=== Environment Variables ===
|
No reverse proxy / external network is configured — it binds `8188` directly. Put it
|
||||||
ROCM_PATH: Not set
|
behind whatever you like if you need auth/TLS.
|
||||||
HIP_PATH: Not set
|
|
||||||
HIP_PLATFORM: Not set
|
|
||||||
HIP_ARCH: Not set
|
|
||||||
HSA_OVERRIDE_GFX_VERSION: Not set
|
|
||||||
|
|
||||||
=== Testing GFX Version Override ===
|
## Updating
|
||||||
Set HSA_OVERRIDE_GFX_VERSION=11.0.0 to test gfx110x mapping
|
|
||||||
✓ Flash Attention worked with GFX override!
|
|
||||||
|
|
||||||
=== Flash Attention Backend Detection ===
|
* **ComfyUI**: `cd ComfyUI && git pull`, then `docker compose restart`. Custom nodes:
|
||||||
✓ flash backend works
|
the ComfyUI Manager is enabled (`--enable-manager`), or install from inside the
|
||||||
✓ mem_efficient backend works
|
container:
|
||||||
✓ math backend works
|
|
||||||
```
|
```bash
|
||||||
|
docker exec -it comfyui /bin/bash
|
||||||
|
cd /opt/ComfyUI && pip install -r requirements.txt
|
||||||
|
```
|
||||||
|
|
||||||
|
* **ROCm / PyTorch**: bump the `FROM rocm/pytorch:...` line in the
|
||||||
|
[Dockerfile](Dockerfile) and the `image:` tag in `docker-compose.yml`, then
|
||||||
|
`docker compose up -d --build`.
|
||||||
|
|
||||||
|
## What's inside / how it works
|
||||||
|
|
||||||
|
The image is based on AMD's [rocm/pytorch](https://hub.docker.com/r/rocm/pytorch)
|
||||||
|
image (Ubuntu 24.04, ROCm, Python 3.12, PyTorch) where everything is configured to
|
||||||
|
work together — see
|
||||||
|
[AMD's docs](https://rocm.docs.amd.com/projects/radeon-ryzen/en/latest/docs/install/installryz/native_linux/install-pytorch.html#use-docker-image-with-pre-installed-pytorch).
|
||||||
|
|
||||||
|
Only two pieces are added: ComfyUI (nothing special — bring your own clone) and
|
||||||
|
**flash-attention**. The latter "doesn't work" out of the box on AMD's image: it lacks
|
||||||
|
the frontend APIs but ships the Triton backend. Setting
|
||||||
|
`FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE` makes flash-attention install fast and route
|
||||||
|
to Triton, which does the actual work. It's cloned from the `main_perf` branch of
|
||||||
|
[ROCm/flash-attention](https://github.com/ROCm/flash-attention/) — that's what others
|
||||||
|
(vLLM on Strix Halo, [kyuz0](https://github.com/kyuz0)'s setups) use, presumably for
|
||||||
|
Triton support not yet in `main` (see
|
||||||
|
[ROCm/flash-attention#27](https://github.com/ROCm/flash-attention/issues/27)).
|
||||||
|
|
||||||
|
The `scripts/` directory holds the test scripts above; they're baked into the image at
|
||||||
|
`/opt/comfyui-gfx1151-utils/`.
|
||||||
|
|
||||||
|
## What did NOT work (so you don't try it)
|
||||||
|
|
||||||
|
* Custom-built wheel setups (e.g.
|
||||||
|
[pccr10001/comfyui-gfx1151-fa](https://github.com/pccr10001/comfyui-gfx1151-fa)) —
|
||||||
|
never worked for me, presumably non-locked dependencies pulling newer ROCm under old
|
||||||
|
wheels. Huge thanks to them for the pointers though.
|
||||||
|
* `ghcr.io/rocm/therock_pytorch_dev_ubuntu_24_04_gfx1151` — no longer published.
|
||||||
|
* Speed "tuning" env vars (`PYTORCH_TUNABLEOP_ENABLED`, `MIOPEN_FIND_MODE`,
|
||||||
|
`ROCBLAS_USE_HIPBLASLT`) — made things slower and (probably) crashed my X11 display
|
||||||
|
server during SD runs. Don't add them unless you enjoy that.
|
||||||
|
|
||||||
## Acknowledgements
|
## Acknowledgements
|
||||||
|
|
||||||
Big thanks to [pccr10001](https://github.com/pccr10001), [lhl](https://github.com/lhl) and [kyuz0](https://github.com/kyuz0) for setting me on the right path!
|
* [ignatberesnev](https://github.com/ignatberesnev) — original repo this is based on
|
||||||
|
* [pccr10001](https://github.com/pccr10001), [lhl](https://github.com/lhl),
|
||||||
|
[kyuz0](https://github.com/kyuz0) — for setting everyone on the right path
|
||||||
|
|||||||
+12
-10
@@ -1,11 +1,18 @@
|
|||||||
services:
|
services:
|
||||||
comfyui:
|
comfyui:
|
||||||
|
# Built locally from the Dockerfile in this repo (docker compose build)
|
||||||
image: comfyui-gfx1151:v0.8
|
image: comfyui-gfx1151:v0.8
|
||||||
container_name: comfyui
|
container_name: comfyui
|
||||||
# Shared memory allocation, should not be > than available RAM, feel free to lower it
|
build:
|
||||||
#shm_size: 8G
|
context: .
|
||||||
|
# Bump this in the Dockerfile (FROM rocm/pytorch:...) and here when updating
|
||||||
|
# Shared memory for PyTorch / ComfyUI internals.
|
||||||
|
# Must NOT be larger than available RAM. This is NOT VRAM.
|
||||||
|
shm_size: 8G
|
||||||
volumes:
|
volumes:
|
||||||
- /data/@sd/ComfyUI:/opt/ComfyUI
|
# Clone ComfyUI into this directory first (see README):
|
||||||
|
# git clone https://github.com/comfyanonymous/ComfyUI
|
||||||
|
- ./ComfyUI:/opt/ComfyUI
|
||||||
ports:
|
ports:
|
||||||
- "8188:8188"
|
- "8188:8188"
|
||||||
environment:
|
environment:
|
||||||
@@ -15,14 +22,9 @@ services:
|
|||||||
- /dev/dri:/dev/dri
|
- /dev/dri:/dev/dri
|
||||||
group_add:
|
group_add:
|
||||||
- video
|
- video
|
||||||
|
- render
|
||||||
cap_add:
|
cap_add:
|
||||||
- SYS_PTRACE
|
- SYS_PTRACE
|
||||||
security_opt:
|
security_opt:
|
||||||
- seccomp:unconfined
|
- seccomp:unconfined
|
||||||
networks:
|
ipc: host
|
||||||
- proxy-network
|
|
||||||
|
|
||||||
networks:
|
|
||||||
proxy-network:
|
|
||||||
name: proxy-network
|
|
||||||
external: true
|
|
||||||
|
|||||||
+10
-3
@@ -1,6 +1,13 @@
|
|||||||
#!/bin/bash
|
#!/bin/bash
|
||||||
|
# Alternative to docker compose: build and run manually.
|
||||||
|
# Clone ComfyUI into ./ComfyUI first:
|
||||||
|
# git clone https://github.com/comfyanonymous/ComfyUI
|
||||||
|
|
||||||
docker run -it \
|
set -e
|
||||||
|
|
||||||
|
docker build -t comfyui-gfx1151:v0.8 .
|
||||||
|
|
||||||
|
docker run -it --rm \
|
||||||
--cap-add=SYS_PTRACE \
|
--cap-add=SYS_PTRACE \
|
||||||
--security-opt seccomp=unconfined \
|
--security-opt seccomp=unconfined \
|
||||||
--device=/dev/kfd \
|
--device=/dev/kfd \
|
||||||
@@ -12,5 +19,5 @@ docker run -it \
|
|||||||
-p 8188:8188 \
|
-p 8188:8188 \
|
||||||
-v $(pwd)/ComfyUI:/opt/ComfyUI \
|
-v $(pwd)/ComfyUI:/opt/ComfyUI \
|
||||||
--shm-size 8G \
|
--shm-size 8G \
|
||||||
--name comfyui-gfx1151 \
|
--name comfyui \
|
||||||
ignatberesnev/comfyui-gfx1151:v0.2
|
comfyui-gfx1151:v0.8
|
||||||
|
|||||||
Reference in New Issue
Block a user