Where the RX 9070 stands today
TRELLIS.2 ROCm source fork
Identical ISA to the tested 9070 XT, so the same GPU_ARCHS value and dependency set apply. A fork of the upstream repository that ships HIP builds of FlexGEMM, CuMesh, nvdiffrast and nvdiffrec, plus a setup script that detects CUDA or ROCm and installs the matching dependencies. Validated on an RX 9070 XT 16 GB.
Open that guide →16 GB GDDR6. The upstream TRELLIS.2 project documents Linux and a 24 GB NVIDIA card as its standard path, so every route below is community work rather than a supported configuration. What follows separates what has actually been reported on this card from what is inferred from a shared instruction set.
What is actually reported
- Shares the gfx1201 target with the tested RX 9070 XT; compiled HIP kernels built for one apply to the other.
- AMD lists the RX 9070 alongside the 9070 XT as supported by its Windows PyTorch package.
Memory and weight formats
Weight size is not peak VRAM. The figures below describe how much of the frame buffer the weights themselves occupy; the texture stage and working buffers sit on top of that. With 16 GB, these are the formats worth trying in order.
- Q8 · ≈ 9.5 GB · 16 GB practical
- Q4 · ≈ 6 GB · 12 GB practical, 8 GB tight
- GGUF K-quants (Q4_K_M – Q6_K) · ≈ 6 GB – 8 GB · 8 GB reported workable
Full detail on the two quantized options lives in the Q8 guide and the Q4 guide. If your workflow goes through ComfyUI rather than a native binary, the GGUF guide covers loader compatibility, which is the part that actually decides whether a quantized file opens.
What to watch for on this card
- Compatibility is inferred from the XT, not separately reported. Treat the first run as a validation run.
- Lower clocks and compute units mean longer generation times at the same settings, not lower memory use.
RDNA 4 in the full matrix
The most actively worked-on target. A dedicated TRELLIS.2 ROCm fork validates the RX 9070 XT 16 GB on Linux, and the ComfyUI Windows effort names RDNA 4 as its first working architecture.
| GPU | VRAM | ISA | ROCm forkLinux | ComfyUI LinuxLinux | ComfyUI WindowsWindows 11 | trellis.cpp ROCmLinux · Windows | trellis.cpp VulkanLinux · Windows |
|---|---|---|---|---|---|---|---|
| RDNA 4 gfx1200 · gfx1201 | |||||||
| RX 9070 XT | 16 GB | gfx1201 | ●Tested | ○Expected | ◐Reported | ○Expected | ○Expected |
| RX 9070 | 16 GB | gfx1201 | ○Expected | ○Expected | ○Expected | ○Expected | ○Expected |
| RX 9060 XT | 16 GB | gfx1200 | ○Expected | –Unknown | –Unknown | ○Expected | ○Expected |
- ●Tested
A named project or maintainer reports this exact card completing the pipeline end to end.
- ◐Reported
Community reports exist but come from work in progress, partial runs, or a single tester.
- ○Expected
Inferred from a shared instruction set with a tested card. No direct report yet.
- ×Blocked
A known dependency, kernel, or runtime gap prevents this path today.
- –Unknown
No usable report either way. Treat as untested rather than as a failure.
Every AMD route on this page is community or third-party work. Neither Microsoft nor AMD lists TRELLIS.2 as an officially supported workload, and AMD's own documentation still limits Windows to PyTorch rather than the full ROCm stack.
TRELLIS.2 ROCm source fork
A fork of the upstream repository that ships HIP builds of FlexGEMM, CuMesh, nvdiffrast and nvdiffrec, plus a setup script that detects CUDA or ROCm and installs the matching dependencies. Validated on an RX 9070 XT 16 GB.
- OS
- Linux
- Stack
- ROCm 7.2 · PyTorch rocm7.2 · Python 3.10+
- Python runtime
- Required
ComfyUI + ROCm on Linux
The community ComfyUI wrapper plus a patch set that fixes the hardcoded GPU architecture flag, package paths and checkpoint downloads. Reported working end to end on a 7900 XTX for both shape-only and textured runs.
- OS
- Linux
- Stack
- ROCm 7.2.2 · PyTorch 2.11+rocm7.2 · Python 3.10–3.12
- Python runtime
- Required
ComfyUI + ROCm on Windows
Work in progress. RDNA 4 is confirmed running and testers are being recruited; RDNA 3 and RDNA 3.5 are next in line. The maintainer warns about silent bugs that can change the final output without raising an error.
- OS
- Windows 11
- Stack
- PyTorch 2.9.1 + ROCm 7.2.1, or PyTorch 2.12 + ROCm 7.14 · Python 3.12
- Python runtime
- Required
trellis.cpp — ROCm build
A native C++/GGML reimplementation of the whole TRELLIS.2 pipeline. Prebuilt ROCm archives are published for Linux and Windows, which removes the CUDA-only Python wheels from the problem entirely.
- OS
- Linux · Windows
- Stack
- Prebuilt ROCm/HIP binaries · no Python runtime
- Python runtime
- Not required
trellis.cpp — Vulkan build
The path with the fewest prerequisites: a Vulkan-capable driver replaces the entire ROCm stack. The project reports Vulkan as its fastest backend on some integrated GPUs.
- OS
- Linux · Windows
- Stack
- Vulkan driver only · no ROCm, no Python runtime
- Python runtime
- Not required
The complete cross-architecture view, including the runtime stacks and version numbers each route was tested against, is on the AMD GPU compatibility page.