Where the Ryzen AI Max+ 395 stands today

Architecture
RDNA 3.5
Compile target
gfx1151
Memory
96 GB
Best evidence
Tested
Start here

trellis.cpp — Vulkan build

The project singles out Vulkan as its fastest backend on some integrated GPUs, and it needs no ROCm install at all. The path with the fewest prerequisites: a Vulkan-capable driver replaces the entire ROCm stack. The project reports Vulkan as its fastest backend on some integrated GPUs.

Open that guide →

Up to 96 GB allocated from unified system memory. The upstream TRELLIS.2 project documents Linux and a 24 GB NVIDIA card as its standard path, so every route below is community work rather than a supported configuration. What follows separates what has actually been reported on this card from what is inferred from a shared instruction set.

What is actually reported

  • trellis.cpp names Vulkan as the fastest backend on some integrated GPUs including Strix Halo.
  • RDNA 3.5 is named as a next validation target in the ROCm ComfyUI Windows test call.
How to read thisEvery AMD route on this page is community or third-party work. Neither Microsoft nor AMD lists TRELLIS.2 as an officially supported workload, and AMD's own documentation still limits Windows to PyTorch rather than the full ROCm stack.

Memory and weight formats

Weight size is not peak VRAM. The figures below describe how much of the frame buffer the weights themselves occupy; the texture stage and working buffers sit on top of that. With 96 GB, these are the formats worth trying in order.

  • f16 / bf16 (default) · ≈ 16.5 GB · 24 GB comfortable
  • Q8 · ≈ 9.5 GB · 16 GB practical
  • Q4 · ≈ 6 GB · 12 GB practical, 8 GB tight

Full detail on the two quantized options lives in the Q8 guide and the Q4 guide. If your workflow goes through ComfyUI rather than a native binary, the GGUF guide covers loader compatibility, which is the part that actually decides whether a quantized file opens.

What to watch for on this card

  • Memory capacity does not buy compute. Expect substantially longer runs than a discrete card at the same settings.
  • Allocation of system memory to the GPU is a firmware or driver setting; check it before blaming the model.

RDNA 3.5 in the full matrix

Under active validation for the ComfyUI route. The trellis.cpp project separately calls Vulkan its fastest backend on some integrated GPUs, which makes it the more realistic entry point here today.

Compatibility matrixReviewed 2026-08-15
GPUVRAMISAROCm forkLinuxComfyUI LinuxLinuxComfyUI WindowsWindows 11trellis.cpp ROCmLinux · Windowstrellis.cpp VulkanLinux · Windows
RDNA 3.5 gfx1150 · gfx1151
Ryzen AI Max+ 39596 GBgfx1151UnknownReportedReportedExpectedTested
  • Tested

    A named project or maintainer reports this exact card completing the pipeline end to end.

  • Reported

    Community reports exist but come from work in progress, partial runs, or a single tester.

  • Expected

    Inferred from a shared instruction set with a tested card. No direct report yet.

  • Blocked

    A known dependency, kernel, or runtime gap prevents this path today.

  • Unknown

    No usable report either way. Treat as untested rather than as a failure.

Every AMD route on this page is community or third-party work. Neither Microsoft nor AMD lists TRELLIS.2 as an officially supported workload, and AMD's own documentation still limits Windows to PyTorch rather than the full ROCm stack.

Tested route

TRELLIS.2 ROCm source fork

A fork of the upstream repository that ships HIP builds of FlexGEMM, CuMesh, nvdiffrast and nvdiffrec, plus a setup script that detects CUDA or ROCm and installs the matching dependencies. Validated on an RX 9070 XT 16 GB.

OS
Linux
Stack
ROCm 7.2 · PyTorch rocm7.2 · Python 3.10+
Python runtime
Required
Open the guide →
Tested route

ComfyUI + ROCm on Linux

The community ComfyUI wrapper plus a patch set that fixes the hardcoded GPU architecture flag, package paths and checkpoint downloads. Reported working end to end on a 7900 XTX for both shape-only and textured runs.

OS
Linux
Stack
ROCm 7.2.2 · PyTorch 2.11+rocm7.2 · Python 3.10–3.12
Python runtime
Required
Open the guide →
Work in progress

ComfyUI + ROCm on Windows

Work in progress. RDNA 4 is confirmed running and testers are being recruited; RDNA 3 and RDNA 3.5 are next in line. The maintainer warns about silent bugs that can change the final output without raising an error.

OS
Windows 11
Stack
PyTorch 2.9.1 + ROCm 7.2.1, or PyTorch 2.12 + ROCm 7.14 · Python 3.12
Python runtime
Required
Open the guide →
Tested route

trellis.cpp — ROCm build

A native C++/GGML reimplementation of the whole TRELLIS.2 pipeline. Prebuilt ROCm archives are published for Linux and Windows, which removes the CUDA-only Python wheels from the problem entirely.

OS
Linux · Windows
Stack
Prebuilt ROCm/HIP binaries · no Python runtime
Python runtime
Not required
Open the guide →
Tested route

trellis.cpp — Vulkan build

The path with the fewest prerequisites: a Vulkan-capable driver replaces the entire ROCm stack. The project reports Vulkan as its fastest backend on some integrated GPUs.

OS
Linux · Windows
Stack
Vulkan driver only · no ROCm, no Python runtime
Python runtime
Not required
Open the guide →

The complete cross-architecture view, including the runtime stacks and version numbers each route was tested against, is on the AMD GPU compatibility page.

Support levels reviewed 2026-08-15. Trellis 3D is an independent project and is not affiliated with AMD, Microsoft, or any of the community projects cited here.