AI Hardware Benchmarking
First-party local-LLM measurements on older AMD GPUs using Linux, Mesa RADV/Vulkan, and llama.cpp. These are separate experiments on separate systems; the numbers are reported as measured, not as projections for untested hardware.
Reading the tables: pp512 is prompt-processing throughput in tokens per second. tg128 is sequential token-generation throughput in tokens per second. Prompt processing and generation are different workloads, so a faster prompt result does not necessarily mean a better interactive response.
AMD Polaris
RX 570 single-GPU compatibility baseline
Measured August 15, 2026 on an X79 S7 running Linux Mint 22.3 x86_64, kernel 7.0.0-28-generic, Intel Xeon E5-1620 v2, and 15.9 GiB RAM. The GPU was an XFX Radeon RX 570 4 GB, reported by Vulkan as RADV POLARIS10. The stack used amdgpu, Mesa RADV, and llama.cpp 6b4344ecc (10440).
llama.cpp enumerated Vulkan0: AMD Radeon RX 570 Series (RADV POLARIS10) (4096 MiB, 4036 MiB free). The backend reported fp16 0, bf16 0, fp4 0, warp size 64, 65536 bytes of shared memory, int dot 0, and no matrix cores. Despite those flags, real Vulkan inference acceleration worked.
For Qwen2.5-1.5B-Instruct Q4_K_M (1.78B parameters, 1.04 GiB):
| Configuration | pp512 | tg128 |
|---|---|---|
| CPU only | 663.63 t/s | 21.09 t/s |
| RX 570 full GPU offload | 986.96 t/s | 104.10 t/s |
That is approximately 1.49× CPU prompt processing and 4.94× CPU token generation. It was the initial proof that the current Linux + Mesa RADV + Vulkan + llama.cpp path could perform useful inference on Polaris. It does not establish Radeon Pro Duo performance.
RX 570 7B VRAM-pressure test
Qwen2.5-7B-Instruct Q4_K_M (7.62B parameters, 4.36 GiB) was intentionally larger than the RX 570's nominal 4 GiB VRAM.
| NGL | pp512 | tg128 |
|---|---|---|
| 0 | 154.45 t/s | 4.95 t/s |
| 20 | 189.95 t/s | 12.25 t/s |
| 24 | 199.18 t/s | 16.74 t/s |
| 26 | 92.25 t/s | 19.96 t/s |
| 27 | 95.45 t/s | 21.39 t/s |
| 28 | 99.24 t/s | 22.54 t/s |
| 29 | 163.42 t/s | 7.30 t/s |
| 30 | 162.77 t/s | 7.30 t/s |
| 31 | 163.20 t/s | 7.30 t/s |
| 32 | 163.60 t/s | 7.30 t/s |
Best measured generation: NGL 28 — 22.54 t/s. CPU-only generation was 4.95 t/s, making this approximately 4.55× faster. The NGL 28 to NGL 29 change caused a sharp, repeatable cliff from 22.54 to 7.30 t/s—approximately a 68% reduction after adding one layer. This is evidence of a memory-placement or execution-regime transition near the VRAM limit, not a conclusively diagnosed driver bug.
RADV nogttspill diagnostic
| RADV mode | pp512 | tg128 | Generation range |
|---|---|---|---|
| Normal RADV | ~99.22 t/s | ~22.42 t/s | 22.27–22.50 t/s |
RADV_PERFTEST=nogttspill | ~170.62 t/s | ~6.87 t/s | 6.86–6.88 t/s |
The results were highly reproducible. nogttspill substantially improved prompt processing while destroying the exceptional sequential generation behavior seen under normal RADV. This disproved the simple explanation that NGL 29 was merely where ordinary GTT spilling began, but the underlying driver/memory behavior was not definitively diagnosed. For interactive dialogue, normal RADV at NGL 28 was clearly preferable.
RX 570 + RX 560 Polaris multi-GPU
Measured August 16, 2026 on a different system: server1200-X79-S7, Linux Mint 22.3 x86_64, kernel 7.0.0-28-generic, AMD Ryzen 7 5800XT, and 48 GB (48092 MiB) RAM. llama.cpp was b10440 / 6b4344ecc; Mesa was 25.2.8-0ubuntu0.24.04.2.
Vulkan0 was an RX 570, RADV POLARIS10, 4096 MiB. Vulkan1 was an RX 560, RADV POLARIS11, 4096 MiB. Both were independently detected by vulkaninfo, llama-bench, and llama-cli. These results must not be combined with the earlier RX 570/X79 NGL experiment.
For Qwen2.5-7B-Instruct Q4_K_M (4.36 GiB, 7.62B parameters), an accidental comma-separated llama-bench command produced separate single-GPU cases. The RX 570 measured pp512 172.09 ± 0.72 and 170.22 ± 0.76 t/s, with tg128 7.86 ± 0.00 and 7.87 ± 0.00 t/s. The RX 560 measured pp512 13.45 ± 0.66 t/s and tg128 6.62 ± 0.02 t/s. The RX 560 was dramatically weaker for prompt processing, while generation was much closer to the RX 570.
| Configuration | pp512 | tg128 |
|---|---|---|
| RX 570 alone | ~171 t/s | ~7.86 t/s |
| RX 560 alone | ~13.45 t/s | ~6.62 t/s |
| Dual 1:1 | 135.81 ± 0.09 t/s | 21.02 ± 0.01 t/s |
| Dual 3:1 | 169.99 ± 0.24 t/s | 23.50 ± 0.05 t/s |
The true dual-GPU syntax was --device Vulkan0/Vulkan1 --split-mode layer --tensor-split 1/1 or 3/1. Equal splitting increased generation from ~7.86 to 21.02 t/s but over-allocated work to the weaker RX 560, reducing prompt throughput. Tuning to 3:1 recovered almost all standalone prompt throughput and produced approximately 23.50 / 7.86 = ~2.99× the RX 570-alone generation rate in this particular test configuration. This is not universal 3× scaling.
Real Skyrim-style llama-cli test
llama-bench used --device Vulkan0/Vulkan1 and --tensor-split 3/1; llama-cli used comma-separated device and split values: --device Vulkan0,Vulkan1 --tensor-split 3,1. The working run also used -ngl 999 --split-mode layer -c 4096 -n 192. The prompt asked the model, acting as Lydia from Skyrim, whether to travel to Whiterun that night and requested a natural two-to-four-sentence response.
| Run | Prompt | Generation |
|---|---|---|
| Skyrim-style run 1 | 143.7 t/s | 22.8 t/s |
| Skyrim-style run 2 | 143.3 t/s | 22.9 t/s |
The runs were extremely consistent and closely matched the synthetic 23.50 t/s generation result. This establishes that the tuned Polaris multi-GPU result translated into a real conversational llama-cli workload.
AMD Hawaii: dual FirePro W8100
Measured August 16, 2026 on server1200 with Linux Mint 22.3 x86_64, AMD Ryzen 7 5800XT, approximately 48 GB RAM, and two FirePro W8100 8 GB cards. Both were RADV HAWAII devices using amdgpu, Mesa RADV, and llama.cpp b10440 / 6b4344ecc. PCIe devices were 04:00.0 and 08:00.0, both Hawaii PRO GL / FirePro W8100. llama.cpp enumerated Vulkan0 and Vulkan1 with 8192 MiB each.
These cards were an older-GCN, higher-capacity experiment—not the final target hardware. The purpose was to test old AMD Vulkan compatibility, genuine multi-GPU layer splitting, aggregate VRAM/model placement, and larger models that do not practically fit on one card.
W8100 7B test
Qwen2.5-7B-Instruct Q4_K_M was 4.36 GiB and 7.62B parameters.
| Configuration | Prompt / pp512 | Generation / tg128 |
|---|---|---|
| Single W8100 interactive | 122.6 t/s | 18.4 t/s |
| Single W8100 benchmark | 122.57 ± 0.99 t/s | 18.51 ± 0.11 t/s |
| Dual W8100 1:1 interactive | 106.5 t/s | 19.1 t/s |
The model fit comfortably on one W8100, so the second GPU brought only a small generation improvement and prompt overhead. Verbose placement confirmed layers 0–14 on Vulkan0 and 15–28 on Vulkan1. Measured VRAM was 2,585,735,168 bytes (~2.41 GiB) on GPU0 and 2,617,847,808 bytes (~2.44 GiB) on GPU1. This demonstrated genuine distribution across both physical GPUs.
W8100 14B aggregate-VRAM test
Qwen2.5-14B-Instruct Q4_K_M was 14.77B parameters, an 8.37 GiB reported model size, and three GGUF shards. It exceeded the practical capacity of one 8 GB W8100 once runtime overhead was included.
| Configuration | Prompt | Generation |
|---|---|---|
| Dual W8100 1:1 interactive | 53.6 t/s | 11.3 t/s |
| Dual W8100 1:1 llama-bench | 53.34 ± 0.26 t/s | 11.37 ± 0.00 t/s |
The synthetic result almost exactly matched the real interactive run. While loaded, GPU0 held 5,013,508,096 bytes (~4.67 GiB) and GPU1 held 4,926,590,976 bytes (~4.59 GiB). This proves aggregate VRAM/model placement worked across two independent W8100 GPU memory pools for this llama.cpp configuration. It does not mean the machine had one contiguous or physically unified 16 GB GPU.
W8100 14B tensor-split tuning
| Tensor split | pp512 | tg128 |
|---|---|---|
| 1/1 | 53.34 t/s | 11.37 t/s |
| 3/2 | 53.51 t/s | 11.19 t/s |
| 2/3 | 52.01 t/s | 11.54 t/s |
| 4/3 | 52.19 t/s | 11.29 t/s |
| 3/5 | 52.84 t/s | 11.59 t/s |
| 4/6 | 50.28 t/s | 11.53 t/s |
| 5/7 | 50.37 t/s | 11.51 t/s |
The best measured generation was 3/5 at 11.59 t/s. The balanced 1/1 baseline was 53.34 pp512 / 11.37 tg128. Prompt processing and sequential generation responded differently; the modest gain demonstrates that tuning matters even between nominally identical GPUs.
What these tests established
- Polaris GPUs worked with current Linux/RADV/Vulkan and llama.cpp, and multiple Polaris devices enumerated independently.
- Heterogeneous RX 570 + RX 560 layer splitting worked, and split tuning mattered substantially.
- Real conversational llama-cli generation closely matched the synthetic Polaris result.
- Dual Hawaii W8100 cards enumerated correctly and performed actual layer splitting.
- Independent W8100 VRAM pools both participated in 14B model placement, with interactive and benchmark results closely matching.
- W8100 tensor split ratios affected performance differently for prompt processing and generation.
Radeon Pro Duo research context
There are two Radeon Pro Duo families. The 2016 model uses dual Fiji GPUs and HBM. The later Radeon Pro Duo 32 GB GDDR5 uses dual Polaris GPUs, with effectively 16 GB attached to each GPU. The target of this research is the later Polaris model, not the Fiji/HBM card.
The RX 570 + RX 560 experiment is relevant because it validates a closely related Polaris-era software path. A planned configuration of two Radeon Pro Duo 32 GB GDDR5 boards would physically contain four Polaris GPU devices and approximately 64 GB of aggregate physical VRAM across those devices. That is not one unified 64 GB GPU, and four-GPU behavior has not been proven.
Still unmeasured are actual Radeon Pro Duo inference performance, four-GPU enumeration and splitting, four-way scaling and tensor-split tuning, PCIe/topology effects, exact prompt and generation throughput, and warm or cold Mantella latency. The Polaris and Hawaii proof-of-concept work reduces software-compatibility and aggregate-VRAM risk, but actual Radeon Pro Duo performance remains an empirical question until those boards are tested directly.