RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8
A user details the technical process of configuring a dual-GPU setup using an RTX 5080 and an RTX 3090 to run large language models like Qwen 3.6. The guide covers hardware requirements, BIOS settings, and the complexities of managing mismatched NVIDIA drivers.
Why it matters
This reflects the growing trend of enthusiasts building high-performance local AI infrastructure to bypass cloud-based limitations.
A year ago, I bought an RTX 5080 for both gaming and AI experiments. Little did I know back then that I would be giving into the joys of local LLM setups. Fast forward 2026, Qwen 3.5, Gemma, Qwen 3.6, I needed more than 16GB. So I got myself a refurbished RTX 3090 with 24GB. I could then run Qwen 3.6 Q4 quants, first at ~30 tok/s, then 50-60 with MTP. Not bad. But still felt limited while my 5080 was barely used. So I began digging what kind of setup could take profit of those 2 cards together. I already had DDR4 sticks and SSD disks ready, I only needed a mobo capable of handling the two cards. Enters the Asus Prime X570-Pro, the "Pro" is important, it is what ensures the 16x PCIe can be splitted in 2x8. The 5080 being the monster it is I bought a good quality PCIe 4 riser to plug it on the second slot.
The content is a technical tutorial and personal experience report with no political or social bias.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in