Article may be outdated

This article is 70 days old. Some details may have changed since publication.

Hacker News·4 min read·hard

RTX 5080 and RTX 3090 Setup: 80 Tok/s on Qwen 3.6 27B Q8

I
iMil
AI Summary

A user details the technical process of configuring a dual-GPU setup using an RTX 5080 and an RTX 3090 to run large language models like Qwen 3.6. The guide covers hardware requirements, BIOS settings, and the complexities of managing mismatched NVIDIA drivers.

Why it matters

This reflects the growing trend of enthusiasts building high-performance local AI infrastructure to bypass cloud-based limitations.

Dive DeeperCreate a free account to unlock

A year ago, I bought an RTX 5080 for both gaming and AI experiments. Little did I know back then that I would be giving into the joys of local LLM setups. Fast forward 2026, Qwen 3.5, Gemma, Qwen 3.6, I needed more than 16GB. So I got myself a refurbished RTX 3090 with 24GB. I could then run Qwen 3.6 Q4 quants, first at ~30 tok/s, then 50-60 with MTP. Not bad. But still felt limited while my 5080 was barely used. So I began digging what kind of setup could take profit of those 2 cards together. I already had DDR4 sticks and SSD disks ready, I only needed a mobo capable of handling the two cards. Enters the Asus Prime X570-Pro, the "Pro" is important, it is what ensures the 16x PCIe can be splitted in 2x8. The 5080 being the monster it is I bought a good quality PCIe 4 riser to plug it on the second slot.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologyai
Political Bias
Center
LeftLean LCenterLean RRight
Confidence: 95%

The content is a technical tutorial and personal experience report with no political or social bias.

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in