DeepSeek V4 Flash on a Single AMD MI300X
This technical post provides a configuration guide for running the DeepSeek-V4-Flash AI model on a single AMD MI300X GPU. It addresses specific hardware compatibility issues, such as FP8 format discrepancies and kernel tuning, to enable production-level deployment.
Why it matters
Optimizing large models for non-NVIDIA hardware is critical for reducing infrastructure costs and diversifying the AI compute ecosystem.
This repository contains the configuration and patches I use to run deepseek-ai/DeepSeek-V4-Flash-0731 on one AMD MI300X in production. It includes the Docker Compose stack, SHA-256-pinned file overlays, reference diffs against upstream, and tuning tables. The checkpoint runs as shipped, without additional weight quantization or offload.
Results from the pinned stack (vLLM ROCm nightly 0.26.1rc1.dev229+g124154a88.rocm723 , AITER 0.1.19 ):
The official vLLM recipe targets NVIDIA and newer AMD hardware. Running the model reliably on MI300X required fixes for its FP8 format, MoE routing at high concurrency, causal speculative verification, CPU-KV synchronization, and several untuned kernel shapes. This repository collects those fixes and pins the versions used in production.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in