DeepSeek-V4-Flash reaches 12.5 tokens per second on RTX 3090 with DDR5
New performance benchmarks indicate that the quantized DeepSeek-V4-Flash-0731 model achieves 12.5 tokens per second on an Nvidia RTX 3090 paired with 128 GB of DDR5 RAM.

What happened?
Users on the r/LocalLLaMA forum have reported performance results for a quantized version of DeepSeek-V4-Flash-0731 (UD-IQ3_S). The model reaches a generation speed of 12.5 tokens per second on a system equipped with an Nvidia RTX 3090 graphics card and 128 GB of DDR5 system memory.
Key facts
| Hastighet | 12,5 tokens/s |
|---|---|
| Grafikkort | Nvidia RTX 3090 |
| Systemminne | 128 GB DDR5 |
| Kvantisering | UD-IQ3_S |
Why it matters
The results demonstrate how advanced quantization methods such as Unsloth/GGUF (IQ3_S) enable the execution of extremely large models on private hardware. It highlights the actual performance limits when combining GPU VRAM with DDR5 RAM for local AI inference.
Who is affected?
This news is relevant to developers, AI researchers, and hobbyists who run large language models locally on their own hardware. It is particularly pertinent to those optimising performance on consumer graphics cards in combination with high-capacity system memory.
Impact on the EU
As the model is executed locally on the user's own hardware, it is not directly subject to EU restrictions or geoblocking. The use of open-source software and local models provides European users with full control over their own data.
What else you should know
Tests show that the quantized version requires significant memory resources, specifically 128 GB of DDR5 RAM in addition to the graphics card's VRAM. The IQ3_S quantization technique allows for the compression of large models to fit consumer hardware, though speeds are impacted by system memory bandwidth bottlenecks.
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Vilka berörs av detta?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
AI-verktyg i artikeln
Topics
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Decide whether this affects strategy over 6–12 months or is just noise.
- Discuss with leadership: do we own the right question or does ownership need to move?
- Ask: what risk are we taking by NOT acting on this this quarter?
Generated angle — not editorial analysis of "DeepSeek-V4-Flash reaches 12.5 tokens per second on RTX 3090"