Skip to content
Kodning & Utveckling· NewsAvailable

DeepSeek-V4-Flash reaches 12.5 tokens per second on RTX 3090 with DDR5

New performance benchmarks indicate that the quantized DeepSeek-V4-Flash-0731 model achieves 12.5 tokens per second on an Nvidia RTX 3090 paired with 128 GB of DDR5 RAM.

By the Aheadline editorial team·2 aug. 2026·1 min read·Source: Reddit r/LocalLLaMAVerifierad signalAI-generated
DeepSeek-V4-Flash reaches 12.5 tokens per second on RTX 3090 with DDR5
DeepSeek-V4-Flash reaches 12.5 tokens per second on RTX 3090 with DDR5
By · Policy- & EU-reporter
Last updated

What happened?

Users on the r/LocalLLaMA forum have reported performance results for a quantized version of DeepSeek-V4-Flash-0731 (UD-IQ3_S). The model reaches a generation speed of 12.5 tokens per second on a system equipped with an Nvidia RTX 3090 graphics card and 128 GB of DDR5 system memory.

Key facts

Hastighet12,5 tokens/s
GrafikkortNvidia RTX 3090
Systemminne128 GB DDR5
KvantiseringUD-IQ3_S

Why it matters

The results demonstrate how advanced quantization methods such as Unsloth/GGUF (IQ3_S) enable the execution of extremely large models on private hardware. It highlights the actual performance limits when combining GPU VRAM with DDR5 RAM for local AI inference.

Who is affected?

This news is relevant to developers, AI researchers, and hobbyists who run large language models locally on their own hardware. It is particularly pertinent to those optimising performance on consumer graphics cards in combination with high-capacity system memory.

Impact on the EU

As the model is executed locally on the user's own hardware, it is not directly subject to EU restrictions or geoblocking. The use of open-source software and local models provides European users with full control over their own data.

What else you should know

Tests show that the quantized version requires significant memory resources, specifically 128 GB of DDR5 RAM in addition to the graphics card's VRAM. The IQ3_S quantization technique allows for the compression of large models to fit consumer hardware, though speeds are impacted by system memory bandwidth bottlenecks.

Frequently asked questions

Quick answers about this story

Vad har hänt?
En användare har publicerat prestandatester för den kvantiserade AI-modellen DeepSeek-V4-Flash-0731 UD-IQ3_S på konsumenthårdvara.
När hände det?
Tester och resultat publicerades i forumet r/LocalLLaMA den 31 juli 2024.
Varför spelar det roll?
Det visar hur stora AI-modeller kan köras lokalt på privat hårdvara genom effektiv kvantisering, vilket är avgörande för lokal AI-utveckling.
Vilka berörs av detta?
Nyheten berör främst utvecklare och AI-entusiaster som kör stora språkmodeller på egna datorer.
Original source
Reddit r/LocalLLaMA·reddit.com

The link opens in a new window and leads to the publisher's own site.

Verifierad signal

Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.

AI-verktyg i artikeln

Topics

#GPU#AI-benchmarking#Large Language Models (LLM)#Models
[ STAY UP TO DATE ]

Get similar news straight to your inbox

No affiliate linksCancel anytimeGDPR-friendly
[ Frequency ]
[ What do you want to read about? ]

You'll receive updates on 2 topics.

The reader's room

Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.

Sign in to submit a comment or question.

Loading comments…
How this affects you

Read the article through your role

  • Decide whether this affects strategy over 6–12 months or is just noise.
  • Discuss with leadership: do we own the right question or does ownership need to move?
  • Ask: what risk are we taking by NOT acting on this this quarter?

Generated angle — not editorial analysis of "DeepSeek-V4-Flash reaches 12.5 tokens per second on RTX 3090"