FuturPulse analysis · 6 October 2026
A single RTX 4090 has not been benchmarked at 100 tokens per second for Qwen3.8-Flash-Next. As of 6 October 2026, Strata publishes a 94 tokens/s RTX 5070 measurement and a 100–140 tokens/s RTX 3090 estimate.
At a glance
- Strata measured Qwen3.8-Flash-Next Q2_0 at 94 output tokens/s on an RTX 5070 with 12 GB of VRAM.
- Strata estimates, rather than measures, 100–140 output tokens/s on an RTX 3090 with 24 GB of VRAM.
- Qwen says Flash-Next has 125 billion language-model parameters but activates 6 billion per token.
- Our calculation puts the IQ2_XS GGUF download at 68.0 GB, or 78.2 GB with a 15% working margin.
- An RTX 4090 has 24 GB of graphics memory, so the model cannot reside wholly in its VRAM.
Is the 100 tokens/s RTX 4090 claim verified?
No. The 100 tokens/s headline is not backed by a published RTX 4090 run. Sifirincidakika describes it as a developer claim and says the implementation details remain unclear.
The strongest public evidence points elsewhere. Strata’s own project page reports measurements on an RTX 5070 and an AMD Radeon RX 9070 XT, not an RTX 4090. It says an RTX 3090 “should” produce roughly 100–140 output tokens/s.
That distinction matters. A measurement records a named machine, model build, prompt size and software version. An estimate scales prior results. It can be useful, but it does not prove that an RTX 4090 reached the quoted rate.
What did Strata actually measure?
Strata’s detailed benchmark measured Q2_0 at 93.0 output tokens/s with a 4K-token prompt. The machine used an RTX 5070, Ryzen 5 7600 processor and 64 GB of DDR5-5200 memory.
“Output tokens” means the words and fragments appearing in the reply. A token is roughly three-quarters of an English word in Strata’s explanation. At 93 tokens/s, a 300-token answer would take roughly three seconds after prompt processing.
The same test recorded 2,171 prompt tokens/s at 32K context for Q2_0. Prompt processing is often called prefill. It measures how fast the system reads your text, code, or chat history before it begins generating.
The test conditions are unusually important here. Strata used 256 generated tokens, speculative decoding, and its own engine settings. Speculative decoding lets a smaller helper propose several next tokens, then lets the larger model check them in a batch.
Strata also says the output result varies by several percent with the text itself. The reason is draft-token acceptance: easier continuations allow more of the helper’s suggestions through. A one-number claim without prompt, output length, quantization and engine version is incomplete.
Why can a 125B model run on one desktop?
Qwen3.8-Flash-Next is not a conventional 125-billion-parameter model that calculates every weight for every word. Qwen’s model card says it has 125B language-model parameters but activates 6B per token, plus 51B parameters in n-gram embedding tables.
It uses a mixture-of-experts design. That means many specialist sub-networks exist, but only a small selection works on each token. Qwen lists 512 experts, with 10 routed experts and one shared expert activated per token.
This reduces compute work, but it does not remove the storage problem. The model still needs a large weight collection available somewhere. Strata addresses that by placing frequently used experts in GPU memory and leaving the rest in system memory.
That arrangement explains both the opportunity and the compromise. An RTX 4090 can accelerate the most-used work, but its memory cannot hold the whole model. The CPU, RAM, PCIe connection and sometimes SSD therefore affect the result.
Qwen’s technical report describes the architecture as sparse and says its extra n-gram tables are held off the accelerator. That is a key reason this model is more desktop-friendly than a similarly sized dense model.
How much memory and storage do you need?
An RTX 4090 has 24 GB of graphics memory, according to Nvidia’s RTX 4090 specification appendix. That is enough for an active cache of experts, not for every compressed weight and a long context cache.
Strata’s practical guidance is 64 GB of system RAM for every main full-model size. It recommends IQ2_XS at that capacity. The project says 48 GB may fit Q2_0 or IQ2_XS, while 32 GB favors its code-focused build.
The figures below add the two published GGUF shards for each format. This is a download-planning estimate, not a hardware test.
| Compressed model choice | Published GGUF pair | Our 15% disk plan | What the trade-off means |
|---|---|---|---|
| Q2_0 | 66.4 GB | 76.4 GB | Smallest full-model download and Strata’s fastest option. |
| IQ2_XS | 68.0 GB | 78.2 GB | Strata’s recommended 64 GB-RAM choice. |
| IQ3_XXS | 75.8 GB | 87.2 GB | Larger download with a quality-focused trade-off. |
| IQ3_S | 83.6 GB | 96.1 GB | Largest listed option, with the slowest listed output speed. |
Our calculation from the two file sizes listed in ISTA-DASLab’s GGUF repository. It is a storage estimate, not a claim about required VRAM.

Storage is not the same as active memory. A large GGUF download can sit on an SSD, while Strata moves or maps parts of it into RAM and uses the GPU for a hot cache. Long chats also increase memory demand because the key-value cache stores information the model needs to continue a conversation.
How does the RTX 4090 compare with tested hardware?
The best direct comparison is not RTX 4090 versus a datacenter GPU. It is measured consumer hardware against a clearly marked estimate. Strata measured Q2_0 at 94 output tokens/s on a 12 GB RTX 5070, while its AMD RX 9070 XT result was 60 tokens/s.
Its estimated 24 GB RTX 3090 result is about 140 tokens/s at a 4K-token prompt for Q2_0. The same table gives about 100 tokens/s at 128K context. Strata labels all of those non-tested GPU figures as estimates with a ±20% range.
That leads to the useful interpretation of the RTX 4090 claim. A 24 GB card gives the engine more room to cache experts than the tested 12 GB RTX 5070. But the published material has not shown whether its software path, CPU pairing and prompt settings turn that into 100 tokens/s.
Long context changes the answer. On the RTX 5070, Q2_0 fell from 93.0 tokens/s at 4K context to 73.7 tokens/s at 128K. The number advertised in a headline therefore should not be read as a guarantee for large codebases or lengthy agent histories.
What would a practical RTX 4090 setup look like?
A sensible target is a desktop with an RTX 4090, 64 GB of RAM, an SSD and a current graphics driver. That matches Strata’s basic capacity guidance, although it does not validate the 100 tokens/s number on that card.
Start with IQ2_XS if general chat and coding matter. Strata recommends it for 64 GB machines. Choose Q2_0 when speed matters more than preserving the larger compressed representation.
- Use a short context first, because long contexts cut output speed and consume more memory.
- Keep the browser and memory-heavy tools closed during startup, because Strata says it loads tens of gigabytes into RAM.
- Record your exact model format, context size, engine version, CPU and RAM before comparing results.
- Separate prompt speed from output speed, because fast prefill does not guarantee fast replies.
For a developer, 60 tokens/s is already beyond normal reading speed. The difference between 94 and 100 tokens/s is therefore less important than stable operation, enough context, and whether a compressed model still meets the task’s quality needs.
How does local use compare with an API?
Local inference removes per-token billing after the hardware and electricity costs. It also puts model inputs on the user’s machine by default. That can matter for private code or files, but it transfers updates, security and reliability work to the operator.
Cloud pricing is not a direct substitute for Flash-Next because providers offer different Qwen models. Still, OpenRouter’s current model list shows why token pricing changes the calculation for frequent use: Qwen-hosted APIs charge separately for input and output tokens, the chunks of text a model reads and writes.
For occasional tasks, an API avoids an 80 GB-class download and desktop setup. For continuous local work, the relevant question is not “can it hit 100?” It is whether the model performs well enough at the context size and output rate your workflow actually needs.
Who is making each claim, and why?
The public benchmark standard is clear. Strata’s community-reporting guide asks for GPU, VRAM, CPU, RAM, storage, software versions, drivers, exact model files and context length. An RTX 4090 result meeting that standard would settle the headline.
What we could not verify?
No public result currently identifies an RTX 4090, its CPU, RAM speed, driver, Strata version, quantization, prompt length and output length in one reproducible 100 tokens/s test. The same missing details prevent a fair comparison with other 4090 reports.
There is also no public quality evaluation here that compares Q2_0, IQ2_XS, IQ3_XXS and IQ3_S on the same task. The model publisher can establish official quality differences, while Strata maintainers or independent users can settle hardware speed through a complete benchmark report.
The next useful update is therefore not another headline. It is a dated RTX 4090 run with the full configuration, several context lengths and separate prompt and output measurements.

