All signals
DeepSeek V4.1 Flash Processes One Trillion Tokens in First 24 Hours, On Pace for 2.8 Trillion in 48 Hours at Five Times Lower Cost Than Comparable Models
Sources: HeadsUpAI September 12, 2026 reporting OpenRouter data; DataNorth AI September 10, 2026; Dataconomy September 11, 2026; Eesel AI, TechJack Solutions, Bitrue, AIToolsReview, TechPillow all September 10-12, 2026; DeepSeek official API documentation and model card on Hugging Face, fetch-verified September 11, 2026 per TechJack Solutions; DeepSeek official announcement September 10, 2026.
DeepSeek V4.1 Flash processed one trillion tokens in its first 24 hours, on pace for 2.8 trillion tokens in 48 hours, according to OpenRouter data reported by HeadsUpAI September 12. Ninety percent of this volume consisted of cache reads priced at approximately 0.006 dollars per million tokens, a rate five times cheaper than comparable models like GLM-5.3 Flash. DeepSeek released V4.1 Flash on September 10, 2026, as a 552-billion-parameter mixture-of-experts model with native vision, a one-million-token context window, and pricing built around low cached-input costs, according to DataNorth, Dataconomy, and multiple technical analysis sites. During off-peak hours, DeepSeek prices the model at 0.003 dollars per million input tokens on a cache hit, 0.15 dollars per million on a cache miss, and 0.60 dollars per million output tokens. Peak rates, defined as Monday through Friday from 01:00 to 04:00 UTC and 06:00 to 10:00 UTC, double those figures. The architecture splits 40 Transformer layers into a 20-layer causal encoder and a 20-layer decoder, activating 8 billion parameters per token during input processing and 16 billion during output generation. DeepSeek said the design combines Compressed Sparse Attention 2, hierarchical sparse indexing, and FP4 KV caching to reduce the global KV cache to 890 bytes per token, about one quarter the size of V4 Flash's, while persistent cache storage falls to roughly one eighth. The model was pre-trained on 45 trillion tokens of multimodal data, with context extended to one million tokens at the 34-trillion-token mark. DeepSeek said V4.1 Flash comprehensively surpassed V4 Pro in performance, cost, speed, and total time, and that V4 Pro requests will route to V4.1 Flash starting 04:00 UTC September 14, 2026, until a future V4.1 Pro release. The weights are published on Hugging Face under the MIT license. The one-trillion-token 24-hour throughput figure, if sustained, would make V4.1 Flash one of the highest-volume open-weight model d