文章详情顶部

DeepSeek Tests Whether Its Cheaper Flash Model Can Displace Pro

DeepSeek opened a short internal test of an intermediate V4.1 Flash build and asked users whether it could fully replace the online V4 Pro. At the same time the company cut Flash-series prices. The moves highlight a broader shift from raw capability toward higher “intelligence density”—delivering strong results at lower cost and latency for everyday and agent workloads.

NextFin News — DeepSeek has put its own pricing ladder under pressure. The company opened a limited internal test of an intermediate V4.1 Flash checkpoint and, in the accompanying feedback form, asked participants whether the new model could comprehensively replace the live V4 Pro.

The test version, reachable under a model name that expires on September 10, is described as using a new structure, adding native multimodal support, and improving capability, speed and cost relative to prior Flash builds. During the trial it is billed at the same rates as the existing V4 Flash and is limited to 20 concurrent requests per account. Early developer reports emphasize markedly higher generation speed on coding and retrieval tasks.

Almost simultaneously, DeepSeek announced that Flash-series prices would fall at noon Beijing time on September 10. Off-peak rates move to 0.02 yuan per million tokens for cache-hit input, 1 yuan for cache-miss input and 4 yuan for output; peak rates remain twice the off-peak levels. The largest reduction, on cache-hit input, reaches 60 percent. The cut applies to both the main Flash model and the vision-experimental variant.

The price gap between Flash and Pro has always been material. At previous peak rates, Flash output sat at 9 yuan per million tokens while Pro commanded a substantially higher figure. Developers have long accepted paying a multiple for the harder cases. The open question is whether a faster, stronger Flash can absorb enough of the everyday and mid-complexity load that the higher tier is reserved for genuinely difficult work.

That question maps onto a concept DeepSeek itself has emphasized: intelligence density. Earlier reasoning-oriented releases demonstrated that longer chains of thought could raise accuracy on hard problems. Later technical discussion turned to whether the same quality could be obtained with fewer tokens and less wall-clock time. A model that reaches a useful answer more quickly and with less intermediate computation improves both user experience and unit economics.

Agent workloads make the economics sharper. A single user request can trigger many model calls as the system reads files, invokes tools, checks intermediate results and continues. Low per-token prices do not automatically produce low task costs if the number of steps multiplies. Frameworks that let the model emit longer programs of tool use, retain relevant reasoning across steps, and avoid re-processing already examined material therefore become as important as the model weights themselves. DeepSeek’s open Harness work sits in this layer: the model decides, the harness supplies tools, session state and execution environment.

Similar recalibrations appear at other frontier labs. When a mid-tier model approaches the practical performance of a previous flagship at a fraction of the cost, the higher tier must justify itself with longer-horizon, higher-stakes tasks that users are willing to wait for and pay for. Flagship models increasingly compete on the ability to carry multi-hour or multi-day agentic work to a verifiable conclusion rather than on every individual benchmark score.

For DeepSeek the near-term test is concrete. If the V4.1 Flash intermediate build, once refined, can handle a large share of what users currently route to Pro, the company can push volume onto the cheaper tier while concentrating remaining research and serving capacity on harder problems. Generation-speed improvements already demonstrated on the V4 stack, together with lower Flash prices, reinforce the same direction: make capable inference frequent rather than occasional.

The intermediate checkpoint is temporary and the full replacement claim remains unproven. Subsequent results will determine how much of the Pro workload can actually migrate. The strategic signal, however, is already clear. After establishing that strong reasoning can be widely available, the next competitive axis is how densely that intelligence can be packed into each second of latency and each unit of cost. Flash is being asked to carry more of the load; Pro, if the test succeeds, will be measured by the problems that still require it.

本文系作者 Chelsea_Sun 授权钛媒体发表,并经钛媒体编辑,转载请注明出处、作者和本文链接
本内容来源于钛媒体钛度号,文章内容仅供参考、交流、学习,不构成投资建议。
想和千万钛媒体用户分享你的新奇观点和发现,点击这里投稿 。创业或融资寻求报道,点击这里

敬原创,有钛度,得赞赏

赞赏支持
发表评论
0 / 300

根据《网络安全法》实名制要求,请绑定手机号后发表评论

登录后输入评论内容

快报

更多

17:54

2026年家用服务机器人产业生态技术交流会将于9月23日在杭州召开

17:53

3连板华脉科技:光缆产品近期价格虽有所上涨,但受多重因素影响涨价不具备可持续性

17:50

硅片价格偏弱运行,低价货源有所增加

17:48

OpenAI或调整定价策略,与开源模型展开竞争

17:47

上海宣布42项产检医保全包,产检分娩住院全报销

17:47

南向资金今日净买入逾37亿港元,百度获买入居前

17:46

ST龙元:法院决定对公司进行预重整

17:45

东兴证券:A股股票将于9月15日起连续停牌直至终止上市

17:45

信达证券:A股股票将于9月15日起连续停牌直至终止上市

17:44

十四届全国人大社会建设委员会原副主任委员孙绍骋被“双开”

17:42

黄益平:AI经济影响暂不明显属索洛悖论,范式或生变

17:41

上海亚虹:筹划控制权变更事项,股票停牌

17:41

掉期市场显示:交易员提高了对欧央行明年进一步加息的押注

17:40

监管通报:有第三方借空壳公司及虚假审计报告诱骗投资者购买高风险债券

17:38

财政部决定发行2026年第七期和第八期储蓄国债(电子式)

17:37

三峡能源在海南昌江成立海上风电公司

17:37

湖南黄金:重大资产重组获湖南省国资委批复

17:36

恒瑞医药:瑞普泊肽注射液上市申请获国家药监局受理

17:35

诺奖得主萨金特:AI投资者不能忽视来自亚洲的竞争,特别是中国

17:35

镭萌科技完成千万级天使轮融资,时尚传媒集团与弘颐资管联合领投

扫描下载App