NextFin News — For much of the past two years the main competitive move in large-language-model APIs was to cut the price per million tokens. Low rates drew in developers, built usage share and helped models become default choices inside applications. That phase is now giving way to a more varied set of pricing tools—time-of-day rates, cache-aware discounts, model tiers and early experiments with fees tied to business results.
On August 17 DeepSeek’s new schedule took effect. The company introduced peak and off-peak windows: higher rates during Beijing daytime hours of 9:00 –12:00 and 14:00 –18:00, and rates set at half that level at all other times. For its V4 Pro model, peak pricing reached 9 yuan per million input tokens on a cache miss, 27 yuan per million output tokens, and 0.3 yuan on a cache hit. Off-peak figures were half those amounts. Earlier public numbers for the same model, released only days before, had listed cache-miss input at 3 yuan, output at 6 yuan and cache-hit input as low as 0.025 yuan. The largest relative increase therefore fell on cached input during peak hours. A similar adjustment applied to the lighter V4 Flash model, with peak rates of 3 yuan, 9 yuan and 0.1 yuan respectively.
The exact figures matter less than the structure they introduce. Time-of-day differentials push non-urgent workloads into quieter periods and improve utilization of fixed serving capacity. Cache-aware pricing rewards systems that reuse context instead of reprocessing it. Separate rates for flagship and lightweight models let high-volume, simpler tasks run at lower cost while complex reasoning still commands a premium. Other providers, both in China and overseas, are testing similar approaches—prepaid packages, base-plus-usage combinations, and, in a smaller number of cases, commercial deals that try to link fees to measurable outcomes.
Several forces are pushing the change across the market. Once models sit inside agent workflows, token use rises sharply. A single user request can trigger long chains of intermediate steps, tool calls and revisions. Serving those chains requires ongoing spending on compute, memory bandwidth and energy; the marginal cost of an extra token is no longer trivial. At the same time, performance gaps between leading open-weight systems and proprietary offerings have narrowed. Global usage trackers show Chinese models taking a larger share of publicly observed API traffic. When substitutes of similar quality exist at different price points, list rates and total cost of ownership become decisive for many buyers.
Higher-priced overseas suppliers have responded with selective cuts on certain models and tiers. The market as a whole therefore shows prices moving in both directions: upward adjustments by providers that once competed mainly on cost, and downward adjustments by incumbents trying to protect volume. The earlier period of across-the-board ultra-low rates is ending. What is replacing it is a more precise effort to match price to actual patterns of demand and cost.
Enterprise customers face parallel internal changes. Traditional software budgets often favor fixed licenses or seat counts—capital spending or predictable operating expense that fits annual planning cycles. Token billing introduces variable cost that is harder to forecast and harder to allocate across teams. Some organizations respond by preferring private or hybrid deployments that restore a degree of cost control. Others accept usage-based charges but insist on clearer monitoring, alerts and charge-back tools. A minority are testing outcome-linked arrangements—fees tied to hours saved, cases resolved or losses avoided—yet these remain hard to standardize. Results depend heavily on how a model is prompted, integrated and supervised; the same system can deliver very different economic value in different hands.
Providers are adjusting their product line-ups in response. Lightweight or “flash” models handle routine, high-volume tasks. Flagship models are kept for harder reasoning problems. Cache discounts and peak pricing improve the economics of serving. A few firms are exploring revenue-sharing or commercial-license structures that shift the conversation away from pure token counts and toward the economic value the customer actually extracts. The common thread is an attempt to move beyond pure commodity competition on price per million tokens.
The core commercial question has therefore changed. It is no longer simply who posts the lowest sticker price. It is whether a given model, at a given price, latency and reliability level, produces enough measurable value to justify its place in a customer’s AI budget. DeepSeek’s move to peak rates and higher absolute prices is one provider’s attempt to answer that question more precisely. The same pressure is visible across the industry. As agent usage grows and the true cost of continuous inference becomes harder to ignore, pricing is becoming less about winning volume and more about aligning cost structures with real patterns of demand and willingness to pay.
The transition will not be uniform. Some workloads will stay highly price-sensitive and will keep chasing the lowest available rates. Others, especially those embedded in critical business processes, will accept higher prices in exchange for reliability, support and predictable performance. The market now taking shape is one in which multiple price points and multiple commercial models coexist—an outcome that looks less like the pure discounting of the past two years and more like the differentiated pricing long familiar in other infrastructure markets.






快报
根据《网络安全法》实名制要求,请绑定手机号后发表评论