2026 itvalue文章顶部

Cheaper AI Inference Drives Record Infrastructure Spending

The cost of running equivalent AI models has fallen sharply since 2022, yet major technology firms continue to increase capital spending on data centers, memory, packaging, and power. Efficiency gains are expanding total demand through multi-step agent systems, shifting physical bottlenecks and concentrating value in slower-to-scale assets.

NextFin News — The price of machine intelligence keeps falling. The bill for the physical systems that produce it keeps rising.

Analyses of provider pricing and capability benchmarks show that the cost of reaching performance levels once associated with GPT-3.5 or early GPT-4 has dropped by factors of tens to hundreds since late 2022. In some tracked series the decline exceeds 280-fold for a fixed capability threshold. Commodity models delivering comparable results now often cost well under a dollar per million tokens. Repeated price cuts of 20 to 80 percent on mid-tier offerings have become routine.

Meanwhile the largest cloud and technology companies have raised their capital expenditure forecasts repeatedly. Combined outlays for 2026 now sit in the range of $700 billion to more than $800 billion for the top providers, with further upward revisions expected. The money is flowing into data-center construction, advanced processors, high-bandwidth memory, networking gear, power infrastructure, and cooling systems.

The two trends are linked. When the unit cost of a capability declines, the volume of economically viable uses expands. This is the pattern first noted by the nineteenth-century economist William Stanley Jevons in the case of more efficient steam engines and rising total coal consumption. In the present case, lower inference costs have made continuous, multi-step systems practical.

These agent architectures do not stop at a single answer. They plan, call external tools, check intermediate results, correct errors, and iterate. Usage data from aggregation platforms show agent-driven token volume growing many times faster than human-initiated traffic and, in recent measurements, exceeding it. Completing a given task this way can require five to thirty times more tokens than a simple chat exchange, and in some configurations the multiplier is higher. Research estimates project that agent activity could multiply overall token throughput by large factors before the end of the decade.

As a result, falling unit prices have not reduced total demand for the resources that generate intelligence. They have increased it. The binding constraints have moved along the supply chain. Early shortages focused on high-end processors. Attention then shifted to high-bandwidth memory and the advanced packaging methods that integrate logic dies with memory stacks. Packaging capacity has expanded several-fold yet remains heavily booked. High-bandwidth memory output is largely committed through 2026 and beyond. Power availability and grid connections have become limiting factors in multiple regions. Cooling systems are shifting toward liquid solutions as rack densities rise.

Value tends to settle where supply expands most slowly. General model intelligence itself has become more competitive and price-sensitive. Returns that prove harder to replicate quickly appear in the physical assets whose production cannot be accelerated by software alone, or in the operational systems that turn inexpensive cognition into reliable commercial results. The latter include feedback loops that improve with each action, sustained customer relationships, and the processes that maintain accuracy across long workflows.

Undifferentiated model access or simple intermediary layers sit between these poles. As base capabilities grow cheaper and more widely available, the ability to charge a lasting premium for mere connectivity shrinks. Buyers increasingly measure systems by concrete cost reductions or output gains.

Earlier technological transitions followed a similar path. Once a scarce human capacity became abundant and inexpensive, scarcity did not disappear. It relocated. Machine cognition is now moving in that direction. The elements that cannot be scaled at the same speed are the points at which economic returns are concentrating. How durable those points prove will depend on the pace of supply response and on how effectively the expanded cognitive capacity is converted into outcomes that justify the physical investment.

本文系作者 Chelsea_Sun 授权钛媒体发表,并经钛媒体编辑,转载请注明出处、作者和本文链接
本内容来源于钛媒体钛度号,文章内容仅供参考、交流、学习,不构成投资建议。
想和千万钛媒体用户分享你的新奇观点和发现,点击这里投稿 。创业或融资寻求报道,点击这里

敬原创,有钛度,得赞赏

赞赏支持
发表评论
0 / 300

根据《网络安全法》实名制要求,请绑定手机号后发表评论

登录后输入评论内容

扫描下载App