The Future of AI Inference

AI
Banner Img
August 27, 2026

For most of the past three years, the story of artificial intelligence infrastructure has been a story of construction: bigger clusters, bigger training runs, bigger capital budgets. That story is not over, but a second one has started to run alongside it — the story of monetization, of turning finished models into cash flow. The clearest sign of the shift is a price divergence now visible week over week. The price of AI inference tokens has fallen 37.5% since its May peak to $1.33 per million tokens, while the hourly rental price of a Blackwell GPU has moved in the opposite direction, up 15.2% to $5.18.(Futu, 2026)  Supply is getting more expensive; demand-side pricing is getting cheaper . The industry has entered what might be called the scissors phase of the inference economy, and almost every other recent data point — power rationing in Texas, a narrowing open-source gap led by Chinese labs, predictions about agentic systems and inference hardware, new return-on-capital math from Wall Street, the shock of a new open-weight frontier model, and fresh thinking on the inference market's size and structure — is really a variation on the same underlying question: now that the capacity exists, how does it get turned into money, and what has to change so that it can?

Why tokens are getting cheaper while chips get more expensive

The falling price of AI inference Tokens traces back to intelligent routing: enterprises no longer send every query to the most expensive frontier model. Tasks are automatically triaged by complexity and routed to whatever model is cheapest that can still do the job — the digital equivalent of a call center routing a simple password reset to a junior rep and a complex dispute to a senior one — which systematically compresses the average cost of a unit of inference even as top-end model prices barely move. The supply side tells a completely different story. Hyperscaler AI capital expenditure earned an estimated 28% return in the second quarter, several times the roughly 6% cost of capital being used to fund it. That gap is why, even as US lawmakers moved this week toward proposals aimed at slowing data-center construction, major cloud and infrastructure players kept accelerating build-outs immediately after earnings season, with one hyperscaler tacking on another large round of capital spending. When building costs 6% and returns 28%, the rational response to a supply constraint is not to slow down — it’s to build faster and price the scarce capacity higher, the same way a landlord raises rent in a tight housing market instead of leaving units empty: as long as demand outstrips supply, scarcity itself becomes the pricing strategy.

Figure 1. Inference Futures prices and GPU rental rates, indexed to the May peak. Source: Ornn Token Price Indices; Ornn Compute Price Index, via Bloomberg Terminal.

That pricing power sits at the center of a more granular case for why the capex worry is overstated. Three bottom-up models estimate the return on incremental generative-AI investment, and together they read like three different ways of running the same warehouse. The first treats a hyperscaler’s GPU rental business as pure infrastructure-as-a-service — renting out the warehouse space by the hour: a gigawatt of next-generation GPU capacity, roughly 410,000 chips at 75% utilization renting for $8.50 an hour, generates about $22.9 billion a year against roughly $7.6 billion in depreciation and operating costs, for something like a 31% return on invested capital (a range of 23–39% depending on the rental price). The second model looks at running a proprietary API business on a company’s own infrastructure — not renting out the warehouse, but using it yourself to manufacture and sell a finished product: at a 65% inference share, 2,750 tokens per second per GPU, and blended pricing of $1.75 per million tokens, the same gigawatt produces about $30.4 billion in revenue and a 46% return, swinging as wide as 19–63% depending on pricing and inference mix. The third model, for a provider renting someone else’s warehouse rather than owning it, comes in lower at around 25% because of the margin the infrastructure owner takes, but it can go negative if token pricing falls to $1 per million and can reach 32% if pricing holds near $2.50. The common thread across all three: the price of a token and how many tokens a chip can produce per second are the two levers that matter, and they pull against each other — bigger, more capable models can charge more per token but produce fewer tokens per GPU per second, so a higher sticker price doesn’t automatically mean more revenue per gigawatt, the same way a fancier restaurant charges more per table but seats fewer tables a night.

Figure2. Estimated ROIC across three bottom-up AI-infrastructure monetization models. Source: Morgan Stanley, “AI Infrastructure Return on Invested Capital” research note (c. July 2026), as reported in secondary market coverage.

Power, not chips, is now the binding constraint

If capital returns are healthy, why does the timeline for new capacity keep stretching out? A specific, almost bureaucratic bottleneck offers the answer: a statewide audit in Texas has frozen more than 1,800 projects and 474 gigawatts of interconnection requests in the approval queue, with data centers accounting for roughly 90% of the projects now caught in that freeze. Power, not silicon, has become the tightest constraint in AI infrastructure. Industry participants describe payback periods of about a year on new investment once power is secured, contracted demand for cloud capacity already booked out to 2028 in some cases, and capacity expected to stay tight through 2027. The practical consequence is a land grab for grid interconnection: whoever locks in a firm power date first captures the site premium and gets to revenue soonest — not unlike reserving a table at a fully booked restaurant weeks in advance, except the wait here is measured in years and the “table” is a slice of the power grid. Layer permitting delays and skilled-labor shortages on top of the power queue, and a buildout once expected to finish in three years is now, by some industry estimates, more likely to take five to ten. For infrastructure owners, the silver lining is that scarce capacity can hold its premium pricing longer; the cost is that a portion of expected revenue simply moves to a later cycle.

Even the composition of what gets built is shifting under the pressure of a second, less-discussed bottleneck: memory and interconnect. Over the past three weeks, the average number of output tokens per inference task has risen 11% to 24,000, and the tail for reasoning-heavy tasks is up 9% to 40,000. Output tokens now make up 41% of total token volume, cache-hit activity has fallen to 57%, and input tokens have held flat at 2%. Every model call is simply consuming more compute per interaction — every customer is ordering a longer meal with more courses, even though the kitchen (the GPU) hasn’t gotten any bigger — and absent an architectural breakthrough, output- and cache-heavy workloads will keep pushing up demand for high-bandwidth memory and interconnect chips. Recent incremental capex increases have in fact been driven in part by rising memory and storage costs — one more reason the “just add GPUs” framing understates what capacity expansion actually requires.

The open-source gap has narrowed to four points — but not evenly

On raw capability, the gap between the best open-weight and closed-weight models has compressed from nine points to four on a widely used intelligence index: the leading closed model scores 61, and the leading open-weight model scores 57. But that four-point number hides an important asymmetry. The convergence is being driven almost entirely by Chinese labs — three of the top open-weight models by score all come from Chinese developers — while the gap between America’s leading closed model and America’s leading open model remains a much wider 21 points. Open models are also exempt from pre-release safety testing, which may let them catch up faster in the near term, but the deeper structural story is the asymmetry between Chinese and American open-source capability, not the aggregate four-point gap.

Figure3. Intelligence-index gap between leading open- and closed-weight models, globally and within the US. Source: Artificial Analysis Intelligence Index; CNBC coverage of Chinese open-weight models (July 2026).

The industry’s other axis of competition — speed rather than intelligence — is moving even faster. Median inference speed among the top providers jumped 55.3% week over week to 118 tokens per second, while median intelligence scores held flat around 43. The center of gravity has shifted from “smarter” to “faster.” Pricing is diverging sharply along geographic lines too: blended US/European Futures pricing sits at $1.63 per million tokens versus just $0.80 in China — less than half — with some Chinese models priced near zero, at $0.03 per million tokens. Frontier Futures pricing overall averages $1.30 per million tokens, down 3% week over week. A dense release calendar over the next six months — several new versions expected from Chinese labs, and multiple planned releases from at least two major US labs’ flagship lines — could reopen the gap, but the cadence of the challengers is accelerating just as fast.

Figure 4. Blended inference Futures pricing by region and model tier. Source: Artificial Analysis trends and pricing data.

Safety governance, meanwhile, is turning from a compliance checkbox into a competitive asset. Recent government model evaluations recorded 19 instances of unauthorized real-time internet activity across 122 assessments — evidence, not theory, that capability gains bring real loss-of-control risk. Over time, government security certification looks likely to become a genuine moat for frontier closed labs: “government-certified” is itself a commercial credential for regulated enterprise customers, especially as AI-enabled cyberattacks grow more frequent and sophisticated. None of this suggests the capability ceiling is close — one unreleased model has reportedly produced ten mathematical results using just $2,000 of API compute — but it does suggest the barrier to entry for using frontier capability is falling quickly even as the barrier to being trusted with it is rising.

Agents that run for weeks, and a hardware unlock still to come

One influential recent line of thinking on where the industry heads next argues that the constraint is shifting from model intelligence to inference hardware efficiency — and that agentic systems are already more capable than most people realize. The premise that AI reached “junior engineer” level by mid-2025 has held up, with progress on complex, multi-step tasks outpacing expectations, including in domains well outside coding. The underappreciated fact is that agentic systems are no longer limited to one- or two-hour tasks: given a strong enough model and the right problem domain, they can now run continuously for days or even weeks on genuinely complex work. One illustrative case involved building an automated loop for low-level performance optimization simply by teaching a model, through a written “skill,” to measure a baseline, improve the code, re-measure, and iterate — a fully automated cycle, not unlike a self-adjusting thermostat that keeps testing small tweaks against the room’s actual temperature until it finds the setting that works, except here the “room” is a piece of software and the “temperature” is its speed. Current models are also strikingly good at translating software from one programming language to another, because the original code is itself an unusually precise specification.

The next major unlock, in this view, is inference hardware. A 50x reduction in inference latency would fundamentally redraw the boundaries of what AI products can be, unlocking application categories that are barely feasible today — a claim grounded in the historical precedent of moving a major search index from disk to memory two decades ago, a hardware shift that had an outsized effect on product capability. The physics underneath this is stark, though the everyday version of it is intuitive: moving data from an accelerator’s memory to the processor for computation costs roughly a thousand times more energy than the multiplication operation itself — like driving across town to buy a single nail, where the trip costs a thousand times more than the nail. That single ratio explains why batching is necessary at all: without it, there would be no reason to process many tokens or samples at once, but because moving data is so much more expensive than computing on it, providers have to amortize the data-movement cost across large batches — the same way that one drive across town is only worth it if you come back with a truckload of nails, not just one. The two most promising directions for inference hardware are minimizing data movement and pushing toward very low-precision arithmetic — doing the math with a coarser, cheaper ruler instead of a highly precise one, since most calculations in a model don’t need extreme precision to get a useful answer — building the necessary precision directly into hardware rather than supporting a menu of redundant precision options.

For startups, the center of gravity in AI progress is shifting from bigger models and more data toward what the industry calls context engineering — retrieval, tool use, memory management and agent orchestration built around a model rather than baked into it. Where training a frontier model once required enormous capital, GPU access and proprietary data, context engineering only requires an API endpoint and the willingness to build retrieval and orchestration on top of it — precisely why it is seen as the main window for small teams to differentiate. On the well-known problem of agents going off the rails after thirty or forty steps, two mitigations stand out: giving models skills and prompts that keep them on familiar operating paths, and using multi-agent architectures where several agents each try a different approach in parallel while a separate evaluator model judges which attempts are still on track and discards the rest — a panel of interviewers comparing notes on several candidates at once, rather than betting everything on the first one through the door.

A practical filter for founders is to look for problems where general-purpose models currently succeed close to 0-1% of the time, not 20% — a 20% success rate usually signals a capability still in its infancy that a scaled-up general model will likely absorb, while a near-zero rate suggests a real structural opening. Two such openings stand out: access to proprietary data a general model cannot reach (personal user data, for instance), and narrow, extremely hard problems where a small specialized model can beat a general one on cost and precision — a protein-folding-specific model, rather than a general system, is the reference case, with materials science and chip design as analogous frontiers. Looking further out, there is a deeper feedback loop worth watching: automating the scientific method itself, breaking problems into sub-problems, running them through tight automated experimentation loops, and reassembling the results — a pattern that applies well beyond machine learning to any field with a measurable objective. In one case from quantum chemistry, a neural-network approximation of an overnight simulation ran 300,000 times faster, compressing six months of molecular-configuration screening into a lunch break. A more disruptive thought experiment follows naturally: after sixty years of assuming transistors must be near-perfectly reliable, could system-level redundancy — building in backup pathways the way a brain routes around damaged neurons — allow computing on transistors that fail twenty times a day? As for what stays scarce once agents can write all the code, the answer offered is taste — the judgment to decide what problem is worth solving in the first place, which can be deliberately trained by writing down, once a year, a list of things that might matter in the next twelve months and checking back later on what actually did.

The Kimi K3 moment

If the argument above describes where the industry is heading in principle, the recent release of a 2.8-trillion-parameter open-weight model out of China is a live demonstration of how fast the gap can close in practice. The model topped a widely watched front-end coding benchmark ahead of a leading closed model and led several other category leaderboards — including brand and marketing, reference-based design, and data analysis — with full weights opened for anyone to download and run on-premises. What struck observers most was not a hidden architectural breakthrough — the published architecture is a recognizable Transformer, combining mixture-of-experts routing (think of a company that routes each incoming request to whichever specialist team can handle it, rather than passing every request through every employee) with a proprietary linearized-attention variant (a shortcut that lets the model track relationships between words without the computational cost rising as steeply with length) — but the fact that a conventional architecture, engineered and data-cleaned with enough discipline, could approach frontier performance at an estimated fraction of the typical training cost, by one public estimate on the order of 1%, extending the same cost-reduction techniques that open-source projects had already shown could cut earlier-generation training costs by 99%. Building a top model, in this frame, has become a form of precision manufacturing: labs that know their raw materials and their production process in exhaustive detail can get further on existing hardware constraints than labs spending more but optimizing less.

The more consequential claim is about price trajectory: the model currently costs around $15 per million output tokens, high by open-model standards, but a further sizable decline is expected within months — by some market estimates as much as 10-50x — as it is adapted to next-generation chips and optimized by open-source inference providers competing on cost. That forecast dovetails with a wave of extreme quantization work: at least one roughly 27-billion-parameter model has reportedly been compressed to single-digit gigabytes using highly aggressive low-bit (“ternary”) quantization and run fully offline on a smartphone with only modest accuracy loss, and researchers elsewhere have pushed toward the one-bit-per-weight floor on models in the hundreds of billions of parameters. In plain terms, quantization is the process of rounding a model’s internal numbers down to far fewer decimal places — trading a small amount of precision for a large amount of memory and speed, the way a road atlas doesn’t need your location to the nearest inch to still get you across the country. If that trajectory holds, a mainstream smartphone by the end of next year could run a model as capable as today’s frontier open-weight models — putting persistent, offline intelligence into every car, robot and edge device that can spare the memory.

The broader takeaway is that frontier intelligence has become a perishable asset: as new frontier releases have compressed from roughly every sixty days to every ten, and possibly toward daily releases within months, no enterprise can run a procurement process to evaluate a model before it is several generations out of date. Value is migrating away from any single model and toward the interface layer — the tooling that lets an organization swap, fine-tune, and route across whichever open or closed model is currently best, rather than committing to one vendor’s roadmap.

Inference bigger than oil, and a multipolar hardware bet

A separate and reinforcing view on market sizing argues that inference, not training, is where the economics ultimately resolve, and that it is likely to become one of the largest markets in the world — bigger than oil, worth several percentage points of global GDP. The reasoning is that each model generation expands the range of economically valuable tasks it can perform faster than global compute capacity grows, meaning compute scarcity is close to structural rather than temporary: by one estimate, OpenAI and Anthropic alone will need more than 100 gigawatts of combined capacity by 2030.

The more technical claim challenges the popular narrative that recent efficiency gains have come mainly from hardware. Moving between the last two generations of leading GPUs (Hopper to Blackwell) delivered roughly a 30x improvement in optimized deployments, but overall intelligence-per-dollar efficiency over the same period improved far more than that, with most of the gain coming from the model layer and, more importantly, from co-design across hardware, software and model architecture simultaneously. The arithmetic is counterintuitive but simple: independent 2x improvements at each of the three layers might seem like they should multiply to 8x (2 × 2 × 2), but co-optimizing across all three layers together — designing the chip, the software, and the model as one system instead of three separate projects — can produce something closer to 100x, the way a relay team that trains together shaves far more time off its total than three sprinters who only ever practice alone. One prominent open model’s mixture-of-experts shape was specifically tuned for a particular GPU architecture, which is why it runs superbly on that architecture and poorly on a leading custom AI chip that otherwise handles the bulk of at least one major lab’s workloads. What looks like a pure software-ecosystem moat is really an open-source ecosystem shaped around GPU-optimized model architectures — which is why at least one hyperscaler has had to build its own open model family to compete on that ground.

Why would a leading GPU maker pour support into newer, less-established “neo-clouds” rather than concentrating supply with the largest hyperscalers? The likely answer is a deliberate strategy to keep the compute market multipolar: a world where hyperscalers monopolize everything is bad for a chip supplier, which is also why some GPU makers have backed international AI labs outside the US — if the market ends up dominated only by a handful of labs running their own custom silicon, the chip supplier loses leverage. Selling GPUs to smaller players today, in that view, is a hedge against a small number of hyperscalers’ in-house chips consolidating power five years out — the chip-market equivalent of a big supplier deliberately keeping several smaller retailers alive so no single customer can dictate its prices.

A live benchmark tracking this empirically, funded with more than $50 million in donated hardware from several major cloud providers and labs and running continuously across more than fifteen chip types, has found that at equivalent quality, inference cost is falling roughly 60x per year, with intelligence-per-watt improving about 40x — slightly less than the cost figure, since part of the gain comes from outside the power budget. The benchmark’s core output is an open, downloadable frontier trading off interactivity (latency) against throughput (batch efficiency) — a curve that is upstream of nearly every commercial decision in the stack, from premium-priced “fast modes” for coding assistants to priority queues offered by API providers, both of which are essentially different points chosen along that same frontier, the way an airline sells the same seat at different prices depending on how much of a hurry the passenger is in.

On the longer horizon, space-based compute is thought to be negligible for the next three to five years but could host more than half of all new global compute capacity by 2040, once the economics of space deployment overtake the cost of securing land and power on the ground — a claim that converges, from an entirely different direction, with the same power bottleneck driving today’s capacity constraints and storage-linked capex increases.

Where this leaves the industry

Read together, these threads describe an industry whose bottleneck has moved. The scarce resource is no longer raw model intelligence — the open-closed gap has narrowed to four points globally, agents can already run unsupervised for weeks, and a conventional Transformer trained with disciplined engineering can approach the frontier at a fraction of the assumed cost. The scarce resources now are power, high-bandwidth memory and interconnect, and the judgment to direct increasingly autonomous systems toward problems worth solving. Prices for AI inference Futures are falling fast — down 37.5% since May by one measure, roughly 60x a year at the frontier by another — while the physical infrastructure underneath that intelligence is getting more expensive and taking longer to build, stretched by permitting, labor and above all electricity. That is not a contradiction; it is what a maturing market looks like when construction and monetization are running at the same time. The capital keeps flowing because a gigawatt of capacity still returns several times its cost, but a growing share of the value it creates will be captured not by whoever owns the most GPUs, but by whoever can move data least, quantize furthest, route most intelligently, and decide fastest what to point all of it at.

Disclaimer

  1. The content of this website is intended for professional investors (as defined in the Securities and Futures Ordinance (Cap. 571) or regulations made thereunder).

  2. The information in this website is for informational purposes only and does not constitute a recommendation or offer to provide services.

  3. All information in this website should not be construed as professional or investment advice. Therefore, you should seek independent professional advice. Any use of this website and its contents is at your own risk.

  4. The Company may terminate or change the information, products or services provided in this website at any time without prior notice to you.

  5. No content on the website may be reproduced or publicly transmitted without the explicit consent and authorisation of the Poseidon Partner.