AI is entering a new operational phase.
Inference workloads accounted for roughly half of all AI compute in 2025 and could reach two-thirds in 2026, reshaping how AI infrastructure is designed, operated, and optimized.
Many of the infrastructure assumptions that enabled rapid AI experimentation were built around flexibility, short-lived workloads, and easy access to compute. Sustained inference changes the economics. Infrastructure efficiency now directly shapes token costs, making utilisation, latency, and operational control critical at scale.
AI Natives are already adapting to these pressures, making infrastructure decisions that prioritise efficiency, predictability, and performance as inference workloads become more demanding.
Much of the wider market is now beginning to encounter the same challenges. This shift reflects a broader change in how AI infrastructure needs to be designed, operated, and optimized as inference workloads grow.
When infrastructure choices become advantage
AI infrastructure is no longer one-size-fits-all.

Abstraction has been one of the defining characteristics of the first wave of AI adoption. By hiding infrastructure complexity, it enabled organizations to deploy AI quickly and focus on applications rather than underlying systems.
For AI-native companies, however, the economics look different. Latency, utilization, and throughput are no longer purely technical concerns. They directly shape margins.
Infrastructure stops being a platform beneath the product and becomes part of the product itself.
A growing number of teams now sit between those two worlds. Their AI systems are already in production, workloads are scaling, and the trade-offs that simplified early adoption are becoming harder to ignore:
- Cost volatility becomes visible.
- Performance ceilings emerge.
- Scaling introduces friction that was not obvious at smaller loads.
Many organizations are discovering that infrastructure choices optimized for speed and convenience can become harder to sustain efficiently as inference scales.
Part of the challenge is that inference places very different demands on infrastructure than training workloads:

At a small scale, these differences are manageable. At a large scale, inefficiencies compound quickly. A general-purpose abstraction layer that adds 15% overhead may be acceptable at hundreds of jobs. At tens of thousands, it becomes a material cost line.
This is why greater infrastructure control is re-emerging as a genuine operational requirement, as a way to regain efficiency, predictability, and visibility as workloads mature.
Infrastructure built for AI
This difference starts at the architectural level. As Hamish Jackson-Mee, VP of Product and Design at Nscale, puts it:
“Hyperscalers started with traditional cloud. We are building for AI, which means we can be focused and selective. No historical bloat, no tech debt.”
Infrastructure designed around AI workloads behaves differently from infrastructure adapted from general-purpose cloud environments. There are fewer inherited assumptions, fewer architectural compromises, and less operational drag between layers.
AI-native companies still rely on abstraction, but increasingly need control over where it applies and how infrastructure adapts as workloads evolve.
A composable stack makes that possible.
Teams can:
- Start with managed infrastructure
- Optimise specific workloads where necessary
- Move deeper into the stack without rebuilding everything else
.png)
At scale, the ability to adapt infrastructure without repeated replatforming becomes a meaningful competitive advantage.
The teams managing this transition most effectively aren't the ones with the most compute. They're the ones who made better infrastructure decisions earlier.
This article was originally published as part of Nscale's Full Stack AI newsletter, where we share perspectives on AI infrastructure, engineering, and emerging industry trends.
.png)

.png)

.png)
