Subscribe to our Daily Briefings
HomeArtificial IntelligenceAI Inference Software Is Driving the Next Compute Investment Boom

AI Inference Software Is Driving the Next Compute Investment Boom

Featured Summary:

  • AI inference software is taking a larger role in the AI compute buildout
  • Nvidia is expanding Dynamo as open-source engines widen developer choice
  • Billions of dollars are being committed to inference platforms and specialised chips
  • Inference infrastructure is projected to overtake training in 2027

Nvidia released Dynamo 1.0 for production use in March 2026. The open-source framework manages requests, GPU resources and memory across large inference systems and supports TensorRT-LLM, vLLM and SGLang.

Nvidia says Blackwell systems using Dynamo served as many as seven times more requests in cited benchmarks.

ByteDance, Meituan, Pinterest, Baseten, CoreWeave and Together AI are among the companies Nvidia lists as using Dynamo in production.

Afritech Biz Hub Daily Briefings — get the week’s Africa business, tech, and finance signals. Sign up here.

Coding agents and other multi-step applications can generate dozens or hundreds of model calls during a single session, increasing demand for scheduling, routing and cache management.

More commercial AI use will put greater pressure on how efficiently existing accelerator fleets are run.

Open-Source Software Is Keeping AI Inference Competitive

Nvidia’s Dynamo works with vLLM, SGLang and TensorRT-LLM rather than tying users to one serving engine. Nvidia documentation also lists support for Nvidia and AMD GPUs and Intel XPUs.

vLLM runs across Nvidia CUDA, AMD ROCm and Intel XPU environments, with a separate backend for Google Cloud TPUs. That allows the same serving software to operate across several accelerator platforms.

U.S. export controls limit sales of some advanced accelerators, while large-scale cloud capacity remains concentrated among a small number of providers.

Open inference frameworks do not remove those restrictions, but they give companies more flexibility when deploying models across different hardware.

Nvidia still benefits from its installed GPU base and software ecosystem. Developers, however, have several serving engines available when choosing how to run inference workloads.

AI Inference Software Is Drawing New Investment

Baseten raised $1.5 billion in June 2026 after reporting revenue growth of about twentyfold over the previous year.

The company said the proceeds would support additional compute capacity, software development and hiring.

Together AI raised $800 million in July at an $8.3 billion valuation. Its infrastructure is used to train and serve open models, with spending directed at production workloads.

Etched raised $700 million in August to develop specialised inference chips and has reported more than $1 billion in customer contracts. Its chips are designed for running large models at lower cost and energy use.

Recent funding has reached model-serving platforms, compute capacity and inference-specific hardware. Investors are also paying more attention to cost per token and how efficiently each unit of compute is used.

Inference Revenue Is Projected to Pass Training in 2027

S&P Global projects inference-related AI infrastructure revenue to rise from $101 billion in 2025 to $532 billion in 2030, a compound annual growth rate of 39%.

Training and fine-tuning infrastructure is also expected to grow, from $127 billion to $336 billion, with inference projected to become the larger market in 2027.

S&P includes hardware, software and accelerated cloud services in its inference category. It also notes that serving software is being developed more closely with servers and runtime environments as companies deploy more AI systems in production.

Nvidia says coding agents such as Claude Code and Codex can generate hundreds of API calls in a single session.

In one 42-call Claude Code example, cache reads reached 891,000 tokens, nearly 12 times the amount written.

That level of repeated model execution gives inference a larger role in the infrastructure companies are building for commercial AI.

AI Adoption Is Turning Inference Into a Recurring Compute Cost

Microsoft 365 Copilot has passed 30 million paid seats, while Azure and other cloud-services revenue rose 43% in the June quarter.

Microsoft reported $90 billion in quarterly revenue and said Azure revenue exceeded $100 billion for the full fiscal year for the first time.

Google’s first-party model APIs are processing about 22 billion tokens a minute, up from 16 billion one quarter earlier.

Nearly 500 Google Cloud customers each processed more than one trillion tokens over the previous year, while Cloud revenue rose 82% in the second quarter. Google said demand for its models continued to exceed available supply.

Meta expects capital expenditure of $130 billion to $145 billion in 2026 after spending $31.1 billion in the second quarter. Stanford’s 2026 AI Index puts organisational AI adoption at 88%, up from 78% a year earlier.

AI services are creating more recurring compute demand across cloud platforms and enterprise software.

Serving those workloads will matter more to margins as usage grows. Better accelerator utilisation can reduce the amount of new capacity companies need to add.

Gideon Omojaunfo
Gideon Omojaunfo
Gideon Omojaunfo covers Africa’s business, technology and financial markets, with a focus on macroeconomic policy, capital flows and FX regimes. His analysis examines structural reform, digital infrastructure and investment risk across the continent.
RELATED ARTICLES

Most Popular

Recent Comments