For much of the artificial intelligence boom, one question dominated the technology industry:
Who can build the biggest and most capable AI model?
A different question is becoming increasingly important:
Who can afford to run it?
The AI industry’s hardware race is shifting toward inference—the process of using a trained artificial intelligence model to answer questions, generate images, write code and perform other tasks.
The change is attracting enormous amounts of investment.
AI chip startup Positron has raised $875 million in new funding at a valuation of about $5 billion, according to current reporting, just months after being valued at roughly $1 billion.
The dramatic increase illustrates how investors are betting that the next major semiconductor opportunity will come not only from training AI models but from running them efficiently.
Training Is Only the Beginning
Training a frontier AI model requires enormous computing resources.
Thousands of advanced processors work together for extended periods, analysing massive datasets and adjusting billions or trillions of internal parameters.
That process can cost enormous sums.
But once training ends, the financial burden does not disappear.
Every interaction with the finished model requires additional computing.
Ask an AI assistant to write an email: inference.
Generate an image: inference.
Translate a document: inference.
Ask an AI coding agent to build software: inference.
As AI services gain hundreds of millions of users, those individual requests accumulate into enormous computing workloads.
Why Inference Is Becoming Critical
Imagine spending billions of dollars building an extremely sophisticated factory.
The factory may produce an excellent product.
But if each unit costs too much to manufacture, the business still fails.
AI faces a similar economic challenge.
Developers need to reduce the cost of generating intelligence.
That means producing more tokens—the units AI models process and generate—using less time, less electricity and fewer expensive computing resources.
Small improvements can become extremely valuable at enormous scale.
Positron Bets on Memory
Positron is among the startups attempting to capture this opportunity.
Its forthcoming Asimov processor uses what the company describes as a memory-first architecture.
The design reportedly supports as much as 2.3 terabytes of memory per chip.
That matters because modern AI inference can be heavily constrained by how quickly processors can access model data.
A processor capable of performing enormous numbers of calculations is less useful if it spends time waiting for information to arrive from memory.
Increasing memory capacity and bandwidth can therefore improve performance for certain inference workloads.
Nvidia Isn’t Standing Still
Startups are not entering an empty market.
Nvidia remains the dominant force in AI computing and is investing aggressively in inference.
AMD is also pursuing the market.
Cloud giants are developing their own custom processors.
Google has TPUs.
Amazon has Trainium and Inferentia.
Microsoft has invested in custom AI silicon.
Meta and other large technology companies are pursuing specialised hardware strategies.
The result is an increasingly diverse AI chip ecosystem.
Why Specialised Chips Are Emerging
GPUs became central to the AI revolution because they are extremely good at performing many calculations simultaneously.
But no single architecture is necessarily optimal for every AI workload.
Training giant frontier models has different requirements from running smaller models on millions of devices.
Real-time voice agents have different needs from image generators.
Autonomous vehicles have different constraints from data-centre chatbots.
This creates opportunities for specialised processors.
Some optimise for memory.
Others focus on energy efficiency.
Some target extremely low latency.
Others are designed for particular model architectures.
Speed Isn’t the Only Metric
The AI chip industry increasingly has to optimise several variables simultaneously.
Performance
How quickly can the system produce an answer?
Cost
How much computing infrastructure is required?
Energy
How much electricity does each request consume?
Memory
How large a model can the hardware handle efficiently?
Latency
How quickly does the user receive the first response?
Scalability
Can the system serve millions of simultaneous requests?
The winner may not simply be the processor with the highest benchmark score.
It may be the architecture offering the best economics.
AI Agents Raise the Stakes
The rise of autonomous AI agents could make inference efficiency even more important.
A traditional chatbot may perform one model request when a user asks a question.
An agent completing a complicated task could potentially make dozens or hundreds of model calls.
It may search.
Plan.
Write code.
Check its work.
Use another tool.
Analyse the result.
Try again.
That multiplies inference demand.
If agents become a normal part of everyday computing, the number of AI operations performed on behalf of each person could rise dramatically.
Why It Matters
This is why inference may become one of the largest technology markets created by artificial intelligence.
Training produces the model.
Inference produces the product people actually use.
And unlike training, which happens periodically, inference happens continuously.
Every user.
Every prompt.
Every generated token.
Every day.
The Bigger Picture
The AI industry spent its first major investment cycle building intelligence.
Its next challenge is industrialising it.
That means making intelligence abundant enough and inexpensive enough to embed inside search engines, operating systems, vehicles, robots, workplaces and everyday devices.
The companies that solve that problem could become enormously valuable.
Some may build models.
Others will build chips.
Others will supply memory, networking, electricity and data centres.
The AI economy is becoming an infrastructure economy.
What Happens Next
Expect the inference market to become one of the fiercest battlegrounds in technology.
Startups such as Positron will attempt to prove specialised architectures can outperform general-purpose alternatives on important workloads.
Nvidia and AMD will defend their positions.
Cloud providers will continue developing custom silicon.
And AI developers will increasingly choose infrastructure based not only on raw capability but economics.
The first phase of the AI race asked:
How intelligent can we make these systems?
The next phase adds another question:
How cheaply can we make that intelligence available everywhere?

