AGP Picks
View all

The Next AI Infrastructure Challenge Is Before the First Token

As the Industry Separates Prefill from Decode, Lumai Says the Next Step Is to Rethink the Compute Architecture Powering Each Stage

OXFORD, United Kingdom, Sept. 10, 2026 (GLOBE NEWSWIRE) -- AI infrastructure has spent the past several years optimizing for the moment a model generates an answer. But as AI applications evolve, the harder problem is increasingly likely to occur before even the first token is generated.

Frontier AI companies project a roughly 1,000x increase in effective compute demand over the next five years. Delivering this using conventional digital accelerators would require an estimated $100 trillion in infrastructure investment and around 1,000 GW of additional electrical capacity. Data centers have limited power budgets, so improving how AI infrastructure addresses incoming tasks matters when every watt counts. Understanding that AI inference is not a single workload helps identify opportunities for optimization.

Prefill, which processes the input context before generation begins, and decode, which generates output tokens, have fundamentally different computational characteristics. Prefill is dominated by highly parallel matrix multiplication and is compute-bound, while decode is driven more heavily by memory bandwidth and data movement. As leading AI infrastructure providers move toward disaggregating prefill and decode into separate infrastructure pools, Lumai sees the next logical step as specializing the hardware powering each stage.

“If the workloads are fundamentally different, it makes sense to stop asking the same hardware to do both jobs,” said Phil Burr, Head of Product at Lumai. “Disaggregating prefill and decode is an important step. But the bigger opportunity is to match the compute architecture to the workload.”

The Next Bottleneck May Be Processing Context, Not Generating Tokens

Much of the AI infrastructure conversation has focused on generation: how quickly can a model produce the next token?

But AI applications are changing the economics of inference. Longer context windows, retrieval-augmented generation, multimodal inputs and agentic workflows are increasing the information that must be processed before token generation begins. A single user request can now trigger multiple model calls, each carrying accumulated context from everything that came before it. That makes prefill as important as decode. Each long prompt, retrieved knowledge set, codebase, document collection, or multimodal context must first be processed before the model can generate its response.

Today, a growing share of inference compute runs on hardware built to do two jobs adequately, rather than one well. As context length grows or an agentic workflow adds another interaction or call, more of the available power is consumed by prefill, widening the gap between the inference capacity a watt could deliver and what it actually delivers. The problem can grow unnoticed until it becomes large enough to constrain both infrastructure capacity and economics.

Prefill, therefore, sits directly on the critical path for time-to-first-token and makes its efficiency increasingly important to the overall economics of inference.

The Cost of Processing Context

More of the work is happening before the first token is generated. At scale, this creates three infrastructure challenges.

1. Power becomes the binding constraint. Data centers effectively convert a fixed megawatt budget into inference capacity. The efficiency of the prefill tier determines the quantity of useful compute that power can deliver before a model generates its first token.

2. GPU capacity ends up stranded. GPUs are highly capable inference processors, but as the same fleet absorbs more prefill computation, that capacity increasingly processes context instead of generating tokens - for which GPUs are better suited.

3. Inference economics becomes a constraint. As context grows and agentic applications trigger more model calls, the cost of processing that context can make long-context and agentic applications increasingly difficult to price competitively. That can ultimately limit not only the economics of existing applications, but what will get built in the first place.

“The real issue is what happens at scale,” said Burr. “Power is fixed, GPU capacity is finite, and every interaction adds more context to process. We have gotten to the point where the economics of inference comes down to the efficiency of the hardware running it.”

Energy Efficiency Becomes Increasingly Important

As AI demand grows, power is increasingly constraining how much inference infrastructure data centers can deploy.

Prefill requires large amounts of highly parallel matrix computation, making its energy efficiency increasingly important. Running prefill on general-purpose GPUs uses capacity that could otherwise support token generation in the decode stage, where the GPU’s memory bandwidth is better utilized. Separating the workloads creates an opportunity to improve both GPU utilization and energy efficiency - and to consider compute architectures designed specifically for prefill.

Lumai Uses Light to Address the Prefill Problem

Lumai Iris Nova uses light rather than electricity to perform the matrix multiplications that define the prefill workload, completing each vector-matrix multiplication in a single optical cycle.

The significance is not merely that optical computing can perform matrix multiplication. It is that optical computing creates the possibility of treating prefill as a purpose-built infrastructure workload rather than running it on general-purpose silicon.

With prefill more efficient, longer context windows and increasingly complex agentic workflows can become more practical and financially viable without proportionally increasing compute and power requirements.

Moving prefill onto purpose-built hardware frees GPU capacity for token generation, enabling existing infrastructure to be more productive, without wholesale replacement. In Lumai testing, Iris Nova ran billion-parameter LLMs in real time and demonstrated approximately 10× more compute per watt than GPU-based equivalents on prefill workloads.

“Our view is not that every accelerator needs to be replaced - it is that we should stop asking every accelerator to solve every problem,” said Burr. “The future AI data center will be increasingly heterogeneous, with different technologies optimized for different stages of inference.”

For a more detailed view on prefill, download the Lumai white paper “The Half of AI Inference Nobody Had Optimized

To learn more about Lumai’s groundbreaking optical AI technology, visit lumai.ai. To request an evaluation of the Lumai Iris Nova Inference Server, visit lumai.ai/eval.

About Lumai

Lumai, the optical compute company, is building the next-generation AI infrastructure for the Inference Era. Spun out of world-leading optics research at the University of Oxford in 2021, Lumai’s mission is to unlock sustainable intelligence at global scale – delivering materially faster inference, significantly higher execution efficiency, and up to 90% lower energy consumption than conventional GPU architecture.

Lumai is the recipient of the Falling Walls Award for Science Breakthrough of the Year 2025. The company was part of Intel Ignite’s first London cohort and won ‘Best Overall Technology’ at the OCP Future Technologies Symposium. Lumai’s CEO, Dr. Xianxin Guo, is an alumnus of the Royal Academy of Engineering’s prestigious Shott Accelerator program – and CTO Dr. James Spall was recognized in the 2025 Photonics 100.

For more information, visit lumai.ai or follow the company on LinkedIn.

Media Contact:
Stephanie Olsen
Lages & Associates
(949) 453-8080
stephanie@lages.com

A photo accompanying this announcement is available at https://www.globenewswire.com/NewsRoom/AttachmentNg/396a17e8-a275-4d2f-bef0-2b218aaa3955


Primary Logo

Lumai Iris Nova

Lumai’s Iris Nova is purpose-built to accelerate AI prefill workloads, freeing GPU capacity for token generation and making existing infrastructure more productive without requiring wholesale replacement. In Lumai testing, Iris Nova ran billion-parameter LLMs in real time and demonstrated approximately 10× more compute per watt than GPU-based equivalents on prefill workloads.

Legal Disclaimer:

EIN Presswire provides this news content "as is" without warranty of any kind. We do not accept any responsibility or liability for the accuracy, content, images, videos, licenses, completeness, legality, or reliability of the information contained in this article. If you have any complaints or copyright issues related to this article, kindly contact the author above.

Share this page:

Advanced Search Options

Search for:

Search scope:

Type:

Search in:

Date range:

The last

Sort by:

Sign up for:

Illinois Tech Journal

The daily local news briefing you can trust. Every day. Subscribe now.

By signing up, you agree to our Terms & Conditions.