Guest Column | September 21, 2026

Token Costs Will Decide Clinical Supply AI

By Eshaan Jain, Senior Consultant, Mphasis

Modern Data Center-GettyImages-2220588451

The price of running a large language model has fallen faster than almost any other technology cost curve on record. Epoch AI's research found that the cost to match GPT-4-level performance on Ph.D.-level science questions dropped from $37.50 per million tokens in March 2023 to $0.18 per million tokens by February 2025, roughly a 208-fold decline in about two years.1 Every AI vendor pitching a clinical trial team right now, for supply forecasting, risk-based monitoring, or medical writing, will show you a version of that chart.

What almost none of them show you: total electricity demand from data centers grew 17% in 2025 alone, with AI-focused facilities growing even faster than that, well ahead of the 3% growth in global electricity demand overall, according to the International Energy Agency.2 AI-specific data center power use is projected to triple by 2030. Per-unit cost is falling. Total infrastructure demand is rising faster. Both things are true at once, and a clinical trial technology program that only tracks the first number is planning against half the picture.

I research token optimization and inference-cost economics outside my day job. Clinical trial technology is exactly the kind of environment where this tension shows up early, because most AI use cases in a trial run continuously, not once, and clinical supply forecasting is usually the first one deployed at real scale.

Where The Volume Hides Across A Trial

A single API call is cheap. A supply forecasting model rerun every night, across every SKU, at every depot, for the life of a multiyear trial, is not one call, it is thousands, and the number climbs every time a new site or cohort is added. A risk-based monitoring agent scanning electronic data capture (EDC) entries across every site, every day, for protocol deviations or data anomalies runs continuously by design. A medical writing assistant that reprocesses an entire clinical study report draft every time a reviewer asks a question, instead of caching what it has already analyzed, pays the same processing cost repeatedly for the same document.

None of these individual calls look expensive on a vendor's price sheet. The aggregate, run at the scale a real multisite trial operates at, is a different number, and it rarely appears in the business case that got any single pilot funded.

Put rough numbers on the supply use case specifically, since it is usually the first to scale. A midsize trial running 60 sites and 150 SKUs, with a forecast rerun daily per SKU per site, generates 9,000 forecast calls a day before anyone adds monitoring, medical writing, or patient-facing tools on top of it. At pilot scale, with five sites and a handful of SKUs, that number is small enough to round to zero on a budget line. At full trial scale, across every AI use case running simultaneously, the workload is what shows up as a surprise the following quarter.

But the workload is not just a function of site and SKU count. Clinical supply forecasts also change when enrollment accelerates, sites activate or close, protocol amendments alter dosing or treatment duration, or inventory is repositioned between depots. An AI forecasting system that responds to those events in real time can generate materially more inference than a static forecast that runs on a fixed schedule.

Falling Prices Do Not Mean Falling Budgets

This is the trap in reading the Epoch AI chart in isolation. A 208-fold price drop sounds like a reason to run more inference, not less, and that is exactly what happens: usage expands to fill the price drop, often faster than the price drop itself. A supply forecast that ran once a week at the old price can now afford to run daily, then hourly, then continuously, and the same expansion happens with monitoring agents, document generation tools, and every other AI use case a trial adopts once the per-call price looks cheap enough to stop counting.

The technology itself is not the problem. Budget for the workload you will actually run at scale, across every AI use case in the trial, not the workload you ran in the pilot with 10 SKUs at one site or one monitoring agent watching one protocol.

There is a CO2 side to this, too, and it tends to get even less attention than the dollar figure. Every one of those calls, whether it is a supply forecast, a monitoring scan, or a document draft, runs on hardware that draws power from a grid with its own emissions profile, and the IEA numbers above show that grid draw is growing faster for AI workloads than for electricity demand overall. A sponsor that reports its own carbon footprint, and many now do, will eventually need an answer for the AI layer's contribution to that number. It’s better to have the answer before someone asks for it.

What A Disciplined Program Tracks

Three practices separate programs that stay ahead of this from ones that get an unpleasant financial conversation 18 months in, and they apply the same way whether the use case is supply forecasting, monitoring, or medical writing.

Match the model to the task. A classification decision: does this shipment need a reorder flag? Does this EDC entry look like a protocol deviation? It does not require the same model as an open-ended reasoning task, such as drafting a clinical study report section. Running a smaller, cheaper, task-specific model for the high-volume decisions and reserving the larger model for genuinely complex judgment calls can cut the aggregate bill without touching accuracy on the decisions that matter.

The same principle applies to supply decisions such as whether a depot needs replenishment, whether inventory should be repositioned, or whether a site is likely to exhaust stock before its next scheduled shipment. These are high-volume decisions where the value comes from consistency and speed, not from using the most computationally expensive model available.

Cache and reuse instead of recomputing. If a contract has not changed, do not reprocess it. If a forecast input has not moved since yesterday, do not rerun the full model. If a document section has already been analyzed, do not analyze it again because a different reviewer asked a similar question.

Track cost per decision alongside accuracy, from day one of every pilot, in every use case, not just the first one that gets funded. A forecasting or monitoring agent that is 2% more accurate but costs 10 times more per decision than the version it replaced is not automatically worth deploying at scale, and you cannot make that call if nobody tracks the cost side of the ledger alongside the accuracy side.

For supply forecasting, that metric could be cost per forecast cycle, cost per SKU-depot decision, or compute cost per avoided stockout or shipment. Those measures give supply leaders a way to compare AI's infrastructure cost with the operational outcome it is supposed to improve.

The Decision This Leaves You With

Before scaling any clinical trial AI pilot from one site to your full trial footprint, start with the use case most likely to be running first, usually supply forecasting, and model its compute cost at full scale, not just the per-call price on the vendor's slide. Compare that number to the cost it is meant to prevent, using a number like the 50% investigational medicinal product waste rate McKinsey has documented in trials with poor forecasting.3 If the infrastructure cost at scale still comes in well under the waste it prevents, scale it, and use what you learned to model the next use case in line. If nobody has run that comparison for even the first one, the rest of the AI road map is being built on a number nobody has checked.

References:

  1. Epoch AI, "LLM Inference Price Trends," analysis of GPT-4 level performance pricing on PhD level science benchmarks, March 2023 to February 2025: epoch.ai/data-insights/llm-inference-price-trends
  2. International Energy Agency, "Data centre electricity use surged in 2025," 2025 data center electricity demand report: www.iea.org/news/data-centre-electricity-use-surged-in-2025-even-with-tightening-bottlenecks-driving-a-scramble-for-solutions
  3. McKinsey & Company estimate on investigational medicinal product waste, as reported by Clinical Leader, "Digitalized Drug Forecasting Minimizes Waste In Clinical Trial Supply Chain": www.clinicalleader.com/doc/digitalized-drug-forecasting-minimizes-waste-in-clinical-trial-supply-chain-0001, and Suvoda, "Strategies to Optimize Clinical Supply": www.suvoda.com/insights/blog/strategies-to-optimize-clinical-supply

About The Author:

Eshaan Jain is a senior consultant at Mphasis and serves as the lead product owner for Salesforce/Vlocity CPQ and CLM at T-Mobile, engaged through Mphasis’s consulting services. He researches token optimization, inference cost economics, and CO2-aware AI infrastructure. He is an IEEE senior member and holds professional membership with Forbes Tech Council, ACM, IEEE, Isaca, and AAAI.