Guest Column | August 14, 2026

Fixing The Data Problem Before AI Can Help Clinical Trial Supply

By Eshaan Jain, Senior Consultant, Mphasis

Pharmaceutical industry research, development data-GettyImages-1354838510

Clinical trial supply teams keep buying forecasting tools and getting the same shortages and overages they had before. I've watched this pattern play out in other supply chains for 15 years. The tools work. The data feeding them usually doesn't.

At Amazon, I co-built a machine learning system that extracted structured terms out of unstructured contracts across a $40 billion+ annual last-mile logistics portfolio, hitting 95% accuracy. Before that system existed, nobody could answer basic questions like which carriers had which liability terms or which lanes had which delivery windows, because the answers lived in PDFs nobody could query at scale. The routing and capacity decisions built on top of that data were only as good as the extraction underneath it. Once extraction hit 95%, the decisions changed: teams reallocated volume by actual contractual capacity instead of the closest guess an ops manager remembered from the last renewal.

I now lead Salesforce CPQ and CLM transformation for one of the largest telecom providers’ quote-to-cash system, as lead product owner on the account through Mphasis. The same problem showed up there in a different form. Clause data, product catalog data, and quote history sat in systems that didn't talk to each other. Structuring that data cut contract creation time in half. The automation built on top of it – approval routing, pricing guardrails, clause suggestion – only started working once the underlying data had a consistent shape.

Clinical trial supply has the same structure and the same failure point.

Where The Data Breaks First

Enrollment data lives in the clinical trial management system (CTMS). Shipment and inventory data live in the interactive response technology (IRT)/randomization and trial supply management (RTSM) system and depot management platforms. Country-specific labeling and packaging requirements live in regulatory documents, often as PDFs or spreadsheets maintained by regional teams. Drug accountability logs live at the site level, sometimes on paper.

A forecasting model asked to predict demand for investigational product needs all four of those sources talking to each other in something close to real time. Most programs don't have that. They have a demand plan built at protocol design, based on assumed enrollment curves, revised manually every few weeks when actual enrollment diverges from the plan. By the time someone catches the divergence, a depot is either overstocked with product approaching expiry or a site is waiting on a resupply that should have shipped two weeks earlier.

The bottleneck is data structure. Enrollment, inventory, and label data don't share a common format across CTMS, IRT, and depot systems. The forecasting math sitting on top of that data is well understood, but it is only as useful as the data feeding it..

What Changes When The Data Is Fixed

Three operational decisions shift once the inputs are structured and current.

Demand forecasting moves from a static protocol-level curve to a per-site model, updated as enrollment velocity changes. A site enrolling faster than assumed gets flagged for resupply before it runs short. A site enrolling slower gets its allocation trimmed before excess product ships and expires. The decision leaders face is not whether to forecast. It's how often the forecast refreshes, and who has authority to act on it without a full protocol amendment cycle.

Investigational medicinal product (IMP) inventory policy can shift from fixed buffer stock per depot to buffer sizing tied to real enrollment variance and shelf life. Most programs still carry buffer sized to the worst historical case, because nobody trusts the forecast enough to carry less. Once the forecast runs on live enrollment and shipment data instead of a protocol design assumption, supply teams can carry a smaller, more targeted buffer and free that working capital without raising stockout risk. That only works if leadership changes the buffer policy. Buying the forecasting tool alone won't do it.

RTSM allocation logic can respond to real enrollment variability instead of the fixed randomization ratios and block sizes set at protocol design. When one region enrolls faster than another, the system should reallocate available supply toward the region that needs it, within whatever constraints the protocol allows. Today that reallocation is usually a manual escalation. It could be a rule an agent executes and logs, with a human reviewing the exceptions instead of every transaction.

The Order Matters

I've seen enterprise programs buy a forecasting or agentic AI tool first and spend the next 18 months trying to retrofit clean data into it. The programs that get value in a reasonable timeframe do the reverse. They standardize site-level enrollment reporting, depot inventory identifiers, and label and packaging metadata first, in a format a model can consume, and bring in the AI layer once that foundation holds.

That data structuring work is unglamorous and rarely gets its own budget line. It also decides whether the forecasting tool a supply team buys next year produces a number worth acting on.

The question clinical supply leaders should answer this quarter: can enrollment, inventory, and label data across CTMS, IRT, and depot systems be queried in one place today? If the answer is no, that's the project. The AI comes after.

About The Author:

Eshaan Jain is a Senior Consultant for Mphasis and through his work serves as lead product owner for Salesforce/Vlocity CPQ and CLM at T-Mobile.. He is an IEEE senior member and Forbes Tech Council member who writes on enterprise AI, applied machine learning, and agentic systems for IEEE, Elsevier, and the Forbes Tech Council.

eshaanjain26@gmail.com | linkedin.com/in/eshaanjain2