Prashant AkhawatBuilding AI-Native Enterprises
CXO Intelligence Series · Edition 09 The Decision Layer · No. 01 9 September 2026 Enterprise AI · Organizational Design · Life Sciences

The Agentic AI-First Life Sciences Enterprise

Nine in ten large companies use AI. Thirty-seven percent can find it in their earnings. The gap is not a technology gap, and life sciences is where it costs the most.

Enterprise AI Series · CXO Intelligence Series · All writing

The argument in 30 seconds

  • Pharmaceutical R&D returns rose to 7.0 percent in 2025. Strip out GLP-1 assets and the same analysis returns 2.9 percent. The machine that produces assets did not improve.
  • Nearly nine in ten organizations use AI and 80 percent report personal productivity gains, but only 37 percent can attribute any EBIT impact, a figure that did not move in a year.
  • The single largest differentiator between the two groups is not technology. Among value captors, 73 percent fundamentally redesigned workflows. Among everyone else, 25 percent did.
  • Accelerating a task inside a process built for human throughput does not shorten the process. The saved time pools in front of the next queue.
  • FDA and EMA jointly issued ten guiding principles in January 2026 that make autonomy a risk and context-of-use decision. The design requirement is already written.
  • One hundred days is enough to find out whether your organization can redesign a workflow. It is not enough to transform one, and the gate must be able to close.

01The productivity mirage

The pharmaceutical industry spent 2025 telling itself a good story about research productivity. Deloitte's sixteenth annual analysis of the twenty largest global biopharma companies put the projected internal rate of return on late-stage R&D at 7.0 percent, up from 5.9 percent the year before and the third consecutive year of improvement.1 After a decade of watching that number fall toward zero, three years of recovery reads like a turning point.

Read the same report one column across. The average cost to develop an asset rose from $2.23 billion to $2.67 billion in a single year, an increase of roughly twenty percent. Average forecast peak sales per asset rose from $510 million to $598 million. The return improved because the numerator grew faster than the denominator, not because the denominator was brought under control.

Read one column further. GLP-1 assets account for an estimated 38 percent of projected pipeline sales, and obesity has displaced oncology as the largest contributor to pipeline value for the first time in the report's sixteen-year history.

Then Deloitte does the arithmetic that settles the question. Exclude GLP-1 assets from the analysis and the rate of return is 2.9 percent.

Exhibit 1The headline number and the number underneath it
R&D IRR, top 20
7.0%
▲ from 5.9% in 2024
Same IRR, excluding GLP-1
2.9%
Deloitte's own sensitivity
Cost per asset
$2.67bn
▲ 19.8% in one year
Peak sales per asset
$598m
▲ from $510m
One molecule class is carrying the industry's headline productivity number. Meanwhile the cost of producing an asset climbed twenty percent in twelve months. Source: Deloitte, Measuring the Return from Pharmaceutical Innovation, sixteenth annual edition, published 4 May 2026. Cohort of the twenty largest global biopharma companies.1
Returns went up. Strip out one molecule class and the return is 2.9 percent. That is not a productivity recovery. It is a portfolio concentration.

This matters for what comes next, because the industry's answer to the cost side is now, almost universally, artificial intelligence. Sanofi has stated an ambition to become the first R&D-driven biopharmaceutical company powered by AI at scale, with more than seventy thousand employees inside that ambition and named platforms running across research, clinical, manufacturing and internal operations.2 Its chief executive, Paul Hudson, has been unusually direct about the measurement problem, arguing that return on AI investment isn't about flashy demos or model size.3

He is right, and the aggregate evidence says the industry has not yet taken his point.

02The adoption-earnings gap

McKinsey's State of AI survey for 2026, fielded in May and June across 1,719 respondents in 97 countries, found that nearly nine in ten organizations now use AI regularly in at least one business function, and eighty percent of respondents report that AI has improved their individual productivity.4

Thirty-seven percent attribute any enterprise-level EBIT impact to it. That figure is unchanged from the previous year. Six percent qualify as high performers, meaning they can attribute at least five percent of EBIT to AI and describe the value as significant.

Exhibit 2The enterprise AI funnel

Share of surveyed organizations, McKinsey State of AI 2026 (n = 1,719, fielded 4 May to 8 June 2026)

Use AI in at least one function
~89%
Report improved individual productivity
80%
Scaling AI across the enterprise
44%
Attribute any EBIT impact
37%
High performers (5%+ of EBIT)
6%
0%25%50%75%100%
Adoption is nearly universal. Attribution is not. The drop between the second and fourth rows is the whole problem: personal productivity gains that never reach the income statement. Source: McKinsey & Company, The State of AI, 2026 global survey.4

You will have seen the more dramatic version of this. An MIT research group reported in August 2025 that around 95 percent of corporate generative AI pilots produced no measurable return, and the number moved markets. It deserves a caveat most people repeating it do not give. The study rested on 153 surveyed leaders, 52 interviews and a review of some 300 public initiatives, and it measured formal pilot programs rather than enterprise AI value as a whole.5 I would not build a board argument on it alone.

The load-bearing evidence is the McKinsey series, because it is longitudinal. The 37 percent did not move between 2025 and 2026, across a year in which adoption, spend and model capability all rose sharply. A single alarming survey can be dismissed. A flat line through a year of that magnitude cannot.

Gartner's first Hype Cycle for Agentic AI, published in April 2026, adds the operational picture. Only 17 percent of organizations have actually deployed AI agents, while more than 60 percent expect to within two years, which Gartner describes as the most aggressive adoption intent among the emerging technologies it tracks. Agentic AI sits at the peak of inflated expectations, and Gartner's analysts note that fully autonomous agents are not ready for the majority of enterprise use cases.6 Gartner separately expects more than forty percent of agentic AI projects to be canceled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls.7

Exhibit 3Two gaps, one cause

Deployment against intent, and the single variable that separates value capture from activity

Have deployed AI agents
17%
Expect to within two years
60%+
Redesigned workflows: high performers
73%
Redesigned workflows: everyone else
25%
0%25%50%75%100%
The top pair is the intent gap. The bottom pair explains it. A three-fold difference in workflow redesign, and it is an organizational decision rather than a technical one. The companies capturing value did not buy better models. They rebuilt the work. Sources: Gartner, Hype Cycle for Agentic AI, April 2026 (top pair)6; McKinsey, The State of AI, 2026 (bottom pair).4

03What absorbs the gains

Consider what actually happens when a medical writer's first draft of a clinical study report goes from 180 hours to 80, a result Merck and McKinsey reported from a generative AI pilot that also cut errors roughly in half.8 That is a genuine, large, verified improvement in a real pharmaceutical task.

Now follow the document. It enters a review cycle whose cadence was set when drafts took four weeks. It waits for a cross-functional review meeting that convenes on a fixed calendar. It queues behind quality control checks sized for the old arrival rate. It reaches a decision forum that meets monthly. The hundred hours you saved are still saved. They are now sitting in a queue.

This is the mechanism behind the funnel. When you accelerate a step inside a process whose sequence, handoffs, controls and review geometry were all designed around the constraint of human throughput, the released time pools in front of the next constraint. The work gets faster. The outcome does not.

McKinsey's own data isolates the variable. Among the six percent of organizations capturing real value, 73 percent report having fundamentally redesigned workflows. Among everyone else, 25 percent do.4 This is the single largest differentiator in the study, and it is not a technology variable at all.

04The retrofit trap

I want to name the failure mode precisely, because it is diagnosable and most enterprise AI programs are inside it right now.

A retrofit inserts AI into a process whose shape was determined by human capacity. The agent gets a task. The task sits where it always sat. The sequence, the batch cadence, the approval geometry and the evidence trail are all inherited. The program reports a percentage time saving on the task and cannot report a cycle-time saving on the process, because the process was never the unit of change.

Three things a retrofit structurally cannot fix.

Batch cadence. Work that arrives continuously is still processed in batches sized for meetings. A monthly review committee imposes a floor of roughly fifteen days of average latency on anything that passes through it, regardless of how fast the upstream work becomes.

Review geometry. Controls designed for human error are aimed at the wrong failure mode. A human reviewer checking whether a machine fabricated a citation is performing a different task from a human reviewer checking whether a colleague was careless, and the review process usually has not been told.

Evidence design. In a regulated enterprise, the evidence that a decision was sound is generated as a byproduct of the process. Change who makes the decision without changing how evidence is produced and you get a defensible outcome with an indefensible file.

Exhibit 4Retrofit against redesign, and which gates survive
RETROFIT AI is placed inside a sequence built for human throughput Data lock and inputs Agent drafts 180h → 80h hands off Queue for monthly review committee QC and rework Filing Saved time pools here Task is faster. Cycle time is unchanged. REDESIGN The sequence itself is rebuilt around continuous machine throughput Data lock and inputs Agents draft, check and reconcile continuously Reviewers sample exceptions in flow, not in batch Evidence written at the point of decision Standing decision authority Filing Two gates removed, not two tasks accelerated The queue and the batch QC step no longer exist to be optimized
The difference is which gates survive. A retrofit accelerates the box in the middle and leaves every queue intact. Redesign deletes the queue by moving review into the flow and generating regulatory evidence where the decision is actually made.

Four questions that tell you which one you have

Programs rarely describe themselves as retrofits, so the label is not useful diagnostically. These four questions are, and they can be asked in a steering committee in under ten minutes.

Did any meeting get canceled? Genuine redesign removes coordination, because coordination existed to synchronize humans working at human speed. If the calendar is unchanged, the process is unchanged.

Did the unit of work change size? Batch sizes are fossils of throughput constraints. A submission assembled in one pass rather than in sequential functional packages is a different process. A submission assembled in the same packages, faster, is the same process.

Can you point to a role that no longer exists in the workflow, and to a new one that does? Not a job eliminated. A role in the flow. If nobody's position in the sequence moved, nothing structural happened.

Does the audit evidence look different? In a GxP environment this is the sharpest test of all. If the validation package for the new process is the old package plus an appendix about a model, the decision architecture was not touched.

A program that answers no to all four has bought capability and changed nothing. It will report impressive task-level metrics and will appear in next year's survey as one of the sixty-three percent.

05The two clocks

There is a timing argument underneath all of this that most boards have not made explicit, and it is the reason the redesign work cannot wait for the technology to settle.

At Davos in January 2026, Novartis chief executive Vas Narasimhan told CNBC-TV18 that the pharmaceutical sector is likely to see tangible benefits from artificial intelligence over the next seven to ten years.9 Coming from the chief executive of one of the industry's most committed AI adopters, that is a sober and probably correct assessment of when the science pays off.

Now set a second clock beside it. The EU AI Act's high-risk obligations, deferred under the omnibus agreement of May 2026, bite in December 2027 for standalone systems and August 2028 for AI embedded in regulated products such as medical devices.10 FDA and EMA published joint guiding principles in January 2026.11 Gartner's cancellation forecast lands at the end of 2027.

The scientific payoff is seven to ten years away. The governance and design deadlines are eighteen to thirty months away. Those two clocks run at different speeds, and the organizations that treat the first as permission to defer the second will meet the second unprepared.

Exhibit 5The two clocks
TWO CLOCKS, RUNNING AT DIFFERENT SPEEDS NOW 2027 2028 2033–2036 100-day proof Starts whenever you decide Dec 2027 EU high-risk, standalone Aug 2028 EU high-risk, embedded Seven to ten years Tangible scientific payoff, per Narasimhan CLOCK ONE · THE SCIENCE Discovery, trial design and molecular prediction compound slowly. Nothing a board does this quarter changes when this arrives. CLOCK TWO · THE ORGANIZATION Governance, autonomy design and workflow redesign are due first, and everything a board does this quarter changes them. The error Reading clock one as permission to defer clock two. The science arriving in 2033 will land inside whatever operating model exists in 2028.
The scientific horizon and the organizational horizon are not the same horizon. When the molecular payoff arrives, it will land inside whichever operating model the company happens to have built by then. That model is being decided now. Sources: Narasimhan at Davos, January 20269; EU AI Act omnibus agreement, May 202610; FDA and EMA joint principles, January 2026.11

06Where the value is, with the arithmetic shown

Abstraction is cheap in this field, so here are two places in a pharmaceutical enterprise where the number is knowable.

Regulatory submissions

Leading companies now file eight to twelve weeks after database lock, cutting historical timelines by fifty to sixty-five percent. McKinsey values that acceleration at approximately $60 million of net present value per month for an asset with $1 billion of peak sales, and puts the total at around $180 million for moving from the 2020 median to the 2024 top quartile.8

Take a mid-cap company with six late-stage assets averaging $1 billion in peak sales. Two months of acceleration on each is 6 × 2 × $60 million, or roughly $720 million of net present value. That is not a productivity metric. It is a number a board recognizes.

Two details in the same analysis matter more than the headline. Only thirteen percent of companies automate table and figure formatting at scale, and health authority query workflows remain largely unautomated despite being among the largest workload drivers. The unexploited ground is not the glamorous part of the process.

And the method McKinsey names for capturing it is zero-based design, rebuilding the entire submission process starting from the last patient's last visit and running to filing, rather than improving the steps that exist. An outside firm advising the industry arrived at the same conclusion from the opposite direction.

Protocol amendments

Tufts Center for the Study of Drug Development published new benchmarks in 2024, drawn from 950 protocols and 2,188 amendments across 16 companies and CROs. Across phases I to IV, 76 percent of protocols now carry at least one amendment, up from 57 percent in the 2015 baseline, and the mean number of amendments per protocol has risen 60 percent to 3.3.12

Volume is not the interesting number, though. This is: the average time from identifying the need to amend to the last oversight approval is 260 days, and investigative sites operate with different versions of the protocol for an average of 215 days.

Sit with those two figures, because they are the entire argument of this article expressed in a single operational statistic. Eight and a half months from recognizing a problem to having the correction approved. Seven months during which the same study is running under different instructions at different sites.

Almost none of that is drafting time.

And here the honest reading cuts against the obvious conclusion. In the 2024 data, 77 percent of amendments were judged unavoidable. Only around a quarter are the product of a design decision a better-informed process would have avoided.

Executive callout · Where to point the agents

If most amendments are unavoidable, then the prize is not preventing them. An agent that drafts amendments faster attacks a small share of 260 days and leaves the rest untouched, because the rest is oversight queueing, version reconciliation and sequential approval across functions and sites.

The redesign question is therefore not how to write amendments faster. It is why a correction everyone agrees is necessary takes 260 days to authorize, and what the process would look like if the impact assessment across every affected document and the site-level version reconciliation were produced continuously rather than assembled in sequence.

The first is automation of a task. The second is redesign of a decision. Only the second touches the 260 days.

Exhibit 6Four value pools, and what each actually requires
Value poolVerified benchmarkScaleRedesign, not automation
Submission acceleration$60m NPV per month of acceleration for a $1bn asset; leaders file 8–12 weeks post database lock~$720m NPV
6 assets × 2 months
Rebuild from last patient visit to filing, not step-level speed-ups
Amendment latency76% of protocols amended, 3.3 each; 260 days from need identified to final oversight approval; 215 days of site version divergence260 days
per amendment cycle
Continuous impact assessment and version reconciliation, not faster drafting
Clinical study reports40% end-to-end cycle-time reduction; first draft 180h → 80h with errors halvedCycle time,
not headcount
Change review cadence, or the saved hours simply queue
Health authority queriesLargely unautomated despite being a major workload driver; 13% automate table and figure formatting at scaleUnexploitedStanding response capability, not ad hoc task forces
Every row has a fast version and a redesigned version, and only the redesigned version reaches the income statement. Benchmarks from sources 8 and 12. The $720m figure applies a published per-unit benchmark to a hypothetical portfolio to demonstrate method; it does not describe any specific company.

07The stack the work actually needs

Most enterprise AI programs buy two layers and skip five. They buy models and they buy an orchestration framework, which are layers three and four below, and then they discover that value lives in layers five, six and seven, which cannot be purchased from anyone.

Exhibit 7The Agentic Enterprise Stack
SEVEN LAYERS, AND THE MONEY IS IN THE WRONG THREE L7 Operating model and accountability Who owns the decision. Whose P&L carries it. How it is reviewed. L6 Control and assurance Autonomy limits, validation, audit evidence, GxP defensibility L5 Redesigned workflow The work as it should be, not the work as it was, with agents in it L4 Agent fabric Planning, memory, tool use, orchestration, agent-to-agent protocol L3 Model and reasoning Foundation and domain models, retrieval, evaluation harness L2 Knowledge and ontology Molecule, study, site, batch, submission. Meaning an agent can act on. L1 Data and evidence substrate Provenance, lineage, immutability, the record that survives an inspection Designed Cannot be bought. Where the 73% live. Bought or built Necessary. Never sufficient. Where most budgets stop. Read bottom up to build it. Read top down to fund it.
Layers one through four are procurement decisions with vendors, budgets and reference architectures. Layers five through seven are leadership decisions with no vendor to blame. The 73 percent who capture value are the ones who did the top three.

A note on reading the diagram in both directions. You build bottom up, because an agent reasoning over data with no provenance produces conclusions you cannot file. You fund top down, because a program that has not answered the layer seven question, who owns the decision and whose P&L carries it, will produce very good pilots and no earnings impact.

A note on buying layers three and four

Layers one to four are real work and should be bought where buying is sensible. The caution is about what the market currently is. Gartner's assessment is that of the thousands of vendors marketing agentic products, roughly 130 are genuinely building agents; the rest are repackaging robotic process automation and chatbots under a new label.7 In a regulated environment this is more expensive than ordinary vendor disappointment, because a system that cannot explain its own reasoning cannot produce the evidence layer six requires.

Two procurement questions cut through most of it. Ask a vendor to show how the system decides to stop and escalate, and ask what artifact it produces when it does. A genuine agent has an answer to both because both are design decisions it had to make. A wrapped workflow engine has an answer to neither, because escalation in such a system is a branch in a script rather than a judgment, and the artifact is a log line.

The second question matters more than it sounds. In life sciences the escalation artifact is the audit trail. A system that escalates well and documents the escalation badly has moved your risk rather than reduced it.

08Autonomy is a consequence decision, not a capability decision

The most common error in agentic design is treating the autonomy level as a function of what the model can do. It is not. It is a function of what happens when the model is wrong and how quickly that can be undone.

A literature surveillance agent that misses a paper is caught by the next sweep. A batch release decision that is wrong reaches patients. Those two facts, not the benchmark scores, set the autonomy level.

Exhibit 8The Agent Autonomy Spectrum
FIVE LEVELS, ASSIGNED PER DECISION A0 Assisted Human acts. Agent suggests. A1 Supervised Agent drafts. Human approves each. A2 Delegated Agent acts in bounds. Human handles exceptions. A3 Governed autonomy Agent closes the loop. Sampled review, monitoring. A4 Self-directed Agent sets sub-goals within a mandate. WHERE THE WORK SITS TODAY Batch release decision Clinical dosing change Health authority responses Clinical study report draft Protocol design analysis Deviation triage Literature surveillance Safety case intake and coding The level is set by two questions, and neither is about the model. 1. What is the consequence if this decision is wrong, and who bears it? 2. How fast, and at what cost, can the decision be detected and reversed? Irreversible and patient-facing stays low, whatever the benchmark says. Reversible and sampled can move right today.
Autonomy is assigned per decision, never per system. The same agent may operate at A3 on literature surveillance and A1 on anything touching a dossier. A program that declares one autonomy level for the whole enterprise has not thought about it.

09What the regulators already decided

This is the part of the argument that changed in January 2026, and a surprising number of AI programs have not caught up with it.

On 14 January 2026, FDA's Center for Drug Evaluation and Research and Center for Biologics Evaluation and Research published, jointly with the European Medicines Agency, ten Guiding Principles of Good AI Practice in Drug Development, spanning nonclinical research, clinical trials, manufacturing and post-market safety surveillance.11 This sits alongside FDA's January 2025 draft guidance on using AI to support regulatory decision-making, which remains in draft and sets out a risk-based credibility assessment framework organized around a defined context of use.13

Read the ten principles as an architecture specification rather than a compliance checklist and something becomes obvious. They are not about models. Principle one is human-centric by design. Principle two is a risk-based approach. Principle four is a clear context of use. Principle eight is risk-based performance assessment. Principle nine is lifecycle management. Every one of those is a decision about who decides what, under which conditions, with what evidence.

Two regulators on two continents have converged on the same answer. Autonomy is a risk and context decision, not a capability decision.The transatlantic alignment is the news. The content is what many programs still have not built.
Exhibit 9Where the ten principles land on the stack
FDA + EMA, TEN PRINCIPLES, 14 JANUARY 2026 L7 · L6 · L5 Designed layers Operating model, control and assurance, redesigned workflow L4 · L3 · L2 · L1 Bought layers Agent fabric, models, ontology, data substrate EIGHT PRINCIPLES LAND HERE 1 · Human-centric by design 2 · Risk-based approach 3 · Adherence to standards 4 · Clear context of use 5 · Multidisciplinary expertise 8 · Risk-based performance assessment 9 · Life cycle management 10 · Clear, essential information TWO PRINCIPLES LAND HERE 6 · Data governance and documentation 7 · Model design and development practices Eight of ten regulatory principles govern the layers you cannot buy. A program that has bought models and orchestration has addressed two principles and believes it has addressed ten. The remaining eight are the same work as capturing the value. Compliance and returns are the same build.
The regulators wrote a specification for the layers with no vendor. Assignment of principles to layers is the author's reading of the published principles, not a regulatory classification. Source: FDA CDER and CBER with EMA, Guiding Principles of Good AI Practice in Drug Development, 14 January 2026.11

In Europe the timeline moved and the requirement did not. Under the digital omnibus agreement reached on 6 May 2026, high-risk obligations for standalone Annex III systems shift to 2 December 2027, and for AI embedded in regulated products under Annex I, including medical devices, to 2 August 2028. General-purpose AI obligations have been in force since 2 August 2025, prohibited practices since 2 February 2025, and Article 50 transparency obligations remain on the original schedule from 2 August 2026.10

The deadline moved. The design requirement did not. A company that treats the deferral as relief will build the wrong architecture twice.

10The formula

Every enterprise AI business case I have reviewed multiplies hours saved by a loaded rate. That calculation is why the funnel in Exhibit 2 collapses. It measures the thing that improved and ignores the thing that absorbed the improvement.

Exhibit 10A value equation that survives a CFO
V = ( T × R × A ) − C T Throughput value released Cycle time, unit cost and quality escape, priced. Not hours saved. R Redesign depth, 0 to 1 How much of the workflow was rebuilt rather than accelerated in place. A Autonomy achieved, 0 to 1 A0 to A4, set per decision by consequence and reversibility. C Assurance cost Validation, monitoring, audit evidence. Real, recurring, and rising with A. R is a multiplier, not an addend. Set it near zero and no value of T or A rescues the program.
The term that decides the outcome is the one with no budget line. T is estimated by finance, A is debated by risk, C is quoted by quality, and R is the term nobody owns. The 73 versus 25 split in Exhibit 3 is R showing up in survey data.

11The 100-day proof

The reasonable objection to everything above is that redesigning an enterprise is not a quarter's work. It is not. But proving whether your organization can do it is.

Exhibit 11One hundred days, with a gate that can close
THE 100-DAY PROOF DAY 1 – 30 DAY 31 – 60 DAY 61 – 100 Instrument the truth Redesign one workflow Run it in parallel Pick three workflows by decision density, not by AI feasibility. Map the real critical path, including every queue and standing meeting. Baseline cycle time, unit cost and quality escape rate. Take the highest-value one and draw it as if the org chart did not exist. Assign an autonomy level to every decision in it. Build the control and evidence layer before the agents. Operate the redesigned workflow alongside the existing one on the same real work. Measure the three baselines, not user satisfaction. Name the executive who will own the redesigned decision. Day 100 gate The redesigned workflow beats the incumbent on cycle time and quality, or the program stops. No extension for promising early signals. The point of the gate is that it can close.
One hundred days does not transform an enterprise. It tells you whether the enterprise can be transformed. A program that cannot be stopped at day 100 was never a test.

Two design rules make the difference between this being a real exercise and another pilot.

Select by decision density, not by AI feasibility. The instinct is to pick the workflow where the technology obviously fits. That instinct produces demonstrations. Pick instead the workflow where the most consequential decisions per unit of elapsed time are being made by people using incomplete information, because that is where redesign has somewhere to go.

Build layer six before layer four. Every organization builds the agents first and the assurance layer afterward, and then discovers that the assurance layer it needed would have required different agents. In a GxP environment this is not a sequencing preference. It is the difference between a system you can validate and a system you rebuild.

12The board's five questions

Nothing in this argument is a technology decision, which is inconvenient, because technology decisions can be delegated and these cannot.

Who owns the redesigned decision?

Not who owns the tool. When a protocol design analysis runs at A2 and recommends against a stratification the therapeutic area head wanted, someone has to own the outcome. If the answer is a committee, the workflow has not been redesigned. It has been decorated.

What autonomy level did you authorize, and who reviews it?

A named executive should be able to say, per decision class, what level the organization operates at, what evidence supports that level, and what the review cadence is. Very few can. That gap is what Gartner is measuring when it forecasts forty percent cancellation for inadequate risk controls.

What did you stop doing?

Redesign that removes nothing is addition. If the review committee still meets, if the QC step still runs on the same schedule, if the same approvals are still collected, then R in the value equation is close to zero and the program will land in the sixty-three percent who cannot find AI in their earnings.

Which of the ten principles have you actually built?

Eight of the ten FDA and EMA principles govern layers five to seven. A program that can produce a model card and a data lineage diagram has addressed two of them. Ask which of the other eight exist as working mechanisms rather than as policy documents.

Which clock is your plan running on?

If the answer is that the company is waiting for the science to mature, note that the governance and design deadlines arrive five to eight years before the science does, and that the operating model in place when the science lands is the one being decided this year.

Executive callout · The position

The life sciences industry does not have an AI adoption problem. Nine in ten organizations already use it, and the individual productivity gains are real and measurable.

It has an enterprise design problem. The work was shaped around the constraint of human throughput, and that constraint has been substantially removed in a growing set of tasks without anyone changing the shape of the work around it.

Design is a leadership act. It cannot be purchased, delegated to a platform team, or deferred until the regulation is final. And on current evidence it is worth roughly the difference between six percent of companies and everyone else.

The next article in this series takes the same argument to a place with an unusual advantage. Most health systems are trapped retrofitting AI onto twenty years of accumulated digital process debt. A few are building capacity now, at scale, on a compressed timeline, which is the one condition under which you can design the operating model assuming AI is a first-class workforce rather than bolting it on afterward.

Selected sources and further reading

  1. Deloitte, Measuring the Return from Pharmaceutical Innovation, sixteenth annual edition, 4 May 2026. Cohort of the twenty largest global biopharma companies, expanded from twelve since 2010. IRR 7.0% (2025) vs 5.9% (2024); 2.9% excluding GLP-1 assets; cost per asset $2.67bn vs $2.23bn; peak sales per asset $598m vs $510m. Deloitte UK
  2. Sanofi, Scaling AI in Healthcare, VivaTech 2026. Ambition to be the first R&D-driven biopharmaceutical company powered by AI at scale; 70,000+ employees; platforms including AI Foundry, eStudy, FUSION, ClinShow and BEAM. sanofi.com
  3. Paul Hudson, Chief Executive Officer, Sanofi, on measuring return on AI investment. STAT News, 17 November 2025. statnews.com
  4. McKinsey & Company, The State of AI, 2026 global survey, fielded 4 May to 8 June 2026, 1,719 respondents across 97 nations, 36% from organizations above $1bn revenue. Near nine in ten regular use; 80% individual productivity; 44% scaling; 37% EBIT attribution; 6% high performers; 73% vs 25% on workflow redesign. mckinsey.com
  5. MIT NANDA, The GenAI Divide: State of AI in Business 2025, August 2025. Cited with limitations stated. 153 surveyed leaders, 52 interviews, ~300 public initiatives; unit of analysis is the formal pilot rather than enterprise AI value. Methodology and reporting have both been challenged. report · critique
  6. Gartner, Hype Cycle for Agentic AI, 15 April 2026. 17% of organizations have deployed AI agents; 60%+ expect to within two years; agentic AI at the peak of inflated expectations. gartner.com
  7. Gartner, Over 40% of Agentic AI Projects Will Be Canceled by End of 2027, press release, 25 June 2025. Also: ~130 of thousands of agentic vendors assessed as genuine; 33% of enterprise software to include agentic AI by 2028 from under 1% in 2024; 15% of day-to-day work decisions autonomous by 2028 from 0%. gartner.com
  8. McKinsey & Company, Rewiring Pharma's Regulatory Submissions with AI and Zero-Based Design. Leaders file 8–12 weeks post database lock, 50–65% faster; ~$60m NPV per month of acceleration for a $1bn asset; ~$180m for a twelve-week acceleration; 40% CSR cycle-time reduction; Merck pilot 180h → 80h first draft with errors halved; 13% automate table and figure formatting at scale. mckinsey.com
  9. Vas Narasimhan, Chief Executive Officer, Novartis, speaking to CNBC-TV18 at the World Economic Forum, Davos, January 2026. report
  10. EU AI Act digital omnibus, political agreement 6 May 2026, confirmed by Council 13 May 2026. Annex III standalone high-risk deferred to 2 December 2027; Annex I embedded systems, including medical devices, to 2 August 2028. GPAI obligations in force since 2 August 2025; prohibited practices since 2 February 2025; Article 50 transparency from 2 August 2026. Gibson Dunn · Pinsent Masons
  11. US FDA (CDER and CBER) with the European Medicines Agency, Guiding Principles of Good AI Practice in Drug Development, 14 January 2026. Ten principles spanning nonclinical research, clinical trials, manufacturing and post-market surveillance. fda.gov · coverage
  12. Getz, K. et al., New Benchmarks on Protocol Amendment Practices, Trends and their Impact on Clinical Trial Performance, Therapeutic Innovation & Regulatory Science, March 2024; data collected 2022 from 950 protocols and 2,188 amendments across 16 companies and CROs. Prevalence up from 57% to 76%; mean amendments per protocol up 60% to 3.3; 260 days from need identified to last oversight approval; 215 days of site version divergence; 77% judged unavoidable. The $535,000 median Phase III direct cost is from the earlier 2015 Tufts study, whose prevalence and volume figures are superseded and are not used here. Springer Nature Link
  13. US FDA, Considerations for the Use of Artificial Intelligence to Support Regulatory Decision-Making for Drug and Biological Products, draft guidance, January 2025, docket FDA-2024-D-4689. Still in draft as of publication. Risk-based credibility assessment framework organized around a defined context of use. fda.gov

On the illustrative figure. The $720 million calculation in section 06 applies a published per-unit benchmark, $60 million of NPV per month of submission acceleration for a $1 billion peak-sales asset, to a hypothetical portfolio of six late-stage assets. It demonstrates a method of sizing a value pool and does not describe any specific company. The benchmark is sourced; the portfolio is not.

On source vintage. Where an older benchmark has been superseded, this article uses the current one and says so. The 2024 Tufts protocol amendment study replaces widely quoted 2015 figures that are still in general circulation and are materially different.

On Exhibit 9. The assignment of the ten FDA and EMA principles to stack layers is the author's analytical reading, offered as a planning aid. It is not a regulatory classification and carries no regulatory status.

About the author

Prashant Akhawat is Chief Technology and AI Officer at Ninestars Information Technologies, where he leads the AOTM agentic AI platform. He previously served as Chief Operating Officer of Telerad Tech and led technology across the Telerad Group, shipping a clinical AI product line into live radiology workflow across breast, neuro, chest and musculoskeletal imaging. He holds degrees from BITS Pilani and IMI Delhi and writes on enterprise AI architecture and regulated-industry governance at akhawat.com.

The Decision Layer

A twenty-part series published under the CXO Intelligence Series, on where enterprise value is won and lost once machines can do the work. Each article is sourced, dated and written for executives who have to decide something.

Next: No. 02, on the one condition under which a health system can design for an AI workforce instead of retrofitting one.

CXO Intelligence Series · Edition 09 · The Decision Layer No. 01 · 9 September 2026 · akhawat.com
Suggested citation: Akhawat, P. (2026). The Agentic AI-First Life Sciences Enterprise. CXO Intelligence Series, Edition 09.