The most valuable AI call you'll make is often 'not this one'
When a pharma AI idea is first proposed: write the value case, sort it into four buckets, and use generative AI only when a rule will not do. Three sourced cases and one illustrative workflow, with the GxP points for each.
Regulatory status as of September 30, 2026. Check current guidance before acting. The framework and decisions are my assessments, not regulatory determinations.
Someone presents a use case and the room asks how to build it. A better first question is whether to build it.
The most valuable call is often “not this one.” The wrong tool adds build cost and validation cost, and in GxP work a wrong answer can be a patient-safety event or an inspection finding.
Here is an illustration, not a documented case: checking structured electronic batch-record data for required signatures, required fields and predefined limits. I would normally use validated rules for these checks; generative AI may add unnecessary complexity and risk. These checks support the required record review and do not replace investigation, quality-unit responsibilities or, where applicable, QP certification.
What this covers
- 1Value case
- 2Four buckets
- 3Five questions, in order
- 4Keep the AI part small
- 5Boundary creep
- 6Four worked examples
- Value case. Name the result, against a baseline, before you choose a tool.
- Four buckets. Put a process fix or a rule ahead of generative AI.
- Five questions, in order. Ask them with the person who knows the process.
- Keep the AI part small. Keep exact facts on rules, and use AI only for the wording.
- Boundary creep. A tool cleared for one use is not cleared for the next.
- Four worked examples. Three sourced cases and one illustration.
Value case
My recommendation: measure against a baseline taken before the build, and net that against the full lifecycle cost.
Each metric needs a baseline, a target, an owner, and a check date.
Lifecycle cost includes:
- build and integration
- validation
- licences and operation
- human review
- monitoring, maintenance, and change assessment
- retirement
Pick one or two of these, not all of them:
- Cost: cash saved when outside spend falls (agency, CRO, or contractor).
- Capacity: skilled hours released, at a stated loaded rate.
- Count released hours as capacity, including hours moved to other work. Do not also count them as cash savings unless spending actually falls.
- Speed: cycle time, or time to submission.
- Quality: error rate, first-pass acceptance, or review cycles.
- Compliance risk: overdue-item count and age, missed-escalation rate, on-time product quality review completion, or recurrence of comparable findings.
- A count of findings avoided is a modeled benefit only.
- Adoption: active users, task completion rate, or how often they use it. Savings count only if people use the tool.
Four buckets

My bucket-routing framework; apply all five intake questions. The bucket you pick does not decide regulatory applicability or validation scope.
Bucket 1: Process fix. The cause is a broken process, not a missing tool.
- Automating a broken process can preserve its defects.
- Fix the process first. Assess and control the change through the quality system.
Bucket 2: Light automation. Triggers, routing, notifications, and scheduling.
- The rules are simple and stable.
- Configuration may be sufficient; validation and audit-trail controls depend on intended use and documented risk assessment.
Bucket 3: Heavier automation without generative AI. Rules engines, scripts, statistical process control, and threshold checks.
- A subject-matter expert can still write the rule.
- Prefer deterministic output: same input, same output.
- In my experience, “let’s use AI” ideas often belong here.
Bucket 4: Generative AI. All three must hold. If one is missing, go back to Bucket 3 or earlier.
- The input is messy language or unstructured content.
- The judgment is genuinely hard to write as an explicit rule.
- The data exists, and you can check whether answers are right.
- Variable output can add evaluation and monitoring work.
I name the kind of model that was built: rules, machine learning, or generative AI.
Five questions, in order
Stop at the first bucket that fits.
- Does a process fix remove the pain? If yes, stop at Bucket 1.
- Can the rule be written down? If a subject-matter expert can state it in plain language, use Bucket 2 or 3, not Bucket 4.
- Is the data there and clean, and can you check the answers? Generative AI will not repair missing or bad data, and without criteria you can defend and references you can trust you cannot show fitness for intended use.
- How bad is a wrong answer, and who checks the output? In my view, for GxP uses, accountable people are trained, authorised personnel, not the model.
- Does the value still cover the full cost (see the value case)? My recommendation: choose the least-cost approach that meets the requirements.
- If a cheaper bucket hits the same targets, take it.
- If the answer is not this one, the quality work still has to be done and resourced.
Keep the AI part small
For document work in Bucket 4, keep structured data and logic on a fixed backbone, and use AI only for the wording.
A bounded AI step can simplify evaluation, but in my assessment validation should still cover the intended workflow and its risks.
For the assisted-drafting workflows discussed here, I recommend authorised review of each draft. Review does not replace validation or determine regulatory acceptability by itself.
“The model decides” is a different system from a checked draft whose fixed parts come from validated source data.
Boundary creep
My recommendation: document the impact assessment and set controls before a new GxP use.
- An assistant assessed for internal questions is not assessed for batch-record entries. Treat that new use as a change.
FDA’s April 2, 2026 warning letter to Purolea Cosmetics Lab cited the firm for using AI agents to create drug product specifications, procedures, and master production or control records, without the quality unit checking those documents.
Four worked examples
Three are sourced, and Example 4 is illustrative.
Example 1: Pfizer’s ‘Charlie’ content platform (my classification: regulated, not GxP)
Bucket 4. The cited interview does not provide a quantified review-effort baseline.
Pfizer built Charlie, named after co-founder Charles Pfizer (Digiday, February 22, 2024; Tracker entry).
- Content types: digital media, emails and digital presentations that sales teams use with physicians.
- Training data for marketing content generation: approved Pfizer content filed by treatment category and product (Digiday says Pfizer also uses a range of sources).
- Answers are checked against previously published and validated Pfizer content.
- Digiday describes a red, yellow, green system that flags assets the medical review team may want to look at more closely (its example is that a headline used many times may need less attention).
- My reading: Digiday’s example suggests that a new claim, or a known claim in a new context, gets more attention than one used many times before.
The cited sources do not specify Pfizer’s detailed implementation.
My view: an AI draft should go through medical, legal and regulatory (MLR) review before anything is published.
My assessment: determining whether content introduces a new claim, or changes an approved claim’s context, can require semantic judgment; literal text novelty alone does not establish the need for generative AI. I would validate the whole workflow, from inputs to published output.
Example 2: Sanofi’s annual product quality reviews (GxP)
Bucket 4. Jacob’s presentation reports a PQR lead-time baseline, but not a complete lifecycle-cost baseline.
Sanofi’s digital and AI page describes automating these reports with generative AI, starting with the annual Product Quality Reports (PQR), then MSAT reports (MSAT is Sanofi’s manufacturing sciences, analytics and technology function) and C&Q (commissioning and qualification) documentation, with information automatically extracted, synthesized and analyzed to draft reports.
A June 19, 2025 magazine article set a target of a 70 percent reduction in creation time for approximately 3,500 PQRs a year (Tracker entry). The Tracker entry cites an Aug 2024 Outsourced Pharma interview (about eightfold faster PQR writing); it does not cover the 70 percent target or the 3,500 figure.
The tool is named GenAIR in a Built In job posting (since removed, with its description still available), in a LinkedIn post by Laure Gurcel and on slide 10 of Yves Jacob’s May 2025 DGRA congress presentation.
My proposed design:
- Data pulls, trend charts, and out-of-trend flags stay on fixed processing, from validated process data.
- Generative AI handles the free text: themes in deviation narratives, and drafted sections.
The cited sources do not specify the complete implementation. The Sanofi page describes extraction, analysis and drafting; it does not say how work is split between fixed processing and generative AI.
EU GMP Chapter 1 section 1.10: product quality reviews are normally documented annually. Section 1.11: the manufacturer and, where different, the marketing authorisation holder evaluate the results and decide whether corrective and preventive action or revalidation is needed, under the pharmaceutical quality system.
In my assessment, authorised quality personnel should approve the review and be accountable for that evaluation.
My assessment: Generative AI (my proposed design: deterministic backbone). A PQR is expected under EU GMP Chapter 1, section 1.10, so I would validate whatever the model extracts and drafts.
- As of September 2026, draft Annex 22 on artificial intelligence was still under development. It proposed limits on generative AI in critical GMP applications.
- EMA’s June 30 and July 1, 2026 workshop sought expert input on safeguards while EMA was still considering the implications of the consultation results.
- As of September 2026, the 2011 Annex 11 is still the applicable text; a revised draft sits on the same consultation page.
Example 3: Lilly patient narratives with Yseop, Merck, and Vertex
Bucket 4 (Yseop describes its platform as hybrid natural language generation). Public sources report some before/after baselines; they do not establish a complete, independently verified net value case.
Yseop’s vendor case study (June 13, 2023) describes drafting plus medical review for Eli Lilly patient narratives. Yseop describes its platform as hybrid natural language generation, blending symbolic, machine learning and large language model techniques; the case study does not say which were used for Lilly’s narratives. It does not say which parts produced the numbers, or how the internal review ran.
These figures are reported in Yseop’s vendor case study; I have not independently verified them with Lilly.
- 2,300 narratives completed
- approximately 10,000 hours of writing and reviewing time eliminated (vendor figure, not verified as cash savings or net of review time)
- 53 percent less third-party spending on patient narratives compared to the prior year
Merck reports that its generative AI platform for clinical study reports (CSRs) cut time to a fully human-reviewed first draft from an average of 180 hours to 80. Errors in data, messaging, citations, terminology, and typography fell by 50 percent (Merck, June 25, 2025; Tracker entry).
Merck says the platform is designed to operate with rigorous oversight by qualified medical writers.
DataIQ’s 2025 award profile says Vertex’s informed consent form (ICF) Autodrafter cut ICF preparation time by 90 percent, had supported more than ten trials, and included human review from the start (Tracker entry).
My assessment: Assisted clinical documents, with different public builds.
- I would validate inputs, generation, review, approval, interfaces, and records. Do not stop at generation.
- In my assessment an approving person is necessary, but approval does not shrink the boundary you have to validate.
Example 4: Training and CAPA overdue tracking (illustrative)
Bucket 2 or Bucket 3. Illustrative only, not a specific organization.
Illustrative assumptions only:
- baseline: 40 overdue items, comprising overdue training assignments and overdue open CAPAs, measured before the build
- target: fewer than 10 items meeting the same definition 90 days after go-live
- owner: head of quality systems
- check date: 90 days after go-live
Improvement depends on follow-up and completion, not tracking alone; lifecycle costs should also be assessed.
- Due dates, status, and assignees are structured.
- Flag an open record past its due date, and escalate after a set number of days.
- A missed escalation can be a compliance risk.
My assessment: A report and a workflow trigger. No generative AI.
- Generative AI would vary output that should stay fixed, with no gain.

Good questions for your next AI project initiation meeting
- Ask “which bucket?” before “how do we build this?”
- Use the five questions with the person who knows the process.
- Write down where each tool stops and look again when its use grows.
Rules for what must be exact, AI for the language, a person for the decision.