Keeping GxP Data Integrity Intact as AI Enters the Regulated Lab
In regulated pharma, a record is only as good as your ability to prove where it came from. That principle has a name, ALCOA+, and for decades it governed lab notebooks and batch records. In 2025 the regulators extended it to AI, and the question an inspector asks about a model's output is now the same one they ask about any GxP record. Can you trace it back to trustworthy data? Answering it is a data governance job.
How the Revised Annex 11 and New Annex 22 Reshape GxP Data Integrity
For years, data integrity in regulated pharma had a settled vocabulary. Records had to be attributable, legible, contemporaneous, original, and accurate, plus complete, consistent, enduring, and available, the principles known as ALCOA+. The systems that held those records, the LIMS, the electronic batch records, the quality management system, were validated once and audited against 21 CFR Part 11 and EU Annex 11.
That extension is now on paper. On 7 July 2025 the European Commission, working with the inspectors of the PIC/S scheme, published draft revisions of Annex 11 on computerised systems, a new Annex 22 written specifically for artificial intelligence, and an updated Chapter 4 on documentation. Consultation closed that October, with final versions expected in the middle of 2026. Annex 11 alone grew from five pages to nineteen, and for the first time it treats cybersecurity, cloud and data-platform qualification, and identity and access management as core requirements.
The through-line is simple. Regulators now expect that every AI decision can be traced back to the data behind it, and that the audit trail proving it is reviewed rather than filed away. That is not a new burden invented for AI. It is ALCOA+ applied to a harder case.
What ALCOA+ Demands and Where AI Strains It
ALCOA+ is easy to state and hard to prove at scale. A record has to say who created it and when, stay readable and unaltered, and remain retrievable years later. The revised Annex 11 sharpens the expectation around the one control that ties it together: the audit trail must be reviewed regularly, and it cannot be disabled without a documented justification.
AI strains three of those principles in particular. A model output is hard to call attributable when no one can say which data and which version produced it. It is hard to call original or accurate when the training and reference data behind it were never traced. And it is hard to keep consistent when the same term means different things in the systems the model reads from. None of that is solved by the model. It is solved by governing the data around it.
Regulators expect a full audit trail linking every AI decision to the underlying data. That is ALCOA+, applied to a model instead of a spreadsheet.
What Annex 22 Actually Allows for AI
Annex 22 is worth reading carefully, because it is more specific than the headlines suggest. For critical GxP decisions, it permits only static, deterministic models, ones with fixed parameters that produce the same output for the same input and do not keep learning after deployment. Adaptive models and generative AI are not permitted for those critical decisions. That single rule reframes what AI in regulated manufacturing even means.
The rest of Annex 22 is a data governance specification in all but name. A model needs a documented intended use with defined inputs, outputs, and performance criteria. Its training data must be representative and free from systematic bias, and fully documented. It needs independent validation datasets, parallel deployment alongside the existing process, and continuous monitoring for performance and input drift. Every one of those requirements is a claim about data you have to be able to evidence, not a property of the model itself.
So whether the AI is a genuinely assistive tool grounded in your documents or a validated model scoring a process, the compliance work lands in the same place. You have to know what data went in, where it came from, who owns it, and how it is classified, and you have to keep that provable over time.
Data Integrity for GxP AI Is a Data Governance Problem
Line up what Annex 11, Annex 22, and ALCOA+ ask for, and it reads like a checklist a governed data platform was built to answer. A data catalog gives you the inventory of every system in scope and what data lives where. A business glossary keeps a term meaning one thing across the lab, the plant, and the report, which is ALCOA+ consistency made operational. Column-level lineage traces a value from the instrument or source system through every transformation to the record it lands in, which is the data-flow audit trail an inspector asks to see.
Around that sit the controls that make a record defensible. Ownership assigns a named, accountable person to each asset. Change history captures who altered what and when. Approval workflows and lifecycle states mean only the current, approved version of a document or definition is in use, the same discipline an electronic batch record demands. Classification flags what is regulated or sensitive. And model metadata, the intended use, inputs, outputs, versions, and training-data provenance Annex 22 asks for, is cataloged alongside the data it depends on.
One honest boundary matters here. A governed data platform is not the validated system of record. It does not run your batch record or your instrument, and no data governance product is "21 CFR Part 11 compliant" on its own, because compliance is something you validate for a system in your own environment. What it does is supply the traceability, controlled vocabulary, and audit-ready lineage that data integrity depends on, and, because it can run fully on-premise or in your own tenant, you can qualify it inside a validated environment rather than trusting a shared cloud.
How to Build Audit-Ready Lineage for GxP AI
The practical path is the same one that makes any data estate trustworthy, pointed at a regulated bar.
Catalog the systems in scope. Scan the LIMS, ERP, warehouse, and BI layer into one catalog so you have a single, current inventory of what data exists and where it lives. Dawiso connects to more than 40 sources, and for systems with no native connector, a REST API and MCP server bring them in.
Fix the vocabulary. Define each regulated term once in the glossary, linked to the actual columns behind it, so a specification limit or a batch attribute means the same thing everywhere it appears.
Trace the lineage. Generate column-level lineage from your SQL and transformations so any value in a report, or any input to a model, can be traced back to the source record. This is the evidence an audit trail review needs.
Govern the model like data. Catalog each model's intended use, inputs, outputs, versions, and training-data provenance next to the data it draws on, and classify what is regulated. That is the documentation Annex 22 asks for, kept in the same place as everything else.
Keep it in a validated environment. Run the platform on-premise or in your own tenant, with a separate metadata store, role-based access, and audit-ready change history, so it can be qualified as part of your validated estate.
The AI question in regulated pharma is not really about the model. It is about whether you can prove the data behind it holds up. Get the governed context and the lineage right, and the same foundation answers ALCOA+, Annex 11, and Annex 22 at once. The pattern is the same one that makes audit-ready lineage work in banking, applied to the regulated lab.
FAQ
Is Dawiso a validated GxP or 21 CFR Part 11 system?
Does the new EU GMP Annex 22 ban AI in pharma manufacturing?
How does data governance support ALCOA+ and audit trail review?
See it in action
Dawiso Interactive Data Lineage
Trace every value from source to report with column-level lineage, the audit-ready backbone GxP data integrity depends on.