What’s new: Openlayer named in the 2026 Gartner Market Guide® for AI Evaluation and Observability Platforms. Learn More

GPAI Documentation Requirements Every Provider Needs (August 2026)

Published August 5, 202618 min read

The August 2026 deadline for high-risk system obligations is getting most of the attention, but GPAI obligations became enforceable a full year earlier. If you're a provider of a general-purpose AI model and your technical documentation, training data summary, or copyright compliance policy isn't in place yet, that's not a preparation gap. It's an open enforcement exposure. Here's what the EU AI Act requires you to have documented right now.

TLDR:

  • Fine-tuning an open-source GPAI model and shipping it in a product likely reclassifies your team as a provider under the EU AI Act, with full Article 53 documentation obligations.
  • Article 53 requires four live artifacts: technical documentation, a copyright compliance policy, a training data summary, and downstream transparency disclosures.
  • Systemic-risk models (trained on compute exceeding 10²⁵ FLOPs) face Article 55 obligations including structured red-teaming, a 15-day incident reporting window, and energy consumption disclosure.
  • Non-compliance with baseline GPAI obligations carries fines up to €15M or 3% of global turnover; systemic-risk failures escalate to €35M or 7% under Article 99.
  • Openlayer connects evaluation, observability, and governance into a single workflow so that test results, drift alerts, and incident logs exist as traceable audit artifacts, not records reconstructed after the fact.

What qualifies as a GPAI model under the EU AI Act

The EU AI Act draws a clear line between AI systems built for a specific task and those designed to serve a wide range of purposes. A GPAI model, under the Act's definition, is any AI model trained on large amounts of data, exhibiting general competence across diverse tasks, and capable of being integrated into a variety of downstream applications or systems.

The distinction matters because it determines which obligations apply. A model purpose-built to screen medical images sits in a different regulatory category (see the guide on high-risk AI systems under the Act) than a foundation model a developer can fine-tune for customer service, code generation, or document summarization.

There are two tiers within the GPAI category worth knowing.

  • Models that meet the baseline GPAI definition are subject to documentation, transparency, and copyright compliance obligations regardless of their downstream use. Any model deployable across multiple task types falls here by default.
  • Models that pose systemic risk carry a heavier set of requirements. The threshold the Act sets is training compute exceeding approximately 10²⁵ FLOPs. At that scale, regulators treat the model as posing risks that extend beyond any single deployment context, and providers face additional adversarial testing, incident reporting, and cybersecurity obligations on top of the baseline set.

One clarification worth making explicit: the GPAI classification attaches to the model itself, not to how any particular deployer uses it. A provider releasing a general-purpose foundation model through an API carries GPAI obligations even if most downstream applications are narrow. The intended use of the downstream product does not change the upstream provider's classification.

Who counts as a GPAI provider

The EU AI Act draws a clear line: if your organization places a general-purpose AI model on the market or puts it into service, you are a GPAI provider subject to the regulation's documentation and transparency obligations. This applies whether you built the model from scratch, fine-tuned an existing foundation model, or released it through an API that third parties build on top of.

There are two tiers within this classification, and the distinction matters for what you must produce.

  • All GPAI providers must meet baseline documentation requirements covering training data, model architecture, capabilities and limitations, and known risks. These apply regardless of model scale.
  • Providers of models with systemic risk face a second layer of obligations. A model reaches systemic risk designation when it is trained on compute exceeding approximately 10²⁵ FLOPs. At that threshold, providers must conduct adversarial testing, report serious incidents to the European AI Office, and maintain enhanced technical documentation available for audit on demand.

One classification question that catches organizations off guard: fine-tuning. A team that takes an open-weight foundation model and fine-tunes it for a specific vertical is almost certainly a GPAI provider under the Act, even if the base model was built elsewhere. The fine-tuned artifact is what goes to market; the organization releasing it owns the compliance obligations that come with it. Treating the original model developer as the sole responsible party is a misreading the Act does not support. The distinction between EU AI Act provider vs deployer obligations determines who owns each requirement.

A second common misread involves internal deployment. Organizations that develop a GPAI model strictly for internal use, with no external release, occupy a different position than providers placing a model on the market. But "internal" has a narrow interpretation here. If the model is accessible to business units outside the team that built it, or if it informs decisions affecting external parties, regulators are unlikely to accept a purely internal classification without scrutiny.

The four baseline obligations under Article 53

Every GPAI provider operating under the EU AI Act must meet four baseline obligations under Article 53 of the EU AI Act, regardless of whether their model carries systemic risk designation. These apply to any general-purpose AI model made available in the EU market.

Here is what each obligation requires in practice:

  • Technical documentation: Providers must produce and maintain records covering model architecture, training data sources and volumes, compute used during training, known limitations, and foreseeable misuse scenarios. This is not a one-time filing; it must stay current and be available for on-demand review by national competent authorities.
  • Copyright compliance policy: Providers must implement and document a policy for complying with EU copyright law, including how the model handles text and data mining exceptions under Directive 2019/790. The policy must be specific enough that an auditor can verify it was applied during training data selection.
  • Training data summary: Providers must publish a sufficiently detailed summary of the content used to train the model. The EU AI Act Office has released a template for this, but the summary must go beyond a list of dataset names; it needs to describe data types, sources, and any filtering or governance steps applied.
  • EU AI Act transparency obligations: Providers must supply downstream deployers with the technical information they need to meet their own EU AI Act obligations. That means sharing documented capability boundaries, known failure modes, and output characteristics, not API documentation alone.

What Annex XI and Annex XII technical documentation must cover

Under the EU AI Act, GPAI providers face two distinct documentation tracks depending on whether their model crosses the systemic risk threshold. The structure matters because the obligations differ in scope, depth, and who gets to inspect them.

There are two main documentation regimes to account for:

Documentation requirementAnnex XI: standard GPAIAnnex XII: systemic risk (>10²⁵ FLOPs)
General description, intended purpose, and task rangeRequiredRequired
Training data sources, volumes, and data governanceRequiredRequired
Compute resources (FLOPs), training duration, and hardwareRequiredRequired
Test and evaluation results, including benchmarksRequiredRequired
Known limitations and foreseeable misuse scenariosRequiredRequired
Public capability summary for downstream deployersRequiredRequired
Adversarial testing and red-team resultsNot requiredRequired
Risk management documentation (catastrophic/irreversible harm)Not requiredRequired
Incident records reported to the AI OfficeNot requiredRequired
Cybersecurity measures for model weights and infrastructureNot requiredRequired

Annex XI: standard GPAI models

Annex XI covers all GPAI models that do not meet the systemic risk threshold. The documentation must include:

  • A general description of the model, its intended purpose, and the tasks it can perform across modalities
  • Information on training data sources, volumes, and data governance practices applied during preparation
  • Compute resources used for training, expressed in FLOPs, along with training duration and hardware infrastructure
  • Test and evaluation results, including benchmarks used and performance characteristics across relevant tasks
  • Known limitations, foreseeable misuse scenarios, and mitigations applied before release
  • A summary of the model's capabilities made publicly available so downstream deployers can assess fit for their use case

Annex XII: Systemic Risk Models

Models trained on compute exceeding approximately 10²⁵ FLOPs face Annex XII requirements on top of Annex XI. The additional obligations include:

  • Detailed adversarial testing results, including red-teaming outcomes conducted before and after model release
  • Documentation of EU AI Act risk management system requirements covering catastrophic or irreversible harm scenarios
  • Records of any incidents reported to the AI Office, including the nature of the incident and corrective actions taken
  • Cybersecurity measures implemented to protect model weights and infrastructure from unauthorized access or extraction

The practical consequence of this two-tier structure is that a provider cannot treat Annex XI as a ceiling. If a model crosses the compute threshold at any point, including through fine-tuning or extended training runs, Annex XII obligations attach retroactively to the documentation record. Providers that begin documenting only after a model is deployed face a gap that auditors will read as an absent control, not a late start.

Systemic risk GPAI models: Article 55 obligations

Models trained on compute exceeding approximately 10²⁵ FLOPs fall into a separate obligation tier under Article 55 of the EU AI Act. The scale of these systems creates systemic risk exposure that standard GPAI documentation requirements do not fully cover, so the regulation layers on additional requirements targeted at this category. The European Commission's official GPAI Q&A confirms this threshold was calibrated to capture the most advanced models available at the time of legislation.

There are four obligation areas Article 55 adds on top of the baseline GPAI documentation requirements.

  • Adversarial testing: providers must conduct model evaluations, including red-teaming, to identify failure modes, capability boundaries, and misuse vectors before deployment. In practice, that requires running structured adversarial probes against the model across the risk domains most relevant to its intended use cases, not a general capabilities benchmark.
  • Incident reporting: serious incidents and corrective measures must be reported to the European AI Office. The applicable reporting window under Articles 61 and 72 is 15 days from the point a serious incident is identified.
  • Cybersecurity protections: providers must put in place protections commensurate with the risk profile of a model at this scale, covering the model weights, inference infrastructure, and any fine-tuning interfaces that downstream deployers access.
  • Energy consumption disclosure: providers must document and publish information about the energy consumption associated with training the model. This is not a voluntary disclosure; it is a documented obligation tied to the systemic-risk classification.

One reading teams sometimes bring to Article 55 is that adversarial testing is satisfied by standard pre-release evaluation. That misreads the obligation. The requirement is for structured red-teaming across failure modes, not a capabilities sweep. A benchmark run that measures performance on standard tasks does not produce the evidence of misuse-vector identification that the European AI Office would expect to see in an audit. The artifact that satisfies Article 55 is a red-team report with named failure modes, not an evaluation leaderboard entry.

Downstream builder obligations: when modifiers become providers

When a downstream company fine-tunes, distills, or substantially modifies a GPAI model, the EU AI Act treats that company as a provider in its own right. The modifier inherits provider-level obligations (technical documentation, transparency disclosures, copyright summaries) and cannot shelter behind the original developer's compliance posture.

The threshold matters here. Minor prompt customization or output formatting does not trigger this reclassification. But retraining on proprietary data, distilling a larger model into a smaller one, or adjusting weights to shift model behavior crosses into substantial modification territory. At that point, the modifier must produce its own Annex XI documentation covering the modified architecture, the additional training data used, and any changes to the model's capabilities or known limitations relative to the base model.

There are three practical implications for teams building on top of foundation models:

  • Downstream fine-tuning records must be maintained from the start of the modification process, not reconstructed after the fact. Training data provenance, compute used, and evaluation results against the modified model are all required artifacts.
  • The original provider's technical documentation does not transfer automatically. Modifiers must either extend it or produce a new document that accurately reflects the modified system's behavior.
  • Deployers who receive a modified GPAI model are entitled to the same transparency disclosures they would receive from any provider. If your organization is the modifier, you become the disclosure source.

The compliance gap most teams miss is the handoff point (worth reviewing against an EU AI Act high-risk compliance checklist): a model leaves a foundation provider with documentation intact, gets fine-tuned internally by a product team, and re-enters deployment without any updated records. The fine-tuned artifact is now a different model under a different provider, but the audit trail treats it as though nothing changed.

The open-source exemption and its limits

The open-source exemption under the EU AI Act sounds straightforward on paper: providers that release model weights publicly under an open license are exempt from most GPAI obligations. But this exemption is narrower than many teams assume, and misreading it creates real compliance exposure.

The exemption covers transparency and copyright obligations only when the model is genuinely open, meaning weights are publicly available and the license permits inspection, modification, and redistribution. But two conditions strip that exemption away entirely:

  • If the open-source GPAI model poses systemic risk (trained on compute exceeding approximately 10²⁵ FLOPs), the full systemic-risk obligation set applies regardless of licensing. Open weights do not change the risk profile the regulator cares about.
  • If a downstream deployer fine-tunes or modifies an open-source GPAI model and deploys it in a product, that deployer may now carry provider-level obligations for the modified system, even though the base model was exempt.

The second condition is where teams most often miscalculate. A team that pulls an open-source base model, fine-tunes it on proprietary data, and ships it in a customer-facing feature has likely crossed from consumer to provider. At that point, GPAI documentation requirements follow the modified model, not the original release.

Implications for Teams Building on Open-Source GPAI

Before assuming the exemption applies, teams should verify three things:

  • The base model's training compute sits below the systemic-risk threshold
  • The deployment does not involve modifications substantial enough to recharacterize the team as a provider
  • The license terms genuinely qualify as open under the Act's definition, not merely permissive by common convention

When any of these conditions is uncertain, the safer posture is to treat general-purpose AI documentation obligations as active and build the audit trail accordingly.

The GPAI Code of Practice: voluntary tool, real consequences

The Code of Practice for GPAI models sits in an interesting regulatory position: participation is voluntary, but the documentation it requires feeds directly into enforcement. Under the EU AI Act, adherence to an approved code of practice creates a presumption of conformity with the Act's GPAI obligations. That presumption matters when regulators come asking.

The Code has gone through multiple iterations since its first publication in early 2025 by a multi-stakeholder drafting body convened by the EU AI Office. The latest version covers four areas providers should be tracking:

  • Transparency and copyright-related rules, which require providers to publish sufficiently detailed summaries of training data and to document policies for honoring copyright opt-outs under Article 53.
  • Risk identification and mitigation for systemic-risk models, covering red-teaming protocols, adversarial testing results, and incident response procedures that feed into the ongoing reporting obligations under Articles 55 and 72.
  • Technical robustness measures, including documentation of evaluation results against recognized benchmarks and any residual risks identified but not fully mitigated.
  • Governance commitments, specifying internal accountability structures and the named roles responsible for compliance decisions across the model lifecycle. In practice, this means naming a designated AI compliance officer with documented authority over deployment approvals and incident escalation, and recording that designation in writing so an auditor can verify who held accountability at each stage.

Providers who choose not to sign the Code face a higher documentation burden: they must prove compliance through alternative means, with no presumption of conformity to fall back on. The practical consequence is that non-signatories carry a heavier audit load, not a lighter one. Regulators will ask for the same evidence either way; the Code just provides a structured framework for organizing it.

GPAI enforcement timeline and penalty structure

GPAI enforcement under the EU AI Act runs on two parallel tracks, and the penalty structure attached to each is steep enough that documentation gaps are not a recoverable oversight.

The first track covers general GPAI obligations: the transparency requirements, technical documentation, copyright compliance, and training data summaries that apply to all GPAI model providers. Non-compliance here falls under Article 99(3): fines of up to €15 million or 3% of global annual turnover, whichever is higher.

The second track applies to prohibited AI practices, which fall under Article 99(6): up to €35 million or 7% of total worldwide annual turnover. This tier is scoped to prohibited practices, not to systemic-risk GPAI documentation failures. Providers of systemic-risk models who fail adversarial testing, incident reporting, or cybersecurity obligations remain subject to Article 99(3) at the same €15M/3% ceiling as baseline GPAI violations — the systemic-risk classification does not automatically escalate the penalty tier.

There are three enforcement milestones worth tracking:

  • Providers who have not published a training data summary, documented training data sources, or published a copyright compliance policy are already exposed under the current enforcement window, not a future one.
  • August 2026: High-risk system obligations take full effect. For providers whose models feed downstream high-risk applications, this deadline compounds the GPAI documentation gap with EU AI Act conformity assessment requirements under Article 43.
  • Ongoing: The European AI Office, designated as the central enforcement body for GPAI models, can request documentation on demand. An incomplete technical summary is not a gap to close before the next audit cycle; it is an open finding from the moment the August 2025 deadline passed.

The enforcement posture here rewards providers who treat documentation as a live artifact, not a pre-submission checklist. Regulators reviewing a GPAI model are not looking for a one-time filing; they are looking for evidence that documentation reflects the model's current state, including updates to training data, capability expansions, and any newly identified limitations.

How Openlayer supports GPAI documentation and governance obligations

Tracking GPAI obligations across documentation, monitoring, and incident reporting is a multi-system problem. Most organizations handle it with a patchwork of spreadsheets, manual audit logs, and point tools that cover individual stages but leave gaps between them. Those gaps are where compliance exposure accumulates.

Two platforms currently hold the most governance-layer mindshare for compliance teams working on GPAI obligations: Credo AI and IBM watsonx.governance. Credo AI covers governance workflow coordination and structured evidence collection well, giving compliance and legal teams a familiar intake environment for policy documentation and review processes. IBM watsonx.governance maps policy requirements to model records and provides a framework for tracking compliance decisions across the model lifecycle. The constraint both platforms share is architectural: each records what an organization intends, not what a deployed model does in production. Neither connects to the live inference pipeline to monitor behavioral drift, block unsafe outputs, or generate incident records traceable to a specific model version and deployment event.

Openlayer closes those gaps by connecting evaluation, observability, and governance into a single workflow that produces audit-ready evidence at each stage of the model lifecycle.

Here is how that maps to the specific obligations GPAI providers face:

Technical Documentation

Openlayer's model registry captures the artifact-level detail that Annex XI requires: training data provenance, model version hashes, evaluation results, and the approval events that authorized deployment. When an auditor asks to see the technical documentation for a specific model version, that record is retrievable and traceable, not reconstructed from memory.

Behavioral Testing and Capability Mapping

Over 175 pre-built tests cover the output quality, safety, and fairness dimensions that high-risk AI model evaluation requires. Test results, pass/fail records, and flagged failure modes are written to the audit trail automatically, so the behavioral evidence regulators expect under Article 53 exists as a structured artifact, not a spreadsheet assembled after the fact.

Production Monitoring and Drift Detection

After deployment, Openlayer monitors live model outputs against the thresholds configured at deployment time. When a metric drifts beyond a defined boundary, the system generates a timestamped alert with the specific metric, the threshold breached, and the model version in question. That alert record becomes the EU AI Act post-market monitoring evidence Annex XI requires, and the monitoring status is visible without manual aggregation across separate tools.

Runtime Enforcement

Logging drift is observation. Openlayer's deployment gates go further: when a groundedness score falls below the configured threshold or a demographic parity gap exceeds the approved limit, inference can be blocked before the output reaches the user. That blocking step, beyond logging the anomaly, is what separates enforcement from observation, and it is the distinction that makes compliance evidence defensible instead of merely documentary.

Incident Reporting Support

When a serious incident occurs, the post-incident investigation record requires the exact input values, the output with confidence scores, the model version hash, and a root cause classification. Openlayer's inference-time logging captures all four, producing audit evidence from LLM traces before the investigation begins, which means the 15-day reporting window under Articles 61 and 72 is not spent reconstructing what the model did.

Credo AI and IBM watsonx.governance cover governance documentation and policy mapping well. But both currently stop at the policy layer: they record what an organization intends, not what a model does in production. That gap matters when preparing for an AI model audit. Openlayer's position is that a governance record without runtime evidence is a policy document, not a compliance artifact. The documentation and the behavioral evidence need to come from the same system, traceable to the same model version, for either to hold up under audit.

Final thoughts on general purpose AI documentation and governance

The GPAI obligations covered here are not a future compliance exercise. Your technical documentation, training data summaries, and copyright compliance policies needed to be in place when August 2025 arrived, and auditors will read any gap from that date forward as an absent control, not a late start. Fine-tuning a base model, deploying through an API, or building on open-source weights each carry their own classification questions worth resolving before an audit does it for you. Reach out to the Openlayer team if you want to see how behavioral evidence and documentation can be kept in sync across your model's full lifecycle.

FAQ

What GPAI documentation does the EU AI Act actually require providers to produce?

All GPAI providers must produce four categories of artifacts under Article 53: technical documentation covering model architecture, training data sources, compute used, and known limitations; a copyright compliance policy specific enough for an auditor to verify it was applied during data selection; a training data summary describing data types, sources, and filtering steps applied; and downstream transparency disclosures giving deployers the capability boundaries and failure modes they need to meet their own obligations. Annex XI specifies the standard documentation track, and models trained on compute exceeding approximately 10²⁵ FLOPs face Annex XII requirements on top of that baseline.

What's the fastest way to close GPAI documentation gaps before an EU AI Office audit?

The most defensible path is connecting evaluation, model registry, and production monitoring into a single workflow so audit evidence is generated automatically at each lifecycle stage, not reconstructed after the fact. A platform like Openlayer captures training data provenance, model version hashes, evaluation pass/fail records, and inference-time incident data as byproducts of normal operation, which means the technical documentation Annex XI requires can be in place before an auditor asks, as part of a broader compliance program. The August 2025 enforcement deadline has already passed, so any gap in your documentation record is an open finding now, not a future risk.

Does fine-tuning an open-source GPAI model make my organization an EU AI Act provider?

Yes, in most cases. When your team retrains on proprietary data, distills a larger model into a smaller one, or adjusts weights to shift model behavior, the EU AI Act treats that modification as substantial enough to reclassify your organization as a provider with its own Annex XI documentation obligations. The original model developer's compliance record does not transfer to your fine-tuned artifact. The open-source exemption only holds when the base model's training compute sits below the systemic-risk threshold, your modifications do not cross into substantial-modification territory, and the license genuinely qualifies as open under the Act's definition.

Credo AI or Openlayer for GPAI general purpose AI compliance programs?

Credo AI is a strong fit if your primary need is policy workflow coordination and structured evidence collection; it handles governance paperwork well and gives compliance and legal teams a familiar intake environment. The architectural gap is that Credo AI does not connect to the model pipeline: it records what your organization intends, not what a deployed model does in production, and it does not block unsafe outputs or monitor behavioral drift against live inference. For GPAI obligations that require continuous post-market monitoring, inference-time incident records, and audit evidence traceable to a specific model version and deployment event, you need a platform that generates that evidence from production behavior, not from policy documentation assembled separately.

How does Openlayer support the 15-day serious incident reporting window under Articles 61 and 72?

When a serious incident occurs, the reporting record requires the exact input values at inference time, the raw output with confidence scores, the model version hash, and a root cause classification. Openlayer captures all four at inference time as part of normal operation, so the investigation record exists before the 15-day window begins, not reconstructed from memory or fragmented logs after it opens. The platform also auto-generates an incident record at the moment a guardrail fires or a metric score breaches its approved threshold, which carries greater evidentiary weight than a log entry assembled after the fact.

Work on the future.

2026 Openlayer. All rights reserved.