AI Due Diligence (2026): What to Check Before You Buy
Co-founder and CEO at Peony. I built the data room platform with a background in document security, file systems, and AI. Founded Peony in 2021 in San Francisco.
Last updated: September 2026
Quick answer: AI due diligence is the buyer-side audit of an AI target across five layers — Data, Model, Infrastructure, Output, Governance — each scored 1-5 for a 5-25 AI Deal Health Score (5-10 walk, 11-17 price-chip, 18-25 healthy). The first test is whether the target is a wrapper on somebody else's hosted model or a proprietary stack. The second is its EU AI Act tier: Article 50 transparency obligations apply from 2 August 2026, and Annex III high-risk obligations are deferred to 2 December 2027.
I'm Deqian Jia, co-founder of Peony, a virtual data room platform used by 6,800+ customers across M&A, fundraising, and diligence workflows. The frames in this post come from a year of fielding questions from corp-dev associates evaluating AI-using targets, PE deal partners running platform diligence on AI-assisted services, in-house GCs scoping AI policy disclosures for buy-side, EU AI Act compliance counsel mapping target tier exposure, and AI-product founders preparing sell-side diligence packets. The question that defines this work in 2026: what specifically does an AI target's data room need to surface — training data provenance, who owns the model weights, inference economics — that a generic tech target does not, and how do you score it?
This post is the AI Due Diligence anchor — the dual-frame deep dive covering both how to audit a target that uses AI (the 5-Layer AI Target Audit) and how to use AI to run faster DD across any target. It anchors against twelve recent deal anchors and integrates the post-Omnibus EU AI Act calendar — Article 50 transparency from 2 August 2026, Annex III high-risk from 2 December 2027 — into the data room scope and gating model.
TL;DR — the 5-Layer AI Target Audit plus EU AI Act Exposure Scoring Matrix defines the playbook:
- The 5-Layer AI Target Audit: Data Layer (training data provenance, PII, scrape risk) + Model Layer (who owns the model weights — proprietary vs licensed, license-cliff terms) + Infrastructure Layer (GPU contracts, cloud concentration, BIS export-control posture) + Output Layer (FTC AI-claims exposure, indemnification posture) + Governance Layer (EU AI Act tier, NIST AI RMF, red-teaming cadence). Each layer scored 1-5. Composite AI Deal Health Score from 5 to 25. Bands: 5-10 walk, 11-17 price-chip, 18-25 healthy.
- The wrapper test: what share of production inference terminates at a third-party hosted model, what percentage of cost of goods sold is somebody else's price card, and what survives a model swap. Bessemer measured roughly 25 percent gross margins at one AI archetype against about 60 percent at another — a 35-point spread at identical revenue (BVP, 13 August 2025). A wrapper is not a bad buy — it is a margin-and-switching-risk buy, and it prices differently.
- The EU AI Act Exposure Scoring Matrix: four tiers (Unacceptable / High-Risk / Limited / Minimal) mapped onto the target's use cases. The deal-cycle relevant timestamps post-Omnibus: Article 50 transparency 2 August 2026; Annex III high-risk 2 December 2027; Annex I embedded 2 August 2028 (deferral enacted by Regulation (EU) 2026/1744, in force 27 July 2026). Article 99 fines — up to EUR 35 million or 7 percent of total worldwide annual turnover — have applied since 2 August 2025 and are group-wide after close (Regulation (EU) 2024/1689, Article 99).
- The key-person question: the 2024-2025 license-and-hire wave (Microsoft-Inflection, Amazon-Adept, Google-Character.AI, Google-Windsurf, Meta-Scale AI) is the market pricing the research team, not the code. Ask who trained the production models and whether they are still employed — and run the IP chain upstream of the buyer too, which is what killed OpenAI's reported roughly $3 billion Windsurf acquisition (TechCrunch, 11 July 2025).
- The training-data and timing risk: the license trail is the single most quantifiable exposure — the Bartz v. Anthropic settlement priced roughly $3,000 per work across about 500,000 books for a $1.5 billion total, with final approval granted 20 July 2026 (TechCrunch, 20 July 2026). There is no published benchmark for what an AI review costs or how long it takes; the Data Layer is the long pole at two to four weeks.
- The AI-for-DD playbook (short): clause extraction, contract redline summarization, Q&A routing, financial-statement reconciliation. Real lift, real audit trail. Avoid black-box "risk scores" with no source citation.
- The Peony spine: NDA gates before AI-specific subfolders are visible, visitor groups so the AI specialist counsel tier sees the model-card archive without the operations tier accessing it, page-level analytics tracking which AI-target documents reviewers spent time on (the engagement signal tells the seller where the price-chip ask will land), dynamic watermarks on every model-card and training-data manifest, and auto-indexing recognizing the typical AI-target document inventory (model cards, eval suites, red-team reports, AI risk registers).
What does "AI due diligence" actually mean in 2026?
AI due diligence in 2026 is a dual-frame discipline combining DD-of-AI (auditing an AI-using target through the 5-Layer AI Target Audit) and AI-for-DD (using LLM tooling with source-citation enforced to speed any DD workflow). The first frame is DD-of-AI — auditing a target company that builds, uses, or sells AI, with structured attention to the AI-specific risk surface that a generic tech audit does not catch. The second frame is AI-for-DD — using AI tools (LLMs, RAG systems, clause-extraction engines) to accelerate the workflow of running any due diligence, AI target or not. Most practitioners use the phrase loosely to mean one or the other. Sophisticated buyers and counsel use both.
The DD-of-AI frame is also no longer a niche discipline, because AI-component deals are now roughly half the tech tape: Bain's Software M&A Report 2026 states that "Software companies acquired a record number of AI assets in 2025, with almost half of tech deals having some AI component in 2025, up from one in four deals in 2024" (Bain, Jan 2026). Every one of those deals needs exactly this audit — and on the sell side, an AI-capability story that survives it; the buyer taxonomy behind that demand shift is in how to sell a SaaS company.
The DD-of-AI frame is the heavier of the two and the focus of this post. It exists because the AI risk surface is structurally different from the standard tech-target risk surface: training data is not source code, model licenses do not behave like software licenses, GPU capacity is not commodity cloud, hallucination is not a software bug, and the EU AI Act is not GDPR. A generic "tech DD" — code-quality review, open-source license scan, GDPR exposure, SOC 2 trail, infrastructure security — misses five categories of liability that show up only when you scope specifically for AI: pirated training corpora, third-party model terms with user-count cliffs, GPU contract concentration risk, FTC AI-claims exposure, and EU AI Act tier obligations. The 5-Layer AI Target Audit is the diligence frame that names and scores each of those five layers. The hub-canonical playbook for the broader workstream sits in M&A solutions and due diligence solutions; private-equity buyers running platform-level AI-DD on add-ons should also map this against the private equity solutions page.
The AI-for-DD frame is increasingly material because the AI-tooling for diligence has crossed a usefulness threshold around 2024-2025. Clause extraction across hundreds of customer contracts (caps, indemnities, change-of-control, MFN, exclusivity) is now a 4-to-6-hour run rather than a 4-to-6-day grind. Q&A routing on inbound buyer questions can be automated against a tagged document store. Redline summarization on amendment stacks is reliable when the source-citation requirement is enforced. The honest test, applied throughout this post: can the AI tool cite the source paragraph for every finding, and would the buyer's counsel sign off on the methodology in a deposition?
The phrase "AI due diligence" in market usage typically means one of three things: (1) the dedicated DD-of-AI workstream when the target builds or uses AI as a material part of revenue, (2) the AI-for-DD application of language models to speed any DD workflow, or (3) the regulatory and policy disclosure scope under the EU AI Act and emerging US frameworks. This post covers all three in sequence. The DD-of-AI workstream is the structural core because it is the workstream where buyer counsel signs off on actual repricing, indemnification, and walk-away decisions.
The 5-Layer AI Target Audit applies to any target where AI is material to revenue, product, operations, or risk profile. That includes the obvious cases (AI labs, AI products, AI infrastructure) and the less obvious ones (any SaaS company that has added an AI feature in the last 18 months, any consumer product with an AI assistant, any regulated business — credit, recruitment, healthcare, biometric ID — that has deployed AI in a covered use case). For targets where AI is not material, the audit reduces to a single-page screening checklist; for targets where AI is the product, it becomes the gating workstream of the entire deal.
How does the EU AI Act change M&A pricing for AI targets?
The EU AI Act changes M&A pricing for AI targets through four mechanisms — tier-driven compliance cost, deferred enforcement uncertainty, GPAI-model obligations that already entered force in August 2025, and the conformity-assessment evidence requirement for High-Risk systems. The deal-cycle relevant dates split post-Omnibus: 2 August 2026 remains the date for the Article 50 transparency requirements, but the Digital Omnibus on AI — proposed 19 November 2025 and enacted as Regulation (EU) 2026/1744, in force 27 July 2026 — deferred the Annex III high-risk obligations to 2 December 2027 and Annex I embedded-AI obligations to 2 August 2028. Deal counsel now prices high-risk conformity readiness against the deferred dates, with transparency duties live from August 2026. (The program-level half of this review — board-oversight evidence and the AI-governance questionnaire — has its own guide: AI governance due diligence.)
| Tier | Enforcement date | Example use cases | Deal-cycle action |
|---|---|---|---|
| Unacceptable | Applied 2 Feb 2025 | Social scoring, real-time biometric ID in public spaces, manipulative techniques | Walk for any EU-touching buyer |
| High-Risk (Annex III) | 2 Dec 2027 (deferred from 2 Aug 2026) | Recruitment scoring, credit scoring, biometric ID, critical infrastructure | Price-chip a readiness reserve; conformity assessment file required |
| High-Risk (Annex I) | 2 Aug 2028 (deferred from 2 Aug 2027) | AI embedded in regulated products — medical devices, machinery, vehicles | Longest runway; conformity route follows the underlying product regime |
| Limited-Risk | 2 Aug 2026 | Chatbots, synthetic-content generators | Article 50 transparency disclosure required |
| Minimal-Risk | No obligation | Internal ops AI, productivity assistants | No AI-specific repricing |
What the fines actually are, and when they started. The penalty chapter is the part most diligence memos skip, and it is the part that was never deferred. Under Article 99 of Regulation (EU) 2024/1689, non-compliance with the Article 5 prohibited practices carries administrative fines of up to EUR 35,000,000 or 7 percent of total worldwide annual turnover for the preceding financial year, whichever is higher; breaches of most operator obligations — including the Article 50 transparency duties — carry up to EUR 15,000,000 or 3 percent; and supplying incorrect, incomplete or misleading information to notified bodies or national competent authorities carries up to EUR 7,500,000 or 1 percent (Regulation (EU) 2024/1689, Article 99). For SMEs including start-ups, Article 99(6) applies the lower of the percentage or the fixed amount, not the higher. Two diligence consequences follow. First, Chapter XII — the penalties chapter — has applied since 2 August 2025 under Article 113(3)(b); it was not touched by the Omnibus. Second, the percentage is of total worldwide annual turnover, not EU revenue and not the target's revenue, so post-close the exposure is measured against the combined group. A EUR 35 million headline is the floor for a large acquirer, not the ceiling. The Digital Omnibus added new Articles 75a(3) and 75c(4), which extend Article 99(3)-(7) to AI Office enforcement — the tiers are no longer only a national-authority instrument.
Mechanism one — tier-driven compliance cost. The Act establishes four tiers. The Unacceptable tier covers prohibited practices (social scoring, real-time biometric ID in public spaces under most conditions, manipulative techniques, exploitation of vulnerabilities) and entered application on 2 February 2025. A target operating in any Unacceptable-tier use case is a walk for any EU-touching buyer. The High-Risk tier covers Annex III use cases — recruitment scoring, credit scoring, education access, law enforcement, biometric identification, critical infrastructure, employment monitoring — and carries the heaviest deal-cycle weight — its obligations now enter application on 2 December 2027, deferred from August 2026 by the Digital Omnibus, a date that still lands inside today's typical hold periods. The Limited-Risk tier covers chatbots and synthetic-content generators, with Article 50 transparency obligations entering force on 2 August 2026 as originally scheduled. The Minimal-Risk tier is everything else, with no obligations beyond voluntary codes of conduct.
Mechanism two — the deferred timeline (uncertainty resolved). The Act itself entered into force on 1 August 2024. The Digital Omnibus on AI, proposed 19 November 2025, was signed on 8 July 2026, published in the Official Journal on 24 July 2026 as Regulation (EU) 2026/1744, and entered into force on the third day following publication — 27 July 2026 (Regulation (EU) 2026/1744). It deferred Annex III high-risk obligations from 2 August 2026 to 2 December 2027 and Annex I embedded-AI obligations to 2 August 2028, citing the delayed availability of standards, common specifications and national competent authorities. Article 50 transparency obligations kept their 2 August 2026 date; the only relief is a four-month transitional under new Article 111(4), giving providers of systems placed on the market before 2 August 2026 until 2 December 2026 to comply with the Article 50(2) marking duty. New prohibitions added by the Omnibus apply from 2 December 2026. The pricing question flipped from "will the deferral pass" to "is the target on track for the deferred dates": a High-Risk target with no conformity-assessment file still carries a compliance-readiness chip sized to use-case complexity — in my own deal exposure a six-figure reserve is common, and I have not found a published benchmark that would let me put a defensible range on it — because December 2027 sits inside the typical hold period. The deferral buys runway, not absolution.
Mechanism three — GPAI model obligations already in force. The governance rules and the obligations for general-purpose AI (GPAI) models became applicable on 2 August 2025. A target that builds or distributes a GPAI model (typically anything trained on broad data and adaptable to a wide range of tasks) has been operating under those obligations since then — documentation requirements, transparency obligations, copyright-policy disclosure, training-data summary publication. The Act will apply from 2 August 2027 to GPAI models that were placed on the market before 2 August 2025; for everything placed after that date, the obligations apply immediately. The audit question: does the target maintain a model card, a training-data summary, and a copyright-compliance policy aligned with GPAI requirements?
Mechanism four — conformity assessment evidence. For High-Risk Annex III systems, the conformity assessment file is the gating evidence pack. It includes the risk-management system documentation, the technical documentation per Annex IV, the data-governance procedures, the post-market monitoring plan, the human-oversight procedures, the cybersecurity measures, and the EU declaration of conformity. A target with a complete conformity assessment file scores 4-5 on the Governance Layer of the 5-Layer Audit; a target with no conformity assessment file and an August 2026 deal-close date scores 1-2 and carries the full compliance-readiness chip.
The EU AI Act Exposure Scoring Matrix is the diligence frame that maps the target's actual use cases against the four tiers and surfaces the conformity-assessment evidence required. Buyer counsel walks the matrix in three passes: pass one inventories every use case the target's AI deploys against (typically 3-12 use cases per target); pass two classifies each use case into one of the four tiers; pass three demands the conformity-assessment evidence for any High-Risk use case and the transparency-disclosure evidence for any Limited-Risk use case. The deliverable is a one-page exposure summary signed by target counsel, with the gap-list and remediation-cost estimate attached. The matrix output then feeds the broader regulatory request list inventoried in the due diligence data room checklist.
What does the 5-Layer AI Target Audit cover?
The 5-Layer AI Target Audit is the proprietary diligence frame that decomposes any AI-using target into five auditable layers, scores each from 1 to 5, and produces a composite AI Deal Health Score from 5 to 25 with three actionable bands. The frame answers a specific buyer question: where exactly is the AI risk in this target, and what is the deal-cycle action that follows?
| Layer | What to ask | Red flags | Score 1-5 | Typical repricing |
|---|---|---|---|---|
| Data | Training data provenance, license trail, PII exposure, scrape risk | Pirated corpora (LibGen / Anna's Archive / Pirate Library Mirror), unlicensed scraped content of copyrighted works, no deletion-on-request log, PII in training corpora | 1-5 | 5-25% on serious findings |
| Model | Proprietary vs licensed, license-cliff terms, fine-tune lineage, model-risk register | Llama 700M MAU cliff approaching, no model cards, fine-tune training data is the in-licensed model's output, no model-risk register | 1-5 | 3-15% on license-cliff |
| Infrastructure | GPU contracts, capacity guarantees, cloud concentration, committed-spend obligations, BIS export-control posture | More than 70% on one hyperscaler, no committed capacity letter, take-or-pay commitments that accelerate or terminate on change of control, no disaster-recovery test, missing BIS classification record | 1-5 | 5-20% on concentration |
| Output | FTC AI-claims exposure, hallucination liability, indemnification posture, insurance rider | Aggressive marketing claims ("first robot lawyer," "replace your accountant"), no incident log, customer indemnification carve-outs missing | 1-5 | 5-15% on claims exposure |
| Governance | EU AI Act tier classification, conformity assessment file, NIST AI RMF program, red-teaming cadence | No tier classification, no conformity assessment file for High-Risk uses, no red-teaming log, no AI risk register | 1-5 | 5-25% on missing conformity |
One honest caveat on that last column. The typical-repricing bands are practitioner estimates from our own deal exposure, not a published dataset — no regulator, bank or research house publishes AI-specific repricing percentages, and anyone who presents one as market data should be asked for the sample. Use them to rank the layers against each other, not to defend a number in a negotiation.
The composite AI Deal Health Score sums the five layer scores. The bands map to deal action: 5-10 walk or restructure to asset-only purchase (the AI-specific risk is large enough that an equity acquisition exposes the buyer to indemnification claims that exceed the deal value); 11-17 price-chip negotiation with escrow or earn-out (the AI-specific risk is real but quantifiable, and the deal closes at a discount with retention mechanics protecting the buyer); 18-25 healthy AI target (no AI-specific repricing required beyond standard rep-and-warranty insurance pricing).
The 5-Layer Audit operates as a structured workstream with its own document checklist and visitor-group access pattern in the data room. The Data Layer audit takes the longest (typically 2-4 weeks for a serious target) because training-data provenance requires per-dataset traceability. The Governance Layer audit is the most time-sensitive because two EU AI Act dates intersect directly with deal-cycle pricing: Article 50 transparency duties are live from 2 August 2026, and the Article 99 fine regime has applied since 2 August 2025 regardless of the Annex III deferral to December 2027. The Output Layer audit is the most heavily lawyered because FTC consent decrees create binding precedent.
The 5-Layer Audit maps onto the data room structure through five top-level folders gated through visitor groups. The Data Layer folder holds training-data manifests, license documentation, opt-out compliance logs, and deletion-on-request records. The Model Layer folder holds model cards, fine-tune lineage documentation, eval-suite results, and model-risk registers. The Infrastructure Layer folder holds GPU contracts, capacity commitment letters, cloud-spend breakdowns, disaster-recovery test logs, and export-control classification records. The Output Layer folder holds the marketing-claims register, customer complaint logs, hallucination-incident database, and customer-contract indemnification provisions. The Governance Layer folder holds the AI risk register, conformity assessment file, post-market monitoring plan, incident response procedures, and red-team logs — best surfaced from a dedicated AI room that keeps the high-sensitivity governance artifacts isolated from the rest of the deal-room scope.
The 5-Layer Audit ties into the broader M&A diligence scope through specific seams. The Data Layer touches IP diligence (copyright posture on scraped training corpora) and privacy diligence (PII in training data, GDPR Article 17 deletion exposure). The Model Layer touches IP diligence (license terms, derivative-work restrictions) and accounting diligence (model amortization treatment, R&D capitalization). The Infrastructure Layer touches commercial diligence (vendor concentration, capacity guarantees) and operational diligence (continuity risk). The Output Layer touches commercial diligence (customer-contract indemnification) and regulatory diligence (FTC, state AG enforcement). The Governance Layer touches regulatory diligence (EU AI Act, NIST AI RMF) and compliance diligence (audit program maturity). The structural diligence anchor across the workstreams is the M&A due diligence process guide.
Is the target a real AI company or a wrapper?
You answer it with three numbers, not with the pitch deck: what share of production inference calls terminate at a third-party hosted model, what percentage of cost of goods sold is that provider's price card, and what the target still owns if the model underneath is swapped out. "Wrapper" gets used as an insult. It should be used as a pricing category. A thin application layer on a hosted foundation model can be an excellent acquisition — distribution, workflow depth and proprietary data all survive a model swap. What does not survive is a gross margin that a supplier can unilaterally reprice. So the wrapper test is not "is this real AI"; it is "who controls this company's cost of goods sold, and what happens to the model if the provider changes its terms."
The market has already moved to the application layer, and the evidence is not ambiguous. ICONIQ's 2026 State of AI bi-annual snapshot, drawing on roughly 300 executives building AI products, found that 49 percent of companies report their primary differentiation comes from application-layer innovation, compared with 14 percent relying primarily on proprietary model development, and that around 70 percent of builders are focused on vertical AI applications (ICONIQ, 2026). Application-layer is now the norm, not the exception. The diligence finding is therefore never "they use OpenAI" — it is the shape of the dependency.
| Question | Wrapper | Hybrid (fine-tune / RAG on proprietary data) | Proprietary |
|---|---|---|---|
| Where inference runs | Third-party hosted API, effectively all production calls | Split — hosted frontier calls plus self-hosted open-weights or fine-tunes | Own weights on own or reserved compute |
| What sets cost of goods sold | The provider's published price card | Blended; the self-hosted share is capex-and-utilization driven | Compute contracts and utilization, negotiated not published |
| Switching cost if the provider reprices | Prompt and eval rework, weeks — but no floor on the margin hit | Meaningful: retraining, re-evaluation, latency and quality regression testing | Low external exposure; high internal fixed cost |
| What is actually defensible | Distribution, workflow, data exhaust, brand — not the model | Proprietary training corpus, evaluation infrastructure, task-specific accuracy | Weights, research team, training pipeline, compute position |
| Evidence the buyer requests | Provider invoices, API call logs by model, contract with rate commitments | The above plus fine-tune lineage, dataset licenses, eval suite, base-model license | GPU contracts, committed-spend schedules, training-run logs, model cards |
| Repricing posture | Price the margin, not the technology; test provider-concentration risk | Price the durability of the data advantage | Price the compute obligation and the key-person risk |
How to test it in the data room. Ask for provider invoices for the last eight quarters, API call logs broken out by model and by product surface, and the model-deprecation playbook — what the engineering team did the last time a provider retired an endpoint or changed a rate limit. That last artifact is the highest-information document in this whole section, because it is the only one that shows behaviour rather than intent. Also ask how many model providers are in production: ICONIQ found builders now use around 3.1 model providers on average, up from about 2.8 six months earlier, which is the market hedging exactly this risk. A target on a single provider with no tested fallback has concentrated a repriceable cost line into one counterparty.
The margin spread is the finding, not the label. Bessemer Venture Partners' State of AI 2025 reported that its "AI Supernovas" archetype runs at only about 25 percent gross margins, "often trading distribution for profit in the short term," while its "Shooting Stars" archetype runs at roughly 60 percent (BVP, 13 August 2025). Two AI companies at identical revenue can sit 35 gross-margin points apart. ICONIQ's surveyed companies expect gross margins to reach around 52 percent on average in 2026 — note the tense, that is a forward expectation from operators, not a realized market average. The framing itself is older than the current cycle: a16z's Martin Casado and Matt Bornstein argued in February 2020 that AI businesses run gross margins "often in the 50-60 percent range — well below the 60-80 percent-plus benchmark for comparable SaaS businesses." Cite that with its 2020 date; it is the origin of the argument, not evidence about 2026.
How do you test an AI target's gross margin and inference costs?
Recompute gross margin with inference in cost of goods sold, by cohort, and then ask who controls the inputs to that line. The single most common restatement in AI diligence is that the target books model spend in research and development rather than in cost of goods sold, which flatters gross margin by anything from a few points to a category. The buyer's finance team rebuilds the bridge from reported gross margin to AI-inclusive gross margin before any multiple is applied.
| Line | What to pull | Red flag |
|---|---|---|
| Reported vs AI-inclusive gross margin | Trial balance mapping for model spend; the R&D-to-COGS reclassification schedule | Inference sits in R&D and management resists reclassifying it |
| Inference cost per active account | Provider invoices divided by monthly active accounts, by cohort and by plan | Cost per account rising faster than price per account |
| Cost per query by model tier | API logs tagged by model and by product surface | Frontier-tier models serving free-tier or trial traffic with no routing policy |
| Third-party model spend as share of COGS | Vendor concentration schedule inside COGS | One provider above half of COGS with no tested fallback |
| Unit-cost trend vs token-price deflation | Realized cost-per-unit-of-work trend, not list-price trend | The plan assumes the list price falls and the workload stays constant |
| Committed-spend amortization | Cloud and GPU minimum-commit schedules, take-or-pay terms, unused-commitment balances | Committed spend expensed as incurred while the obligation runs years past the model |
Do not let the target underwrite its margin plan on token deflation. The deflation is real and it is dramatic: a16z's Guido Appenzeller measured that for an LLM of equivalent performance the cost of inference is falling roughly 10x every year, a factor of about 1,000 over three years — GPT-3-level quality moved from about $60 per million tokens in November 2021 to about $0.06 per million tokens (a16z, 12 November 2024). But the decline is wildly uneven. Epoch AI's analysis of six benchmarks found the rate of decline ranges from 9x to 900x per year depending on the capability milestone — GPT-4-level performance on GPQA Diamond fell from $37.50 to $0.12 per million tokens between March 2023 and December 2024, roughly 40x per year (Epoch AI, 12 March 2025). Epoch's own caveat is the line that belongs in the diligence memo: the fastest price drops in that range occurred in the most recent year, so it is less clear that those will persist. A target serving frontier-capability workloads is at the 9x end of that range, not the 900x end. If the 2027 margin plan requires the 900x end, the plan is underwriting a trend the people who measured it decline to extrapolate.
Size the market context so the concentration finding lands. Menlo Ventures put enterprise generative-AI spend at $37 billion in 2025, 3.2 times the $11.5 billion of 2024, with $12.5 billion of that flowing to foundation-model APIs, and enterprise LLM spend split Anthropic 40 percent, OpenAI 27 percent, Google 21 percent (Menlo Ventures, 9 December 2025). Three suppliers hold most of the API layer. That is the structural reason provider-concentration risk is not a theoretical exercise: there is no long tail to fall back on. The IT-side view of the same cost line — token usage pulls, growth stress-tests, AI cost trajectory — sits in IT due diligence; this section is the canonical margin treatment.
What documents does an AI target have to produce?
One consolidated request list, keyed to the five layers, with a named reviewer per line — that is the artifact that turns the audit into a data room. Scattering the requests across five workstream emails is how confirmatory diligence slips two weeks. The table below is the layer-level request list; the full 174-document cross-workstream companion is the due diligence data room checklist. The same request list is available as a downloadable AI due diligence checklist — tick the lines you have received, filter by layer, disclosure wave or reviewer, and export the outstanding requests to CSV, Markdown or PDF.
| Layer | Documents and artifacts requested | Who reviews it | Disclosure wave |
|---|---|---|---|
| Data | Per-dataset training-data manifest; dataset licenses with effective dates; acquisition receipts and purchase trail; dataset cards; PII inventory and DPIA; opt-out, DMCA and GDPR Article 17 deletion logs; shadow-library screening record | Buyer's data-protection counsel + IP counsel | Confirmatory |
| Model | Model inventory (foundation, fine-tune, embedded open-weights, API-accessed, internal); base-model licenses; fine-tune lineage; model-card archive; eval-suite results and held-out test design; model-risk-management register | Buyer's AI/ML advisor + IP counsel | LOI (inventory) → confirmatory (lineage) |
| Infrastructure | GPU reservation contracts; capacity commitment letters; cloud minimum-commit and take-or-pay schedules with unused-commitment balances; change-of-control clauses in compute agreements; cloud-spend breakdown; DR test log; BIS classification record | Buyer's AI/ML advisor + commercial counsel + lender | LOI (spend) → confirmatory (contracts) |
| Output | Marketing-claims register with substantiation files; customer-complaint log; hallucination-incident database; customer-contract indemnification index; insurance policies and AI riders | Buyer's regulatory counsel + commercial counsel | Confirmatory |
| Governance | EU AI Act tier classification per use case; conformity assessment file and Annex IV technical documentation; post-market monitoring plan; NIST AI RMF program documentation; ISO/IEC 42001 certificate and scope statement if claimed; red-team log; AI risk register | Buyer's regulatory counsel + EU AI Act specialist | LOI (tier) → confirmatory (file) |
Two lines in that table are the ones sellers are least prepared for. The acquisition receipts in the Data Layer are not the same document as the dataset license — the Bartz holding turned on how the copies were obtained, not on what the license said afterwards. And the unused-commitment balance in the Infrastructure Layer is an obligation the buyer inherits that frequently appears nowhere in the reported cost of goods sold.
How do you audit the Data Layer — training provenance, PII, scrape risk?
The Data Layer audit is the heaviest of the five and the one where 2024-2026 case law has materially repriced risk. The 5-Layer Audit gives Data Layer a score from 1 to 5 based on five diagnostic tests — training-data provenance, license documentation, scraped-content posture, PII inventory, and opt-out / deletion-on-request compliance — with each test scored Pass / Yellow / Red and the layer score derived from the worst-case combination.
Diagnostic test one — training-data provenance. Every dataset used to train a target's proprietary model (foundation or fine-tune) needs documented provenance: source name, acquisition date, license terms with effective date, intended-use scope, and the licensee-of-record. The buyer's data-counsel team reads the training-data manifest and traces each dataset to its origin. The acceptable evidence stack: per-dataset license agreement, dataset-card documentation, source-of-truth identifier, and the deletion log for any dataset that has been removed from training (typically post-DMCA, post-GDPR Article 17, or post-customer opt-out). The unacceptable evidence stack: a single line item reading "publicly available web data" with no further documentation, scraped content from sites with terms-of-service prohibiting scraping, or any dataset acquired from a torrent-tracker or shadow-library source.
Diagnostic test two — license documentation. Every licensed dataset needs license documentation effective on the training run date. Commercial-use clauses, derivative-works clauses, attribution requirements, and effective-date language all need explicit review. The buyer's data-counsel team checks for license-mismatch (training was completed under a license that has since been revoked or downgraded) and for attribution-gap (the target's product owes attribution to a dataset source that was never credited). The exposure is quantifiable: the Anthropic Bartz settlement of $1.5 billion — preliminary approval from Judge William Alsup on 25 September 2025, final approval granted on 20 July 2026 by Judge Araceli Martínez-Olguín after Alsup retired — priced the per-work damages at approximately $3,000 per book for an estimated 500,000 books, establishing a baseline-per-work damages number that buyer counsel now multiplies against the target's likely-pirated training-corpus volume.
Diagnostic test three — scraped-content posture. The buyer's data-counsel team asks the binary question: did the target's training corpora include any content scraped from shadow libraries (Library Genesis, Anna's Archive, Pirate Library Mirror, Z-Library)? The answer comes from the training-data manifest, from engineering-team interviews, from the data-acquisition-purchase trail, and from internal Slack archives if the target opens them. The Bartz court ruled in June 2025 that downloading from shadow libraries was not protected fair use even where the downstream training use might be — the piracy itself was the actionable claim. A target with documented shadow-library training-corpus content scores 1 (Red) on this test regardless of post-detection cure unless the cure includes substantial settlement payment, retraining from clean corpora, and customer notification.
Diagnostic test four — PII inventory. Any PII in training corpora creates GDPR Article 17 (right to erasure) exposure, CCPA exposure, and state-law biometric privacy exposure. The buyer's privacy-counsel team requires a PII inventory across all training datasets, the data-protection impact assessment (DPIA) for any dataset containing PII, and the deletion-on-request workflow showing how customer or data-subject deletion requests propagate from production data through to training corpora and any fine-tuned model derived from them. In the data room, the PII inventory itself is staged behind NDA gates at the model-card folder and surfaced through redaction for any identifier-level fields the seller cannot share without further legal review. The technical reality: deleting from training corpora does not delete from a trained model — but the buyer requires documented evidence of the model-retraining or unlearning approach taken to honor deletion requests. A target with no PII inventory and no documented deletion-on-request workflow scores 1-2.
Diagnostic test five — opt-out and deletion compliance. Beyond GDPR Article 17, the buyer's data-counsel team reviews the target's response posture to publisher opt-outs (the NYT case discovery order of January 2026, where US District Judge Sidney Stein affirmed compelling OpenAI to produce a 20-million-log sample of anonymized ChatGPT conversations, set the precedent that AI companies must preserve user logs against future copyright discovery). The target needs documented opt-out compliance for any website using the robots.txt AI-crawler directives, any DMCA notice received, and any GDPR Article 17 request received. The buyer also reviews the user-log preservation policy because litigation-preservation obligations now flow from preservation orders, not just from internal retention schedules.
The Data Layer score derives from the worst-case test result. A target with documented per-dataset provenance, clean license trails, no shadow-library content, complete PII inventory, and documented opt-out compliance scores 5 (Pass-Pass-Pass-Pass-Pass). A target with one Yellow (license-mismatch on one dataset, partial PII inventory) and four Pass scores 3-4. A target with one Red (shadow-library content, no PII inventory) scores 1-2 regardless of how clean the other four tests are. The proprietary insight: shadow-library content posture is the highest-information-content test because it produces the most quantifiable downside risk per the Bartz settlement precedent.
The Data Layer feeds the broader IP-diligence and privacy-diligence workstreams covered in the M&A due diligence process guide and the vendor due diligence checklist. The deal-cycle handoff: data-counsel completes the Data Layer audit, surfaces findings into the data-room Q&A, and the IP and privacy workstreams use the Data Layer findings to size their own indemnification and rep-and-warranty asks.
Do the customer contracts actually let the target train on that data?
Read the customer contracts, not just the vendor contracts — because the data asset you are paying for may be contractually unusable after close. Most AI diligence reads the target's inbound agreements (which models it licenses, which datasets it bought) and stops. The buy-side question that gets skipped is the outbound one: the target's product accumulates customer data, the valuation model treats that accumulation as a compounding moat, and somewhere in the customer contract set is a population of agreements that say the target may not train on any of it. Every one of those contracts is a hole in the asset.
There is no credible published statistic for how common each clause type is. I looked — every "X percent of enterprise contracts contain a no-training clause" figure I could surface traces back to vendor content marketing with no sample, no instrument and no methodology. So do not benchmark this; count it. The count is produced by clause extraction across the full customer contract set with a citation back to the source paragraph for every hit, which is exactly the workload AI extraction is for, and it is a number the seller can produce in days rather than weeks.
| Clause type | What it says | Effect on the data asset post-close | Where it shows up |
|---|---|---|---|
| Express training grant | Customer grants a right to use its data to improve the service, sometimes including model training | Cleanest case; check whether the grant survives assignment and whether it names "model training" explicitly | MSA, order form, product terms |
| Silence | No clause either way | Jurisdiction- and regulator-dependent, and the weakest position to defend to an acquirer's counsel | Older MSAs, self-serve click-through terms |
| No-training / no-derivative | Expressly prohibits using customer data or outputs to train or improve any model | The data is unusable as a training asset; and if a model was already trained on it, that is a live exposure | Enterprise-negotiated MSAs, security riders |
| DPA-restricted (processor-only) | Target processes only on documented instructions; training is not a documented instruction | Training is out of scope even where the MSA is silent | Data processing agreement, Annex/Schedule |
| Deletion on termination | Requires deletion or return of customer data within a fixed window | Deleting from a corpus does not delete from a trained model — ask what the unlearning or retraining plan is | MSA termination article, DPA |
| Change-of-control consent | Requires customer consent to assignment, or grants a termination right on change of control | Can strip the data right and the revenue in the same clause; count these separately | MSA assignment article |
Then check the chain of promises upstream. The finding worth the most in this section is a target that promises its customers "we never train on your data" while routing production traffic to a model provider tier where that is not the default. The promise is only as good as the weakest link between the customer, the target, its sub-processors, and the upstream model provider — and that chain is checkable in diligence from public, dated provider terms without any survey number. Note that the vendor-side mirror of this question, the AI clause inventory for the tools the target buys, belongs to the program-level review in AI governance due diligence. One further distinction worth keeping straight: a data room vendor's own no-training commitment — the promise that your diligence documents are never used to train a model — is a different question entirely, covered in AI in the data room.
How do you audit the Model Layer — ownership, license terms?
The Model Layer audit traces every model in the target's stack to its license posture, identifies any license-cliff trigger, and scores the model-risk-management discipline. Five diagnostic tests anchor the audit: model-inventory completeness, license-term clarity, fine-tune lineage, model-card archive, and model-risk-management register.
Diagnostic test one — model-inventory completeness. The buyer's tech-counsel team requires a complete model inventory across the target's stack: every foundation model, every fine-tune, every open-source model embedded in product, every third-party model accessed via API, and every model used internally for ops, sales, or marketing. For each model: model name, license terms, version, training-data summary, intended-use scope, and the in-production indicator. A target with a complete model inventory scores 4-5 on this test; a target where the engineering team can identify only the top-three models out of a stack of ten scores 1-2.
Diagnostic test two — license-term clarity. Each licensed model needs a clean license-term review. The two most material 2025-2026 license-cliff terms: (1) the 700-million-monthly-active-user threshold that appears at section 2 of every Meta Llama Community License — Llama 2 (18 July 2023), Llama 3 (18 April 2024), Llama 3.1 (23 July 2024) and Llama 4 (5 April 2025) — under which a licensee above 700 million monthly active users must request a separate license from Meta, which Meta may grant at its sole discretion; and (2) the naming and attribution conditions at sections 1.b.i and 1.b.iii, which bind any model improved with Llama Materials or their outputs and any redistribution, regardless of MAU. The Mistral commercial terms need parallel review. The Anthropic AUP needs review for any enterprise customer of the target using Anthropic's models in regulated use cases. A target with documented license-term compliance and a contingency plan for license-cliff scenarios scores 5; a target with an undisclosed cliff approaching scores 1-2.
Two precision points that counsel, not an engineer, should resolve. First, the threshold is measured "on the [model] version release date" — a fixed historical date in the licence text, not a rolling monthly test, and the practical reading of that mechanic is a legal question rather than a diligence assumption. Second, the clause reaches "Licensee or Licensee's affiliates," so post-close the acquirer's user base enters the test for any new Llama version the target adopts. A sub-scale target inside a buyer already above 700 million monthly active users can lose the right to take the next Llama release even though it was comfortably compliant standalone. Ask which Llama versions and weights are in the stack, whether they are fine-tuned or merged, and whether Llama sits in the inference path of the revenue-generating product. And do not write "Llama is open source" in the diligence memo — Meta's instrument is a Community License with a use-based commercial trigger and naming conditions; the accurate phrase is open-weights under a community licence, and the difference is the whole finding.
Diagnostic test three — fine-tune lineage. Every fine-tuned model in the target's product stack needs lineage documentation: what base model was fine-tuned, what data was used to fine-tune, what license attaches to the fine-tune output, and what derivative-work restrictions apply. The restriction is version-specific and most memos get it backwards. The Llama 2 and Llama 3 licences prohibited using Llama Materials or outputs to improve any other large language model (excluding Llama and its derivatives). Llama 3.1 and Llama 4 dropped that prohibition and replaced it with a naming condition at section 1.b.i: any model created, trained, fine-tuned or improved using Llama Materials or their outputs must carry "Llama" at the start of its name. So the question is not "did they train on Llama output" but "which licence version governs that run, and does the resulting model name comply." The audit asks: is every fine-tune in the stack training-data-traceable to a license-compliant source? A target with documented fine-tune lineage scores 5; a target with engineering team admission of cross-model training scores 1-2.
Diagnostic test four — model-card archive. Each production model needs a model card per the NIST AI RMF July 2024 Generative AI Profile (NIST.AI.600-1) and per the EU AI Act GPAI obligations effective 2 August 2025. The model card documents intended use, performance metrics, bias evaluation, safety evaluation, training-data summary, environmental footprint estimate, and limitations. The model-card archive sits in a data room with page-level analytics so the seller can see which model cards specialist counsel actually reads end-to-end versus which they skim. A target with model cards for every production model scores 5; a target with model cards only for the flagship product scores 3; a target with no model cards scores 1-2.
Diagnostic test five — model-risk-management register. Financial-services targets, healthcare targets, and any target deploying AI in regulated use cases need a model-risk-management (MRM) program aligned with the OCC 2011-12 framework or equivalent. The MRM register documents every model, its inherent risk tier, its validation schedule, its monitoring metrics, and its retirement criteria. For non-regulated targets, the equivalent is the NIST AI RMF-aligned risk register. A target with an MRM register reviewed quarterly scores 5; a target with a static one-time risk register scores 3; a target with no register scores 1-2.
The Model Layer score derives from the worst-case test result. The license-cliff test is the most material because the consequence is binary — either the target's product is in license compliance, or the target is operating in license violation with potential injunctive remedies. The proprietary insight: the Llama license-cliff at 700M MAU is the single most underappreciated license-cliff term in 2026 because the threshold is non-trivial to cross but non-trivial to monitor — many fast-scaling targets are not actively monitoring their MAU against the Llama threshold and only discover the exposure in diligence. The cliff documentation handoff feeds the buyer-side Q&A workflow described in the due diligence questionnaire guide.
The Model Layer feeds into the broader tech-DD and IP-DD workstreams covered in the startup due diligence guide (for founder-stage targets where the AI is the product) and the M&A due diligence process guide (for the hub-canonical workstream relationships). The deal-cycle handoff: tech-counsel completes the Model Layer audit, surfaces license-cliff findings into the data-room Q&A, and the IP workstream uses the Model Layer findings to size license-cliff indemnification asks.
How do you audit the Infrastructure Layer — GPU contracts, vendor concentration?
The Infrastructure Layer audit covers the compute supply chain the target depends on: GPU contracts, cloud provider concentration, capacity guarantees, disaster recovery posture, and BIS export-control compliance. Six diagnostic tests anchor the audit: GPU contract terms, capacity commitment letter, cloud concentration ratio, disaster recovery test log, export-control classification record, and committed-spend obligations.
Diagnostic test one — GPU contract terms. Targets running their own model training or large-scale inference workloads need committed-capacity GPU contracts. The buyer's tech-counsel team reads each GPU contract for: committed capacity (number of GPUs and hours), allocation guarantees (priority versus best-effort), force-majeure clauses, termination triggers, pricing structure (fixed versus indexed), and renewal mechanics. A common 2025-2026 finding: targets running on hyperscaler-hosted GPU capacity have only best-effort allocation with no committed-capacity SLA, which becomes a continuity risk in capacity-constrained quarters. A target with documented committed-capacity contracts scoring well above peak demand requirement scores 5; a target running on best-effort allocation only scores 1-2.
Diagnostic test two — capacity commitment letter. Beyond the contract terms, the buyer's tech-counsel team requires a capacity commitment letter from each GPU provider summarizing the committed-capacity scope and any near-term capacity expansion plans. The commitment letter is the buyer's evidence that the target's growth model is compute-supportable through the deal-close horizon and the first 12-18 months of post-close operations. A target with capacity commitment letters covering the next 18 months scores 5; a target with no commitment letter scores 1-2.
Diagnostic test three — cloud concentration ratio. The buyer's tech-counsel team computes the cloud concentration ratio: what percentage of the target's compute spend runs on a single hyperscaler? More than 70 percent on one provider is our own working rule of thumb rather than a published market threshold — I use it because above roughly that line the renewal negotiation stops being a negotiation. The market-structure reason concentration matters is that there is no long tail to fall back on: Menlo Ventures put enterprise LLM spend at Anthropic 40 percent, OpenAI 27 percent and Google 21 percent in 2025 (Menlo Ventures, 9 December 2025), and ICONIQ found builders running around 3.1 model providers on average precisely to hedge it. A target with cloud concentration below 50 percent on any single provider scores 5; above 70 percent scores 2-3; above 90 percent scores 1.
Diagnostic test six — committed-spend obligations. The Infrastructure Layer is the one layer where the buyer inherits a liability that does not appear in the reported cost of goods sold. Minimum-commit and take-or-pay cloud and GPU agreements are off-P&L obligations in substance: the target has promised a supplier a spend floor for a fixed term, and if usage falls short the shortfall is still owed. The buyer's team pulls every committed-spend schedule, the remaining term on each, the unused-commitment balance to date, the true-up mechanics, and — the question that gets asked last and matters most — what happens to the commitment on change of control. Some agreements accelerate, some terminate with a penalty, some are silent and therefore negotiable. Model the unused balance as a deal liability, not as a run-rate cost. A target that discloses committed-spend schedules with change-of-control terms and a utilization history scores 5; a target with a multi-year take-or-pay commitment sized to a growth plan it has already missed scores 1-2.
Diagnostic test four — disaster recovery test log. Targets running production AI workloads need documented disaster-recovery (DR) testing on a defined cadence. The DR test log documents the test scope, the date, the recovery time objective (RTO) achieved, the recovery point objective (RPO) achieved, and any gaps identified. A target with quarterly DR tests showing achievable RTO under 4 hours scores 5; a target with annual DR tests scores 3; a target with no documented DR testing scores 1-2.
Diagnostic test five — BIS export-control classification record. Targets serving international customers need a BIS Export Administration Regulations (EAR) classification record for any AI-related technology or model exported. The January 2026 BIS final rule revising the license-review posture for NVIDIA H200- and AMD MI325X-equivalent chips from "presumption of denial" to "case-by-case review" effective 15 January 2026 changed the compliance calculus. The January 2026 25-percent tariff on covered products (advanced AI chips not destined for the US supply chain) added a parallel duty consideration. BIS stayed the Affiliates Rule on November 10, 2025 for one year (final rule, stay; 90 FR, November 12, 2025), removing one compliance lever but leaving others in place. The stay ends November 9, 2026, and the 50 percent ownership provisions return to the Export Administration Regulations on November 10, 2026 unless BIS extends the suspension, so a deal closing after that date should be diligenced as if the rule is live. The export-control file lives in a permissions-controlled data room with link expiry so outside trade counsel get time-boxed access only. A target with documented BIS classifications, export licenses where required, and a screening program for foreign-affiliate destinations scores 5; a target with no BIS classification record scores 1-2.
The Infrastructure Layer score derives from the worst-case test result, weighted toward the cloud-concentration and GPU-contract tests because those create the most immediate continuity risk. The proprietary insight: GPU-contract concentration is the single most common Infrastructure Layer finding in 2025-2026 because hyperscaler-hosted AI workloads default to best-effort allocation unless the customer specifically negotiates committed-capacity terms — and most early-to-mid-stage AI targets did not negotiate committed-capacity terms during their initial cloud agreement.
The Infrastructure Layer feeds into the broader operational-diligence and commercial-diligence workstreams covered in the vendor due diligence checklist and the due diligence cost breakdown. The deal-cycle handoff: tech-counsel completes the Infrastructure Layer audit, surfaces concentration findings into the data-room Q&A, and the operational workstream uses the Infrastructure Layer findings to size business-continuity indemnification asks.
How do you audit the Output Layer — hallucination, FTC AI-claims exposure?
The Output Layer audit prices the marketing-claim and customer-deliverable risk arising from AI output behavior. Five diagnostic tests anchor the audit: marketing-claim posture, customer-complaint log, hallucination-incident database, customer-contract indemnification, and insurance carrier AI-rider terms.
Diagnostic test one — marketing-claim posture. The buyer's regulatory-counsel team reviews every public marketing claim about AI capability, accuracy, or substitutability. The FTC Operation AI Comply enforcement sweep launched 25 September 2024 cracked down on deceptive AI claims and schemes. The DoNotPay final order announced 11 February 2025 required $193,000 in monetary relief, prohibits DoNotPay from advertising that its service performs like a real lawyer unless it has sufficient evidence to back it up, and requires notice to past subscribers between 2021 and 2023 — DoNotPay's claim of being "the world's first robot lawyer" was the trigger. The FTC Rite Aid case (December 2023) banned facial-recognition use for security or surveillance for five years after finding the technology produced more false-positive results for Black and Latino customers. A target with conservative marketing claims, signed-off by counsel, with substantiation files for each claim scores 5; a target with aggressive marketing claims and no substantiation files scores 1-2.
Diagnostic test two — customer-complaint log. The buyer's regulatory-counsel team reviews the target's complaint log for any complaints alleging AI inaccuracy, AI bias, AI harm, or AI substitution failure. The log inventory: source of complaint (customer, regulator, media), date, complaint substance, target's response, and resolution. A target with a documented complaint log showing low volume and prompt resolution scores 5; a target with a complaint log showing pattern-of-issue or recurring inaccuracy complaints scores 2-3; a target with no complaint log scores 1.
Diagnostic test three — hallucination-incident database. Production GenAI deployments need a hallucination-incident database documenting each material false-output incident: customer impact, root-cause analysis, mitigation deployed, and recurrence indicator. The database is the operational evidence pack that the target takes hallucination risk seriously. When the seller surfaces the hallucination log into the data room, dynamic watermarks burned into every page deter casual leakage to journalists or short-sellers. A target with a documented hallucination-incident database and quarterly review scores 5; a target with no formal hallucination tracking scores 1-2.
Diagnostic test four — customer-contract indemnification. The buyer's commercial-counsel team reads every material customer contract for AI-specific indemnification provisions. Does the target indemnify customers for AI-output errors? Does the target carve out AI-output from liability caps? Does the target preserve the right to inject disclaimers into AI-output? The 2025-2026 baseline: enterprise customers increasingly require AI-output indemnification, but mature targets carve out indemnification specifically for hallucinations attributable to documented model limitations. A target with consistent AI-output indemnification posture across customer contracts scores 5; a target with inconsistent posture or open-ended indemnification exposure scores 2-3.
Diagnostic test five — insurance carrier AI rider. The target's commercial general liability, errors-and-omissions, and cyber insurance policies need AI riders. The 2024-2026 development: insurers added AI-specific exclusions and AI-specific riders, and the rider terms vary widely. A target with documented AI riders on all material policies and with the rider scope reviewed by counsel scores 5; a target with no AI rider scores 1-2.
The Output Layer score derives from the worst-case test result. The marketing-claim posture test is the most material because it determines the FTC and state-AG enforcement exposure surface. The proprietary insight: the FTC Operation AI Comply enforcement pattern through 2024-2026 (DoNotPay, Rite Aid, and the September 2024 Operation AI Comply sweep) plus the FTC's December 2025 Rytr final order set aside under the Trump Administration's AI Action Plan show that the enforcement target shifted from broad AI claims toward specific deceptive-claim patterns — but the underlying risk surface remains the marketing-claim register and the substantiation files.
The Output Layer feeds into the broader regulatory-diligence workstream covered in the hub-canonical M&A diligence playbook and the commercial workstream. The deal-cycle handoff: regulatory-counsel completes the Output Layer audit, surfaces marketing-claim findings into the data-room Q&A, and the commercial workstream uses the Output Layer findings to size customer-indemnification asks.
How do you audit the Governance Layer — EU AI Act tier, NIST AI RMF, red-teaming?
The Governance Layer audit confirms the target's compliance posture under the EU AI Act and the NIST AI RMF, and reviews the red-teaming cadence and AI risk register. Five diagnostic tests anchor the audit: EU AI Act tier classification, conformity assessment file, NIST AI RMF program, red-team log, and AI risk register.
Diagnostic test one — EU AI Act tier classification. The buyer's regulatory-counsel team requires the target's EU AI Act tier classification across every use case (typically 3-12 per target). The classification document maps each use case to one of four tiers: Unacceptable (prohibited, applied 2 February 2025), High-Risk (Annex III, applies 2 December 2027 — deferred from August 2026 by Regulation (EU) 2026/1744), Limited-Risk (Article 50 transparency, applies 2 August 2026), or Minimal-Risk (no obligations). A target with a complete tier classification reviewed by EU counsel scores 5; a target with no tier classification scores 1-2.
Diagnostic test two — conformity assessment file. For any use case classified as High-Risk, the conformity assessment file is the gating evidence pack required by 2 December 2027 (the Digital Omnibus deferral, enacted July 2026; formerly 2 August 2026). The file includes the risk-management system documentation, technical documentation per Annex IV, data-governance procedures, post-market monitoring plan, human-oversight procedures, cybersecurity measures, and EU declaration of conformity. A target with a complete conformity assessment file for every High-Risk use case scores 5; a target with a partial file scores 2-3; a target with no file and a close date approaching the December 2027 deadline scores 1.
Diagnostic test three — NIST AI RMF program and the ISO/IEC 42001 question. The buyer's regulatory-counsel team reviews the target's alignment to the NIST AI Risk Management Framework 1.0 (NIST AI 100-1, released 26 January 2023) plus the Generative AI Profile (NIST AI 600-1, released 26 July 2024) (NIST). The four functions of the AI RMF — Govern, Map, Measure, Manage — each need documented procedures, owners, and review cadence. The GenAI Profile adds focus on confabulation, data privacy, harmful bias, and misuse. A target with documented AI RMF program alignment scores 5; a target with selective alignment scores 3; a target with no AI RMF program scores 1-2.
The asymmetry between the two dominant standards is itself the diligence ask, and it is the thing most memos get backwards. ISO/IEC 42001:2023 — "Information technology — Artificial intelligence — Management system," published December 2023 by ISO/IEC JTC 1/SC 42 — specifies requirements for establishing, implementing, maintaining and continually improving an AI management system, and it is certifiable: a target holding a certificate has a scope statement and a named certification body the buyer can call. The NIST AI RMF is voluntary and non-certifiable, so "we follow the NIST AI RMF" is an unauditable claim until someone reads the procedures. Two traps follow. Write the standard as ISO/IEC 42001, not "ISO 42001." And read the certificate's scope statement rather than its existence — a management-system certificate is not a conformity assessment under the EU AI Act and does not substitute for one; the program-level treatment of that distinction sits in AI governance due diligence. A target claiming "alignment with" 42001 without a certificate has told you something different from what you heard.
Diagnostic test four — red-team log. Production GenAI deployments need a red-team log documenting cadence (quarterly minimum for production GenAI, monthly for safety-critical deployments), test scope, findings, and remediation. The red-team test scope covers prompt injection, jailbreak resistance, data exfiltration, bias evaluation, harmful-content generation, and use-case-specific safety tests. A target with quarterly red-teaming, documented findings, and remediation tracking scores 5; a target with annual red-teaming scores 3; a target with no red-teaming scores 1-2.
Diagnostic test five — AI risk register. The AI risk register is the operational artifact that ties all four prior tests together. It documents every AI risk the target has identified, the risk owner, the mitigation in place, the residual risk score, and the review cadence. A target with an AI risk register reviewed quarterly by the executive AI committee scores 5; a target with a static risk register scores 3; a target with no AI risk register scores 1-2.
The Governance Layer score derives from the worst-case test result, weighted toward the EU AI Act tier and conformity-assessment tests because those drive direct compliance exposure. The proprietary insight: the absence of an AI risk register and a documented red-teaming cadence is the single most common Tier 1 Governance Layer finding in 2025-2026 — many AI-using targets have NIST AI RMF familiarity at the engineering level but have not formalized the governance artifacts at the executive level. Buyer counsel requires the executive-level artifacts.
The Governance Layer feeds into the broader regulatory-diligence workstream covered in the end-to-end M&A diligence sequence and the compliance workstream. The deal-cycle handoff: regulatory-counsel completes the Governance Layer audit, surfaces AI Act and AI RMF findings into the data-room Q&A, and the compliance workstream uses the Governance Layer findings to size remediation-readiness chips.
What the OECD Due Diligence Guidance for Responsible AI adds
When we diligence an AI target, the governance question is never whether someone wrote an AI policy. It is whether a process exists that produces artifacts. The OECD Due Diligence Guidance for Responsible AI, published February 19, 2026, is the cleanest benchmark I have found, because it takes the six-step risk-based framework from the OECD guidelines for multinational enterprises (embed, identify and assess, cease or prevent or mitigate, track, communicate, remediate) and opens every step with a table mapping it to the frameworks a target already claims to follow, including the NIST AI Risk Management Framework and ISO/IEC 42001. That crosswalk is the useful part. It turns a vague responsible-AI claim into a specific list of documents you request in the data room, and the gaps show up fast. An AI target almost always sits in the guidance's second group, enterprises active in the AI system lifecycle, so these three requests cover most of it:
- Step 1, embed. The document assigning AI due diligence oversight to named senior management and the board, plus the written risk-tolerance thresholds defining low, medium, and high severity. A policy without thresholds is a brochure.
- Steps 2 and 4, identify and track. The testing, evaluation, verification and validation record: documented test sets, metrics, tooling, and measured performance improvements or declines across releases. Ask for the trend, not a snapshot.
- Steps 5 and 6, communicate and remediate. The incident log covering significant incidents, errors, and attempts to circumvent safeguards, plus the channel through which affected stakeholders seek remedy.
How much of the value walks out with the research team?
In an AI target the research team frequently is the asset, and the market has priced it that way in public: a run of 2024-2025 transactions bought the people and licensed the technology, deliberately, instead of buying the company. That structure is the strongest available evidence about where AI value actually sits. Five examples, all of them license-and-hire rather than acquisitions — in each case the licence was non-exclusive and the target continued to exist as an independent company:
- Microsoft and Inflection AI (21 March 2024). A reported $650 million, split by reporting into roughly $620 million to licence Inflection's models non-exclusively for Azure and roughly $30 million to waive legal claims over the hiring. Co-founders Mustafa Suleyman and Karén Simonyan and most of a roughly 70-person staff moved to Microsoft. Microsoft bought no assets.
- Amazon and Adept (28 June 2024). Amazon hired co-founder and CEO David Luan and a small group of colleagues into its AGI team and took a non-exclusive licence to Adept's agent technology, multimodal models and some datasets. No price was disclosed — the only figure in circulation is a reported ~$25 million, never confirmed. Adept continued independently.
- Google and Character.AI (announced 2 August 2024). A reported $2.7 billion for a non-exclusive licence plus the return of co-founders Noam Shazeer and Daniel De Freitas and part of the research team to Google DeepMind; the general counsel became interim CEO.
- Meta and Scale AI (announced 12 June 2025). $14.3 billion for a 49 percent stake with expressly no voting power, per Scale AI's own spokesperson. Founder Alexandr Wang left to run Meta's superintelligence effort and stayed on Scale's board; Scale continues as an independent company. This is not an acquisition and Meta does not control it.
- Google and Windsurf (announced 11 July 2025). A reported $2.4 billion in licensing fees and compensation for a non-exclusive licence plus the hire of CEO Varun Mohan, co-founder Douglas Chen and senior researchers into Google DeepMind. Google explicitly did not invest in Windsurf, and Windsurf remained free to license its technology to others; most of the team stayed put (CNBC, 11 July 2025).
Note the direction of the evidence. Buyers repeatedly paid nine and ten figures for a small number of named researchers and a non-exclusive right to technology the seller kept. If that is what the team is worth to a strategic acquirer, it is also what walks out of your target on a resignation letter. And the structure is not a regulatory safe harbour: Amazon's Adept deal drew FTC scrutiny in July 2024 and the Google-Character.AI agreement drew DOJ attention over whether it was built to avoid merger review. A licence-and-hire is not an escape from review — it is a different review.
The key-person test, in six questions. Who specifically trained each production model, and are they still employed today? How many people currently at the company could retrain the flagship model from scratch — if the honest answer is one, that is the finding. Is the training pipeline documented to the point that a new hire could reproduce a run, or is it tribal knowledge in a notebook? What are the vesting cliffs, and where do they land relative to the expected close? Are the non-solicits and non-competes enforceable in the jurisdictions where the researchers actually live? And what do the comparables above imply the buyer would have to pay to replace this team on the open market — because that number, not the code, is the retention package you are negotiating. Structure follows: retention pools, earn-outs tied to named individuals, and staged consideration. The generic bus-factor mechanics — interviewing the team without the key person present, knowledge-transfer plans, retention packages — are covered in technical due diligence; this section is the AI-specific overlay.
One more warning from the same run of deals: run the IP chain upstream of the acquirer, not just the target. OpenAI's reported roughly $3 billion agreement to acquire Windsurf collapsed on 11 July 2025 without closing, and the reported reason was a term in the buyer's own pre-existing arrangements — Microsoft had access to all of OpenAI's intellectual property, and OpenAI did not want its largest backer to receive Windsurf's coding technology as well (TechCrunch, 11 July 2025). Windsurf's residual business went to Cognition, announced 14 July 2025. The deal died on nothing in the target. Before you diligence someone else's IP chain, diligence the licence-backs and IP-assignment rights your own investors and model partners already hold.
Which 2025-2026 deals were AI-repriced, AI-stopped, or AI-talent-absorbed?
Twelve deals in 2024-2026 anchor the modern AI DD reference set. Each illustrates a specific AI-DD-relevant element that buyer counsel cites when negotiating reps, warranties, indemnification, and price.
One — Microsoft-Inflection AI (reported $650 million license + acqui-hire, 21 March 2024). Microsoft paid a reported $650 million to license Inflection's technology non-exclusively and hire most of its roughly 70-person workforce including co-founders Mustafa Suleyman and Karén Simonyan; Microsoft acquired no assets and Inflection remained a separate company. The reported split: roughly $620 million for the non-exclusive model license and roughly $30 million to waive legal rights related to mass hiring. Both figures are source-reported rather than filed. The FTC opened an informal inquiry in June 2024 and the UK CMA opened a probe shortly after. The diligence frame Microsoft-Inflection established: the reverse-acquihire structure as an alternative to a traditional acquisition that would have triggered HSR review. The 2026 deal-cycle implication: any buyer evaluating a reverse-acquihire path against a traditional acquisition must price the post-close FTC inquiry probability.
Two — Google-Character.AI (reported $2.7 billion license-and-hire, announced 2 August 2024). Google paid a reported $2.7 billion for a non-exclusive license to Character.AI's technology plus the return of co-founders Noam Shazeer and Daniel De Freitas (both former Google employees); Character.AI continued to operate independently under an interim CEO. The structure drew DOJ attention over whether it was built to avoid merger review, alongside an FTC staff report concluding such pseudo-acquisitions could constitute unfair competition by depriving rivals of essential engineering talent. The diligence frame Google-Character.AI established: AI safety and youth-protection liability flowing from chatbot product behavior, separate from the antitrust review of the deal structure.
Three — Amazon-Adept AI (June 2024). Amazon hired Adept co-founder David Luan and a group of colleagues into its AGI team — Semafor reported Adept retained only about a third of its roughly 100 staff — and licensed Adept's technology, multimodal models, and some datasets. Amazon disclosed no price; a reported ~$25 million licensing figure (Semafor, 2 Aug 2024) has never been confirmed by either party. Adept investors were paid back through the licensing proceeds. The FTC opened an informal inquiry. The diligence frame Amazon-Adept established: the licensing-fee-only-no-equity transfer structure that is HSR-exempt because no voting securities or assets transfer.
Four — Stability AI Sean Parker-led rescue (June 2024). Sean Parker and an investor group including Greycroft, Coatue Management, Sound Ventures, and Lightspeed Venture Partners closed a rescue round (approximately $80M new equity plus approximately $400M debt forgiveness organized by Parker with creditors). Prem Akkaraju became CEO; Sean Parker became Executive Chairman. By December 2024, the company had erased all debt and delivered triple-digit revenue growth. WPP added a strategic investment in 2025. The diligence frame Stability AI established: the AI restructuring playbook for distressed-AI targets — debt forgiveness plus new-equity plus executive-team turnover, executed without receivership.
Five — Anthropic Bartz v. Anthropic settlement ($1.5 billion, final approval July 2026). Judge William Alsup of the Northern District of California gave preliminary approval on 25 September 2025 to a settlement requiring Anthropic to pay approximately $1.5 billion (approximately $3,000 per book for roughly 500,000 works) to resolve claims that Anthropic downloaded pirated books from shadow libraries (Library Genesis, Pirate Library Mirror) to train Claude. Alsup subsequently retired, and Judge Araceli Martínez-Olguín granted final approval on 20 July 2026 (TechCrunch, 20 July 2026). Judge Alsup's June 2025 summary-judgment ruling: using legally-acquired books to train AI was fair use; pirating books to train AI was not. The diligence frame Bartz established: the per-work damages baseline for any AI target with shadow-library training-corpus exposure, plus the binary-cure framework (retraining from clean corpora, settlement, customer notification).
Six — Universal Music / Concord et al. v. Anthropic (filed January 2026 by major music publishers, $3 billion+ damages alleged). Concord Music Group, UMG, and ABKCO sued Anthropic in late January 2026 alleging Anthropic illegally downloaded more than 20,000 copyrighted songs including sheet music, lyrics, and compositions — described as potentially the largest non-class-action copyright case in US history. Anthropic had earlier entered into a January 2025 partial-settlement on a separate music-publisher case to avoid an injunction. The diligence frame: music-publisher litigation extends the Bartz-precedent damages model to a new content vertical and signals continued copyright-exposure layering.
Seven — New York Times v. OpenAI / Microsoft (20-million-log discovery order, January 2026). US District Judge Sidney Stein affirmed a magistrate judge's order compelling OpenAI to produce the entire 20 million-log sample of anonymized ChatGPT conversations, not just plaintiff-selected logs. The earlier May 2025 demand had been for 1.4 billion logs; the discovery order narrowed to 20 million. The diligence frame NYT v. OpenAI established: user-log preservation as a litigation hold obligation, with consequence for data-retention policy and infrastructure cost. Any AI target with user-log volume in the hundreds of millions or above needs documented preservation-policy design.
Eight — OpenAI Sora 2 launch (30 September 2025) and opt-out reversal. OpenAI launched Sora 2 with an opt-in for likeness (cameo feature) but opt-out for copyright — users could generate videos with copyrighted characters unless rightsholders explicitly opted out. Within days, the Motion Picture Association demanded immediate action. Three days post-launch, OpenAI reversed to an opt-in model for copyrighted characters. By 2026 the Disney $1B partnership enabled licensed Disney characters in Sora 2 output. The diligence frame Sora 2 established: opt-in versus opt-out as the binary content-rights design choice for any generative product, with rapid-reversal capability becoming part of the product-architecture diligence question.
Nine — xAI $20 billion Series E closed January 6, 2026, upsized from a $15 billion target at a reported $230 billion valuation, then merged into SpaceX on February 2, 2026. xAI closed a $20 billion Series E in January 2026 at $230 billion valuation (upsized from $15 billion target) including Valor Equity Partners, StepStone Group, Fidelity, Qatar Investment Authority, MGX, Baron Capital, Nvidia, Cisco Investments, and Tesla (which committed about $2 billion to the Series E in January 2026 and funded it in March 2026, receiving SpaceX stock because the merger closed first, per Tesla's Q2 2026 10-Q). On February 2, 2026, SpaceX completed its acquisition of xAI through a share exchange that valued the combined company at $1.25 trillion (SpaceX at $1 trillion, xAI at $250 billion per deal documents reviewed by CNBC), with a cash election at $75.46 per share for eligible holders rather than a pure all-stock structure. The diligence frame xAI-SpaceX established: the AI-target absorption-into-strategic-acquirer pattern at scale, with valuation discipline maintained through a predominantly share-based structure with a limited cash election rather than a cash-out.
Ten — FTC Operation AI Comply, DoNotPay final order ($193,000, February 2025) and Rytr final order set aside (December 2025). The DoNotPay February 2025 final order required $193,000 in monetary relief, prohibits the company from advertising that its service performs like a real lawyer unless it has sufficient evidence to back it up, and requires notice to past subscribers between 2021 and 2023. In December 2025, the FTC reopened and set aside the Rytr final order in response to the Trump Administration's AI Action Plan — signaling a shift in FTC enforcement posture but not a withdrawal of the underlying Operation AI Comply enforcement framework. The diligence frame: FTC AI-claims enforcement remains active, but the specific enforcement-priority mix is shifting under the new Administration.
Eleven — FTC Rite Aid Corporation facial-recognition five-year ban (December 2023). The FTC banned Rite Aid from using facial-recognition technology for security or surveillance for five years after finding the technology produced more false-positive results for Black and Latino customers and that Rite Aid failed to implement reasonable procedures. The five-year ban runs through approximately December 2028, well into typical M&A deal cycles. The diligence frame Rite Aid established: facial-recognition and biometric-ID use as a board-level governance risk with explicit FTC consent-decree precedent.
Twelve — BIS H200 / MI325X chip licensing rule (January 2026) and Hart-Scott-Rodino threshold increase ($133.9 million effective 17 February 2026). The BIS final rule of 13 January 2026, effective 15 January 2026, revised the license-review posture for NVIDIA H200- and AMD MI325X-equivalent chips from presumption-of-denial to case-by-case review. On 14 January 2026, the Trump Administration imposed a 25-percent value-based tariff on covered products effective 15 January 2026. The HSR Size-of-Transaction threshold was revised from $126.4 million in 2025 to $133.9 million effective 17 February 2026 (a 5.9 percent increase). The diligence frame: BIS export-control posture is now a deal-by-deal modeling exercise rather than a clean-allow or clean-deny analysis, and the HSR threshold delineates which AI deals require pre-merger notification.
The twelve-deal reference set anchors the 5-Layer Audit's specific score-decisions. The Bartz precedent anchors the Data Layer scrape-risk test. The Llama license-cliff anchors the Model Layer license-term test. The xAI-SpaceX absorption anchors the Infrastructure Layer cloud-and-compute concentration test (xAI on its Memphis Colossus build versus SpaceX's existing data-center stack). The Rite Aid and DoNotPay precedents anchor the Output Layer marketing-claim posture test. The EU AI Act calendar — Article 50 from 2 August 2026, Annex III from 2 December 2027, Article 99 fines live since 2 August 2025 — plus NIST AI RMF 1.0 and the Generative AI Profile anchor the Governance Layer tier-classification and AI RMF tests. The HSR $133.9 million threshold establishes the pre-merger notification gate for any AI deal at or above the threshold.
Which AI tools genuinely speed DD vs. which are theater?
The AI-for-DD half of the AI Due Diligence frame is materially smaller than the DD-of-AI half in deal economics but materially larger in day-to-day workflow time savings. Real-DD use cases and theater use cases divide on a single test: can the AI tool cite its source paragraph for every finding, and does buyer's counsel sign off on the methodology?
Real-DD use case one — clause extraction across customer contracts. Modern LLMs can process hundreds of pages of unstructured deal-room content and surface clause types (caps, indemnities, change-of-control, MFN, exclusivity, auto-renewal, termination-for-convenience, governing law). The honest output: a clause-by-clause table with source-citation back to the specific contract paragraph. Inside Peony, AI extraction speeds the licensing-clause review the same way, with citations back to the originating contract page. A diligence team running clause extraction across 200 customer contracts can complete the workstream in 4-6 hours rather than 4-6 days. The methodology check: every finding cites the source paragraph, the LLM does not invent clauses that are not in the source, and the counsel reviewer can verify each finding against the source in 30 seconds. Real lift, real audit trail.
Real-DD use case two — redline summarization on amendment stacks. Long-running customer contracts often accumulate 5-20 amendments. Modern LLMs can summarize the net-effect changes across the amendment stack — what terms moved, what terms were added, what terms were deleted — with source-citation back to each amendment. The methodology check: the summary references the specific amendment for each change, and the counsel reviewer can verify by reading the cited amendment.
Real-DD use case three — Q&A categorization and routing. Inbound buyer Q&A in a data room can be categorized and routed to the right responder using LLM tagging. Peony's smart Q&A handles this routing with category suggestions and assignee defaults. Real lift: 30-50 percent reduction in routing time. The methodology check: the routing decision is logged with the LLM's reasoning, and the human reviewer can override any routing decision.
Real-DD use case four — financial-statement reconciliation. LLMs combined with structured-data tools can reconcile financial statements against bank records, identify timing differences, and surface unreconciled items. The methodology check: every reconciliation finding cites the specific statement line and bank-record entry.
Real-DD use case five — entity matching across customer contracts. Customer-name normalization across hundreds of contracts (handling typos, legal-name versus DBA, parent-versus-subsidiary references) is a structured task that LLMs handle well with appropriate guardrails. The methodology check: every normalization decision is logged and reversible.
Theater use case one — black-box risk scoring. Vendors that produce an "AI risk score" on a 1-100 scale from an LLM with no audit trail and no source citation are theater. The buyer's counsel cannot sign off on the methodology, and the score has no defensible deal-cycle use. The test: can the tool show what specific document features drove the score, and can the counsel reviewer verify the features against the source documents?
Theater use case two — "diligence in 24 hours" promises. Vendor pitches that promise complete diligence in 24 hours without a human reviewer in the loop are theater. The 24-hour timeline is achievable only by skipping the verification step that converts LLM output into defensible findings.
Theater use case three — AI summarization of executive Q&A with no source citation. Summarizing executive Q&A in a data room is real lift only if the summary cites the source answer. Summaries without source-citation are an audit-trail gap that counsel cannot sign off on.
The honest AI-for-DD posture: use LLMs aggressively for the structured-extraction tasks where source-citation is enforced, and avoid the black-box risk-scoring products. The deal-cycle economics: I am not going to put a week-count on the saving, because I have no published benchmark for it and neither does anyone else quoting one. What is observable in a room is narrower and more honest — the extraction workstreams that used to take days take hours, and the hours that come back go to senior counsel and partners for the higher-judgment work in the 5-Layer Audit rather than to first-pass reading. The timeline compression then changes the cost-stack covered in the due diligence cost breakdown — senior-counsel hours redirect to AI-specific judgment work.
What does AI due diligence cost, and how long does it take?
There is no published, independent benchmark for either number, and I would rather tell you that than repeat one. I went looking for a primary source — Big Four published fee ranges, law-firm surveys, consultancy rate cards with an explicit AI-diligence scope, any dataset behind the figures that rank for this query. What exists instead is a closed loop: every specific range on the first page of search results lives on a data room vendor's content-marketing page, none of them discloses a sample, a methodology, a survey instrument or a collection date, and they cite each other. Big Four firms do not publish transaction-diligence fee ranges at all; those sit in engagement letters. So if you have seen "$35,000 to $95,000" or "three to six weeks" for an AI technical review, you have seen a number with no evidence behind it. Do not put it in a budget memo you will have to defend.
What you can do is bound it from your own deal. Cost is driven by five things, in roughly this order of magnitude. Corpus size and messiness — how many training datasets exist and whether provenance is documented or has to be reconstructed from engineering memory, which is the single biggest swing factor. Contract volume — how many customer agreements need clause-level review for training rights, indemnification and change-of-control, and whether extraction can run across them or a human has to read each one. Regulatory surface — an EU-touching Annex III use case with no conformity assessment file means building the file, not reviewing it, and that is a project rather than a diligence line. Specialist count — an AI/ML technical advisor, IP counsel, data-protection counsel and an EU AI Act specialist are four separate rate cards, and the AI workstream is what adds them. Seller readiness — a target with a model inventory, dataset manifest and eval file already assembled costs a fraction of one where diligence has to build those artifacts before it can assess them. Price it as a premium on top of the technology and IT diligence workstream rather than as a standalone budget; the workstream-by-workstream cost stack it attaches to is in the due diligence cost breakdown.
On timeline, use phase durations rather than a total. The Data Layer is the long pole at roughly two to four weeks on a serious target, because per-dataset traceability cannot be parallelized past the point where the target's own engineers become the bottleneck. Model, Infrastructure and Output layers run concurrently once the folders open, and each is gated by document production rather than by review capacity. Governance is bimodal: days if a conformity assessment file and risk register already exist, months if they do not — which is precisely why the Governance Layer tier classification belongs in the LOI-stage disclosure wave, not in confirmatory. The seller controls most of this. A room that opens with the five layer folders already populated removes the request-and-wait cycle that is the actual reason AI diligence runs long.
How does AI DD feed the data room scope and staged disclosure?
The AI DD workstreams (DD-of-AI and AI-for-DD) feed the data room scope and the staged disclosure model through three specific mechanisms — folder structure, visitor-group access patterns, and engagement-signal monitoring. The structural pattern: every layer of the 5-Layer AI Target Audit gets its own top-level folder, every reviewer tier gets its own visitor group, and every reviewer engagement signal gets fed back to the seller's IR team for negotiation calibration.
Folder structure. The data room top-level folders for AI-using targets typically follow the 5-Layer Audit structure: 01_Data_Layer (training-data manifests, license documentation, opt-out compliance logs, deletion-on-request records), 02_Model_Layer (model cards, fine-tune lineage, eval-suite results, model-risk register), 03_Infrastructure_Layer (GPU contracts, capacity commitment letters, cloud-spend breakdown, DR test logs, export-control records), 04_Output_Layer (marketing-claims register, complaint logs, hallucination-incident database, customer-contract indemnification index, insurance riders), 05_Governance_Layer (AI risk register, EU AI Act tier classification, conformity assessment file, NIST AI RMF program, red-team log). The 5-folder structure mirrors the audit structure so reviewer-by-reviewer drill-down maps directly to layer-by-layer audit work, and manage links lets the seller revoke any over-shared layer-folder link the moment scope drifts.
Visitor-group access patterns. Each reviewer tier gets a visitor group with permission gates that surface only the relevant layer-folders. The matrix below is the one I hand sellers at kickoff — five reviewer groups, five folders, and an explicit default for every cell, because the failure mode in AI diligence is not under-sharing, it is a lender or an operating-team member landing in the training-data manifest.
| Reviewer group | 01 Data | 02 Model | 03 Infrastructure | 04 Output | 05 Governance |
|---|---|---|---|---|---|
| Buyer's AI/ML advisor | View, manifests only | Full view | Full view | View | No access |
| IP counsel | Full view (licenses, receipts) | Full view (lineage) | No access | View (claims register) | No access |
| Data-protection counsel | Full view (PII inventory, DPIA) | View (fine-tune data) | No access | View (complaint log) | View (risk register) |
| Deal team / full-deal partner | Full view | Full view | Full view | Full view | Full view |
| Lender | Summary only | Summary only | View (committed spend) | Summary only | Summary only |
Peony delivers per-group permissions of this shape on the Data Room plan at $52 per admin per month, which is the tier that carries granular per-file and per-user permissions with a full audit trail, dynamic watermarking, the advanced NDA with a countersigned PDF for both parties, Screenshield capture prevention and auto-indexing. Business at $30 per admin per month covers the lighter end — screenshot protection, allow/block visitor lists and a simple acknowledge-only NDA — and is where sellers usually sit before diligence opens. The free plan at $0 handles up to 50 documents, which is enough for a teaser-stage AI exposure summary and nothing more. Reviewers are free on every plan, so a fifth group costs nothing to add; 6,800+ customers run rooms on this model. The visitor-group structure is what prevents the operating tier from accidentally accessing model-card archives or AI risk registers that are properly gated to outside specialist counsel.
Engagement-signal monitoring. Peony's page-level analytics tell the seller's IR team which pages each reviewer tier spent time on. The data-counsel team spending 6 hours on the training-data manifest signals data-license findings are landing; the tech-counsel team spending 4 hours on Llama license-cliff documents signals license-cliff findings are landing; the regulatory-counsel team spending 3 hours on the EU AI Act conformity-assessment file signals AI Act tier classification findings are landing. The engagement-signal tells the seller's IR team where the price-chip ask will come from before the buyer's first markup of the share-purchase agreement.
The Peony brand spine. The AI-DD workstream maps onto Peony product surfaces. NDA gates sign each outside specialist counsel onto the seller's NDA template before any AI-specific folder is visible — particularly material for the training-data manifest and the AI risk register, which are both highly sensitive seller-side artifacts. Visitor groups gate each layer-folder to the relevant reviewer tier without manual permission resets. Dynamic watermarks embed firm-name plus reviewer-name plus timestamp on every model-card and conformity-assessment-file page so any leak is traceable. Screenshot protection blocks naive capture on the training-data manifest and red-team logs at the renderer level. Leak protection deters print-screen and download exfiltration on the most sensitive AI artifacts (training-data manifest, fine-tune lineage, red-team logs). Auto-indexing recognizes typical AI-target document inventory (model cards, eval suites, red-team reports, AI risk registers) across the multi-GB archive without manual folder configuration. Peony Data Room at $52/admin/month ships unlimited storage for the typical 8-15 GB AI-target archive across model-card, training-data manifest, GPU-contract, and conformity-assessment files.
The staged disclosure pattern proceeds in three waves. Wave one — the teaser-stage disclosure surfaces a one-page AI exposure summary (5-Layer Audit scores plus EU AI Act tier classification summary) before any deep folder is visible. Wave two — the LOI-stage disclosure opens the AI risk register, the model inventory, and the EU AI Act tier classification in full. Wave three — the confirmatory-DD-stage disclosure opens every layer-folder in full, including the training-data manifest, conformity-assessment file, and red-team log. The three-wave structure protects the seller's most sensitive AI artifacts (training-data manifest, red-team log) until the buyer has demonstrated sufficient seriousness through LOI and exclusivity execution. The pricing math for the seller stays simple at the Data Room plan on the pricing page: one $52/admin/month seat per deal lead, with reviewer seats free.
The broader M&A DD workstream relationships are covered in the M&A due diligence process guide. The data-room-side platform comparison and the redaction-and-permissions diligence patterns are covered in the VDR permissions guide for due diligence. The cost-side benchmarks are covered in the due diligence cost breakdown. For founder-stage targets where AI is the product, the startup due diligence guide adds the early-stage diligence frame. For the vendor-side audit pattern, the vendor due diligence checklist covers the SaaS-and-vendor angle. The fuller checklist of documents the buyer requests across all DD workstreams sits in the due diligence data room checklist.
Related resources
- Technical due diligence (2026) — code, architecture, and team review for software targets, plus the lender-grade independent report
- IT due diligence (2026) — 6-Axis Tech-Stack Fragility Audit (parallel to AI DD's 5-Layer)
- Environmental due diligence (2026) — PFAS + CSDDD + BFPP Defense Stack
- M&A due diligence process guide — hub for the broader DD cluster
- Due diligence data room checklist — 174-document file-side companion
- Due diligence questionnaire (DDQ) — 5-persona template library
- IP due diligence — 5-Asset Encumbrance Matrix
- Operational due diligence — 8-System audit + 3-Axis severity
- Private equity due diligence — 6-strategy hold playbook
- Vendor due diligence checklist — procurement third-party risk
- Due diligence cost breakdown — what diligence really costs
- Startup due diligence guide — investor-to-startup framing
- How to securely share a Claude artifact — sharing an AI-generated diligence artifact without losing the audit trail
- AI governance due diligence — the program-level half: board oversight evidence, ISO/IEC 42001, the AI-governance questionnaire
- AI in the data room — the trust boundary, abstention, permission-scoped AI Q&A, and why a vendor's no-training promise is a different question from your target's
- MCP data room — letting an AI agent query a permissioned room without downloading it
- AI due diligence checklist (tool) — the 5-Layer request list as a tickable, exportable checklist with the evidence, reviewer and disclosure wave on every line
- Best AI virtual data rooms for M&A due diligence — using AI to run diligence (the opposite job to diligencing an AI company)
- 9 Big Data Techniques That Create Business Value — the analytics-method map underneath modern diligence: gradient boosting, graph analysis, embeddings, and LLM extraction, each with honest limits
- Share AI-generated documents securely — the workflow for distributing AI-built CIMs, memos, and live models under watermark, NDA, and audit
- Peony pricing — Data Room at $52/admin/month
Frequently asked questions
What is AI due diligence?
AI due diligence is the buyer-side investigation of a target company's artificial-intelligence stack before an acquisition or investment closes. It examines five things a generic tech audit misses: where the training data came from and whether it was lawfully acquired, who owns or licenses the models, what the compute supply chain and inference costs look like, what liability the model's outputs create, and how the target is governed against the EU AI Act, ISO/IEC 42001:2023 and the NIST AI Risk Management Framework. The output is not a memo — it is a score, a document request list, and a repricing or walk-away recommendation.
How is AI due diligence different from technical due diligence?
Technical due diligence asks whether the software works, scales and is maintainable: code quality, architecture, open-source license scan, security posture, engineering team depth. AI due diligence asks a different set of questions that a code review structurally cannot answer, because the risk does not live in the repository. Training data is not source code, so provenance and acquisition receipts matter more than commit history. Model licenses carry user-count cliffs that software licenses do not. Inference is a cost of goods sold that a supplier can reprice. And hallucination is a liability event, not a bug ticket. For a corp-dev associate evaluating an AI-using SaaS target with three production models, the 5-Layer AI Target Audit (Data / Model / Infrastructure / Output / Governance) covers the five things a generic tech audit misses: scraped training data with no license trail, third-party model terms that flip on a user-count threshold, GPU contracts that drop on capacity, FTC-actionable marketing claims, and EU AI Act tier obligations on the post-Omnibus timeline (GPAI live since August 2025, Article 50 transparency from 2 August 2026, Annex III high-risk deferred to 2 December 2027). Run both — they overlap by maybe a fifth.
How do you tell a real AI company from a wrapper?
You do not tell by the pitch deck; you tell by the invoices and the gross margin. Ask what share of production inference calls terminate at a third-party hosted model, what percentage of cost of goods sold is third-party model spend, and what survives if that provider doubles its price card or deprecates the endpoint. A target that owns proprietary data, evaluation infrastructure, workflow depth and distribution can route around a model swap in weeks. A target whose only asset is a prompt cannot. Bessemer's State of AI 2025 measured roughly 25 percent gross margins at one archetype and about 60 percent at another — a 35-point spread at identical revenue.
Who owns an AI model's weights, training data, and outputs?
Usually three different parties, which is exactly why the question is asked separately. Weights are owned by whoever trained them, unless the target fine-tuned somebody else's base model — then a community or commercial license governs the derivative, and Meta's Llama Community Licenses, for example, require a separate license request from anyone above 700 million monthly active users on the model's version release date. Training data is normally licensed rather than owned, so the acquisition receipts matter more than the manifest. Outputs are governed by the target's customer contracts. Ask for all three chains in writing before you price the intellectual property.
How long does AI due diligence take?
I looked for a published benchmark and there is not one — every specific range circulating on this topic traces back to vendor marketing pages with no sample, no methodology and no collection date. What I can say from running rooms is which phases dominate. Data Layer provenance work is the long pole, typically two to four weeks on a serious target because every dataset needs tracing to an acquisition receipt. Model, Infrastructure and Output layers run in parallel once the folders open. Governance depends entirely on whether a conformity assessment file already exists, or has to be built from nothing.
How should a data room be structured for AI due diligence?
Five top-level folders keyed to the five audit layers, and one reviewer group per specialist tier so nobody sees more than their scope. The buyer's machine-learning advisor needs the model and infrastructure folders; intellectual-property counsel needs training-data licenses and fine-tune lineage; data-protection counsel needs the personally identifiable information inventory; the lender usually needs only the summary. In Peony, that is the Data Room plan at $52 per admin per month — granular per-folder permissions, dynamic watermarks on every model card, an advanced NDA gate and a full audit trail. Business is $30 per admin per month, the free plan covers 50 documents, and reviewers are always free. 6,800+ customers run deals on it.
How does the EU AI Act change M&A pricing for AI-using targets in 2026?
If you're EU AI Act compliance counsel scoping a target's tier exposure for a buyer, the Annex III high-risk obligations now enter application on 2 December 2027 — deferred from 2 August 2026 by the Digital Omnibus on AI (Regulation (EU) 2026/1744, in force 27 July 2026) — while obligations for general-purpose AI (GPAI) models have been in force since 2 August 2025 and Article 50 transparency obligations still apply from 2 August 2026. A target whose product sits in the High-Risk tier (recruitment scoring, credit scoring, biometric ID, critical infrastructure) carries conformity assessment, post-market monitoring, and registration duties that meaningfully shift deal economics — now priced to the December 2027 (Annex III) and August 2028 (Annex I embedded) deadlines, both of which land inside a typical hold period.
What does the 5-Layer AI Target Audit actually cover, and how do I score it?
The 5-Layer AI Target Audit scores Data, Model, Infrastructure, Output, and Governance from 1 (broken) to 5 (clean) per layer, producing a composite AI Deal Health Score from 5 to 25. Bands map to deal action: 5-10 walk away or restructure to asset-only purchase, 11-17 enter price-chip negotiation with escrow or earn-out, 18-25 healthy AI target with no AI-specific repricing required. Each layer has its own diagnostic checklist, document list, and threshold benchmarks anchored against real 2025-2026 deal precedents.
How do you audit the Data Layer — training provenance, PII exposure, and scrape risk?
The Data Layer audit traces every training dataset to a documented source with a license trail, screens for scraped content of copyrighted works (the Bartz v. Anthropic $1.5 billion settlement made this exposure quantifiable), and inventories any personally identifiable information (PII) in training corpora. The buyer requires per-dataset provenance documentation, license terms with effective dates, opt-out compliance records (DMCA, GDPR Article 17), and the deletion-on-request log. The pirated-corpora question is now a binary: did training corpora include shadow library content (LibGen, Anna's Archive, Pirate Library Mirror), and if yes, is there documented post-detection cure?
How do you audit the Model, Infrastructure, and Output Layers?
The Model Layer audit categorizes every model in the stack as proprietary (built and trained in-house), licensed (third-party with license terms), or hybrid (fine-tuned on a base model), then hunts the license-cliff: the Meta Llama Community Licenses terminate free commercial use above 700 million monthly active users measured on the version release date, and Llama 3.1 and Llama 4 add a naming condition on any model improved with Llama Materials or their outputs. The Infrastructure Layer audit covers GPU contract terms (committed capacity, allocation guarantees, force-majeure clauses), cloud provider concentration (more than 70 percent on one hyperscaler is a chip-on-the-table flag), and BIS export-control posture given the January 2026 codification of NVIDIA H200- and AMD MI325X-equivalent licensing. The Output Layer audit prices marketing-claim risk under the FTC Operation AI Comply framework launched September 2024, including the DoNotPay $193,000 final order (February 2025) and the Rite Aid facial-recognition five-year ban. Across all three, buyers request the model-card archive and fine-tune lineage, the GPU reservation contract and cloud spend breakdown, and the marketing-claims register alongside customer-contract indemnification and the insurance carrier's AI rider terms.
How do you audit the Governance Layer — EU AI Act tier, NIST AI RMF, and red-teaming cadence?
The Governance Layer audit confirms the EU AI Act tier classification (Unacceptable / High-Risk / Limited / Minimal), maps the target's risk-management program against NIST AI RMF 1.0 plus the July 2024 Generative AI Profile (NIST.AI.600-1), and reviews the red-teaming cadence. Buyers request the AI risk register, the conformity assessment file (if High-Risk), the post-market monitoring plan, the incident response procedure, and the red-team log with cadence (quarterly minimum for production GenAI). A target with no AI risk register and no red-teaming cadence is the single most common Tier 1 finding and triggers automatic Governance Layer score of 1-2.
Which 2025-2026 deals were AI-repriced, AI-stopped, or AI-talent-absorbed?
The reference set: SpaceX completed its acquisition of xAI on February 2, 2026, a share exchange that valued the combined company at $1.25 trillion, with SpaceX at $1 trillion and xAI at $250 billion per deal documents reviewed by CNBC (February 3, 2026); each xAI share converted into 0.1433 SpaceX shares, and eligible holders could elect cash at $75.46 per share instead, so it was not an all-stock deal. CNBC called it the biggest merger on record, and SpaceX has since gone public (final prospectus filed June 12, 2026). The Microsoft-Inflection ($650 million license + acqui-hire, March 2024), Google-Character.AI ($2.7 billion, August 2024), and Amazon-Adept (undisclosed license + partial staff hire, June 2024) trio established the reverse-acquihire pattern and triggered FTC and CMA inquiries through 2025. The Stability AI Sean Parker-led rescue (June 2024, ~$80M new equity plus ~$400M debt forgiveness) became the playbook for distressed-AI restructuring rather than receivership.
Which AI tools genuinely speed due diligence work, versus which are theater?
If you're a PE deal partner running platform diligence on an AI-assisted services target, the real-DD use cases are clause-extraction across hundreds of contracts (caps, indemnities, change-of-control, MFN), Q&A categorization for question routing, redline summarization on amendments, financial-statement reconciliation against bank records, and trade-area entity matching. Theater: AI 'risk scoring' on a 1-100 scale derived from a black-box LLM with no audit trail, vendor pitches that promise 'AI diligence in 24 hours' without a human reviewer in the loop, and AI summarization of executive Q&A with no source-citation back to the underlying answer. The honest test: can the tool cite its source paragraph for every finding, and does the buyer's counsel sign off on the methodology?
You might also like
Aug 22, 2026
AI Governance Due Diligence (2026): The 10-Artifact Evidence Pack + Corrected EU AI Act Timeline
Aug 21, 2026
Insurance Due Diligence in M&A (2026): The Collateral Trap + Loss-Run Playbook
Aug 21, 2026
ESG Due Diligence (2026): The Post-Omnibus Scope Reset + the Scope 3 Evidence Test

