Where We Help Experience Insights Resources Leadership Book a Meeting →

AI Due Diligence for Private Equity: How to Test a Target’s AI Claims

A practical framework for determining whether an acquisition target’s AI capabilities are real, defensible, scalable, compliant, and economically valuable.

← All insights
The short version

An AI demonstration is not evidence of an AI advantage. Private equity buyers should require six connected proofs: ownership and dependencies, data and intellectual-property rights, task-specific performance, actual human intervention, production unit economics, and post-deployment controls. Each finding should lead to an underwriting decision, deal protection, or funded post-close action.

AI-specific diligence should establish whether a target’s claimed capability works under real operating conditions, what it depends on, whether the company has the necessary rights, and whether the economics support the investment case.

A polished demonstration proves only that a workflow can work under selected conditions. It does not establish representative accuracy, automation at scale, ownership of the underlying capability, regulatory readiness, or attractive margins.

Buyers should require six connected proofs:

  1. Ownership and dependencies: What has the target built, and what does it rent or outsource?
  2. Data and intellectual-property rights: Can the company use the inputs, outputs, customer data, and improvements on which the product depends?
  3. Task-specific performance: Does the system meet its commercial claim on representative data?
  4. Human intervention: How much labor is required to produce, review, correct, or deliver the output?
  5. Production unit economics: What does each successful completed task cost at expected volume and service quality?
  6. Post-deployment controls: Can the company detect and contain failures, security threats, behavioral changes, and regulatory exposure after release?

Weakness in any one area can change the investment case. It can reduce defensibility, compress margins, delay growth, create remediation costs, or make the value-creation plan impractical.

When AI-specific diligence is necessary

Conventional software diligence remains essential. Architecture, security, engineering practices, scalability, technical debt, and product delivery still matter. AI-specific diligence adds another layer when a material product feature or investment thesis depends on probabilistic model behavior, training or retrieval data, third-party models, automated decisions, or AI-generated outputs.

That additional work is warranted when the target claims AI:

  • differentiates its product or supports a valuation premium;
  • automates work previously performed by people;
  • improves accuracy, speed, conversion, retention, or another customer outcome;
  • creates a proprietary data or learning advantage;
  • materially expands gross margin at scale;
  • performs a regulated, safety-sensitive, or high-consequence task;
  • can extend rapidly into new products, customers, or geographies; or
  • supports a central part of the post-close value-creation plan.

The issue is not limited to companies marketed as AI businesses. A software or services company may use AI deeply enough that model costs, data rights, human review, or third-party dependencies materially affect its economics.

AI use alone says little about the maturity of a specific capability. Stanford’s 2026 AI Index reports that organizational AI adoption reached 88% in 2025 and that 70% of surveyed organizations used generative AI in at least one business function. Agent deployment remained in the single digits across nearly all functions. [1] Buyers need evidence of what the target’s AI does reliably in production, not evidence that the company uses AI somewhere.

The six proofs an investment committee should require

ProofCore diligence questionEvidence to requestCommercial implication
Ownership and dependenciesWhat capability belongs to the target?Architecture diagrams, model inventory, vendor contracts, source repositories, deployment recordsDefensibility, switching cost, concentration risk, replacement cost
Data and IP rightsDoes the target have sufficient rights to operate and improve the system?Data lineage, licenses, customer terms, consent records, retention policies, invention assignmentsLegal exposure, product continuity, ability to train and expand
Task-specific performanceDoes the system meet the commercial claim on representative work?Evaluation datasets, scoring methods, error analysis, production samples, version historyRevenue quality, customer value, warranty risk, growth assumptions
Human interventionHow automated is the delivered outcome?Workflow logs, review queues, staffing data, escalation rules, correction recordsGross margin, capacity, service quality, labor exposure
Production unit economicsWhat does a successful completed task cost?Model and cloud bills, usage data, latency, support costs, retry rates, volume forecastsContribution margin, pricing, cash needs, scalability
Post-deployment controlsCan management detect and contain failures?Monitoring dashboards, incident logs, access controls, test cadence, rollback proceduresOperational resilience, security exposure, compliance cost

These proofs must be tested together. A target can produce strong evaluation results but lack the rights to use its data. It can own valuable software while relying on an external model that undermines expected margins. It can report high output quality while concealing a substantial manual review operation.

1. Separate proprietary capability from assembled capability

Third-party models are not inherently a problem. They are often the sensible architectural choice. The diligence question is whether management presents an assembled application as proprietary AI and whether the target controls the components that matter commercially.

Start with a dependency map covering:

  • foundation models and model APIs;
  • open-source models and licenses;
  • cloud and specialized infrastructure;
  • retrieval systems, vector stores, and external data feeds;
  • orchestration, prompts, tools, and workflow logic;
  • fine-tuning, classifiers, evaluation systems, and guardrails;
  • outsourced annotation, review, or operational support; and
  • customer-specific integrations or configurations.

For each component, determine who owns it, who can change it, how it is priced, what happens if access ends, and whether the target can switch without rebuilding the product.

If most performance comes from a third-party model, defensibility may reside elsewhere: workflow integration, proprietary data, distribution, customer switching costs, or operational execution. The buyer should value the company on that basis rather than assigning unsupported value to model ownership.

Dependency disclosure is also a credibility test. In a January 2025 enforcement action, the U.S. Securities and Exchange Commission found that Presto Automation had failed to disclose that third-party technology powered all deployed units of its AI product for a period. The SEC also found that the vast majority of orders processed by a later version using Presto’s own technology required human intervention, contrary to the company’s automation claims. [2] Buyers should verify both the technology supply chain and the labor supply chain.

2. Trace rights across data, models, prompts, and outputs

A data-room folder labeled “AI IP” is not sufficient. The buyer needs a rights map connecting the system’s important inputs, components, and outputs to contracts, licenses, policies, and technical controls.

That map should address:

  • the origin of training, fine-tuning, evaluation, and retrieval data;
  • rights to use customer data for service delivery and model improvement;
  • whether personal, confidential, regulated, or copyrighted material enters prompts or datasets;
  • ownership of prompts, workflow logic, labels, evaluations, and fine-tuned artifacts;
  • rights to AI-generated outputs and customer-specific derivatives;
  • vendor rights to retain submitted data or use it for training;
  • deletion, portability, and retention obligations; and
  • employee and contractor invention assignments.

NIST’s Generative AI Profile recommends documenting data origin and content lineage, testing data and content flows, and documenting reliance on upstream data sources. [3] The U.S. Copyright Office has separately examined the copyrightability of outputs created with generative AI and the use of copyrighted materials in AI training. [7]

Technical diligence does not need to resolve every unsettled legal question. It should identify where the product’s operation or expansion depends on rights the target may not possess. Counsel can assess the legal consequences while the deal team estimates remediation cost, customer impact, and timing.

3. Test the commercial claim, not a generic benchmark

Evaluate the AI system against the task the target sells. A broad model benchmark rarely proves that an application performs well on the target’s customers, documents, edge cases, languages, or operating environment.

NIST’s August 2026 initial public draft of the TEVV-Athlon Framework states that AI measurement approaches should be customized to organizational objectives and the relevant application context. [4] NIST’s Generative AI Profile also recommends evaluating outputs against known ground-truth data and using multiple evaluation methods. [3]

An independent test should follow seven steps:

  1. State the commercial claim. Examples include reducing review time, correctly classifying cases, drafting acceptable outputs, or resolving requests without escalation.
  2. Define success. Specify the quality threshold, latency, failure tolerance, and permitted degree of human review.
  3. Select representative cases. Include routine work, difficult cases, missing information, adversarial inputs, and known failure categories.
  4. Create or verify ground truth. Use qualified reviewers and record disagreement instead of forcing false precision.
  5. Run the production workflow. Test the deployed configuration, not a specially prepared demonstration environment.
  6. Measure the full distribution. Report failure types and severe outliers, not only an average score.
  7. Reconcile results with commercial claims. Determine whether observed performance supports the revenue, retention, staffing, or margin assumptions in the model.

Management’s internal evaluations are useful evidence, but buyers should inspect them for dataset leakage, selective exclusions, stale model versions, small samples, subjective scoring, and methodology changes. If management cannot reproduce its headline metric or explain its denominator, that metric should not support underwriting.

4. Measure the labor behind the automation claim

Hidden human work can affect both credibility and margin.

Human review is not automatically a weakness. It may be the correct control for a high-consequence workflow. The problem arises when the financial model assumes automation while the operating model depends on unreported review, correction, escalation, or data preparation.

Request evidence connecting each production output to the work required to deliver it:

  • percentage of cases reviewed by a person;
  • review and correction time per case;
  • escalation and exception rates;
  • staffing by customer, product, and shift;
  • outsourced labor contracts and locations;
  • rework caused by model failures;
  • service-level commitments; and
  • whether reviewers approve outputs or substantially create them.

The relevant measure is not model accuracy alone. It is the cost and quality of the completed customer outcome. If a nominally automated transaction requires retries, manual correction, and supervisory approval, all of that belongs in the unit economics.

5. Rebuild AI unit economics from production evidence

AI cost models should start with observed production usage, not vendor list prices or management’s estimate of a typical request.

For each material workflow, calculate the cost per successful completed task. Include:

  • model inference or API charges;
  • input and output volume;
  • retrieval, storage, and data-processing costs;
  • cloud compute and specialized infrastructure;
  • retries, fallbacks, and failed runs;
  • monitoring, evaluation, and security tooling;
  • human review and exception handling;
  • customer support and implementation labor; and
  • minimum commitments or volume tiers in vendor agreements.

Test the model under several scenarios: current volume, the base underwriting case, a high-usage case, a model-price increase, a switch to a higher-quality model, and increased review for a new customer or product category.

The U.S. Government Accountability Office’s 2026 review of 13 federal AI acquisitions found that agency officials had difficulty understanding AI-related costs. GAO also identified data-rights and testing clauses as examples of acquisition lessons that agencies should capture and reuse. [5] The same questions matter in commercial diligence because variable costs and contractual constraints may become visible only at higher volume.

The analysis should answer four investment questions:

  • Does gross margin improve, remain stable, or decline as usage grows?
  • Is pricing aligned with the target’s actual cost driver?
  • How exposed is the company to one model or infrastructure provider?
  • What engineering or operational investment is required to reach the underwritten margin?

6. Test controls for production, not only launch

Pre-close testing provides a snapshot. AI behavior can change as models, prompts, datasets, customer inputs, and external dependencies change. NIST states that controlled pre-deployment evaluations are inherently limited and should be complemented by repeated testing, evaluation, validation, and verification after deployment. [9]

Diligence should establish whether the target can:

  • identify which model and configuration produced an output;
  • detect changes in quality, cost, latency, or usage;
  • restrict access to sensitive data and high-impact actions;
  • test updates before release;
  • investigate incidents and reproduce failures;
  • roll back models, prompts, or workflow changes;
  • set consumption and authorization limits; and
  • monitor third-party changes affecting the product.

Security testing should cover AI-specific attack paths as well as conventional application security. OWASP’s LLM and generative-AI security initiative identifies prompt injection, sensitive-information disclosure, supply-chain vulnerabilities, data and model poisoning, excessive agency, misinformation, and unbounded consumption among its principal risk workstreams. [8] Tests should match the application. A drafting assistant and an autonomous agent with access to customer records and payment systems do not present the same exposure.

Regulatory analysis should begin with the target’s role, use case, customers, and geography. The European Union AI Act became generally applicable on August 2, 2026, with exceptions. Under the European Commission’s current timeline, rules for specified high-risk uses apply on December 2, 2027, while rules for high-risk systems embedded in regulated products apply on August 2, 2028. [6] The buyer should identify the target’s role in affected systems and ask counsel to confirm the resulting obligations.

Convert findings into deal decisions

An effective AI diligence report does more than classify findings as red, yellow, or green. It connects each material finding to underwriting, transaction terms, or execution.

Possible responses include:

  • Valuation adjustment: Reduce credit for unsupported differentiation, automation, growth, or margin claims.
  • Model adjustment: Add realistic model, infrastructure, labor, compliance, and remediation costs.
  • Deal protection: Address ownership, data rights, undisclosed dependencies, security incidents, and regulatory representations with counsel.
  • Closing condition: Require resolution when a missing right, contract, control, or dependency threatens business continuity.
  • Integration dependency: Sequence the identity, data, security, vendor, and architecture work required to operate the business safely.
  • 100-day initiative: Fund the evaluations, monitoring, rights remediation, vendor diversification, or workflow redesign needed to support the thesis.
  • Thesis rejection: Walk away when the claimed advantage cannot be reproduced or delivered economically.

AI diligence should not end with a red-flag presentation. The evidence should shape the operating plan. EE Solutions’ Private Capital Technology Decision Playbook provides a broader structure for connecting technical findings to investment and operating decisions. A technology operating partner becomes particularly important when the work must continue from assessment into implementation.

The minimum AI diligence request

Before close, the buyer should obtain enough access to reproduce the target’s central claim. The exact request will vary, but a focused package usually includes:

  • a live production walkthrough using buyer-selected cases;
  • architecture and data-flow diagrams;
  • a complete model, vendor, and open-source dependency inventory;
  • model and infrastructure contracts, pricing, and usage reports;
  • evaluation datasets, methods, results, and version history;
  • production quality, latency, cost, escalation, and intervention metrics;
  • data provenance and rights documentation;
  • customer terms governing data use and outputs;
  • security assessments, incident records, and access controls;
  • monitoring dashboards and change-management procedures;
  • staffing and outsourced-review data; and
  • interviews with product, engineering, data, security, legal, finance, sales, and frontline operations.

If management restricts access for legitimate confidentiality or security reasons, use a controlled test environment, clean-room review, or scoped independent assessment. Do not accept the claim without evidence.

The investment committee does not need certainty about every future AI development. It needs evidence that the underwritten capability exists now, performs the task customers pay for, operates at the modeled cost, and can withstand foreseeable technical, contractual, and regulatory pressure. That is the standard an AI claim should meet before it enters the deal thesis.

Sources

  1. 1
    The 2026 AI Index Report: EconomyStanford Institute for Human-Centered Artificial Intelligence · accessed 2026-09-02
  2. 2
    SEC Charges Restaurant-Technology Company Presto Automation for Misleading Statements About AI ProductU.S. Securities and Exchange Commission · 2025-01-14 · accessed 2026-09-02
  3. 3
    Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence ProfileNational Institute of Standards and Technology · 2024-07-26 · accessed 2026-09-02
  4. 4
    The TEVV-Athlon Framework for Evaluating AI Systems: Initial Public DraftNational Institute of Standards and Technology · 2026-08-07 · accessed 2026-09-02
  5. 5
  6. 6
    AI ActEuropean Commission · accessed 2026-09-02
  7. 7
    Copyright and Artificial IntelligenceU.S. Copyright Office · accessed 2026-09-02
  8. 8
    Top 10 for LLM and GenAIOWASP GenAI Security Project · accessed 2026-09-02
  9. 9
    Challenges to the Monitoring of Deployed AI SystemsNational Institute of Standards and Technology · 2026-03-01 · accessed 2026-09-02
EE Solutions

Need a senior technology team around the decision?

EES works with private capital firms and portfolio companies from technical assessment through execution.

Book a Meeting