Contact us

Enterprise AI Security: What Buyers Actually Ask Before They Sign

September 4, 2026 13 min 05 sec

 

Your SOC 2 report answers a different question.

TL;DR:

  • An enterprise security review does not ask whether your company is secure. It asks what your AI does with the buyer’s data, whether that AI is one feature in your product or the product itself, and it expects the answer to come out of the system.
  • Enterprise questionnaires increasingly carry a separate AI section, and it asks about the model, the data path, and retention behavior rather than about your company’s controls 
  • The wording varies, but the same themes recur: data flow, training use, retention, isolation, human oversight, incident handling 
  • The artifact list is short and specific, and the records of what the model did are what teams most often cannot produce
  • OWASP published the 2026 LLM Top 10 in August 2026, and prompt injection stayed at number one
  • Most of these are architectural decisions. By the time the questionnaire arrives, they have already been made

The deal stopped at a question about your model

An enterprise security review is the point where a buyer’s risk team decides whether your product may touch their data, and AI, whether it’s one feature or the whole product, turns that review into a separate line of questioning. 

The sequence is familiar to anyone who has sold into an enterprise account. Procurement moves, legal moves, and then a vendor security review lands in your inbox as a spreadsheet: the questionnaire a buyer’s risk team uses to vet everything from access control to incident response. Near the end of it there is a block headed “AI and machine learning.” Your SOC 2 report is attached to the reply already. It covers most of the spreadsheet. That block is where it runs out, and the deal stops there while your team tries to work out who owns the response.

What the reviewer is examining has narrowed. Enterprise AI security review questions target the behavior of one functional block in your product, and the boundary of that examination runs along the buyer’s data. By this point reviewers usually assume the organizational controls have already been checked. 

For many enterprise buyers, AI is no longer evaluated as an isolated capability. They are assessing whether adding it changes their own data risk profile, which is why questions about model providers, retention, tenant isolation and auditability now sit alongside the traditional security controls.

Many of the enterprise AI security best practices published today describe governance processes. The questionnaire asks for evidence produced by a running system, so that is where this article starts.

Where SOC 2 stops and the AI section begins

SOC 2 attests that your organization operates the controls it described, and the AI section of a questionnaire asks a separate question about the model itself.

The Trust Services Criteria cover access management, change management, monitoring, and the rest of your control environment. A Type II report covers a defined system description over a defined examination period. If your AI feature shipped after that period, or falls outside the system description, the report says nothing about it, and reviewers check the scope section for exactly this.

Scope is the first gap. The second one is structural. The Trust Services Criteria examine the controls a service organization operates around its systems. A model’s behavior is shaped by its training data and by the context it receives at inference, so evidence about the surrounding control environment says little about what the model does with a specific input.

Two frameworks have filled part of that space, and buyers increasingly ask about both:

  • ISO/IEC 42001 — the certifiable management system standard for AI. Buyers ask whether you hold it, or when you plan to
  • NIST AI RMF — a voluntary risk framework, used as the vocabulary risk teams write their questions in

Neither replaces SOC 2. They sit beside it and answer what it was never scoped to answer. If you need the SOC 2 side in depth, SOC 2 Type II Compliance: What Auditors Actually Expect Your System to Do covers it.

What does an enterprise security team ask about your AI feature?

The AI section of a vendor questionnaire is narrower than it looks. Across the questionnaire sources we reviewed, the same six themes recur even when the wording changes, and each one expects an answer a system can produce.

The Cloud Security Alliance released AI-CAIQ v1.1 on June 23, 2026; the AI sections of the SIG questionnaire from Shared Assessments use a different taxonomy and surface the same concerns.

One detail in AI-CAIQ v1.1 explains the dynamic. Alongside control specifications and self-assessment questions, it carries a separate set of justification questions — a section for evidencing each answer. The evidence request is built into the instrument.

Six themes of AI security questions.

Data path

Where the buyer’s data physically goes once it enters your AI feature.

  • “Does customer data leave our tenant boundary at any point during inference?”
  • “Which subprocessors receive this data, and under what agreement?”
  • “Where is the data physically processed?”

The third question catches teams who route inference through a model provider in a different jurisdiction without recording that in the subprocessor list.

Training use

Whether the buyer’s data improves your product for someone else.

  • “Is our data used to train, fine-tune, or improve your models?”
  • “If opt-out exists, is it default or on request, and how is it enforced technically?”

The second half of that question is where prepared answers fall apart. A policy statement is not enforcement, and reviewers now ask which configuration setting carries it.

Retention and deletion

How long the data stays, and what deletion actually removes.

  • “How long are prompts and outputs retained?”
  • “Does deletion of our account remove the data from model context, caches, and logs?”

Deletion questions have grown teeth because the answer touches four stores that rarely share a retention policy: the primary database, the vector store, the inference logs, and whatever your observability vendor holds.

Isolation and access

Whether one customer can reach another customer’s data, and who inside your company can read any of it.

  • “Can one tenant’s data appear in another tenant’s output?”
  • “Who inside your company can read prompts submitted by our users?”

The second question is about your own staff. Support tooling that surfaces raw prompts for debugging is a common finding, and it is usually discovered during the review by the vendor themselves.

Oversight and output control

What happens when the model is wrong.

  • “What happens when the model produces a wrong or harmful output?”
  • “Is a human in the loop, and at which step?”

For some use cases “no human in the loop” is a complete answer on its own. What reviewers follow up on is the detection: what catches a bad output when nobody is watching.

Testing and incident handling

Whether anyone has tried to break the feature on purpose.

  • “Have you tested this feature against adversarial input?”
  • “What is your procedure when an AI-specific incident occurs?”

An incident response plan that covers infrastructure and says nothing about model behavior reads as an incomplete answer, because the failure modes are different.

What the risk team is doing with these answers. Most of these questions end up as requests for evidence, and the justification section of the questionnaire is where that evidence goes. The reviewer checks the artifact behind the sentence, and an answer arriving without one carries little weight in the scoring. It is why an AI security questionnaire rewards a mediocre diagram over confident prose. 

 

The answers come from two layers of the architecture

Every question above resolves to one of two places: the data governance layer that controls what enters and leaves the model, and the model governance layer that records what the model did.

Enterprise AI governance is usually presented as policy work owned by a committee. In a product company with no security function, it is closer to system design, and it lives in two layers of the AI Compliance Stack: data governance and model governance.

Data and model governance layers.

Data governance. This layer decides what the model is allowed to see and what survives afterward.

  • A tenant boundary enforced at the query, so isolation is a property of the data path
  • Input classification that tags what is sensitive before it reaches the prompt
  • Retention expressed as configuration, with an explicit period set for each store
  • Training opt-out enforced by a provider setting, where one exists, asserted by your code on every call
  • Processing region pinned in the deployment configuration, per customer where the contract requires it

Model governance. This layer records what happened, so questions about a specific decision have an answer months later.

  • Invocation logging that captures the request, the model version, and the outcome
  • Model versioning with a recorded rollback path
  • Human review, where it exists, recorded as a workflow state so it stays queryable
  • Output monitoring that fires on the conditions the buyer asked about

Log the minimum needed to reconstruct the event. Where prompts or outputs carry customer data, use references, redaction or controlled access instead of storing raw content.

We refer to this approach as Compliance-Native Architecture: designing the data path and the evidence model so buyer requirements are enforced by the system itself. None of these questions should come as a surprise. 

 A team that reads them before designing the data path builds an enterprise AI security architecture that answers them by construction, and a team that reads them afterward pays for a migration.

Classification under the EU AI Act runs on its own thresholds. The questionnaire themes described here stay the same, and high-risk classification adds obligations on top of them. 

Data leakage is the question reviewers keep returning to

Across every questionnaire theme, one concern repeats in different wording: whether the buyer’s data can surface somewhere the buyer did not authorize.

AI data leakage prevention comes down to four common paths and the evidence that each one is closed. Each path fails differently, and each needs its own proof.

  1. Training ingestion. The model provider uses submitted content to improve its models. Provider capabilities differ, so closing this path means confirming which controls your provider offers and configuring them. The evidence is the provider’s terms, the account configuration, and a test showing your integration uses it.
  2. Context bleed between tenants. Retrieval pulls a document belonging to another customer into the prompt. Closing it means tenant identity enforced inside the retrieval query, at the storage layer. The evidence is a test that attempts a cross-tenant retrieval and fails.
  3. Logs and traces. Prompts and outputs land in observability tooling with a different retention policy and a wider access list than your database. Closing it means redaction before the log write and a retention window set deliberately for the log store. The evidence is a sample log line and the access list.
  4. Third-party integrations. Data reaches an analytics vendor, an evaluation service, or an internal tool that nobody listed. Closing it means an inventory generated from the code rather than maintained by hand. The evidence is that inventory reconciled against the contractual subprocessor documentation.

The enterprise LLM security question underneath all four is identical: can you demonstrate the path is closed, today, without asking an engineer to remember? A claim that you do not train on customer data carries weight only when a configuration and a test stand behind it.

The artifacts a reviewer expects to receive

A security review closes on documents, and each document has to describe something the system actually does.

Artifact What it proves

Where it comes from

Model card What the model does Model registry
Model lineage Origin of each model version Training pipeline
Invocation logging Which version served which call Runtime logs
Data flow diagram Every path data takes Architecture review
Subprocessor list Who else receives data Vendor inventory
DPA Processing terms and obligations Legal, per vendor
Log retention policy Logs outlive the question Infrastructure config
Incident response procedure AI failures have an owner Security runbook
Penetration test report Someone tried to break it External tester
SBOM / AIBOM Software and model components Build pipeline, where applicable
Training data provenance Origin of training data Data pipeline, where you train or fine-tune yourself
Data residency statement Processing location per region Deployment config

Two rows in that table are frequently collapsed into one, and reviewers separate them.

Model lineage describes where a version came from: the base model, what it was fine-tuned on, what changed against the previous version. It reconstructs later only if the training pipeline recorded its own inputs at the time.

Invocation logging ties a single call to a model version, a timestamp, and a request identifier. Nothing recreates it retroactively. A team that starts logging today has records from today, and last quarter stays unaccounted for.

The rest of the table splits into two kinds. An artifact generated by the system stays accurate between questionnaires. An artifact assembled by hand before a review is stale the week after you send it, and the next buyer receives a document that no longer matches the product.

This is where an external AI security assessment usually starts: someone reads the architecture, maps it against the artifact list, and names what the system can already produce.

Subprocessors and DPAs bring third-party obligations that run past this article. GDPR-Compliant AI: What Engineers Need to Log, Trace, and Explain covers the logging and traceability side in engineering detail.

Artifacts an AI security review expects.

 

 

What kind of testing does the questionnaire actually accept?

An AI feature introduces failure modes that a standard application penetration test was never built to reach, and reviewers now ask for that distinction explicitly.

A conventional test probes your infrastructure, your authentication, and your application logic. LLM penetration testing probes four surfaces that a conventional test has no reason to touch:

  • The input, for instructions embedded in user content
  • The context, for data pulled in by retrieval that the user should never see
  • The tools, for actions the model can trigger beyond its intended scope
  • The output, for content that leaks system instructions or other customers’ data

OWASP published the 2026 edition of the Top 10 for LLM Applications on August 4, 2026, with prompt injection at LLM01, the first position.

An analysis by members of the OWASP GenAI working group describes the 2026 ranking as blending expert voting with incident data at weights of 0.75 and 0.25, against a corpus of 6,639 classified incidents. It also notes that prompt injection holds its position on expert judgment more than on recorded incident volume — useful to know when a reviewer asks why you prioritized it.

AI red teaming appears as its own questionnaire line, separate from penetration testing. Reviewers treat the two differently:

  • A penetration test produces findings against a defined scope, with severities and remediation status
  • A red team exercise produces a narrative of what an adversary achieved, and the interesting part is usually the path through the system

Which one satisfies the questionnaire depends on the buyer’s wording and evidence requirements. A statement that the feature was “tested internally” leaves the reviewer without a comparable artifact.

A readiness sequence that survives the next questionnaire

Readiness is not a document set assembled before a deal; it is the order in which the system is built so the documents fall out of it.

Securing enterprise AI in a company without a security function works when the sequence follows the data.

  1. Map the data path end to end. From the user input through the prompt, the retrieval layer, the model provider, the logs, and every store that keeps a copy. Most teams find a path they had forgotten, and forgotten paths are where findings concentrate.
  2. Fix the tenant boundary and the retention windows in configuration. Set an explicit retention period per store — prompts, logs, vector storage, backups — and enforce each one in configuration. Retention that lives in a policy document and nowhere in the infrastructure fails on the first follow-up question.
  3. Turn on invocation logging and model versioning. Model version, request identifier, timestamp, and a reference to the input. The raw content stays out of the log.
  4. Generate lineage and runtime evidence from what the system already writes. Keep the model card synchronized with the model registry and evaluation records.
  5. Run an LLM penetration test and write the AI incident procedure. The test produces the artifact reviewers ask for. The procedure answers the question that follows it.

In practice, teams that follow this sequence find the documentation much easier, because the evidence is already being generated. Starting from the documents tends to produce a description of a system nobody has built yet.

Related questions worth exploring

Wondering what it costs to postpone this work? The commercial case for doing it before a deal forces it sits in AI Governance and Compliance: The Hidden Cost of Waiting Until Your First Enterprise Deal.

Need the SOC 2 side rather than the AI side? SOC 2 Type II Compliance: What Auditors Actually Expect Your System to Do covers what auditors examine and what evidence satisfies them.

Trying to work out whether your system is high-risk under EU law? Classification changes what enterprise AI risk management has to cover, and it runs on its own thresholds. EU AI Act Risk Categories Explained: Is Your AI System High-Risk? walks through the four tiers.

 

 

Sources and verification

Everything in this article was checked on August 31, 2026 against the primary documents, not summaries of them.

The OWASP Top 10 for LLM Applications came out on August 4, 2026. How that ranking was actually put together is described in a separate analysis by members of the OWASP GenAI working group, which is where the weighting and the incident corpus come from.

The questionnaire structures are the Cloud Security Alliance’s AI-CAIQ v1.1, released June 23, 2026, and the SIG questionnaire from Shared Assessments. Where this article refers to frameworks, it means ISO/IEC 42001, the NIST AI Risk Management Framework, and the SOC 2 Trust Services Criteria.

None of this is legal advice.

Share this post:

Subscribe to our blog

Frequently Asked Questions

We already passed SOC 2 Type II. Why is the buyer asking for more?

Because the report describes a defined system over a defined examination period, and the AI section asks about model behavior today. Two gaps come up most:

  • The AI feature shipped after the examination period, or falls outside the system description, so the report says nothing about it
  • The Trust Services Criteria examine the controls operated around a system, and the buyer is asking what the model does with a specific input

AI SOC 2 compliance, as buyers use the phrase, means the AI feature sat inside the audit scope and the controls described actually reach it.

Our AI feature is just an API call to OpenAI. Do these questions still apply to us?

Yes, and the routing makes some of them sharper. If the provider processes buyer data on your behalf, expect it to appear in your vendor documentation with the corresponding agreement.

The questions you now own: what your code sends, what configuration governs training use and retention on your account, and what your own logs keep afterward.

The buyer asked for ISO/IEC 42001. Do we need to get certified to close this deal?

Often not for the current deal. Some buyers accept a documented plan with a target date; others require certification or another form of assurance.

Certification takes months and requires a management system that exists before the audit. If several buyers have now asked, treat it as a planning question for your next two quarters.

We have no security team. Who is supposed to own the answers internally?

In companies at this stage the answers live with engineering, because that is where the evidence is generated. What usually goes wrong is ownership of the response process itself.

  • Engineering owns the technical answers and the artifacts
  • One named person owns the questionnaire, its deadlines, and the reusable answer library
  • Legal owns the DPA and the subprocessor agreements

The failure mode is a questionnaire that circulates without an owner while the deal ages.

Can we answer "we do not train on customer data" if our vendor's default setting handles it?

The statement is accurate and incomplete. Reviewers ask how the setting is enforced, and a default is a property of an account that somebody can change.

A stronger answer names the configuration, points to the provider’s terms, and describes the code path that asserts it on every request. That converts a claim into something the reviewer can verify.

The questionnaire arrived with 40 AI questions and a two-week deadline. What do we do first?

Sort the questions into three groups before writing anything: answers you can support with an existing artifact, answers that are true but undocumented, and questions where the system does not yet do what is being asked.

Send the first group immediately. Document the second, since it is usually the largest and the fastest to close. For the third, give the buyer a dated remediation plan. A specific plan with owners is treated far better than silence or an optimistic answer that unravels on the follow-up call.

Andrii Svyrydov

Founder / CEO / Solution Architect

Have more questions or just curious about future possibilities?