
TL;DR:
- An enterprise security review does not ask whether your company is secure. It asks what your AI does with the buyer’s data, whether that AI is one feature in your product or the product itself, and it expects the answer to come out of the system.
- Enterprise questionnaires increasingly carry a separate AI section, and it asks about the model, the data path, and retention behavior rather than about your company’s controls
- The wording varies, but the same themes recur: data flow, training use, retention, isolation, human oversight, incident handling
- The artifact list is short and specific, and the records of what the model did are what teams most often cannot produce
- OWASP published the 2026 LLM Top 10 in August 2026, and prompt injection stayed at number one
- Most of these are architectural decisions. By the time the questionnaire arrives, they have already been made
The deal stopped at a question about your model
An enterprise security review is the point where a buyer’s risk team decides whether your product may touch their data, and AI, whether it’s one feature or the whole product, turns that review into a separate line of questioning.
The sequence is familiar to anyone who has sold into an enterprise account. Procurement moves, legal moves, and then a vendor security review lands in your inbox as a spreadsheet: the questionnaire a buyer’s risk team uses to vet everything from access control to incident response. Near the end of it there is a block headed “AI and machine learning.” Your SOC 2 report is attached to the reply already. It covers most of the spreadsheet. That block is where it runs out, and the deal stops there while your team tries to work out who owns the response.
What the reviewer is examining has narrowed. Enterprise AI security review questions target the behavior of one functional block in your product, and the boundary of that examination runs along the buyer’s data. By this point reviewers usually assume the organizational controls have already been checked.
For many enterprise buyers, AI is no longer evaluated as an isolated capability. They are assessing whether adding it changes their own data risk profile, which is why questions about model providers, retention, tenant isolation and auditability now sit alongside the traditional security controls.
Many of the enterprise AI security best practices published today describe governance processes. The questionnaire asks for evidence produced by a running system, so that is where this article starts.
Where SOC 2 stops and the AI section begins
SOC 2 attests that your organization operates the controls it described, and the AI section of a questionnaire asks a separate question about the model itself.
The Trust Services Criteria cover access management, change management, monitoring, and the rest of your control environment. A Type II report covers a defined system description over a defined examination period. If your AI feature shipped after that period, or falls outside the system description, the report says nothing about it, and reviewers check the scope section for exactly this.
Scope is the first gap. The second one is structural. The Trust Services Criteria examine the controls a service organization operates around its systems. A model’s behavior is shaped by its training data and by the context it receives at inference, so evidence about the surrounding control environment says little about what the model does with a specific input.
Two frameworks have filled part of that space, and buyers increasingly ask about both:
- ISO/IEC 42001 — the certifiable management system standard for AI. Buyers ask whether you hold it, or when you plan to
- NIST AI RMF — a voluntary risk framework, used as the vocabulary risk teams write their questions in
Neither replaces SOC 2. They sit beside it and answer what it was never scoped to answer. If you need the SOC 2 side in depth, SOC 2 Type II Compliance: What Auditors Actually Expect Your System to Do covers it.
What does an enterprise security team ask about your AI feature?
The AI section of a vendor questionnaire is narrower than it looks. Across the questionnaire sources we reviewed, the same six themes recur even when the wording changes, and each one expects an answer a system can produce.
The Cloud Security Alliance released AI-CAIQ v1.1 on June 23, 2026; the AI sections of the SIG questionnaire from Shared Assessments use a different taxonomy and surface the same concerns.
One detail in AI-CAIQ v1.1 explains the dynamic. Alongside control specifications and self-assessment questions, it carries a separate set of justification questions — a section for evidencing each answer. The evidence request is built into the instrument.

Data path
Where the buyer’s data physically goes once it enters your AI feature.
- “Does customer data leave our tenant boundary at any point during inference?”
- “Which subprocessors receive this data, and under what agreement?”
- “Where is the data physically processed?”
The third question catches teams who route inference through a model provider in a different jurisdiction without recording that in the subprocessor list.
Training use
Whether the buyer’s data improves your product for someone else.
- “Is our data used to train, fine-tune, or improve your models?”
- “If opt-out exists, is it default or on request, and how is it enforced technically?”
The second half of that question is where prepared answers fall apart. A policy statement is not enforcement, and reviewers now ask which configuration setting carries it.
Retention and deletion
How long the data stays, and what deletion actually removes.
- “How long are prompts and outputs retained?”
- “Does deletion of our account remove the data from model context, caches, and logs?”
Deletion questions have grown teeth because the answer touches four stores that rarely share a retention policy: the primary database, the vector store, the inference logs, and whatever your observability vendor holds.
Isolation and access
Whether one customer can reach another customer’s data, and who inside your company can read any of it.
- “Can one tenant’s data appear in another tenant’s output?”
- “Who inside your company can read prompts submitted by our users?”
The second question is about your own staff. Support tooling that surfaces raw prompts for debugging is a common finding, and it is usually discovered during the review by the vendor themselves.
Oversight and output control
What happens when the model is wrong.
- “What happens when the model produces a wrong or harmful output?”
- “Is a human in the loop, and at which step?”
For some use cases “no human in the loop” is a complete answer on its own. What reviewers follow up on is the detection: what catches a bad output when nobody is watching.
Testing and incident handling
Whether anyone has tried to break the feature on purpose.
- “Have you tested this feature against adversarial input?”
- “What is your procedure when an AI-specific incident occurs?”
An incident response plan that covers infrastructure and says nothing about model behavior reads as an incomplete answer, because the failure modes are different.
What the risk team is doing with these answers. Most of these questions end up as requests for evidence, and the justification section of the questionnaire is where that evidence goes. The reviewer checks the artifact behind the sentence, and an answer arriving without one carries little weight in the scoring. It is why an AI security questionnaire rewards a mediocre diagram over confident prose.
The answers come from two layers of the architecture
Every question above resolves to one of two places: the data governance layer that controls what enters and leaves the model, and the model governance layer that records what the model did.
Enterprise AI governance is usually presented as policy work owned by a committee. In a product company with no security function, it is closer to system design, and it lives in two layers of the AI Compliance Stack: data governance and model governance.

Data governance. This layer decides what the model is allowed to see and what survives afterward.
- A tenant boundary enforced at the query, so isolation is a property of the data path
- Input classification that tags what is sensitive before it reaches the prompt
- Retention expressed as configuration, with an explicit period set for each store
- Training opt-out enforced by a provider setting, where one exists, asserted by your code on every call
- Processing region pinned in the deployment configuration, per customer where the contract requires it
Model governance. This layer records what happened, so questions about a specific decision have an answer months later.
- Invocation logging that captures the request, the model version, and the outcome
- Model versioning with a recorded rollback path
- Human review, where it exists, recorded as a workflow state so it stays queryable
- Output monitoring that fires on the conditions the buyer asked about
Log the minimum needed to reconstruct the event. Where prompts or outputs carry customer data, use references, redaction or controlled access instead of storing raw content.
We refer to this approach as Compliance-Native Architecture: designing the data path and the evidence model so buyer requirements are enforced by the system itself. None of these questions should come as a surprise.
A team that reads them before designing the data path builds an enterprise AI security architecture that answers them by construction, and a team that reads them afterward pays for a migration.
Classification under the EU AI Act runs on its own thresholds. The questionnaire themes described here stay the same, and high-risk classification adds obligations on top of them.
Data leakage is the question reviewers keep returning to
Across every questionnaire theme, one concern repeats in different wording: whether the buyer’s data can surface somewhere the buyer did not authorize.
AI data leakage prevention comes down to four common paths and the evidence that each one is closed. Each path fails differently, and each needs its own proof.
- Training ingestion. The model provider uses submitted content to improve its models. Provider capabilities differ, so closing this path means confirming which controls your provider offers and configuring them. The evidence is the provider’s terms, the account configuration, and a test showing your integration uses it.
- Context bleed between tenants. Retrieval pulls a document belonging to another customer into the prompt. Closing it means tenant identity enforced inside the retrieval query, at the storage layer. The evidence is a test that attempts a cross-tenant retrieval and fails.
- Logs and traces. Prompts and outputs land in observability tooling with a different retention policy and a wider access list than your database. Closing it means redaction before the log write and a retention window set deliberately for the log store. The evidence is a sample log line and the access list.
- Third-party integrations. Data reaches an analytics vendor, an evaluation service, or an internal tool that nobody listed. Closing it means an inventory generated from the code rather than maintained by hand. The evidence is that inventory reconciled against the contractual subprocessor documentation.
The enterprise LLM security question underneath all four is identical: can you demonstrate the path is closed, today, without asking an engineer to remember? A claim that you do not train on customer data carries weight only when a configuration and a test stand behind it.
The artifacts a reviewer expects to receive
A security review closes on documents, and each document has to describe something the system actually does.
| Artifact | What it proves |
Where it comes from |
| Model card | What the model does | Model registry |
| Model lineage | Origin of each model version | Training pipeline |
| Invocation logging | Which version served which call | Runtime logs |
| Data flow diagram | Every path data takes | Architecture review |
| Subprocessor list | Who else receives data | Vendor inventory |
| DPA | Processing terms and obligations | Legal, per vendor |
| Log retention policy | Logs outlive the question | Infrastructure config |
| Incident response procedure | AI failures have an owner | Security runbook |
| Penetration test report | Someone tried to break it | External tester |
| SBOM / AIBOM | Software and model components | Build pipeline, where applicable |
| Training data provenance | Origin of training data | Data pipeline, where you train or fine-tune yourself |
| Data residency statement | Processing location per region | Deployment config |
Two rows in that table are frequently collapsed into one, and reviewers separate them.
Model lineage describes where a version came from: the base model, what it was fine-tuned on, what changed against the previous version. It reconstructs later only if the training pipeline recorded its own inputs at the time.
Invocation logging ties a single call to a model version, a timestamp, and a request identifier. Nothing recreates it retroactively. A team that starts logging today has records from today, and last quarter stays unaccounted for.
The rest of the table splits into two kinds. An artifact generated by the system stays accurate between questionnaires. An artifact assembled by hand before a review is stale the week after you send it, and the next buyer receives a document that no longer matches the product.
This is where an external AI security assessment usually starts: someone reads the architecture, maps it against the artifact list, and names what the system can already produce.
Subprocessors and DPAs bring third-party obligations that run past this article. GDPR-Compliant AI: What Engineers Need to Log, Trace, and Explain covers the logging and traceability side in engineering detail.

What kind of testing does the questionnaire actually accept?
An AI feature introduces failure modes that a standard application penetration test was never built to reach, and reviewers now ask for that distinction explicitly.
A conventional test probes your infrastructure, your authentication, and your application logic. LLM penetration testing probes four surfaces that a conventional test has no reason to touch:
- The input, for instructions embedded in user content
- The context, for data pulled in by retrieval that the user should never see
- The tools, for actions the model can trigger beyond its intended scope
- The output, for content that leaks system instructions or other customers’ data
OWASP published the 2026 edition of the Top 10 for LLM Applications on August 4, 2026, with prompt injection at LLM01, the first position.
An analysis by members of the OWASP GenAI working group describes the 2026 ranking as blending expert voting with incident data at weights of 0.75 and 0.25, against a corpus of 6,639 classified incidents. It also notes that prompt injection holds its position on expert judgment more than on recorded incident volume — useful to know when a reviewer asks why you prioritized it.
AI red teaming appears as its own questionnaire line, separate from penetration testing. Reviewers treat the two differently:
- A penetration test produces findings against a defined scope, with severities and remediation status
- A red team exercise produces a narrative of what an adversary achieved, and the interesting part is usually the path through the system
Which one satisfies the questionnaire depends on the buyer’s wording and evidence requirements. A statement that the feature was “tested internally” leaves the reviewer without a comparable artifact.
A readiness sequence that survives the next questionnaire
Readiness is not a document set assembled before a deal; it is the order in which the system is built so the documents fall out of it.
Securing enterprise AI in a company without a security function works when the sequence follows the data.
- Map the data path end to end. From the user input through the prompt, the retrieval layer, the model provider, the logs, and every store that keeps a copy. Most teams find a path they had forgotten, and forgotten paths are where findings concentrate.
- Fix the tenant boundary and the retention windows in configuration. Set an explicit retention period per store — prompts, logs, vector storage, backups — and enforce each one in configuration. Retention that lives in a policy document and nowhere in the infrastructure fails on the first follow-up question.
- Turn on invocation logging and model versioning. Model version, request identifier, timestamp, and a reference to the input. The raw content stays out of the log.
- Generate lineage and runtime evidence from what the system already writes. Keep the model card synchronized with the model registry and evaluation records.
- Run an LLM penetration test and write the AI incident procedure. The test produces the artifact reviewers ask for. The procedure answers the question that follows it.
In practice, teams that follow this sequence find the documentation much easier, because the evidence is already being generated. Starting from the documents tends to produce a description of a system nobody has built yet.
Related questions worth exploring
Wondering what it costs to postpone this work? The commercial case for doing it before a deal forces it sits in AI Governance and Compliance: The Hidden Cost of Waiting Until Your First Enterprise Deal.
Need the SOC 2 side rather than the AI side? SOC 2 Type II Compliance: What Auditors Actually Expect Your System to Do covers what auditors examine and what evidence satisfies them.
Trying to work out whether your system is high-risk under EU law? Classification changes what enterprise AI risk management has to cover, and it runs on its own thresholds. EU AI Act Risk Categories Explained: Is Your AI System High-Risk? walks through the four tiers.
Sources and verification
Everything in this article was checked on August 31, 2026 against the primary documents, not summaries of them.
The OWASP Top 10 for LLM Applications came out on August 4, 2026. How that ranking was actually put together is described in a separate analysis by members of the OWASP GenAI working group, which is where the weighting and the incident corpus come from.
The questionnaire structures are the Cloud Security Alliance’s AI-CAIQ v1.1, released June 23, 2026, and the SIG questionnaire from Shared Assessments. Where this article refers to frameworks, it means ISO/IEC 42001, the NIST AI Risk Management Framework, and the SOC 2 Trust Services Criteria.
None of this is legal advice.