Skip to main content

    Vendor Risk

    AI Vendor Risk Assessment: The Questions to Ask Before You Buy — and the Evidence Each Answer Needs

    Most AI capability arrives through a vendor: a model API, a feature switched on inside a tool you already pay for, or a product that is itself an agent. The vendor's answers to a security questionnaire are self-reported evidence. An AI vendor risk assessment is the work of deciding which answers you need more than self-report for, what evidence would satisfy you, and what you will do if the vendor cannot provide it.

    This article maps the questions to ask before you buy to the evidence each answer needs, across six areas: data, access, model changes, sub-providers, incidents and exit. A fictional worked review shows the mapping in use.

    By
    AuditPartners
    Added
    Reading time
    10 min
    Status
    Educational guide

    Why a questionnaire is a starting point, not an assessment

    A completed questionnaire tells you what the vendor says. For low-risk uses, self-report is potentially proportionate after a risk review, not automatically sufficient. For a system that will see customer data, make or influence decisions about people, or act in your environment, the NIST AI RMF 1.0 Map function expects you to understand the system's context, its components and who is responsible for which risks, including third-party components. You cannot do that from assertions alone.

    The practical move is to attach an evidence level to each question before sending it. For each answer, decide whether you will accept self-report (the vendor's statement), require documents (a policy, a report, a contract clause), require inspection (you or an independent party look at the configuration, the data flow or the report detail), or require a test (you exercise the control yourself in a trial tenant). The same four-level distinction runs through AI Technical Controls Assessment for internally built systems.

    The question-to-evidence map

    The table lists the questions we ask most often, the minimum evidence that supports an answer for a system that handles personal or confidential data, and what a weak answer looks like. Scale the evidence level down for low-risk uses and up for systems that make consequential decisions.

    AI vendor questions mapped to the evidence each answer needs
    AreaQuestionMinimum evidence for a data-handling systemWeak answer to watch for
    DataIs our data (prompts, files, outputs, feedback) used to train or improve any model, for us or for other customers? Can that be switched off, and is the default off for our tier?Explicit contractual training-use terms covering the vendor and appropriate upstream sub-processor/model-provider commitments; product documentation showing the setting; inspection of the setting in your tenant. Zero retention alone does not establish no training use."We take privacy seriously" with no contractual term; training opt-out only on an enterprise tier you are not buying.
    DataWhere is our data processed and stored, for how long, and is it logged by the vendor or its model provider?Documented regions and retention periods; the model provider's retention terms if the vendor resells a third-party model.A region list for the vendor's own servers that is silent about the upstream model provider.
    AccessWhich vendor staff can access our data or our tenant, under what approval, and is that access logged and available to us?Access-control policy; an independent report covering logical access (for example a SOC 2 Type 2 report from a CPA firm, read in full, not the cover letter); sample of access logs if offered.A bridge letter or a badge image instead of the report; 'engineers may access for support' with no approval or log.
    AccessWhat permissions does the product need in our environment, and can we scope them down?Permission manifest or integration documentation; a trial-tenant test of the minimum scope that still works.Requests for admin or organisation-wide scopes with no granular option.
    Model changesHow will we know when the underlying model, prompt or safety configuration changes, and can we pin a version or test before adoption?Release-notes policy; a changelog you can inspect; contractual notice period; evidence of a previous notification."We continuously improve the model" with no notice mechanism; no way to reproduce last month's behaviour.
    Model changesWhat evaluation does the vendor perform before a model change reaches customers, and will they share the results relevant to our use?Description of the evaluation process; a redacted evaluation summary for the most recent change; the NIST Generative AI Profile lists the kinds of risks such evaluations should cover.Marketing benchmarks with no description of method or of failure cases.
    Sub-providersWhich model providers, hosting providers and other sub-processors sit behind the product, and how are we told when they change?Current sub-processor list; the vendor's own due-diligence summary for the model provider; notice terms.A list that omits the model provider, or a list with no date.
    Sub-providersIf the upstream model provider changes its terms or withdraws a model, what is the vendor's contingency?A documented contingency (alternative model, migration path) and evidence it has been exercised at least once."That has never happened."
    IncidentsHow are security and AI-behaviour incidents (data leakage, harmful output, unauthorised action) detected, and how and when will we be notified?Incident-response policy; notification clause with a time-bound commitment; a sanitised example of a past notification.Notification 'without undue delay' with no definition and no example.
    IncidentsCan we obtain logs of the product's actions in our tenant sufficient to investigate an incident ourselves?Log schema and retention documentation; export tested in a trial tenant.Logs available 'on request to support' only.
    ExitIf we leave, how do we export our data and configurations, in what format, and when is our data deleted from the vendor and its sub-providers?Export documentation; deletion commitment with timeline and certificate; sub-provider deletion flow-down in the contract.Export limited to the UI; deletion that excludes backups or the model provider's logs.
    ExitWhat happens to work in progress if the product is withdrawn or the vendor fails?Continuity or escrow terms proportionate to the dependency; a tested export.No term, on the assumption the vendor will always exist.
    'Minimum evidence' is a judgement for a typical data-handling system and is not a legal or regulatory requirement. A SOC 2 report is an independent CPA firm's examination against the AICPA Trust Services Criteria; read the scope, the period, the carve-outs and the exceptions, because the report describes the vendor's system as the vendor defined it.

    A fictional worked review

    Marlowe Mutual's review of Clauseworks: answers, evidence obtained and outcome
    AreaVendor answer (self-report)Evidence obtainedAssessment
    Data: training use"Customer data is never used for training."Data-processing addendum: covers the vendor only. Upstream model provider's terms (inspected): zero-retention is available via an enterprise API flag, but explicit upstream training-use commitments were not obtained. Vendor's configuration (inspected via a screenshot of their console, not independently verified): flag enabled.No-training claim not established: zero retention alone does not prove no training use. Requirement: explicit contractual training-use terms covering the vendor and appropriate upstream sub-processor/model-provider commitments, plus flow-down maintaining the retention setting and notice of changes.
    Access"Access is restricted and audited."SOC 2 Type 2 report, period to June 2026, read in full: logical access controls tested with one exception (a departed contractor's access removed 19 days late). Vendor support access to customer tenants is not within the report's scope.Supported for the vendor's corporate environment; not supported for tenant access. Requirement: tenant access approval and logging, with Marlowe able to see the log.
    Model changes"We notify customers of major model changes."Changelog inspected: three model changes in twelve months, two announced after the fact. No version pinning.Not supported. Requirement: 30-day notice and a test period for model changes affecting clause classification; or accept the risk with a documented monthly re-test by Marlowe's legal operations team.
    Sub-providers"We use a leading model provider and major cloud host."Sub-processor list obtained; model provider named; list dated. Contingency: vendor states a second provider is 'supported' but has never switched in production.Supported for transparency; contingency untested. Accept with a contractual commitment to disclose provider changes.
    Incidents"Customers are notified without undue delay."Incident policy obtained; no defined notification window; no example notification available.Not supported. Requirement: a 72-hour notification clause and a tested log export; Marlowe tested the export in the trial tenant and it worked.
    Exit"Full export available at any time."Export tested in trial tenant: contracts and flags exported; redline history and custom clause libraries not included. Deletion: 30 days, backups 90 days, sub-provider deletion not addressed.Partially supported. Requirement: export of clause libraries; deletion flow-down to the model provider's logs or confirmation of zero retention.

    The outcome was not a yes or a no. Marlowe proceeded subject to contractual requirements such as explicit training-use terms and upstream commitments, tenant-access logging, incident notification and exit provisions, with an accepted risk and compensating control owned by legal operations, and an entry in its AI system inventory recording the system, its data categories, the sub-providers and a review date six months out. The decision and its basis are now something an internal auditor can test later; see ISO 42001 Internal Audit: An Evidence Request Checklist, request 15.

    Scaling the effort, and what a pre-purchase review cannot do

    Not every vendor deserves this depth. An illustrative tiering approach: self-report is potentially proportionate after a risk review for low-risk uses where the product sees no personal or confidential data and takes no actions; documents and a read of any independent report where it sees data; inspection and tests where it takes actions in your environment or influences decisions about people. Record the tier and the reason in the inventory so the decision can be revisited.

    If you want help structuring a review like Marlowe's, or an independent read of a vendor's evidence, that falls within our AI Technical Controls Assessment work. If you are earlier than that, the free AI Governance QuickScan will tell you whether you have an inventory and owners to attach vendor reviews to in the first place.

    Key takeaways

    • Attach an evidence level (self-report, documents, inspection, test) to each question before you send the questionnaire.
    • The six areas that most often decide the outcome are data use, access, model change notice, sub-providers, incident notification and exit.
    • Read independent reports in full: scope, period, carve-outs and exceptions, and whether tenant access is even in scope.
    • The output is a decision with conditions, accepted risks and owners, recorded in the AI system inventory with a review date.
    • A pre-purchase review is point-in-time; model changes and agent oversight need their own ongoing tests.
    Related serviceAI Technical Controls AssessmentRelated resourceAI System Inventory TemplateFree toolAI Governance QuickScan

    References

    Primary sources cited above. Standards are referenced by number and year; their text is licensed and is paraphrased, not reproduced.

    1. NIST AI 100-1, AI Risk Management Framework (AI RMF 1.0), January 2023 (opens in a new tab) — Map function, including third-party components and responsibilities.
    2. NIST AI 600-1, AI RMF: Generative Artificial Intelligence Profile, July 2024 (opens in a new tab) — Risk categories relevant to evaluating model changes and supplier models.
    3. NIST AI Risk Management Framework programme page (opens in a new tab)
    4. AICPA, 2017 Trust Services Criteria with Revised Points of Focus (2022) (opens in a new tab) — Criteria a SOC 2 examination is performed against.
    5. AICPA, System and Organization Controls (SOC) suite of services (opens in a new tab)
    6. ISO/IEC 42001:2023, Artificial intelligence management system — Requirements (opens in a new tab) — Licensed text; supplier-related control objectives referenced, not reproduced.

    Need this applied to your environment?

    A scoped assessment follows the agreed criteria, procedures and evidence described here. Findings are reviewed and signed off by accountable professionals.