Skip to main content

    AI Controls

    AI Technical Controls Assessment: Following One Change from Work Ticket to Production

    Most AI governance reviews stop at policy: is there an acceptable-use statement, is there a steering committee, is there an inventory. An AI technical controls assessment asks a narrower and harder question: for one change that an AI tool helped produce, can you show who asked for it, who approved it, what was reviewed, what was tested, how it reached production, and what the system could touch once it got there?

    This article follows a single fictional change from work ticket to production, names the control that should exist at each step, shows the evidence that would demonstrate it, and explains why a control that is well designed can still fail in operation.

    By
    AuditPartners
    Added
    Reading time
    11 min
    Status
    Educational guide

    Why trace one change instead of reviewing the programme

    Policies describe intent. A change record describes what happened. When an engineer uses a coding assistant or an autonomous agent to produce a change, the organisation's exposure is concentrated in the delivery pipeline: the ticket that scoped the work, the review that read the diff, the tests that ran, the release that deployed it and the identities and connectors the deployed code can use. Tracing one change through that pipeline is the fastest way to learn whether the controls on paper are the controls in practice.

    This is consistent with how the NIST AI Risk Management Framework 1.0 organises risk work: Govern sets accountability, Map establishes context and what the system does, Measure collects evidence about behaviour, and Manage acts on it. A ticket-to-production trace is a Measure activity that feeds Manage. It is also the approach behind our AI Technical Controls Assessment service and the free AIGC Technical AI Controls Checklist, which lists the same checkpoints in a form a team can self-complete before any external review.

    The fictional change we will follow

    The exact revision matters. An assessment that reasons about "the payout service" in general cannot tell you whether this change was reviewed; an assessment anchored to commit `9f3c2e1` can. Every evidence request in the table below therefore names the ticket, the commit, the pipeline run or the identity.

    Six control points and the evidence at each

    The table distinguishes three kinds of evidence. Self-reported means someone told us. Inspected means we looked at a record or configuration ourselves; inspection of reliable operating records can support operating effectiveness, while configuration alone generally supports design at that point. Tested means we re-performed the control or observed it operate on a sample. Conclusions depend on the reliability and coverage of inspected and tested evidence; self-report tells you where to look. Any permission-denial exercise is performed only in an approved isolated sandbox with synthetic data and recovery controls, never by attempting access to unrelated real data or deploying to production.

    Control points for PAY-4312, with the evidence that supports each level of assurance
    Control pointWhat should be trueSelf-reportedInspectedTested
    1. Approved scopeThe work was requested and authorised before code was written, and the change stayed inside that scope.Engineer says the PM approved it in stand-up.Ticket PAY-4312 shows a requester, an approver with a timestamp before the first commit, and acceptance criteria that match the diff.Compare the diff of `9f3c2e1` to the acceptance criteria; any file outside the scope (here, the connector change) is an exception to explain.
    2. Peer reviewA second person who did not author the change read it and had the authority to block it."We always do code review."Pull request shows a reviewer other than the author, approval recorded before merge, and branch protection requiring one approval.Attempt a merge without approval only in an approved isolated sandbox repository with synthetic data, recovery controls and equivalent protection rules; read review comments to confirm the reviewer engaged with the idempotency logic, not just the formatting.
    3. TestsAutomated tests covering the new behaviour ran and passed on the revision that shipped."CI is green."Pipeline run #2087 is linked to `9f3c2e1`, shows test stage results and includes a test for duplicate-payout prevention.Re-run the test suite on `9f3c2e1`; introduce a deliberate duplicate webhook only in an approved isolated sandbox with synthetic data and recovery controls, and confirm the idempotency key rejects it.
    4. ReleaseOnly the reviewed and tested artefact reached production, through a controlled path."Deploys go through the pipeline."Deployment log shows the artefact digest from run #2087 deployed to production; no manual deploy permissions on the production cluster for engineers.Compare the running container's image digest to the pipeline artefact; attempt a manual deploy only in an approved isolated sandbox with synthetic data and recovery controls, using a test identity with equivalent engineer permissions, and confirm denial.
    5. Role-based accessThe deployed service and the humans around it have the least privilege the change needs."The worker only touches payouts."IAM policy for `payout-worker` lists its permissions; access review record shows the policy was reviewed in the last cycle.Enumerate effective permissions for `payout-worker`; attempt a read of a synthetic out-of-scope table only in an approved isolated sandbox with synthetic data and recovery controls, using a test identity with equivalent permissions, and confirm denial. Do not access unrelated real data.
    6. Connectors and secretsNew integrations were approved, their credentials are managed, and their blast radius is understood."The ledger connector was already there."Connector inventory lists the ledger write connector, its owner and approval date; the secret is stored in the secrets manager with rotation enabled.Confirm the connector credential in production matches the managed secret and that the change in `9f3c2e1` that widened write scope has an approval that post-dates the request.
    The 'what should be true' column is written for this scenario, not copied from any standard. Map it to your own criteria (for example the relevant NIST AI RMF subcategories or ISO/IEC 42001:2023 Annex A objectives) before relying on it.

    Design effectiveness versus operating effectiveness

    A control is designed effectively when, if it operates as described, it would prevent or detect the risk it addresses. A control is operating effectively when evidence shows it actually operated that way over the period and population that matter. The distinction decides what you can claim.

    In our scenario, branch protection requiring one approval is a well-designed peer-review control. Inspection shows it is enabled. But when the assessor reads the review on PAY-4312, the approval came 40 seconds after the pull request was opened and contains no comments, and the reviewer is the engineer's pairing partner who co-wrote the prompt that produced the code. The control existed and technically operated; whether it provided an independent second look is a judgement the assessor must make and document, not a box to tick.

    What AI-generated changes alter about the picture

    The control points above predate AI coding tools; they are ordinary software change controls. AI changes the pressure on them in four ways.

    • Volume and plausibility. Generated code arrives faster and reads cleanly, which makes a 40-second approval more likely, not less. Review evidence needs to show engagement with the logic, not presence of a click.
    • Scope creep inside a single change. Assistants and agents readily touch adjacent files. In PAY-4312 the connector scope widened without a separate request. Scope comparison between ticket and diff becomes a primary control rather than a formality.
    • Provenance. If an agent authored the commit, who is the accountable human? Commit metadata, agent run logs and the ticket should agree on a named owner.
    • Tool permissions as a control surface. The coding agent's own credentials (repository write, CI triggers, cloud access) are part of the inventory. The NIST Generative AI Profile (AI 600-1) discusses human-AI configuration and information-security risks that are specific to generative systems; the practical consequence is that the agent's access must be assessed like any other identity.

    Where the agent acts with autonomy at run time rather than only at build time, the oversight questions get their own treatment in How to Test Human Oversight Controls for AI Agents.

    Writing the finding

    A finding from a trace should name the control point, state the condition observed with the identifier, state the criterion it was compared against, explain the consequence and recommend an action with an owner. Avoid verdict words such as "compliant" or "non-compliant" unless the engagement was scoped to give one.

    If you want to rehearse this before an external review, the AIGC Technical AI Controls Checklist walks through the same six control points, and the AI Governance QuickScan gives a preliminary view of where the governance side has gaps. Neither is an assessment; both tell you where to start.

    Key takeaways

    • Anchor the assessment to an exact revision, ticket, pipeline run and identity; generalities cannot be verified.
    • Classify every piece of evidence as self-reported, inspected or tested, and only conclude on what was inspected or tested.
    • A control that exists and operated can still fail its purpose; document the judgement, especially for review quality.
    • AI-generated changes raise the stakes on scope comparison, review engagement, provenance and the agent's own permissions.
    • One trace supports only that trace; period conclusions need reliable population-based evidence, with samples and periods selected by risk and scope, not one automated-control test.
    Related serviceAI Technical Controls AssessmentRelated resourceAIGC Technical AI Controls ChecklistFree toolAI Governance QuickScan

    References

    Primary sources cited above. Standards are referenced by number and year; their text is licensed and is paraphrased, not reproduced.

    1. NIST AI 100-1, Artificial Intelligence Risk Management Framework (AI RMF 1.0), January 2023 (opens in a new tab) — Govern, Map, Measure and Manage functions.
    2. NIST AI 600-1, AI RMF: Generative Artificial Intelligence Profile, July 2024 (opens in a new tab) — Risks and suggested actions specific to generative AI.
    3. NIST AI Risk Management Framework programme page (opens in a new tab)
    4. NIST SP 800-53A Rev. 5, Assessing Security and Privacy Controls, January 2022 (opens in a new tab) — Examine, interview and test as assessment methods, the basis for the self-reported / inspected / tested distinction used here.
    5. ISO/IEC 42001:2023, Artificial intelligence management system — Requirements (opens in a new tab) — Licensed text; referenced for Annex A control objectives, not reproduced.

    Need this applied to your environment?

    A scoped assessment follows the agreed criteria, procedures and evidence described here. Findings are reviewed and signed off by accountable professionals.