Audit of AI
What is an audit of AI?
An independent examination of an AI system and everything around it — the people accountable for it, the data feeding it, the instructions governing it, the tools it can reach, the vendors behind it and the context it operates in — to establish, on evidence, whether it can be relied on for its intended use.
Why does AI need independent audit?
Because an AI system can keep operating normally while the assurance around it breaks. A model version changes, a data source is reshaped, a vendor updates something invisible, a human review step stops being recorded — the outputs still look right, and nothing fails loudly. Audit exists to establish whether the control still works, on evidence rather than on the policy that describes it.
What are the different types of AI audit?
Six disciplines, each answering a different question about the same system. An AI System Audit asks whether it performs as specified. An Algorithm Audit asks whether its decisions are fair, explainable and defensible. A Data Audit asks whether the data feeding it is reliable, complete and traceable. An AI Compliance Audit asks whether the system meets the obligations that apply to it. An AI Security Audit asks whether it withstands AI-specific threats such as adversarial input, data poisoning and model theft. A Third-Party AI Tool Audit asks what risk the organisation is inheriting from AI it did not build.
How does a Nakoda AI audit engagement work?
Twenty-four steps in six phases. Enter and define covers client enquiry, qualification, proposal and scope, and kickoff. Understand the environment covers discovery, process mapping, current AI assessment, AI maturity assessment and data and technology assessment. Find what matters covers AI ideation, the use-case catalogue, value and feasibility scoring, risk assessment, the business case and prioritization. Turn findings into a plan covers the AI roadmap, pilot planning, execution and validation, and KPI measurement. Assure and decide covers the final audit report, the executive presentation and decision or sign-off. Continue covers the 30, 60 and 90-day follow-up and implementation or scale.
What does Nakoda do at each stage of an AI audit?
In steps 01 to 04 we qualify the need, define the scope, establish objectives and set the engagement. In 05 to 09 we discover the environment, map the process, assess current AI, evaluate maturity and examine data and technology. In 10 to 15 we identify relevant use cases, catalogue them, score value and feasibility, assess risk, build the business case and prioritize the portfolio. In 16 to 19 we build the roadmap, plan pilots, support validation and define measurement. In 20 to 22 we produce the final audit report, present material findings and support executive decision-making. In 23 to 24 we follow up, re-test where required, and support implementation and scale.
What does an AI audit produce?
An AI System Assurance Report: an executive conclusion on whether the system can be relied on for its intended use, the scope and methodology, the material findings with the evidence behind each, and a remediation plan with owners, target dates and a scheduled re-test. The sign-off records one of five outcomes — accept, remediate, restrict, do not use, or re-test.
Can Nakoda build an internal AI audit function?
Yes. Internal AI Audit Function Design establishes permanent capability rather than repeat external engagements: the methodology adapted to AI evidence, the tooling and data access, training to the level required to audit rather than build, the governance structure that gives the function independence, the reporting line from a finding to the board, and a first cohort of audits run under supervision to validate the method before the team operates on its own.
Which frameworks does the audit practice follow?
ISACA's frameworks — COBIT for governance and management of enterprise IT, CRISC for risk and information systems control, and the CISA standard for information systems auditing — alongside the professional standards of chartered accountancy. Alignment means the methodology, evidence standards, documentation requirements and reporting format follow those standards.
What is the difference between audit of AI and AI in auditing?
Audit of AI examines the AI systems an organisation relies on — the system, its data, controls, people, vendors, risk and accountability — to answer whether it can be trusted for its intended use. AI in auditing examines how an audit function uses AI in its own work: enhanced sampling and coverage, continuous monitoring and anomaly detection, document analysis, AI-enhanced risk assessment for audit planning, and finance audit specialisation — and how those outputs are validated, how evidence is handled and how audit judgement is protected. They are different disciplines and they work better together.
When does an AI system need to be re-audited?
When something material changes. A new model version, a changed data source, a vendor altering its service, a change in volume, scope or regulation. Assurance is a statement about a system at a point in time, so a material change is the trigger to establish whether the earlier conclusion still holds. The engagement schedules 30, 60 and 90-day follow-up for exactly this reason.
What an audit of AI examines
Seven areas surround the AI system, and each carries one question the audit has to answer on evidence.
- People — Who is accountable?
- Decision rights: Who is allowed to decide, and who answers for it. Named owner, delegated authority, escalation path, the reviewer who sees it first.
- Data — Can we trust the information?
- Data lineage: Where did the information come from? Source systems, transformations, retention, the point where provenance stops.
- AI system — Does it behave as intended?
- Control effectiveness: Did the control actually work? Version in production, performance across the input range, monitoring, degradation.
- Instructions — What are we asking it to do?
- Control boundary: What the system has been told it may and may not do. System instructions, guardrails, refusal conditions, who may change them.
- Tools — What can it access?
- Authority: What the system can reach, and what it can do there. Connected systems, write permissions, transaction limits, the stop.
- Vendors — What risk are we inheriting?
- Third-party risk: What we accepted when we bought it. Model provider, contractual accountability, change notice, what we can see.
- Context — When does yesterday's assurance stop applying?
- Change management: What changed since we last looked. Model versions, data sources, volumes, regulation, the trigger to reassess.
The six audit disciplines
- AI System Audit — Does it perform as specified?
- Accuracy across the full input range, consistency, degradation as conditions change, monitoring, edge cases and the human oversight designed to catch failures. An inaccurate system that is confident is more dangerous than one that fails loudly. Control effectiveness: Did the control actually work?
- Algorithm Audit — Are its decisions fair, explainable and defensible?
- Decision logic tested for bias across protected and proxy characteristics, whether a decision can be explained to the person it affects, and whether oversight is proportionate to consequence. Bias scales before anyone notices it. Disparate impact: Does the outcome differ by group without a reason that holds?
- Data Audit — Can we trust what feeds it?
- Quality, integrity, completeness and lineage across every source the system depends on — and the point at which provenance stops being recoverable. Bad data does not always look bad. Data lineage: Can we trace where the information came from?
- AI Compliance Audit — Does it meet the obligations that apply to it?
- The system mapped against its applicable obligations — Dubai AI Seal, UAE PDPL, DIFC data protection, sector rules — with documented evidence of conformance and of every gap. An obligation you have not mapped is one you cannot evidence. Conformance: Which rules apply here, and can we show we meet them?
- AI Security Audit — Can it withstand AI-specific attack?
- Attack surface, detection of adversarial input, training pipeline security, access control around model weights and training data, and monitoring for anomalous production behaviour. Traditional security controls do not see every AI attack. Adversarial input: Input crafted to make the system answer wrongly.
- Third-Party AI Tool Audit — What risk are we inheriting?
- Externally sourced AI evaluated on accuracy, bias, data handling, security posture, vendor governance and contractual accountability, rated as a portfolio. Most organisations use more AI than they built. Third-party risk: What we accepted when we bought it.
The engagement — 24 steps in six phases
01 Enter & define — First, why does this need auditing at all?
We establish what prompted the question and what should actually be examined — before anything is looked at. Audit scope: What is being examined, and what for. Result: Audit opened.
- 01 Client enquiry — Something prompted the question. We start there.
- 02 Qualification — Is an independent audit the right instrument for this?
- 03 Proposal & scope — What will be examined, against what, and what is excluded.
- 04 Kickoff — Owners named, access agreed, the file is opened.
02 Understand the environment — First, we understand what we're actually auditing.
AI rarely exists in isolation. We map the business, the processes, the systems and the environment around it. AI inventory: Everything in scope, and what each part decides. Result: Baseline established.
- 05 Discovery — What the organisation actually does, before any AI is discussed.
- 06 Process mapping — The workflow the AI sits inside, end to end.
- 07 Current AI assessment — Every AI system in scope, and what each one decides.
- 08 AI maturity assessment — How well the organisation governs what it already runs.
- 09 Data + technology assessment — The infrastructure underneath: sources, pipelines, platforms, controls.
03 Find what matters — Now we determine what deserves attention.
An audit that reports two hundred things reports nothing. The question is which findings, risks and opportunities actually matter. Materiality: Would this change what leadership does? Result: Priorities established.
- 10 AI ideation — Where AI could matter, and where it demonstrably could not.
- 11 Use-case catalogue — Every candidate written down before any is judged.
- 12 Value + feasibility scoring — What it is worth, and whether it can actually be done here.
- 13 Risk assessment — What could go wrong, how badly, and who carries it.
- 14 Business case — The cost, the return, and the assumptions both rest on.
- 15 Prioritization — The short list leadership can actually act on.
04 Turn findings into a plan — A finding is useful only when it changes what happens next.
The priorities move into lanes with owners and dates, and the first of them is proven on a pilot before anything scales. Remediation: What has to change before the control can be relied on. Result: Plan validated.
- 16 AI roadmap — Sequenced by dependency and risk, not by enthusiasm.
- 17 Pilot planning — The smallest test that would change our mind.
- 18 Execution / validation — The pilot runs and produces evidence, not opinion.
- 19 KPI measurement — The measure agreed before the pilot, applied after it.
05 Assure & decide — The evidence becomes a decision.
Everything gathered assembles into one report, the material part of it is put to leadership, and leadership decides. Executive sign-off: Someone with the authority accepts the conclusion. Result: Decision recorded.
- 20 Final audit report — Evidence, findings, conclusion and remediation in one file.
- 21 Executive presentation — Only what is material, to the people who decide.
- 22 Decision / sign-off — Accept, remediate, restrict, do not use, or re-test.
06 Continue — The audit doesn't end when the report is signed.
The system is re-examined at 30, 60 and 90 days, and what holds becomes part of how the organisation operates. Continuous assurance: Does the assurance still hold after the system changes? Result: Continuous assurance.
- 23 30/60/90-day follow-up — What changed, and does the assurance still hold?
- 24 Implementation / scale — What worked becomes part of the operating model.
What Nakoda does at each stage
- Steps 01–04 — Frame the engagement
- We establish whether an independent audit is the right instrument, and what it has to cover. Qualify the need. Define the scope. Establish the objectives. Set the engagement.
- Steps 05–09 — Establish the baseline
- We map what exists before we judge any of it. Discover the environment. Map the process. Assess current AI. Evaluate maturity. Examine data and technology.
- Steps 10–15 — Identify and prioritize
- We separate what is material from what is merely true. Identify relevant use cases. Catalogue them. Score value and feasibility. Assess risk. Build the business case. Prioritize the portfolio.
- Steps 16–19 — Turn findings into action
- We give every finding an owner, a date and a measure. Build the roadmap. Plan the pilots. Support validation. Define measurement.
- Steps 20–22 — Assure and report
- We put the material findings, and only those, in front of the people who decide. Produce the final audit report. Present material findings. Support executive decision-making.
- Steps 23–24 — Continue
- We come back, because the system will not have stood still. Follow up. Re-test where required. Support implementation. Support scale.
Testing a control
Claim: Human review is required before a payment decision is actioned. Evidence: of 1,250 decisions, 1,201 carried a documented review and 49 had no documented review. Conclusion: control effectiveness partial, raised as F-03 — Human oversight exists but is not consistently evidenced.
- F-01 (Context) — Model version changed without reassessment. Classified: Restrict. Material.
- F-02 (Data) — Supplier master feed lineage not captured end to end. Classified: Remediate. Material.
- F-03 (People) — Human oversight exists but is not consistently evidenced. Classified: Remediate. Material.
- F-04 (Vendors) — No contractual notice of provider model change. Classified: Escalate.
- F-05 (AI system) — Decision path not reconstructable beyond 30 days. Classified: Remediate.
- F-06 (Instructions) — Instruction changes recorded but not reviewed. Classified: Accept.
Remediation
- R-01 (0–30 days) — Enable review logging on the payment decision workflow. Owner: Finance Systems.
- R-02 (0–30 days) — Hold the model version pending reassessment. Owner: AI Platform.
- R-03 (30–90 days) — Restore lineage capture on the supplier master feed. Owner: Data Engineering.
- R-04 (30–90 days) — Define the reassessment trigger for model and vendor change. Owner: Risk.
- R-05 (30–90 days) — Re-test the oversight control on a full quarter. Owner: Internal Audit.
- R-06 (3–6 months) — Change-notice clause in the AI provider contract. Owner: Procurement.
- R-07 (3–6 months) — Decision reconstruction for the payment workflow. Owner: AI Platform.
- R-08 (6–12 months) — Instruction change control in the release process. Owner: AI Platform.
AI System Assurance Report
Conclusion: Conditional. Reliable for the intended use, once F-01 to F-03 are addressed.
- 03 Material findings
- 4 / 7 Domains effective
- 124 Evidence items reviewed
- 08 Recommendations
- 90 days Re-test
- Executive conclusion — What the AI can be relied on for, and what it cannot.
- Scope & methodology — What was examined, how, and against what.
- Material findings — Evidence, risk, recommendation, re-test — per finding.
- Remediation & re-test — Owners, target dates, and what gets tested again.
Internal AI Audit Function Design
For organisations that want permanent capability rather than repeat engagements.
- Method — The audit methodology, adapted to AI evidence.
- Tools — Tooling and the data access an audit actually needs.
- People — Training to the level required to audit, not to build.
- Governance — The structure that gives the function independence.
- Reporting — The line from a finding to the board.
- First audits — A first cohort run under supervision, to validate the method.
Professional basis: COBIT, CRISC, CISA. Methodology, evidence standards, documentation and reporting follow ISACA frameworks and the professional standards of chartered accountancy.
AI in auditing
The same audit desk, doing more of the work that used to be sampled.
- Enhanced sampling and coverage
- Population coverage: How much of the ledger was actually looked at. From statistical sample to full population, or risk-stratified where it cannot be.
- Continuous monitoring
- Anomaly detection: Finding it as it happens, not at year end. From point in time to continuous, with thresholds tuned against alert fatigue.
- Document analysis
- Extraction accuracy: Whether the machine read the contract correctly. From manual review queue to extraction validated against a manual benchmark.
- Risk assessment for internal audit
- Audit universe: What the annual plan decided to look at. From interviews and last year's plan to transaction, control and external signals, ranked.
- AI in finance audit
- Materiality: The threshold above which an error would change a decision. From two methodologies, separately to financial audit discipline applied to ai systems, and back.
Output is evidence, not a substitute for professional judgement.



