Skip to content
8 min read

An AI management system is daily operating discipline

See what an AI management system requires each day: ownership, risk decisions, change records, evidence, audits, and corrective action under ISO 42001.

An AI management system is daily operating discipline
Table of Contents

ISO/IEC 42001 does not ask you to produce a handsome policy binder and return to shipping as before. It asks you to run an AI management system: a repeatable way to decide which AI uses are acceptable, who owns each decision, what evidence supports it, and what happens when reality disagrees with the plan.

That changes ordinary work. A product manager cannot add a model provider with only a procurement ticket. An engineer cannot change a prompt, evaluation set, or fallback without considering impact. Leadership cannot approve an AI policy once and disappear. The standard turns AI governance into an operating rhythm, and an auditor tests whether that rhythm works.

The standard asks for a system, not a binder

The standard requires connected management practices, not a prescribed pile of documents. Clauses 4 through 10 follow the familiar management-system sequence: understand context, establish leadership, plan, provide support, operate controls, evaluate performance, and improve. Annex A supplies reference controls, while Annex B explains how those controls can be implemented. The organization still chooses controls based on its risks and scope.

ISO's public overview describes ISO/IEC 42001 as a Plan-Do-Check-Act management-system standard. That wording matters. Plan means setting scope, objectives, risk criteria, and treatments. Do means running the chosen processes. Check means monitoring, internal audit, and management review. Act means correcting failures and changing the system. A policy without the last three verbs is not an operating system.

The standard uses two kinds of language that teams often blur. Requirements in the main clauses state what the organization must achieve. Annex A controls are a reference set considered during risk treatment, not a universal checklist that every company implements word for word. You may exclude a control when it is not relevant, but you need a defensible reason and must not undermine required outcomes.

Certification is also narrower than many sales decks imply. ISO says certification is voluntary, independent certification bodies perform it, and ISO itself does not certify companies. A certificate says the management system conforms within its stated scope at the time of assessment. It does not prove that every model output is correct, that a product complies with every AI law, or that no harmful incident can occur.

Day to day, the system should answer five questions without a scavenger hunt:

  • Which AI systems and uses fall inside scope?
  • Who can accept a risk or approve a material change?
  • Which controls apply, and why?
  • Where is the evidence that people followed them?
  • How does the company detect and correct failure?

Leadership has recurring duties behind those answers. Top management must establish the AI policy, assign authority, integrate the management system into business processes, provide resources, communicate the importance of effective management, and support continual improvement. Delegating coordination to a compliance lead is sensible. Delegating accountability is not. A useful responsibility map names the person who owns the AIMS, owners for individual systems and controls, risk acceptance levels, escalation paths, and who reports performance to leadership.

The AI policy should fit the organization's purpose and provide a frame for objectives. It also needs commitments relevant to the standard, including meeting applicable requirements and continually improving the system. Staff must be able to find and understand it, and the organization must make it available to interested parties where appropriate. A copied promise to use AI ethically will not guide a product owner choosing whether to automate an account restriction. The policy needs enough specificity to shape decisions without pretending to anticipate every model change.

Clause 6 turns that policy into objectives. Make each objective measurable where practicable and state what will be done, which resources are needed, who owns it, when it will finish, and how results will be evaluated. Reduce AI risk is not an objective. Close all overdue high-risk treatments within 30 days, with exceptions accepted by the risk owner, gives a team a threshold and an action. So does require every material AI release to carry a linked evaluation record. Those are examples, not thresholds imposed by ISO.

Track objectives in the same operating review that can allocate staff and money. A dashboard without an owner or response rule only displays drift. When a target is missed, record the reason, decide whether the action, resource, measure, or target should change, and follow the decision to completion. That evidence later feeds performance evaluation and management review, closing the loop between leadership's stated intent and daily behavior.

If staff need a consultant to reconstruct those answers before an audit, the system is documentation theater.

Scope starts with an inventory people actually maintain

Your scope determines every obligation that follows, so write it around business activities, products, teams, locations, and interfaces rather than an aspiration such as responsible AI. Clause 4 expects the organization to understand internal and external issues, interested parties, their requirements, and the boundaries of the management system. A startup may scope a customer-support product and the engineering team that operates it. It should not quietly ignore a shared data pipeline, a contracted model provider, or the human review team when those parts affect the product.

The practical foundation is an AI system inventory with an accountable owner. It must include purchased tools and embedded features, not only models your engineers train. A recruiting team using a ranking service, a sales team pasting account data into a public assistant, and a support workflow that drafts replies all create decisions and impacts worth managing. Discovery can begin with interviews, expense records, identity logs, code searches, and vendor lists, but a one-time scan will decay quickly.

Use one record per defined AI system or use case. The distinction matters: one foundation model can support several uses with different affected people, data, failure modes, and oversight. Treating the vendor name as the inventory item hides that difference. A minimal record can look like this:

id: ai-017
use_case: Draft support replies
business_owner: Head of Support
technical_owner: Platform Lead
provider: Contracted model API
inputs: Ticket text and approved account fields
outputs: Reply draft for agent review
affected_parties: Customers and support agents
decision_role: Advisory, human sends every reply
risk_tier: medium
impact_assessment: IA-017-v3
last_reviewed: 2026-07-15
next_review: 2026-10-15
status: production

The dates are examples, not an ISO-mandated quarterly schedule. Set review frequency from risk, change rate, contractual duties, and legal requirements. A low-impact internal summarizer may need review after material change. A system influencing access to work, credit, care, or education deserves tighter triggers and closer monitoring.

Assign a process owner to keep the inventory alive and require teams to update it through workflows they already use: purchasing, architecture review, release management, and employee onboarding. The owner should chase gaps, but business owners remain accountable for their uses. Central governance cannot know how every output affects customers.

Risk assessment and impact assessment answer different questions

AI risk assessment manages uncertainty against organizational objectives, while AI system impact assessment examines consequences for people, groups, and society. ISO/IEC 42001 requires processes for both. Combining them into one score often erases the human consequence behind an operational label.

Start by defining assessment criteria before teams score anything. Specify impact dimensions, likelihood scales, acceptance thresholds, who can accept residual risk, and when reassessment occurs. Otherwise every owner calls a favored project medium risk. The method should cover the intended use, reasonably foreseeable misuse, data issues, opacity, security, safety, bias, availability, environmental concerns where relevant, and dependency on people or suppliers. Only include dimensions that can be applied honestly, but do not remove an awkward impact because measurement is difficult.

An impact assessment should identify affected parties and effects across the system lifecycle. Ask who benefits, who can be denied something, who carries the work of correcting an output, and who may never know AI influenced the result. Consider normal operation, failure, misuse, and retirement. Record assumptions and consultation, especially when the team lacks direct knowledge of an affected group.

A common failure looks like this. A support assistant is classified low risk because a human approves every message. In production, agents handle a queue that rewards speed, accept most drafts, and cannot see which account fields reached the provider. The nominal control, human review, exists. Its design fails because the reviewer lacks time, context, and a defined rejection route. The assessment must examine how people actually work, not how the diagram says they work.

Risk treatment follows the assessment. A team can avoid the risky use, change it, add controls, share part of the risk contractually, or accept residual risk through an authorized owner. Acceptance should be a signed decision with an expiry or review trigger, not a permanent label. New providers, material model or prompt changes, new data categories, expanded users, serious incidents, and changed legal duties are sensible reassessment triggers.

Keep an explicit chain from risk to treatment to evidence. If hallucinated policy advice is a risk, the treatment might restrict retrieval to approved documents, require citations to source passages, route low-confidence cases to a specialist, and sample outputs weekly. The evidence then includes retrieval configuration, test results, routing logs, and sample-review records. A vague treatment such as improve accuracy gives nobody an action and gives an auditor nothing meaningful to test.

The statement of applicability records control decisions

The statement of applicability, often shortened to SoA, is the bridge between assessed risks and Annex A. It records which reference controls are necessary, whether they are implemented, and why controls are included or excluded. It should reflect your actual design, not copy the annex into a spreadsheet with every row marked applicable.

Annex A covers policies for AI, internal organization, resources, impact assessment, system lifecycle, data, information for interested parties, use of AI systems, and third-party relationships. Those groups are prompts for a complete treatment discussion. They do not replace the main-clause requirements, and they do not prevent you from creating controls outside Annex A when your risk treatment needs them.

A useful SoA row contains the control reference, applicability decision, justification, implementation status, owner, linked risks, and evidence location. The justification should say something testable. Applicable because we use AI says nothing. A better entry explains that the company provides customers with information about AI-assisted decisions because the system affects their requests and because the impact assessment identified transparency and contestability needs.

Exclusion needs equal care. Suppose a control concerns development resources, but your company only buys a finished AI service. Declaring it irrelevant may be premature if your engineers still configure prompts, connect data, test behavior, set thresholds, or build fallback logic. You may not train the model, yet you develop an AI-enabled system. Scope follows the work and impact, not the invoice category.

Do not ask control owners to upload proof into a special audit folder each quarter. Link the SoA to evidence produced by normal work: approved policies, role descriptions, training records, architecture decisions, evaluation runs, release tickets, supplier reviews, user notices, monitoring reports, incident records, and management minutes. An evidence link that points to an empty template is worse than no link because it signals that staff optimize for appearance.

Review the SoA after changed risks and at a planned interval. The review is a decision meeting, not a formatting exercise. Owners should confirm whether each control still addresses the linked risk, whether its implementation operates, and whether monitoring found exceptions. Record changes and follow unresolved gaps into corrective action.

Ordinary changes need explicit AI gates

Put owners behind every system
Oleg maps AI work to accountable people before policies turn into unused paperwork.

Operational control under clause 8 means teams plan and control the processes needed to meet the management system's requirements. In a software company, the cleanest implementation adds AI-specific gates to purchasing, design, development, release, support, and retirement. A parallel governance ceremony creates delay and gets bypassed.

Define what counts as a material AI change. Provider model upgrades, system-prompt changes, new tools, evaluation changes, retrieval-source changes, threshold adjustments, new input data, new affected groups, and a shift from advisory output to automated action can all change risk. Cosmetic interface work usually cannot. Put the definition in the release process and allow owners to explain why a change is non-material.

A change ticket should force a compact decision, not a long essay:

Change: CHG-2841
System: ai-017
Purpose: Add order-status context to reply drafts
New data: Order ID, shipment state, delivery estimate
Affected risks: R-31 disclosure, R-44 incorrect commitment
Evaluation: EVAL-017-22 passed acceptance criteria
Human oversight: Agent still reviews and sends
Notices affected: Customer AI notice unchanged
Rollback: Remove context connector and restore prompt v18
Approvers: Product owner, privacy owner, technical owner
Monitoring: Review first 200 eligible drafts and all escalations

This artifact creates traceability across intent, data, risk, testing, approval, rollback, and monitoring. The sample number is illustrative. Your company should choose monitoring volume using risk and expected traffic, then record why the sample can reveal the failure modes you care about.

Approval authority must match consequences. A technical lead can approve a low-impact prompt correction within accepted risk. A business owner and privacy or legal owner may need to approve a new personal-data flow. Leadership should reserve decisions that exceed risk criteria or change the management system's objectives. If the CEO approves every prompt, work stalls. If engineers can accept every social or legal impact, accountability has vanished.

Emergency changes still need control. Permit a shorter approval path when harm or outage demands it, require the owner to state the emergency, preserve rollback, and run retrospective assessment within a defined period. Never label routine deadline pressure an emergency. Auditors notice when the exception path becomes the main road.

Data, suppliers, and people remain your responsibility

Buying an AI service transfers tasks, not accountability for your management system. The organization must control externally provided processes, products, or services that affect conformity. Day to day, that means assessing suppliers before use, setting contract and information requirements, monitoring relevant changes, and maintaining an exit or fallback plan proportional to risk.

A supplier review should cover the evidence you need to operate safely: service description, intended and prohibited uses, data handling, security responsibilities, model-change notice, performance information, incident communication, subcontractors where relevant, deletion and retention, audit or assurance evidence, and termination support. You may not get every term from a large provider. Record the gap, assess the resulting risk, add your own control, choose another supplier, or accept the residual risk through the right authority. A questionnaire marked complete does not remove the gap.

Data governance is broader than training-data quality. Inventory input, training, validation, test, retrieval, feedback, and monitoring data as applicable. Record provenance, permitted purpose, quality criteria, preparation, access, retention, and known limitations. When users paste data into an assistant, the input channel itself needs rules and technical boundaries. A policy telling staff not to paste secrets will fail if the product workflow rewards exactly that behavior.

People also need defined competence and awareness. Train them for their role. Engineers need the change criteria, evaluation method, and logging rules. Product owners need impact assessment and acceptance authority. Reviewers need enough time, interface context, escalation routes, and feedback loops. General annual AI awareness can support those controls, but it cannot replace job-specific instruction and observed competence.

Information for interested parties should match the effect on them. A user may need to know that AI participates, what the output means, its main limitations, how to seek human review, and how to report a problem. The exact notice depends on context and applicable duties. Publishing a generic responsible AI page while the affected screen offers no explanation may satisfy communications planning on paper but fails the person making a decision.

Retirement deserves a control too. Decide how to disable the system, retain required evidence, delete or return data, notify users or customers, migrate dependencies, and watch for downstream processes that still call the old service. Teams that inventory only launches discover abandoned API keys and undocumented automations much later.

Evidence should fall out of the work

Find shadow AI before audit
A Team & AI Audit traces tools, workflows, and handoffs that your formal inventory misses.

ISO/IEC 42001 requires documented information where the standard calls for it and where the organization decides it is necessary for effective operation. It does not require a document for every sentence. The useful question is whether another competent person can see what decision was made, under which criteria, by whom, using what evidence, and whether the result worked.

Design records at the point of action. Identity systems can show training completion. Version control can preserve policy and configuration changes. Evaluation pipelines can retain test version, inputs, metrics, acceptance criteria, result, and approver. Ticketing systems can connect incidents and corrective actions. Meeting notes can preserve management decisions. Keep access, retention, integrity, version, and approval controls appropriate to the record.

A lean operating calendar might assign recurring work like this:

  • Product owners review inventory status and open exceptions monthly.
  • Control owners inspect their evidence and failed checks on a risk-based schedule.
  • The AI governance owner reports objectives, incidents, supplier changes, and overdue actions to leadership.
  • Internal auditors sample processes independently of the people who operate them.
  • Top management holds a formal management review at planned intervals and records decisions.

ISO does not mandate those frequencies. Choose intervals that can catch drift before the impact becomes unacceptable. A rapidly changing customer product may need weekly operational monitoring and quarterly governance review. A stable, low-impact internal use may justify less. Event-driven reviews sit beside the calendar because a serious incident should not wait for next quarter.

Metrics need decisions attached to them. Count overdue high-risk treatments, evaluation failures, unresolved incidents, unreviewed material changes, supplier-notice response time, and control exceptions if those measures fit your objectives. Accuracy alone can hide unequal error distribution, unsafe rare cases, or reviewers who accept outputs automatically. Define the threshold, data source, owner, review forum, and required response for each measure.

Management review is not a status presentation delegated away from leadership. Top management must review whether the system remains suitable, adequate, and effective, considering matters such as changes in context, performance, audit results, nonconformities, risk, opportunities, and improvement. The record should capture decisions on resources, changes, and actions. Attendance without decisions is weak evidence.

Audits test behavior and corrective action changes it

Map your AI management gaps
The five-day Team & AI Audit finds unowned AI uses, weak gates, and costly process gaps.

Internal audit tests whether the management system conforms to the organization's own requirements and ISO/IEC 42001, and whether it operates effectively. The audit program should consider process importance, changes, and previous results. Independence matters: a person should not audit decisions they made or controls they operate. A small company can cross-audit between functions or use an external auditor, but it still owns the program.

Good auditors sample a trail rather than admire policies. Pick an inventory entry, trace it to assessment and applicable controls, select a material release, examine evaluation and approval, inspect monitoring, then interview the people who performed the work. Compare their explanation with the written process. Sample a supplier change and an incident. The gaps between declared and actual behavior are where the useful findings live.

A nonconformity is failure to meet a requirement. An incident may cause or reveal a nonconformity, but the terms are not interchangeable. A harmful output can occur despite staff following a well-designed process, in which case the team still responds and reassesses risk. Conversely, a missing approval is a nonconformity even when no customer reports harm. Blurring the terms causes teams to ignore process failures until damage appears.

Correction contains the immediate issue. Corrective action removes or reduces its cause so it does not recur. If an unapproved model version reaches production, rolling back is correction. Corrective action asks how deployment bypassed the gate, whether permissions or pipeline checks should change, whether similar systems share the weakness, and how the company will verify the fix. Retraining one engineer by default is usually lazy root-cause analysis.

Keep a record with the finding, violated requirement, evidence, containment, cause analysis, action owner, due date, completion proof, and effectiveness check. Close the action only after evidence shows the change worked. A revised procedure nobody follows is not an effective corrective action.

External certification audit follows the same broad logic but does not replace internal audit. ISO's own explanation says organizations may seek independent confirmation through a certification body. Treat the external auditor as a test of the system, not the designer of it. If your process exists only because an auditor requested a template, ask which risk or requirement it addresses before making it permanent.

A small company can keep the system lean

A small company can conform without creating an AI governance department. ISO/IEC 42001 scales to the organization, but small does not mean informal. One person may hold several roles, provided authority is clear and conflicts are managed. A seven-person company still needs someone to accept risk, someone competent to operate controls, and an audit arrangement with enough independence to challenge them.

Begin with boundaries and decisions, not templates. Establish the scope, inventory uses, identify interested parties and obligations, set the AI policy and measurable objectives, define assessment criteria, assess risks and impacts, choose treatments, complete the SoA, then put operational gates and evidence into existing work. Run the processes long enough to produce records. Complete internal audit, management review, and corrective actions before asking a certification body to judge readiness.

Do not buy certification on a deadline set by sales before understanding the scope. The standard can improve governance without certification, and certification does not replace customer due diligence or legal analysis. If a customer asks for ISO 42001, clarify whether it requires a certified AIMS, what scope must appear on the certificate, which accreditation expectations apply, and by when. Those details can change the work substantially.

For an AI-heavy engineering company, the expensive part is rarely writing policies. It is finding shadow uses, deciding which releases need review, building evaluations that represent real failure, and making owners preserve evidence while shipping. That is where experienced technical leadership earns its keep. At oleg.is, the Team & AI Audit is the practical entry point when a founder needs to map those gaps alongside team cost and delivery constraints.

An effective AI management system makes an uncomfortable decision visible before deployment and a failed assumption visible after it. If your register is current, approvals follow stated authority, evidence appears during normal work, audits find reality, and corrective actions change the process, the standard has become part of operations. If all activity starts six weeks before the certification audit, you have a calendar event, not a management system.

Frequently Asked Questions

What does ISO 42001 require from a company?

ISO/IEC 42001 requires a managed system for AI policy, objectives, risks, impacts, controls, resources, operations, monitoring, audits, management review, and improvement. The exact controls and evidence depend on the company's scope and assessed risks.

Is ISO 42001 certification mandatory?

No. ISO describes certification as voluntary, although a contract, customer, regulator, or procurement program may make it commercially necessary. An independent certification body performs the audit; ISO does not certify organizations.

Does ISO 42001 certify an AI model as safe?

No. Certification assesses the organization's AI management system within a defined scope. It does not guarantee every model output, prove universal legal compliance, or eliminate the chance of harm.

Do we need an inventory for ISO 42001?

You need enough controlled information to know which AI systems and uses the management system covers. A maintained inventory is the clearest way to connect owners, impacts, risks, controls, suppliers, changes, and evidence.

What is an ISO 42001 statement of applicability?

The SoA records the necessary controls, why they apply, whether they are implemented, and why any Annex A controls are excluded. It should trace back to risk treatment rather than repeat generic annex text.

Are all ISO 42001 Annex A controls mandatory?

Annex A is a reference control set considered during risk treatment, not a list applied identically by every organization. Exclusions need sound justification, and the selected treatment must still meet the main requirements and address risk.

How often should AI risk assessments be reviewed?

ISO 42001 does not impose one universal interval. Set a risk-based schedule and reassess after material changes, serious incidents, new data or affected groups, supplier changes, and changed obligations.

Can a startup implement ISO 42001 without a compliance team?

Yes. People can hold several roles and existing engineering workflows can carry approvals and evidence. The startup still needs clear authority, competent owners, leadership review, and an internal audit arrangement with sufficient independence.

How long does ISO 42001 implementation take?

The standard sets no fixed implementation period. Timing depends on scope, existing governance, number and risk of AI uses, evidence already available, and how many gaps require operational change before an audit.

Does ISO 42001 satisfy the EU AI Act?

ISO 42001 can organize governance and evidence that support legal compliance, but it does not replace the EU AI Act or legal analysis. Map each applicable legal duty to processes and controls instead of treating a certificate as automatic conformity.

Related Posts