# The ISO 42005 AI impact assessment in practice

> Use the ISO 42005 AI impact assessment to connect ISO 42001 governance with system evidence, approval gates, monitoring, and a practical template.

ISO/IEC 42005:2025 gives teams a usable method for examining how one AI system can affect people and society. It does not replace ISO/IEC 42001, certify a product, or turn an AI risk register into an impact assessment. It fills the operational gap between a company policy that says impacts will be assessed and the evidence needed to approve, restrict, monitor, or stop a specific deployment.

That distinction matters in a startup. I have seen teams write a broad responsible AI policy, attach a spreadsheet of model risks, and assume the governance job is done. Then a customer asks who can be harmed, which use was approved, what evidence supported the decision, and what change forces a new review. The spreadsheet cannot answer. A good assessment can, because it ties an identified impact to affected people, deployment context, controls, an owner, an approval decision, and a review trigger.

## ISO/IEC 42005 adds a system level method

ISO/IEC 42005 adds guidance for assessing the impacts of a particular AI system and its reasonably foreseeable applications across its life cycle. ISO and IEC published the first edition on 28 May 2025. The standard applies to organizations that develop, provide, or use AI systems, regardless of their size or sector. Its focus is the effect on individuals, groups, communities, organizations, and society, not merely whether a model meets a technical score.

The word guidance deserves attention. ISO/IEC 42005 is not a management system requirements standard, and organizations do not obtain a standalone ISO/IEC 42005 certificate in the way they seek certification of a management system against ISO/IEC 42001. A customer or contract can still require you to follow it. A regulator can also expect an impact assessment under a law with its own scope. Those are separate sources of obligation. Buying the standard does not create a legal duty, and following it does not prove compliance with every law.

The standard separates impact from risk more clearly than many internal templates do. A risk describes uncertainty around an objective. An impact is a consequence experienced by someone or something. A hiring model can carry a risk of inconsistent recommendations, but the impact may be delayed employment, lost income, or unequal access for a particular applicant group. If the record stops at model drift or inaccurate output, it describes a technical failure while leaving the human consequence unnamed.

Benefits belong in the assessment too. Faster review, better access, reduced waiting time, and more consistent service can be real benefits, but a team must identify who receives them and under which conditions. A benefit claimed by the operator can coexist with harm imposed on an affected group. Recording both prevents the assessment from becoming a one sided defense of a launch decision.

ISO/IEC 42005 also treats foreseeable misuse and unintended use as part of the context. That is more useful than limiting the review to the product manager's happy path. If a support summarizer will predictably be copied into employee performance reviews, saying that performance management is outside the intended use does not remove the impact. The assessment should either control that use or explain why the organization accepts it.

## ISO/IEC 42001 owns the management system

ISO/IEC 42001 owns the organizational process, while ISO/IEC 42005 explains how to perform and document the system assessment that the process calls for. ISO/IEC 42001 specifies requirements for establishing, operating, maintaining, and continually improving an artificial intelligence management system, usually shortened to AIMS. It covers organizational context, leadership, planning, support, operation, performance evaluation, and improvement.

Clause 6.1.4 of ISO/IEC 42001 addresses AI system impact assessment in planning. Clause 8.4 brings the assessment into operations. Annex A.5 provides related controls, and Annex B.5 gives implementation guidance. The exact applicability depends on the organization's role, context, and documented control choices, but the management system must make the process repeatable rather than treat an assessment as a presentation prepared for one audit.

ISO/IEC 42005 supplies the deeper method: establish timing and scope, assign responsibilities, perform the assessment, analyze results, record and report them, approve the outcome, then monitor and review it. Its content prompts cover the system description, capabilities, purpose, intended and unintended uses, data and model information, geography, languages, deployment constraints, interested parties, benefits, harms, failures, and reasonably foreseeable misuse. That is the companion relationship in practical terms.

Keep one process, not parallel paperwork. The impact assessment should feed the AI risk assessment, risk treatment plan, operational controls, monitoring plan, internal audit evidence, and management review. The risk register may point to the assessment record, while the assessment points back to the risks and controls it creates or changes. Duplicating every field in both places produces conflicting versions within a month.

The NIST AI Resource Center published a crosswalk between ISO/IEC 42005 and the NIST AI Risk Management Framework. It maps responsibility and approval topics to GOVERN, system context to MAP, assessment and analysis to MEASURE, and monitoring to both governance and measurement activity. That mapping is useful because it shows that impact assessment is a process, not a document at the end of MAP. The crosswalk labels its ISO source as DIS 42005, however, so use it as a routing aid and verify clause references against the published edition you license.

ISO/IEC 23894 plays another role. It gives guidance on AI risk management across organizations that develop, deploy, or use AI. Use it to improve risk methods, not to replace the human and societal inquiry in ISO/IEC 42005. ISO/IEC 42001 tells the organization to run a managed process, ISO/IEC 42005 develops the impact assessment method, and ISO/IEC 23894 helps with the wider risk discipline.

## You need an assessment before consequences become expensive

You need an AI impact assessment when an AI system can materially affect people, rights, access, safety, work, money, or public participation, even if no customer has yet requested a named standard. The sensible threshold depends on severity, scale, reversibility, dependency, vulnerability, and the operator's power over the affected people. A spelling assistant and an automated credit decision should not pass through the same depth of review.

There are four practical triggers. First, ISO/IEC 42001 implementation can require a defined and operated assessment process within the scope of the AIMS. Second, legislation may mandate a particular assessment. Third, a contract, procurement questionnaire, board policy, insurer, or enterprise customer may demand evidence. Fourth, your own launch controls should require an assessment when the downside justifies it. The fourth trigger often arrives before the others and costs less than investigating harm after deployment.

The EU AI Act provides a good example of why teams must not equate ISO guidance with law. Article 27 requires a fundamental rights impact assessment before certain deployers use specified high risk systems. Its scope includes bodies governed by public law, private entities providing public services, and deployers of specified systems concerning creditworthiness and life or health insurance, subject to the regulation's detailed boundaries and exceptions. It asks for the deployment process, duration and frequency, affected categories, specific risks of harm, human oversight, and measures for when risks materialize. It does not require every private company using any AI tool to file the same assessment.

GDPR Article 35 can separately require a data protection impact assessment when processing is likely to create high risk for people's rights and freedoms. Examples named in the article include systematic and extensive automated evaluation that supports decisions with legal or similarly significant effects, large scale processing of sensitive categories, and large scale systematic monitoring of public areas. The EU AI Act says its fundamental rights assessment can complement an existing data protection impact assessment where obligations overlap. Complement does not mean rename. Privacy, broader fundamental rights, and the wider social impacts covered by ISO/IEC 42005 have different scopes.

Outside a legal mandate, use a short screening record for every inventoried AI system. Escalate to a full assessment when the system makes or materially shapes consequential decisions, acts with meaningful autonomy, handles sensitive data, reaches vulnerable people, operates at scale, is hard to contest, can cause difficult to reverse outcomes, or changes a worker's livelihood. Record why low impact systems were screened out. Silence looks like omission; a dated rationale shows a decision.

## Scope follows the deployment, not the model name

A useful assessment scopes the system as deployed, including the people, workflow, data, interfaces, operator behavior, and environment around the model. Assessing a foundation model in the abstract will not tell you what happens when a sales team lets it rank leads, when a clinic drafts patient instructions with it, or when an engineering agent can merge code. The same model can sit inside systems with radically different impacts.

Start with a boundary statement that another person can test. Name the business process, system owner, operator, affected population, decision or action, deployment regions, supported languages, upstream data, third party components, downstream integrations, and human override. State what the assessment excludes and why. If the boundary says only customer service assistant, but the tool can also cancel accounts through an action connector, the boundary is false.

Treat intended use and foreseeable use as different fields. Intended use comes from the approved purpose. Foreseeable use comes from incentives and access. A recruiter under a deadline may turn a writing assistant into a candidate scorer even if the interface never advertises that feature. Product permissions, training, logging, and review must address what people will predictably do, not what the policy hopes they do.

Interested parties are more than users and buyers. Include people subject to decisions, workers who must correct outputs, people represented in training or retrieval data, operators, customers, vendors, regulators, and communities that bear indirect effects. Consultation should fit the stakes. A low impact internal assistant may need interviews with a few operators. A system that affects access to essential services needs input from people who understand the affected groups and the practical route to appeal.

Do not collapse a system portfolio into one assessment called use of generative AI. Shared components can support shared evidence, but each deployment needs its own context and decision. You may reuse the model card, vendor review, security evaluation, and data provenance record. You cannot reuse the conclusion if the affected people, action authority, or consequence changes.

## A template must end in a decision

An ISO 42005 template should produce an owned decision with evidence, conditions, and review triggers, not a polished narrative that nobody operates. The following YAML is a compact record structure for a startup. It is not an official ISO form, and adopting it does not establish conformity. It gives engineering, legal, security, product, and leadership one object they can inspect.

```yaml
assessment_id: AIIA-2026-014
status: draft | approved | approved_with_conditions | rejected
system:
  name: support-response-agent
  version: 3.2
  owner: vp-customer-operations
  purpose: draft replies and propose account actions
  intended_uses:
    - answer product questions
    - propose refunds below the approved limit
  foreseeable_unintended_uses:
    - infer customer intent from protected traits
deployment:
  regions: [US, EU]
  languages: [en, es]
  affected_groups: [customers, support_staff]
  human_override: required_before_account_action
evidence:
  - evals/response-quality-3.2
  - security/vendor-review-2026-04
impacts:
  - group: customers
    outcome: an incorrect action delays access to paid service
    severity: high
    likelihood: possible
    reversibility: partial
    control: require_agent_confirmation_and_manager_review
decision:
  approver: ai-governance-owner
  conditions: [weekly_exception_review]
  review_triggers: [model_change, new_action, new_region, serious_incident]
  next_review: 2026-10-01
```

The identification block should include the assessment owner, contributors, approver, dates, system version, and links or references to controlled evidence. Do not assign approval to a committee without a named accountable role. Committees discuss; one authorized person accepts conditions or stops release. Separation helps when possible: the person trying to ship should not be the sole person deciding whether the evidence is enough.

The system and context block should describe purpose, functionality, capabilities, limits, users, affected groups, locations, languages, operating environment, data, model, suppliers, integrations, and human involvement. Write observable facts. Replace supports human review with the exact point where a person sees the output, the information they receive, the time available, the authority to reject it, and what the system does after rejection.

The impact block should contain benefits and harms for each affected group. For every material impact, record source or causal path, existing controls, severity, likelihood or uncertainty, duration, scale, reversibility, distribution, and supporting evidence. Do not force every harm into a single numeric score. A rare irreversible harm and a common minor burden can land on the same arithmetic total while requiring different decisions. Keep the dimensions visible.

The decision block should state approved, approved with conditions, rejected, or retired. Conditions need owners and due dates. A condition such as improve monitoring is not executable. Record the event to measure, source, threshold, review cadence, person receiving the alert, and action when the threshold is crossed. Link the decision to release management so a failed condition can actually block deployment.

The review block should list time based and event based triggers. Useful events include a model or material prompt change, a new data source, expanded autonomy, a new user group, a new region or language, a supplier change, a complaint pattern, a control failure, an incident, and evidence that an assumption was wrong. The next review date is a backstop, not permission to ignore change until the calendar catches up.

## Run the assessment as an engineering workflow

Run the assessment before the architecture becomes costly to change, then update it at release gates and during operation. A late legal review can identify a serious impact, but it cannot cheaply restore a human appeal path after every downstream system assumes an automated decision is final. Product discovery is the right time to test purpose, necessity, affected groups, and less harmful alternatives.

A five stage workflow is enough for most teams:

1. Screen the proposed system, assign an owner, and decide whether a full assessment is needed. Preserve the screening rationale either way.
2. Set the deployment boundary and collect evidence from product, engineering, security, privacy, legal, operations, suppliers, and affected people where appropriate.
3. Identify benefits, harms, failures, and foreseeable misuse for each affected group. Test controls against the actual workflow.
4. Analyze results and make a recorded approval decision. Convert every condition into owned delivery work with acceptance evidence.
5. Monitor indicators and triggers, review the assessment after material change, and preserve previous versions.

Do not begin with a blank workshop. Send contributors a system description and a first pass at affected groups before the session. Use meeting time to challenge assumptions, trace causal paths, and decide controls. Asking twelve people to brainstorm risks on demand produces familiar topics such as bias and privacy while missing the exact account action, delay, or denial that hurts someone.

Engineering evidence should be specific enough to rerun. Record evaluation dataset version, sample policy, metrics, subgroup slices, observed failure examples, test environment, model and prompt version, action permissions, and test date. Evidence can include qualitative research and complaints as well as automated tests. A clean benchmark does not disprove an operational impact that appears only when workers trust the output under time pressure.

Make the assessment record part of change management. A pull request or release ticket can reference the assessment ID and declare whether a review trigger fired. The reviewer should see the decision status and open conditions without hunting through a governance folder. This small integration catches scope changes earlier than a quarterly meeting.

## Evidence quality matters more than document length

Evidence quality matters more than document length because an assessment decision is only as sound as the facts behind it. A thirty page report can still hide an untested assumption, while a shorter record can expose its evidence, limits, and open questions well enough for an approver to act. Treat every material claim as something another person should be able to inspect.

Separate facts, estimates, judgments, and unknowns. The production log may show how often an operator overrode a recommendation; that is a fact for the measured period. A product manager's belief that customers will appeal a wrong decision is a judgment. Expected volume before launch is an estimate. The effect on a group that has not been consulted is unknown. Combining all four under one confidence score makes weak evidence look settled.

Use an evidence table with six fields: evidence ID, claim supported, source owner, system version, collection date, and limitation. Add retention and access rules where the material contains personal, confidential, or security sensitive information. The assessment should reference the controlled item rather than paste raw data into a widely shared document. This keeps access narrow without breaking traceability.

Vendor documents need the same skepticism as internal claims. A model card can describe the supplied component, but it rarely proves that your prompts, retrieval data, permissions, interface, operators, and downstream actions produce acceptable outcomes. Record which supplier claim you rely on, test whether it holds in your deployment, and keep a fallback for evidence the supplier will not provide. Procurement cannot transfer accountability for your use.

Negative evidence deserves a place in the record. Keep failed evaluations, disagreement among reviewers, complaints that do not fit the initial taxonomy, and groups for which the sample is too small to support a conclusion. Deleting awkward results makes the report cleaner and the decision worse. The approver needs to see the limits, especially when uncertainty sits next to severe or irreversible harm.

Define acceptance evidence before building a control. If the condition is that staff can override an agent, acceptance should test that an authorized operator can see the basis for the proposed action, reject it, choose a safe alternative, and leave a trace. A screenshot of an override button proves the interface contains a button. It does not prove that the control works inside the operating pressure described in the assessment.

When evidence conflicts, do not average it into false certainty. Investigate why a controlled evaluation and field complaints tell different stories. The evaluation may omit a language, the complaint channel may capture only severe cases, or operators may use the system outside the tested workflow. Preserve both observations, narrow the claim, and make the approval conditions match what is actually known.

## A risk register cannot carry the whole assessment

A risk register cannot carry the whole assessment because it usually removes the narrative link between system behavior and a person's outcome. Registers are good for prioritization, ownership, treatment status, and aggregation. They are poor at preserving deployment context, competing benefits, affected group input, causal chains, and the rationale behind approval.

Consider a support agent that drafts replies and proposes refunds. The risk register says inaccurate output, medium likelihood, high impact, treatment: human review. The team launches. Operators handle four conversations at once and accept proposals quickly. The interface shows no source, the refund rule changed last month, and customers cannot tell whether a person reviewed the action. Some valid requests are denied, repeat contacts grow, and staff copy the same output because the queue rewards speed.

The register was not false; it was too compressed. The impact assessment would ask which customers are more likely to struggle with appeal, what information the reviewer sees, how much time they have, whether a rejected proposal teaches the system anything, how the customer contests a denial, and who watches for patterns. Human review is not a control until the workflow gives the human context, time, authority, and feedback.

The popular recommendation is to score likelihood times severity and approve anything below a threshold. It is popular because a colored matrix fits executive reporting. It is wrong as the sole decision method for AI impacts. Evidence can be thin, distribution can be uneven, and reversibility can matter more than the product of two guessed numbers. Use scoring to sort attention, then keep the reasoning and uncertainty visible.

The opposite mistake is writing a long ethics essay with no control mapping. An auditor, engineer, or incident lead needs to trace each material impact to prevention, detection, response, evidence, owner, and status. If that trace takes an afternoon, the document will drift away from the system.

## Approval and monitoring make the record credible

Approval makes the assessment operational only when the approver has authority to impose conditions or stop the system. The record should show what was known, which uncertainties remained, whose input was considered, what tradeoff was accepted, and why. A signature on an incomplete template proves that someone clicked approve; it does not prove a reasoned decision.

Use approval with conditions sparingly. It fits a limited gap that can close by a fixed date while existing controls keep deployment within tolerance. It does not fit missing evidence about a potentially severe and irreversible impact. Teams often abuse conditional approval because launch dates feel fixed. If every release carries old conditions forward, the status is approval by exhaustion.

Monitoring must follow the causal path in the assessment. For the support agent, model accuracy alone is insufficient. Watch overrides, denial reversals, repeat contacts after automated action, complaints, action frequency, time available for review, subgroup differences where lawful and appropriate, and cases that lack evidence. Pair outcome indicators with control indicators. A stable model can still cause new harm after a policy, interface, customer mix, or operator incentive changes.

Preserve version history. Each record should show the system version, assessment version, previous decision, changed assumptions, new evidence, and approver. Do not overwrite the approved record with current text. When an incident occurs, the team needs to know what decision makers saw at launch and whether a trigger was missed.

Set a review cadence based on the speed of change and severity of impact. A frequently updated agent with action permissions may need event driven review plus a short calendar interval. A stable internal classifier with bounded effects may justify a longer interval. The assessment should explain the choice. Planned intervals satisfy a process; relevant triggers protect people.

## Small teams should keep the evidence close to delivery

Small teams can implement ISO/IEC 42005 without creating a separate compliance department, but they cannot remove accountability. One person can hold several roles; the record still needs to show when that person acts as system owner, assessor, control owner, or approver. For higher impact decisions, bring in independent review or affected party expertise rather than pretending a two person team contains every perspective.

Keep a short controlled template, a central AI system inventory, and evidence in the systems where work already happens. Reference test runs, release tickets, incident records, vendor reviews, and decisions instead of copying them into a giant document. Access control and retention still matter. The goal is one traceable chain, not one enormous file.

At oleg.is, I use the Team & AI Audit to separate governance work that protects the business from ceremony that merely consumes engineering time. The useful test is concrete: can the team show which deployment was assessed, who may be affected, what evidence supports its controls, who approved it, and which event reopens the decision?

Start with the deployment that has the greatest combination of consequence, reach, and weak reversibility. Give it a stable assessment ID, write its real boundary, and name one accountable approver. If your current process cannot turn a failed condition into a blocked release, fix that connection before producing another policy.
