# How high-risk AI systems are classified under the EU AI Act

> Learn how high-risk AI systems are classified under the EU AI Act, including Annex III, HR and credit edge cases, and a practical self-assessment.

High-risk classification under the EU AI Act follows intended purpose, not the model name, vendor label, or presence of a human reviewer. A system that ranks job candidates can be high-risk even if a recruiter clicks the final button. The same model used to rewrite a job description may sit outside Annex III because it does not evaluate a person or shape a selection decision.

That distinction sounds clean until a product combines intake, summarisation, scoring, and recommendations behind one interface. Founders then ask the wrong question: "Is our AI in the high risk category?" Classify each AI system and intended purpose in the actual workflow. The legal answer can change when a customer switches on a feature, changes who receives the output, or uses a harmless assistant to make consequential decisions.

This is an operational reading of Regulation (EU) 2024/1689, the EU AI Act, and the European Commission's May 2026 draft classification guidelines. The regulation controls; the draft guidelines explain the Commission's current interpretation but are not binding. Employment, lending, data protection, consumer protection, and discrimination rules still apply when an AI system falls outside the high risk category.

## How the two classification routes work

Article 6 creates two separate routes into the high risk regime. Run both. A negative answer on one route says nothing about the other.

The Article 6(1) route covers an AI system that is itself a regulated product or is intended as a safety component of one, when a third party must assess the product's conformity under the Union laws listed in Annex I. Annex I points to product regimes for machinery, toys, lifts, pressure equipment, medical devices, in vitro diagnostic devices, motor vehicles, civil aviation, and other regulated products. Both conditions must hold: the product connection and the required assessment by a third party.

This route is often misread as "AI used in hardware." That is too broad. Demand forecasting for a manufacturer of medical devices does not enter the high risk category because the company makes regulated products. Conversely, software can qualify when it is itself a medical device or performs a safety function whose failure endangers people or property. The product's regulatory role matters, not whether the AI runs in a physical box.

The Article 6(2) route covers intended purposes listed in Annex III. These are often called independent use cases because classification does not depend on a conformity assessment under product law. Recruitment screening and consumer credit scoring usually enter here.

One system can satisfy both routes. Document each result because the applicable conformity process, registration details, and transition date can differ. Do not stop after finding the first convenient answer.

## Annex III names uses, not industries

Annex III contains eight areas, but an entire company or department does not enter the high risk category. The listed use case inside the area is the unit that matters.

1. Biometrics covers permitted remote biometric identification, certain biometric categorisation by sensitive or protected attributes, and emotion recognition. Verification that solely confirms one person's claimed identity is excluded from the remote identification entry.
2. Critical infrastructure covers safety components used to manage or operate critical digital infrastructure, road traffic, and supplies of water, gas, heating, or electricity.
3. Education and vocational training covers admission or assignment, evaluation of learning outcomes, assessment of an appropriate education level, and monitoring prohibited behaviour during tests.
4. Employment covers targeted job advertising, recruitment and selection, decisions on work terms, promotion or termination, certain task allocation, and monitoring or evaluation of worker performance and behaviour.

The remaining four areas concern access to services and the exercise of public power:

5. Essential services covers eligibility for public benefits, creditworthiness and credit scores for natural persons, life and health insurance risk and pricing, classification of emergency calls and dispatch priority, and emergency healthcare triage.
6. Law enforcement covers specified permitted uses such as assessing the risk of victimisation, tools similar to polygraphs, assessing the reliability of evidence, certain offending or reoffending risk assessments, and profiling during investigation or prosecution.
7. Migration, asylum, and border control covers specified permitted risk assessments, application examination, tools similar to polygraphs, and identification in that context, excluding verification of travel documents from the final entry.
8. Justice and democratic processes covers assistance to judicial authorities in applying law to facts, similar alternative dispute resolution, and systems intended to influence voting or an election or referendum outcome. Pure campaign administration whose output does not reach voters is excluded from that election entry.

The list is narrower than labels such as "fintech AI" or "HR automation." An anomaly detector for accounts payable at a bank is not credit scoring. A tool for planning shifts that allocates work only from store opening hours and required staffing levels differs from one that assigns shifts based on each worker's speed, mood, or reliability score.

The Commission can update Annex III through delegated acts within the areas already listed, subject to the Act's conditions. A classification record therefore needs a source version and review date. A spreadsheet that says only "outside high risk" will age badly.

## Intended purpose beats the sales description

The provider's intended purpose anchors classification, including the context and conditions of use described in instructions, technical documentation, promotional material, and other supplied information. A carefully narrow contract cannot rescue broad product behaviour and sales claims that encourage a listed use.

Start by drawing the decision workflow. Identify the person affected, the decision, each input, each AI output, who sees it, and what happens when the output is adverse. Then define system boundaries. A hiring suite may contain a writer for job descriptions, an interview scheduler, a CV parser, a candidate ranker, and an attrition predictor. Treating the suite as one black box hides which functions enter Annex III and which do not.

Human involvement does not remove a use from Annex III. The words "assist" and "support" appear throughout consequential workflows. If a recommendation determines which applications receive attention, adds scrutiny, causes delay, changes price, or strongly frames a decision, it can materially influence the outcome even when a person retains formal authority.

Watch downstream use as well. A credit bureau can be the deployer that produces a score while another organisation uses it to decide on a mortgage or housing. The Commission's draft guidance treats that separation as irrelevant to whether the scoring purpose falls within Annex III point 5(b). Classification follows the purpose of the score, not corporate boundaries.

A model built for general purposes does not carry one fixed risk class into every application. The downstream system provider decides how the model is integrated and presented. A language model that extracts dates from forms may qualify for the Article 6(3) filter; the same underlying model ranking applicants by predicted success can enter the high risk category. "We only call an API" is an architecture fact, not a classification argument.

## The Article 6(3) filter is narrow

An Annex III system can avoid classification in the high risk category only when it poses no significant risk of harm to health, safety, or fundamental rights, including because it does not materially influence the decision, and at least one of four conditions applies.

- It performs a narrow procedural task.
- It improves the result of a human activity that has already been completed.
- It detects patterns or departures from earlier decisions without replacing or influencing the completed human assessment and with proper human review.
- It performs a preparatory task for an Annex III assessment.

Those conditions are not four magic labels. The opening test still requires no significant risk and no material influence. Calling candidate scoring "preparatory" fails when the score determines the shortlist a recruiter actually reads.

A narrow procedural hiring tool might detect exact duplicate files, convert a CV to a standard format, or route an application to the role the applicant explicitly selected. It crosses the line when it hides documents, infers suitability, ranks candidates, or marks employment gaps as risk signals. A proofreading tool used after a manager has completed an appraisal can improve a finished human activity. A tool that introduces new performance reasons or recommends a rating does something else.

The filter has a hard stop: an Annex III system that profiles natural persons always belongs in the high risk category. Profiling generally means automated processing of personal data to evaluate personal aspects, including performance at work, economic situation, health, preferences, reliability, behaviour, location, or movements. Do not reduce that question to whether the output is called a "profile." A risk category, propensity estimate, or score of candidate fit may perform profiling without using the word.

The provider must document a filter decision before placing the system on the market or putting it into service, register the system under Article 49(2), and produce the assessment to a competent authority on request. The exemption removes Chapter III obligations for systems in the high risk category; it does not erase GDPR duties or make the use lawful under other rules.

## HR tools turn on influence, not automation

Recruitment systems enter the high risk category when their intended purpose includes targeted job advertising, analysing or filtering applications, or evaluating candidates. The category reaches earlier than the offer decision. Choosing who sees an opportunity and who reaches an interview can shape access to employment.

Candidate ranking is the easy case. Suppose a system scores written answers, ranks applicants, and sends the top group to interview. A recruiter may inspect every recommended candidate and still never inspect the bottom group. The AI has already allocated attention and materially influenced selection. The Commission's draft employment examples place this pattern in the high risk category.

Scoring during a background check is similar. An output such as "low, medium, or high risk" based on work history, education, financial data, or online information evaluates a person. Human review after the score does not undo its effect when flagged candidates move slower or disappear from the active queue. If personal data is used to predict reliability or suitability, the profiling override also blocks the Article 6(3) filter.

Several nearby functions deserve separate treatment:

- Generating a generic job description without evaluating people falls outside the recruitment use case.
- Scheduling an interview from calendars usually falls outside if it neither prioritises candidates nor infers suitability.
- Parsing a CV into fields may fit the filter when it preserves all information and performs only clerical structuring.
- Summarising an application becomes risky when the summary selects favourable facts, omits material facts, or feeds the only view a recruiter reads.
- Matching workers to tasks enters point 4(b) when it relies on individual behaviour or personal traits; allocation from location, availability, and objective operational constraints may fall outside that wording.

Worker management extends beyond employees with standard contracts. Annex III refers to contractual relationships involving work and access to freelance work. Promotion recommendations, dismissal flags, productivity scoring, behaviour monitoring, and task allocation based on personal traits need review. Payroll arithmetic and expense categorisation do not enter the high risk category merely because HR uses them.

The most common failed argument is "the manager decides." Ask what evidence the manager sees without the AI, whether an adverse output changes the queue, and how often staff depart from it. Formal approval is weak evidence when the interface has already filtered the world.

## Credit tools depend on the person and purpose

Annex III point 5(b) covers AI intended to evaluate a natural person's creditworthiness or establish that person's credit score. Consumer loans and mortgages are the clear cases. The draft guidelines also say a score produced by one organisation can qualify when third parties use it for lending, housing, healthcare, or telecommunications decisions.

The express exception for AI used to detect financial fraud depends on purpose. It does not exempt a combined model that labels an applicant as possible fraud and uses the same output to decide whether that person deserves credit. Split the systems, outputs, permissions, and documentation if the purposes truly differ. A renamed decline score remains a creditworthiness assessment when it predicts repayment or drives approval.

Corporate credit can fall outside point 5(b) when the system evaluates a legal entity from company data and does not assess a natural person's personal finances. The Commission's draft goes further in examples: analysis of business data for an owner of a small business that is not a legal entity can remain outside, and assessment of an owner backing a company loan can remain outside where the company is the primary beneficiary. Treat these as draft interpretive examples, not a blanket exemption for sole traders. If personal income, spending, payment behaviour, or household circumstances drive the result, document why the system is not evaluating the person before relying on that position.

Some adjacent tools fall outside because their intended purpose differs. Customer support that explains application terms, helps complete a form, or handles a complaint after the decision does not itself assess creditworthiness. Customer segmentation for marketing remains outside if it plays no part in credit assessment. Collateral valuation can remain outside when it evaluates only the asset. Internal monitoring of credit exposure after a loan, for prudential purposes, is also treated as outside point 5(b) in the draft guidance.

Boundaries break when teams reuse features. A marketing affordability segment copied into underwriting has acquired a role in credit assessment. A fraud flag that automatically reduces an approval score now influences creditworthiness. A model used for early warnings after a loan may no longer fit the narrow internal monitoring example if it cuts a person's available credit or sets renewal terms.

Credit scoring will often evaluate economic situation, reliability, or behaviour through personal data, which is profiling. In that case the Article 6(3) filter cannot apply even if the team describes the score as one small input. The correct work is classification and compliance, not creative relabelling.

## A defensible assessment takes eight passes

A useful assessment produces a traceable decision for each pair of system and purpose. It should let a product manager, counsel, and engineer reach the same system boundary without reconstructing months of meetings.

1. Confirm that the component meets the AI Act definition of an AI system. Ordinary software that executes rules defined solely by people may fall outside, but do not improvise this test. Use Article 3(1) and the Commission's guidelines on the AI system definition.
2. Check territorial scope and exclusions. Record where the provider and deployer operate, where the system is marketed or used, and whether its output is used in the Union. Research, personal use outside work, military and national security uses, and certain open source situations have specific rules, not broad slogans.
3. Write one statement of intended purpose per function. Include the affected person, decision, output, user, context, and prohibited uses. Compare the statement with the interface, sales material, contracts, and observed use.
4. Test Article 6(1). Name the Annex I product law, the safety function, and the requirement for conformity assessment by a third party. If any element is absent, record it instead of writing "not applicable."

At this point, the system boundary and the first legal route should be explicit. The next four passes test the sensitive use, allocate responsibility, and keep the result current:

5. Test every Annex III entry. Cite the point, such as 4(a) for candidate evaluation or 5(b) for an individual's creditworthiness, and map each legal phrase to product evidence.
6. If Annex III matches, test Article 6(3). Explain the absence of significant harm and material influence, identify one filter condition, and test for profiling. A filter conclusion needs evidence from workflow behaviour, not an adjective.
7. Assign roles and obligations. Identify the provider, deployer, importer, distributor, authorised representative, and any party that rebrands or substantially modifies the system. Record who owns conformity work, instructions, logs, monitoring, human oversight, notices, and incident handling.
8. Set review triggers. Reassess when features, input data, output meaning, user group, geography, instructions, integration, model behaviour, or downstream decisions change. Include an owner and a date.

Do not ask a vendor only whether its product is "EU AI Act compliant." Ask for documentation of intended purpose, system boundaries, instructions for use, known and prohibited uses, validation evidence, logging behaviour, and the provider's classification rationale. Your organisation still owns the way it deploys the system.

## Keep a classification record engineers can review

The decision record should live beside product and risk documentation, under version control where practical. A legal memo that engineers never see cannot constrain a feature flag or integration.

This compact template forces the important facts into the open:

```yaml
system_id: hr-candidate-assist
version: 2.4
owner: product-risk
provider: Example Provider Ltd
deployer: Example Employer Ltd
ai_system_basis: "Article 3(1) assessment reference"
intended_purpose:
  affected_person: job applicant
  decision: invitation to interview
  output: ranked candidate list
  user: recruiter
  context: EU engineering recruitment
system_boundary:
  included: [cv_parser, suitability_score, ranking]
  excluded: [calendar_scheduler]
article_6_1:
  annex_i_law: null
  safety_component: false
  third_party_assessment: false
annex_iii:
  match: "4(a) recruitment and selection"
article_6_3:
  significant_risk_absent: false
  material_influence_absent: false
  condition: null
  profiling: true
classification: high-risk
evidence: [workflow-map-v3, recruiter-ui-v5, model-card-2.4]
review_triggers: [new-output, new-data, new-user, new-market]
next_review: 2026-11-15
```

The expected output shape is deliberately boring: one classification, the legal route, evidence references, an owner, and review triggers. If Article 6(3) places the system outside high risk, fill every filter field and add the Article 49(2) registration reference. If the system sits in the high risk category, connect this record to the broader risk management and conformity plan rather than stuffing compliance evidence into one file.

Test the record against the running product. Can recruiters sort by a supposedly hidden score? Does a support agent copy a fraud label into a credit decision? Does the calendar scheduler prioritise "strong" candidates despite its harmless name? I have seen teams classify a slide deck while production followed different rules. Regulators, affected people, and your own incident team will care about production.

## Responsibility follows control over the system

Buying an AI product does not transfer every duty to the vendor. The AI Act assigns roles from conduct, so the contract label is only a starting point. A company that develops a system or has it developed and places it on the market under its own name is a provider. The organisation using it under its authority is usually a deployer. Importers and distributors have separate duties in the supply chain.

A deployer can become the provider under Article 25 when it puts its name or trademark on an existing system in the high risk category, makes a substantial modification, or changes the intended purpose so the system enters that category. A substantial modification is not every configuration change. It is a change after market launch that the provider did not foresee or plan in its initial conformity assessment and that affects compliance or changes the intended purpose.

This matters in ordinary product work. A vendor supplies a tool intended only to extract fields from CVs. The customer adds prompts, weights, and an automated cutoff, then markets the resulting workflow internally as candidate scoring. The customer has not merely deployed the documented parser. It has created a new intended purpose and should assess whether it now carries provider duties for the resulting system.

The opposite mistake is to treat every customer configuration as a new system. Setting an allowed threshold, choosing documented inputs, or connecting the product as instructed may stay within the provider's assessed purpose. Record the provider's permitted configuration range and compare your changes against it. Do not decide from the size of the code diff. A single flag can activate candidate ranking; a large infrastructure migration can leave purpose and compliance untouched.

Contracts should make the operational split explicit. Require the provider to supply the classification rationale, instructions, technical information needed for deployment, expected input characteristics, design for human oversight, logging controls, change notices, and incident contacts. Require the deployer to follow use restrictions, train reviewers, monitor overrides and complaints, preserve relevant logs, and report serious or unexpected behaviour. These clauses do not override the regulation, but they prevent each party from assuming the other owns the work.

Procurement teams also need a rule for components. A provider of models for general purposes, an HR application provider, an integration partner, and the employer can all sit in one chain. The HR application provider normally owns classification of the completed system for ranking candidates that it markets. The employer owns its deployment context and can take on more responsibility through modification or repurposing. The model provider's documentation supports the downstream assessment but cannot classify uses it did not build.

Run change control through the classification record before release, not during an annual legal review. Product tickets that add new data, new recommendations, new affected groups, or new downstream actions should require a classification check. Sales promises need the same gate. If a salesperson offers "automatic rejection of weak applicants" while the instructions describe clerical CV parsing, the evidence of intended purpose has already split in two.

Finally, preserve disagreements. When counsel, product, and engineering interpret material influence differently, record each factual assumption and the decision owner. That record is more useful than false consensus because it tells the next reviewer what new evidence could change the result.

## Classification starts the work, it does not finish it

A result in the high risk category triggers a compliance programme, not an automatic ban. Providers face requirements for risk management, data and data governance, technical documentation, record keeping, deployer information, human oversight, accuracy, resilience, and cybersecurity. They also need the applicable conformity assessment, quality management, registration, monitoring after market launch, corrective action, and processes for serious incidents.

Deployers must follow instructions, assign capable human oversight, monitor operation, and control input data where they supply it. Workplace use brings advance information duties toward affected workers and representatives. Systems that make or assist decisions about natural persons bring notice duties. Public bodies, private entities providing public services, and deployers of creditworthiness or life and health insurance systems have specified duties under Article 27 to assess effects on fundamental rights. GDPR, collective labour rules, consumer credit law, and national employment law may impose separate controls.

The current application calendar gives teams time, but not permission to postpone classification. Following the Digital Omnibus change that entered into force in July 2026, the rules for Annex III systems in the high risk category apply from 2 December 2027, while Article 6(1) rules linked to products apply from 2 August 2028. Check the consolidated legal text and Commission material when planning because standards, final guidance, and delegated acts can change the implementation details.

For a portfolio with dozens of AI features, classification belongs in the same inventory as ownership, cost, access, and production risk. A Team & AI Audit at oleg.is can map that portfolio and the engineering work around it, but legal conclusions should still be reviewed by qualified EU counsel.

Do the boundary work before buying a compliance platform. If your team cannot state what the system does to whose decision, software will only organise an unknown. The strongest first deliverable is a versioned map that connects intended purpose to production evidence, with no gap where "human review" used to be.
