Europe is about to approve “human oversight” of AI without testing it under pressure. The windows are closing now.

TL;DR: The EU’s draft human oversight standard mandates that oversight be verified but leaves testing under realistic conditions optional. I propose four wording changes and explain how to submit them before the comment windows close (some shut within days!). Below: the finding, the fix, how EU standardisation works, and what no standard can fix because Article 14 left it out.

A mistake waiting to happen

Imagine a triage tool used in a health emergency: it passes every requirement in the current draft, verified the way the draft permits — maybe in a quiet room, several clinicians, a few minutes per patient. The system passes with flying colours and is quickly deployed in the field.

Header image: Lera Kogan /​ Unsplash

Two paramedics find themselves as the first responders of a mass casualty incident, a collision between a bus and an 18-wheeler truck with 24 injured. The tool tags a 29-year-old with chest pain and stable vitals as “green”, but it’s wrong. Internal bleeding that the paramedic would have caught had she been in a quiet room gets missed. The paramedic has 30 seconds, a queue of patients and a screen that tells her “low priority”. She doesn’t overthink it.

While verification is mandatory, testing with representative users under realistic conditions is recommended or noted as possible, not required. Automation bias rises sharply under time pressure and task load (Parasuraman & Manzey 2010[1]), which is precisely the condition the verification never checked.

Nothing malfunctioned. The oversight the standard verified was not the oversight that would play out in real life, at the roadside.

What the Draft says

The draft — prEN 18229-3 — is the European standard that will spell out what Article 14′s ‘human oversight’ means in practice (how it got there is explained below).

Notably, JTC 21 (the body in charge of the draft) went further than instructed in the standardisation request. The current draft addresses many of the issues that are not delineated in Article 14, such as having reaction timeframes for intervention as a key element of human oversight, interface design, and deployer-side obligations. These are all important advances, yet there is still a gap.

The document itself defines the convention: “shall” is a requirement, “should” a recommendation, “can” a possibility. It mandates the result (oversight must be verified), but every test that would reveal whether human oversight survives contact with a stressed human — verification under realistic conditions, automation-bias testing, training content — is recommended or noted as possible, not required. In a conformity assessment (usually done in-house, a topic for another piece), testing under real-life conditions is essentially optional. The draft even describes the test itself as an example: realistic conditions, flawed outputs mixed in. It just doesn’t require it.

Call to Action

The draft is still in enquiry. Now is the time to comment via your national standards body (the organisation that represents your country in CEN-CENELEC and consolidates comments for JTC 21). This is the only public window before the formal vote.

The CEN-level enquiry closes on 22 October, but national bodies close earlier and dates vary: Netherlands (NEN) 22 September, Ireland (NSAI) 24 September, Sweden (SIS) 7 October, Norway (Standard Norge) 8 October, Spain (UNE) 9 October, France (AFNOR) 12 October. Check your national body’s date today and treat it as the real deadline.

The fix is small: change “can/​should” to “shall” in 5.3.2.3.2, 5.3.4.5, 5.5.2 and 5.6.2 (the numbers are the draft’s own section headings), so that testing with representative users under realistic conditions becomes a requirement.

You can access your national body’s enquiry portal via CEN-CENELEC’s page here.

Ready-to-paste comment text

National body portals ask for clause, comment and proposed change, one entry per clause. The draft is in English and JTC 21 works in English, so submit your comment in English even on a national portal (no need to translate the quotes). Each block below is one complete entry, paste as is.

Example screenshot of a national body’s form filled with one proposed change (in this case it’s the Spanish portal, UNE). Yours will look similar:

Screenshot of UNE's enquiry comment form for prEN 18229-3, with fields for section (5.3.2.3.2), comment type (technical), comment, and proposed change, filled in with the text from this post.
  • 5.3.2.3.2

    • Comment: The draft mandates outcomes (5.3.2.3, 5.3.4.4, 5.6.1) but leaves optional the only methods that expose failure under time pressure and task load. Under clause 0.2, only “shall” carries presumption of conformity (AI Act Art. 40); “can” and “should” are invisible to a conformity claim. Automation bias rises sharply under exactly these conditions (Parasuraman & Manzey, 2010). The clause itself notes verification can address “high-volume or high-fatigue environments”, yet leaves the verification optional.

    • Proposed change: “Verification can be performed by testing the retrospective review interface with representative designated persons under realistic oversight conditions” → “Verification shall include testing the retrospective review interface with representative designated persons under realistic oversight conditions.”

  • 5.3.4.5

    • Comment: The draft mandates outcomes (5.3.2.3, 5.3.4.4, 5.6.1) but leaves optional the only methods that expose failure under time pressure and task load. Under clause 0.2, only “shall” carries presumption of conformity (AI Act Art. 40); “can” and “should” are invisible to a conformity claim. Automation bias rises sharply under exactly these conditions (Parasuraman & Manzey, 2010). 5.3.4.4 requires providers to verify that designated persons are able to complete the key oversight tasks within the reaction timeframe, yet makes the only method named for doing so optional.

    • Proposed change: “Usability testing can be conducted” /​ “The interface should be tested” → “Usability testing shall be conducted” /​ “The interface shall be tested”.

  • 5.5.2

    • Comment: The draft mandates outcomes (5.3.2.3, 5.3.4.4, 5.6.1) but leaves optional the only methods that expose failure under time pressure and task load. Under clause 0.2, only “shall” carries presumption of conformity (AI Act Art. 40); “can” and “should” are invisible to a conformity claim. Automation bias rises sharply under exactly these conditions (Parasuraman & Manzey, 2010). The clause’s own example describes the test (“a mix of correct and intentionally flawed AI outputs under realistic conditions”), and 5.5.1 states that “the testing in 5.5.2 can be used to assess the adequacy of these measures”, yet the testing itself is only a “can”. The draft therefore requires the 5.5.1 measures but no evidence that they reduce over-reliance.

    • Proposed change: “During development the provider can conduct tests” /​ “Scenarios should be designed to compare” → “During development the provider shall conduct tests” /​ “Scenarios shall be designed to compare” (automation-bias testing).

  • 5.6.2

    • Comment: The draft mandates outcomes (5.3.2.3, 5.3.4.4, 5.6.1) but leaves optional the only methods that expose failure under time pressure and task load. Under clause 0.2, only “shall” carries presumption of conformity (AI Act Art. 40); “can” and “should” are invisible to a conformity claim. Automation bias rises sharply under exactly these conditions (Parasuraman & Manzey, 2010). 5.6.1 requires that designated persons be able to understand the system’s capacities and limitations, and refers to “mandatory training materials”, yet 5.6.2 makes the minimum content of those materials only recommended. The draft mandates the understanding but not the method to make it happen.

    • Proposed change: “the documentation or training materials provided should include” → “the documentation or training materials provided shall include”.

How the machinery works (skip if you know EU Standardisation)

The Structure

The legislation-to-standards pipeline follows this structure:
The EU AI Act (Article 14) → C(2025)3871 → prEN 18229-3

  • The Act. The legislation that was proposed by the European Commission, adopted by the European Parliament and the Council. Hard to change once passed.

  • The Standardization Request [C(2025)3871]. The European Commission’s formal instructions to the relevant bodies for harmonised standards. The relevant bodies are:

    • CEN (European Committee for Standardization), entrusted with general standardisation across multiple industries such as construction, healthcare, energy and information technology

    • CENELEC (European Committee for Electrotechnical Standardization), responsible for European standardisation of electrical engineering.

    These are independent bodies working jointly through CEN-CENELEC JTC 21 (Joint Technical Committee) on the standardisation of, amongst other items, human oversight.

  • prEN 18229-3. A draft European Standard or prEN. This is the text that CEN-CENELEC JTC 21 produces in response to the request. Standards are voluntary, but following a harmonised one gives providers “presumption of conformity” with Article 14 (the system is presumed to comply with EU rules).

What Article 14 states about Human Oversight

In short, Article 14 requires that AI systems be effectively overseen through their interface, that risks are prevented or minimised and that oversight measures are commensurate with those risks and are provided to the deployer.

The Article has no specifics regarding reaction timeframe, interface requirements or automation bias training.

This can be excused by understanding that such legislation is horizontal in nature — intended to act as a framework within which artificial intelligence systems operate — yet naming these issues would have cost nothing.

The Commission’s 2025 Standardization Request

From the Commission Implementing Decision[2]:

2.5. Human oversight
The harmonised standards and standardisation deliverables in this area shall set up specifications for human oversight. Those specifications shall comprehensively cover all elements referred to in Article 14 of the Regulation (EU) 2024/​1689.

While the mandate says to “comprehensively cover all elements of Article 14”, there are no design-in conditions, so the vagueness of the Act is inherited in the mandate (as I previously addressed in this piece).

The Standard currently in development

The European Standard draft on human oversight, prEN 18229-3 (Part 3 of CEN-CENELEC’s AI trustworthiness framework) is currently in development, alongside other standard requests under C(2025)3871.

A standard gains legal effect in two steps:

  1. Getting adopted as an EN by the CEN-CENELEC members (it becomes a European Standard and no longer a work in progress)

  2. Citation by the Commission in the Official Journal. This is what triggers the ‘presumption of conformity’ with Article 14 and makes it a harmonised standard.

Institutional residue

Even with mandatory testing, issues remain. The draft (and the Article it originates from) assumes a designated individual will take responsibility — and, if anything fails, the burden — of working with these systems, as pointed out by Elish (2019)[3]. Humans become “moral crumple zones”, soaking the blame for mishaps that actually originated in the development and management structure of the AI system, not in human operators.

Oversight might not fail individually, but institutionally: through problems in staffing, load and authority (as pointed out by Green, 2022)[4].

The Article brings up institutional mechanisms once, in paragraph 5, for high-risk AI systems that use remote biometric identification systems:

[...] no action or decision is taken by the deployer on the basis of the identification resulting from the system unless that identification has been separately verified and confirmed by at least two natural persons with the necessary competence, training and authority.


The requirement for a separate verification by at least two natural persons shall not apply to high-risk AI systems used for the purposes of law enforcement, migration, border control or asylum, where Union or national law considers the application of this requirement to be disproportionate.

This last section adds an institutional mechanism, but immediately makes a carve-out for law enforcement, migration, border control and asylum. It is made optional in precisely the situations that have high-stakes and intense cognitive load, when human oversight is predicted to fail more often (Green 2022[4], Parasuraman & Manzey 2010[1]).

To Reiterate

A standard can’t fix what the original article waived. For now, we must focus on fixing the four “can/​should”s to “shall”s (5.3.2.3.2, 5.3.4.5, 5.5.2 & 5.6.2).

Comment via your national body using the text above in the “Ready-to-paste comment text” section of this piece.

  1. ^

    Parasuraman, R., & Manzey, D. H. (2010). Complacency and bias in human use of automation: An attentional integration. Human factors, 52(3), 381-410. https://​​doi.org/​​10.1177/​​0018720810376055

  2. ^

    Commission Implementing Decision C(2025)3871 on a standardisation request to CEN and CENELEC as regards high-risk AI systems in support of Regulation (EU) 2024/​1689 (the AI Act), repealing Implementing Decision C(2023)3215. https://​​ec.europa.eu/​​transparency/​​documents-register/​​api/​​files/​​C(2025)3871_0/​​de00000001072818?rendition=false

  3. ^

    Elish, M. C. (2019). Moral crumple zones: Cautionary tales in human-robot interaction (pre-print). Engaging Science, Technology, and Society (pre-print). https://​​dx.doi.org/​​10.2139/​​ssrn.2757236

  4. ^

    Green, B. (2022). The flaws of policies requiring human oversight of government algorithms. Computer Law & Security Review, 45, 105681. https://​​doi.org/​​10.1016/​​j.clsr.2022.105681

  5. ^

    CEN-CENELEC (2026). prEN 18229-3: AI trustworthiness framework — Part 3: Human oversight. Draft European Standard, Enquiry stage, CEN-CENELEC JTC 21, July 2026.

No comments.