Article 14 was drafted for a system a person watches. A CV screener returns a ranked list, a human reads it, the human decides. Oversight in that picture is a role someone occupies.
An agent does not sit still for that. It plans across steps, calls tools, acts on what comes back, and reaches the consequential moment somewhere in the middle of a run that nobody was watching in real time because watching it would defeat the purpose of having it. The oversight Article 14 describes cannot be staffed onto that. It has to be built into it.
Which is the useful thing about reading the provision closely. Article 14 does not actually ask for a person watching. It asks for five capabilities to be available to a person, and every one of them is an architectural property. Read that way it becomes a design specification, and a fairly precise one.
What does Article 14 of the EU AI Act require?
Article 14(1) requires that high-risk AI systems be designed and developed so that they can be effectively overseen by natural persons during the period in which they are in use, including through appropriate human-machine interface tools. Article 14(2) sets the objective: preventing or minimising risks to health, safety or fundamental rights.
Three drafting choices in those two paragraphs do most of the work.
The obligation is on design and development, so it falls on the provider before the system reaches anyone. Oversight is not something a deployer can add later to a system that does not support it, which is why Article 14(3) offers providers two routes: measures built into the system where technically feasible, or measures identified by the provider as appropriate for the deployer to implement. Either way the provider identifies them, before placing on the market.
"Effectively" is doing more than it appears to. The standard is not that oversight is possible but that it works.
And Article 14(3) requires oversight measures commensurate with the risks, the level of autonomy, and the context of use. That is the sentence agent builders should read twice. Level of autonomy is named as a scaling factor in the operative text, which means a more autonomous agent owes more oversight machinery for the same use case. The Act anticipated the argument that autonomy is a reason for lighter controls and rejected it in advance.
When does Article 14 apply to AI agents?
Article 14 binds high-risk AI systems. Following the Digital Omnibus, Regulation (EU) 2026/1744, it applies from 2 December 2027 to standalone systems classified as high-risk under Article 6(2) and Annex III, and from 2 August 2028 to systems classified under Article 6(1) and Annex I.
For most agents the route in is Annex III rather than Annex I. Employment, meaning CV screening and candidate evaluation, and essential services, meaning creditworthiness and eligibility assessment, are where agent deployments most often land. Scheduling, coding assistance and internal workflow automation generally sit outside.
The Article 6(3) derogation is worth knowing precisely, because teams reach for it too readily. An Annex III system is not high-risk where it does not pose a significant risk of harm, and one of four conditions is met: narrow procedural task, improving the result of a previously completed human activity, detecting deviations from prior decision-making patterns without replacing human assessment, or performing a preparatory task. But the last subparagraph of Article 6(3) overrides all four: a system that performs profiling of natural persons is always high-risk. For an agent that builds a picture of a person across steps, that override is easy to trip and hard to argue out of.
Article 6(4) also requires a provider relying on the derogation to document the assessment before placing the system on the market, and to register under Article 49(2). The derogation is a documented position, not a silent one.
Alongside Article 14 sits Article 26(2), the deployer's obligation to assign human oversight to natural persons who have the necessary competence, training and authority, as well as the necessary support. Two duties on two parties: the provider builds the capability, the deployer staffs it with someone who can actually use it. A capable system overseen by someone without authority to stop it satisfies neither.
The full deadline picture after the Omnibus matters here, because the transparency duties that already bind are frequently confused with these.
What are the five capabilities a human overseer must have?
Article 14(4) requires that the system be provided to the deployer so that the persons assigned to oversight are enabled, as appropriate and proportionate: to understand its capacities and limitations and monitor operation including detecting anomalies; to remain aware of automation bias; to correctly interpret output; to decide not to use the system or to disregard, override or reverse its output; and to intervene or interrupt through a stop button or similar procedure allowing a halt in a safe state.
They are cumulative. A system enabling three of the five does not satisfy Article 14(4), and the qualifier about appropriateness and proportionality governs how each capability is exposed rather than whether it exists.
Translated into agent architecture, each becomes something specific.
Understanding capacities and limitations, and monitoring operation. Documentation plus a live view of what the agent is doing now. Article 14(4)(a) explicitly extends to detecting and addressing anomalies, dysfunctions and unexpected performance, which is a monitoring requirement, not a documentation one.
Automation bias awareness. The only one of the five that is about the human rather than the system, and the hardest to evidence. More on it below.
Correct interpretation of output. Explanations at the decision point rather than in a manual. Article 13(3)(b)(iv) requires the provider to supply the technical capabilities relevant to explaining output, and 13(3)(f) requires describing how deployers collect, store and interpret logs under Article 12.
Deciding not to use, or to disregard, override or reverse. Note that reverse is in the text alongside disregard and override. Disregarding an output is trivial when a human acts on a recommendation. It is meaningless when the agent has already acted, which makes this capability a reversibility requirement for any agent that executes rather than recommends.
Intervening or interrupting. A stop that halts execution and leaves the system in a safe state. Not an alert. Not a flag for tomorrow's review.
For biometric identification systems under Annex III point 1(a), Article 14(5) adds a further requirement: no action or decision may be taken on the basis of the identification unless separately verified and confirmed by at least two competent, trained and authorised natural persons, subject to a carve-out where Union or national law considers it disproportionate for law enforcement, migration, border control or asylum. Outside those carve-outs it is a binding requirement rather than good practice, and Article 12(3)(d) requires logging who those verifying persons were.
Can you satisfy Article 14 with a human review dashboard?
No, and the reason is 14(4)(b). Article 14(4) requires that overseers remain aware of the possible tendency to automatically rely or over-rely on system output, which the provision names as automation bias, and singles out systems providing information or recommendations for decisions taken by natural persons.
That paragraph is unusual. Every other capability in 14(4) describes something the system must let a person do. This one describes a cognitive failure the arrangement must guard against, and no interface satisfies it merely by existing.
A dashboard is where automation bias goes to compound. A reviewer facing a queue of agent outputs, most of which are fine, converges on approval. The measurement problem is that the queue looks healthiest at exactly the moment oversight has stopped working: throughput up, escalations down, no complaints. Those are the metrics of a rubber stamp and of a well-functioning review, and they do not distinguish between them.
What does distinguish them is override rate and time to decision. If reviewers essentially never override, either the agent is perfect or nobody is really reading. If median review time is four seconds, you know which. Those two numbers are the closest thing to evidence that 14(4)(b) has been addressed, and it is worth noting that Singapore's IMDA arrived independently at the same pair of indicators in its agentic framework, which is some confirmation that they are the right ones.
Article 26(2)'s requirement of authority points the same way. A reviewer who is measured on throughput does not have authority in any sense that matters, whatever the org chart says. If overriding the agent is career-costly, oversight is decorative.
The honest summary: a dashboard is necessary and nowhere near sufficient. What makes it sufficient is a checkpoint the agent cannot pass without a decision, an overseer with authority and time, and measurement showing the decisions are real.
What does "intervene or interrupt" mean for an agent mid-task?
Article 14(4)(e) requires the ability to intervene in the operation of the system, or to interrupt it through a stop button or similar procedure that allows the system to come to a halt in a safe state. The last five words are the requirement. Stopping is easy. Stopping safely, mid-plan, is a design problem.
Consider an agent five steps into a seven-step task. It has read a record, called an external API, written to a database, sent an email, and is about to issue a refund. The stop arrives now. What state is safe?
The database write can be rolled back. The refund has not happened. The email has been sent and cannot be recalled, so the safe state has to include a compensating action rather than a rollback: a follow-up message, a flag on the account, a note in the record explaining that a process was interrupted. That is what "halt in a safe state" means for an agent, and it is not something a kill signal delivers on its own. It requires knowing, per action, whether the action is reversible and what compensates for it if it is not.
Which is why Article 14(4)(d) and (e) are really one requirement viewed twice. Reversibility is the precondition for meaningful interruption. An agent whose actions cannot be undone can be stopped but cannot be safely stopped, and the difference is the whole of 14(4)(e).
Three design consequences follow.
The stop has to reach the executor, not the queue. An interrupt that prevents new tasks from starting while the current run completes is not an interrupt within the meaning of the provision.
Irreversible actions belong behind checkpoints, because the alternative is a stop capability that cannot deliver a safe state for exactly the actions where safety matters. This is the point at which Article 14 and DIFC's requirement for evidence of a human-intervention algorithm converge on the same architecture from opposite drafting traditions.
And the interruption must be recorded, including the state it halted in and what compensating actions ran. Otherwise you can perform the capability but not evidence it, and the audit trail fields that carry this are the ones conventional observability does not produce.
A closing point about the calendar. December 2027 sounds distant and is not, because none of the above is documentation. A checkpoint before consequential actions, an interrupt that reaches running execution, a compensating-action model for irreversible steps, and override metrics that mean something are architecture, and architecture retrofitted into a shipped agent is the expensive version of this work. India's regime asks a version of the same question through grievance redress and asks it sooner, from 13 May 2027.
Teams that build it now are not complying early. They are building the only version that will work.
Truveil scores AI agents against the EU AI Act and five other frameworks, and produces the evidence described here.