Telonic

GovernanceSEP 30, 2026Ishaan Mirchandani, Anthony Ombrore

Designing agents that answer to governance at enterprise scale

How an enterprise should decide which decisions an AI agent takes on its own, which it takes within limits, and which always stay with a person. And why the record of what the agent did, and why, matters more than the instructions it was given.

On this page
  1. The situation
  2. What the evidence shows
  3. What that means
  4. How we design for it
  5. What we don't know yet
  6. Sources

The situation

A UAE motor insurer has to tell a claimant within three days which documents are missing, and settle within fifteen days of receiving a complete file [1]. At 11pm a policyholder messages on WhatsApp: "Has my claim been approved?" An agent can read the claim's status from the claims system, list the two documents still outstanding and explain what the policy says about the excess. The one thing it must never do is answer that question with a yes or a no of its own making.

Every enterprise that deploys an agent draws a line like this. A developer decides whether the agent may restructure a payment plan. A hotel decides whether it may issue a goodwill voucher, and up to what value. The hard part is not agreeing that a line exists. It is deciding where it goes, making sure it holds when nobody is watching, and being able to show afterwards exactly what the agent did and on what basis.

What the evidence shows

The frameworks scale oversight to risk. None asks for a person on every decision.

The NIST AI Risk Management Framework, a voluntary US standard, defines risk as a composite of the probability of an event and the magnitude of its consequences, and asks organisations to document how humans oversee a system and to keep mechanisms to "supersede, disengage, or deactivate" it [2]. ISO/IEC 42001, a certifiable management standard, has a single control on logging, which requires event logs to be kept at least while the system is in use [3]. Neither says where a person must sit.

The EU AI Act is the most specific. For systems it classes as high risk, Article 14 requires that the people overseeing the system can monitor it, override or reverse its output, and stop it. It is the only binding text located for this article that names automation bias, the well-documented tendency to accept a machine's output without checking it, as something overseers must be helped to resist [4]. Its criteria for classing a use as high risk include how far the system acts on its own and "the extent to which the outcome produced involving an AI system is easily corrigible or reversible" [5].

Two points of precision matter for a Gulf reader. The high-risk list is organised by topic, and customer service and claims handling are not on it; only life and health insurance risk assessment and pricing are [6]. And the Act reaches a company outside the EU that uses an AI system only where the system's output is used inside the EU [7], so a UAE or Saudi enterprise serving Gulf customers is outside its scope. Its high-risk obligations now apply from December 2027 [8].

Gulf law is keyed to the effect of a decision, not its subject.

UAE federal data protection law gives a person the right to object to decisions "resulting from automated processing, including profiling, particularly those decisions which have legal impact on or adversely affect" them, and requires the controller to "include the human element in reviewing automated processing decisions at the request of the Data Subject" [9]. The law has been in force since January 2022, but its executive regulations had not been issued as of mid-2026 and the federal Data Office never became fully operational, so there is no enforcement apparatus yet [10]. It does not apply in the DIFC or ADGM, which have their own regimes.

The DIFC's rules are the most developed in the region. Regulation 10 permits a system to process personal data commercially only if its purposes are human-defined or human-approved, or, where the system sets its own purposes, it does so "solely within the limits of human-defined constraints"; and Article 38 lets an individual object to a solely automated decision with legal or "other seriously impactful consequences" and require manual review [11][12]. These bind DIFC entities only. In Saudi Arabia, the implementing regulations of the Personal Data Protection Law, enforced since September 2024, require explicit consent for solely automated decisions and an impact assessment where decisions rest on automated processing; this rests on a law firm's summary, as the regulator's English text was not reachable [13].

For insurers licensed onshore in the UAE, the Central Bank's February 2026 guidance note sets out three oversight tiers. "Human-out-of-the-loop", where the system acts without direct human involvement, "should only be utilised for low-risk, non-material processes", and consumers "should be able to request human review or explanation of AI generated decisions" [14]. It is guidance, not regulation, and it applies to licensed financial institutions, not to developers or hotels.

The person checking the machine is a weaker control than it looks.

Human in the loop means a person approves each decision before it takes effect. Human on the loop means the system acts and a person monitors the results. Both depend on the person engaging, and decades of research say that is where it fails.

In a controlled flight-simulation study, participants given a highly but imperfectly reliable automated aid missed 41% of the events it failed to flag, and followed its wrong recommendation 65% of the time even when other instruments contradicted it [15]. In a study of 554 lay participants predicting pretrial outcomes, of the 304 given a risk score only 23.7% did better than the score alone, 64.1% did worse, and their confidence was negatively correlated with their performance [16]. Explaining the machine's reasoning does not fix this: in a study of 1,626 participants, explanations increased agreement with the recommendation whether it was right or wrong [17]. OWASP's 2025 list of the top ten risks in AI agent applications includes human-agent trust exploitation, where confident, polished explanations mislead operators into approving harmful actions [18].

The public failures share one shape. In the Dutch childcare benefits scandal, the civil servants who reviewed flagged applications could not see the information used to assign the risk score [19]. In Australia's Robodebt scheme, staff had previously reviewed around 20,000 high-risk files a year; the automated system sent out notifications at around 20,000 a week [20]. In each case the reviewer lacked at least one of three things: time, information or authority. Where any of the three is missing, the human check is a formality.

Gates that fire often are ignored.

A gate that interrupts every routine action trains people to click through it. In a study of 1.27 million clinical reminders across 112 clinicians, the likelihood that a reminder was accepted fell by 30% for each additional reminder in the same encounter [21]. So a gate should sit only where it is needed. One vendor's analysis of agent traffic on its own platform, not peer reviewed, found that only about 0.8% of agent actions appeared to be irreversible [22]. If that holds in customer communication, concentrating human attention on the irreversible actions is affordable in a way that reviewing everything is not.

The agent's own sense of when to stop cannot stand alone.

Confidence-based escalation means the agent hands to a person when its estimate of being right falls below a threshold. That estimate is a real signal, but a biased one. A 2026 preprint testing eight language models on tasks with real human decision data found that the confidence at which each model chose to escalate ranged from about 53% to above 100%, with two models escalating everything, and that in 66% of model and condition combinations the model overestimated its own accuracy [23]. On a benchmark of 2,000 agent safety tests, no agent scored above 60%, and the authors concluded that "reliance on defense prompts alone may be insufficient" [24]. A threshold has to be measured for the specific model before deployment, and it belongs alongside rules the agent cannot argue with, not in place of them.

What a usable record contains is already settled elsewhere.

A flight data recorder logs what the aircraft did and what the crew did on one time base, so the cockpit voice recorder can be laid over it. Cockpit recorders on new large aircraft now retain 25 hours rather than two, because two-hour recordings kept being overwritten before anyone knew they were needed, including after a 2017 incident at San Francisco [25]. European investment firms must record conversations about client orders "even if those conversations or communications do not result in the conclusion of such transactions", for five to seven years [26]. US securities rules require records kept in write-once storage or with "a complete time-stamped audit trail" of every change and who made it [27]. The AI industry is converging on the equivalent field list: the OpenTelemetry conventions for generative AI, an open observability standard, name the model, the instructions, the input and output messages, each tool called with its arguments and result, and the agent and conversation identifiers [28]. A tamper-evident log chains its entries so each carries a cryptographic fingerprint of the one before, which makes any later edit detectable [29].

The record decides disputes. In February 2024 a Canadian tribunal ordered Air Canada to compensate a passenger after its website chatbot told him he could apply for a bereavement fare within 90 days of buying a ticket, contradicting the airline's policy page. Air Canada argued the chatbot was, in effect, a separate legal entity responsible for its own actions. The tribunal called that "a remarkable submission" and held that "it should be obvious to Air Canada that it is responsible for all the information on its website" [30]. The passenger had a screenshot. A 2025 UK regulator review of 23 home and travel insurers, with no reference to AI, found that in many cases the minutes of key claims committee meetings "lacked sufficient detail to prove meaningful discussion, challenge or decision-making" [31]. The standard of evidence does not change with the kind of agent that made the decision.

What that means

The working hypothesis for this article was that decisions should be sorted by consequence and reversibility rather than by topic, and that governance has to be enforced by what the agent is able to do and by a complete record, not by instructions alone. The evidence supports both, with two corrections.

First, consequence and reversibility are the right axes, but topic still matters. Legislators use consequence and reversibility to decide what to regulate, then regulate by topic [5][6]. Some decisions are reserved whatever their reversibility: by law, such as a medical judgement; by a regulator's expectation, such as a claims decision, which the Central Bank's guidance would not class as a low-risk, non-material process; or by the business itself, such as anything already with a regulator. A sorting scheme needs a rule that removes those from the agent entirely.

Second, the judgement of consequence must be made in advance, per type of action, by people. The agent's runtime sense of risk is too variable to carry it [23][24]. Confidence-based escalation is a second line, not the first.

The enforcement hypothesis holds, with an addition. Capability limits and a complete record are the floor. But the human approval gate is itself a control that fails in predictable ways, so it has to be designed for the reviewer: fire rarely, arrive with facts rather than a persuasive summary, and go to someone with the authority to change the outcome. A record prevents nothing on its own; it makes accountability possible afterwards and, if it is read, makes the system better. The trade-off is real. Fewer gates mean more decisions taken without a person; more gates mean a person who stops reading. The gate should sit where the irreversible actions are, and almost nowhere else.

How we design for it

Actions are sorted before deployment, and the sorting is enforced in what the agent can do. Each action an agent can take in a customer's systems falls into one of four zones. Actions it takes and records: reading a claim's status, listing missing documents. Actions it takes within limits the business sets: a rebooking within fare rules, a voucher up to a value. Actions it prepares for a person to approve. And actions it is never given. In insurance, the agent never assesses liability, sets a settlement amount, or approves or declines a claim, and this holds because those decision actions are not exposed to the agent at all, not because it has been told not to use them. The business sets, and can change, which conversations go to a person, on what triggers, and which topics are off limits. In the second zone, the customer's rules, such as approval limits and entitlements, are enforced in the architecture rather than written as instructions the agent might ignore, and actions the business marks as sensitive wait for a person's approval.

The agent's confidence is a second trigger, and the handover gives the person facts, not a case. A conversation hands to a person when the agent is not confident enough; given the evidence on overconfidence, the threshold is measured per model before deployment and layered on the rules above. The agent asks a clarifying question when it is unsure rather than guessing. The person who takes over receives the full conversation and a summary, and given what explanations do to reviewers, the brief is designed to show what the customer said, what the agent did and where its uncertainty lies, not an argument for an outcome.

The record is built for a dispute in two years, not a dashboard tomorrow. Every call can be recorded, subject to consent, and transcribed as it happens; every conversation is summarised when it ends and its outcome classified; and the agent answers only from the customer's approved material, so an answer can be traced to its source. A full audit log records every action the agent took and why, in a form where any later change is detectable, with version control so every change to an agent is recorded and reversible, and retention windows the customer sets. The model version and the identity of any person who approved a gated action belong in the same record.

Review is for learning, not blame. European aviation pairs its recorders with a reporting regime built on a "just culture", in which people are not punished for honest errors and safety information may not be used "to attribute blame or liability" [32]. Every conversation can be searched. Every conversation is scored against the customer's own standard, so that review is systematic rather than sampled and feeds back into the agent's design.

What we don't know yet

No published study compares a deployment gated by consequence and reversibility with one gated by topic on the harm that actually resulted. The sorting described here follows from the evidence, but it has not been tested against the alternative.

There is no audited data on how often insurance claims reviewers or customer-service supervisors override an automated recommendation, and nobody has published how large a sample of conversations must be reviewed to hold a given error rate. What exists is litigation allegation and industry custom, not measurement.

If human accept or correct decisions are used to tune when an agent escalates, and those humans are themselves subject to automation bias, the escalation signal inherits the bias. No study appears to have measured this.

Gulf-specific evidence is thin. There is no peer-reviewed study of oversight behaviour in Arabic-language or Gulf customer service, and no published court or regulator decision in the UAE or Saudi Arabia concerning an AI system's statements to a customer could be located. Three regulatory questions were open at the time of writing: when the UAE's executive regulations will be issued, whether the DIFC's 2026 consultation on Regulation 10 has been enacted, and whether Saudi Arabia's Insurance Authority has adopted the shorter settlement window for individual claims it proposed in late 2025 [33].

Sources

  1. 1.Insurance Authority Board of Directors' Decision No. (25) of 2016 Pertinent to Regulation of the Unified Motor Vehicle Insurance Policies, UAE (in force 2017; now in the CBUAE Rulebook).⁠https://rulebook.centralbank.ae/en/rulebook/insurance-authority-board-directors-decision-no-25-2016-pertinent-regulation-unified-motor
  2. 2.NIST, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1, 2023.⁠https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf
  3. 3.ISO/IEC 42001:2023, Information technology, Artificial intelligence, Management system, Annex A control A.6.2.8, 2023.⁠https://www.iso.org/standard/42001
  4. 4.Regulation (EU) 2024/1689 (EU AI Act), Article 14, 2024.⁠https://artificialintelligenceact.eu/article/14/
  5. 5.Regulation (EU) 2024/1689, Article 7(2), 2024.⁠https://artificialintelligenceact.eu/article/7/
  6. 6.Regulation (EU) 2024/1689, Annex III, 2024.⁠https://artificialintelligenceact.eu/annex/3/
  7. 7.Regulation (EU) 2024/1689, Article 2(1)(c), 2024.⁠https://artificialintelligenceact.eu/article/2/
  8. 8.Regulation (EU) 2026/1744 (Digital Omnibus on AI), Official Journal, 24 July 2026.⁠https://eur-lex.europa.eu/eli/reg/2026/1744/oj
  9. 9.UAE Federal Decree-Law No. 45 of 2021 on the Protection of Personal Data, Article 18, 2021 (English translation, UAE legislation portal).⁠https://uaelegislation.gov.ae/en/legislations/1972
  10. 10.Chambers and Partners, Data Protection & Privacy 2026: UAE, Trends and Developments, 2026; Morgan Lewis, UAE Establishes Federal Authority for Artificial Intelligence and Data, June 2026.⁠https://practiceguides.chambers.com/practice-guides/data-protection-privacy-2026/uae/trends-and-developments ; https://www.morganlewis.com/pubs/2026/06/uae-establishes-federal-authority-for-artificial-intelligence-and-data
  11. 11.DIFC Data Protection Regulations, Consolidated Version No. 2, Regulation 10 (Personal Data Processed through Autonomous and Semi-Autonomous Systems), in force 1 September 2023.⁠https://www.difc.com/business/registrars-and-commissioners/commissioner-of-data-protection/regulation-10
  12. 12.DIFC Data Protection Law, DIFC Law No. 5 of 2020 (consolidated July 2025), Article 38.⁠https://www.difc.com/business/registrars-and-commissioners/commissioner-of-data-protection
  13. 13.Clyde & Co, Saudi Arabia issues Implementing Regulations to the Personal Data Protection Law, September 2023 (secondary source; regulator text not reachable).⁠https://www.clydeco.com/en/insights/2023/09/saudi-arabia-issues-implementing-regulations
  14. 14.Central Bank of the UAE, Guidance Note on Consumer Protection and the Responsible Adoption and Use of Artificial Intelligence and Machine Learning by Licensed Financial Institutions in the U.A.E., February 2026, sections 7(a) and 7(c).⁠https://rulebook.centralbank.ae/en/rulebook/guidance-note-consumer-protection-and-responsible-adoption-and-use-artificial-intelligence
  15. 15.Skitka, L. J., Mosier, K. L. and Burdick, M., Does automation bias decision-making?, International Journal of Human-Computer Studies 51(5), 1999; figures as restated in Skitka, Mosier and Burdick, Accountability and automation bias, IJHCS 52(4), 2000, p. 702.⁠https://www.sciencedirect.com/science/article/abs/pii/S1071581999902525 ; https://lskitka.people.uic.edu/IJHCS2000.pdf
  16. 16.Green, B. and Chen, Y., Disparate interactions: an algorithm-in-the-loop analysis of fairness in risk assessments, ACM FAT* 2019.⁠https://www.benzevgreen.com/wp-content/uploads/2019/02/19-fat.pdf
  17. 17.Bansal, G. et al., Does the whole exceed its parts? The effect of AI explanations on complementary team performance, ACM CHI 2021.⁠https://idl.cs.washington.edu/files/2021-AIExplanationsTeamPerformance-CHI.pdf
  18. 18.OWASP GenAI Security Project, OWASP Top 10 for Agentic Applications, December 2025, item ASI09.⁠https://genai.owasp.org/2025/12/09/owasp-top-10-for-agentic-applications-the-benchmark-for-agentic-security-in-the-age-of-autonomous-ai/
  19. 19.Amnesty International, Xenophobic Machines, October 2021, p. 26.⁠https://www.amnesty.nl/content/uploads/2021/10/20211014_FINAL_Xenophobic-Machines.pdf
  20. 20.Royal Commission into the Robodebt Scheme, Report, Volume 1, Overview, pp. xxiv and xxvi, Australia, July 2023.⁠https://robodebt.royalcommission.gov.au/
  21. 21.Ancker, J. S. et al., Effects of workload, work complexity, and repeated alerts on alert fatigue in a clinical decision support system, BMC Medical Informatics and Decision Making 17:36, 2017.⁠https://link.springer.com/article/10.1186/s12911-017-0430-8
  22. 22.Anthropic, Measuring AI agent autonomy in practice, February 2026 (vendor-published analysis of its own platform traffic).⁠https://www.anthropic.com/research/measuring-agent-autonomy
  23. 23.DosSantos DiSorbo, M. and Ju, H., Act or escalate? Evaluating escalation behavior in automation with language models, arXiv preprint 2604.08588, 2026 (not yet peer reviewed).⁠https://arxiv.org/abs/2604.08588
  24. 24.Zhang, Z. et al., Agent-SafetyBench: evaluating the safety of LLM agents, arXiv 2412.14470, 2024.⁠https://arxiv.org/abs/2412.14470
  25. 25.US Federal Aviation Administration, 25-Hour Cockpit Voice Recorder Requirement, New Aircraft Production, final rule, Federal Register, 2 February 2026 (FR Doc. 2026-02110), and the December 2023 proposed rule; 14 CFR 121.344 and 121.359.⁠https://www.federalregister.gov/documents/2023/12/04/2023-26144/ ; https://www.ecfr.gov/current/title-14/chapter-I/subchapter-G/part-121/subpart-K/section-121.359
  26. 26.Directive 2014/65/EU (MiFID II), Article 16(7), 2014.⁠https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX:02014L0065-20240328
  27. 27.US Securities and Exchange Commission, Rule 17a-4(f), 17 CFR 240.17a-4, as amended 2022.⁠https://www.ecfr.gov/current/title-17/chapter-II/part-240/section-240.17a-4
  28. 28.OpenTelemetry, Semantic conventions for generative AI systems, attribute registry, 2024 to 2026.⁠https://opentelemetry.io/docs/specs/semconv/registry/attributes/gen-ai/
  29. 29.Crosby, S. A. and Wallach, D. S., Efficient data structures for tamper-evident logging, USENIX Security 2009.⁠https://static.usenix.org/event/sec09/tech/full_papers/crosby.pdf
  30. 30.Moffatt v. Air Canada, 2024 BCCRT 149, Civil Resolution Tribunal of British Columbia, 14 February 2024, paras 27 to 29 and 40 to 44.⁠https://decisions.civilresolutionbc.ca/crt/crtd/en/item/525448/index.do
  31. 31.UK Financial Conduct Authority, Home and travel claims handling arrangements: good practice and areas for improvement, July 2025.⁠https://www.fca.org.uk/publications/good-and-poor-practice/home-travel-claims-handling-arrangements
  32. 32.Regulation (EU) No 376/2014 on the reporting, analysis and follow-up of occurrences in civil aviation, Articles 2(12), 15 and 16, 2014.⁠https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX:32014R0376
  33. 33.Middle East Insurance Review, Saudi Arabia: Insurance Authority proposes shorter claims settlement period, 25 November 2025; Atlas Magazine, Saudi Arabia to speed up insurance claims processing, November 2025 (secondary sources reporting the Insurance Authority's draft amendment; regulator text not reachable).⁠https://www.atlas-mag.net/en/article/saudi-arabia-to-speed-up-insurance-claims-processing

Get new research as it's published

Occasional emails when we publish.

Subscribe

Read next