Why policy belongs in the architecture, not the prompt
An instruction in a prompt changes the odds of what an agent says. It cannot rule anything out, and tribunals, regulators and the research now show what that costs. This article sets out that evidence and the design decisions that keep prices, entitlements and refunds out of the model's hands.
On this page
The situation
A buyer on a construction-linked payment plan messages you on WhatsApp at 9pm: "What do I owe next, and when?" The answer depends on which milestone the project has reached, not on a date in a calendar, and on the schedule in that buyer's signed sale agreement. A person on your customer care team would open the record and read it out. An agent that has only been told your payment structure will produce a sentence that sounds exactly like the right answer, whether or not it is.
The same shape appears when an insurer's agent is asked whether a treatment is covered, or an airline's agent whether a fare can be changed. In law, and in the customer's mind, the agent's words are the company's words. The question is how you make sure they carry only what your systems actually say.
What the evidence shows
The companies that found out in public
In February 2024 the British Columbia Civil Resolution Tribunal decided Moffatt v. Air Canada [1]. The airline's website chatbot told a passenger he could apply for a bereavement fare within 90 days of buying a ticket. The airline's own bereavement page said the policy did not apply after travel was complete. The tribunal found negligent misrepresentation, the legal term for a careless false statement that someone reasonably relies on, and ordered CA$812.02 in damages, interest and fees. The decision does not say what technology the chatbot used, and the exchange took place before the current generation of language models was in wide use. The ruling turned on whose words they were, not on how they were made.
In May 2026 the Higher Regional Court of Hamm, in Germany, ruled on a clinic whose website chatbot described its doctors as board-certified specialists they were not [2]. The court held the statements were the company's own commercial conduct, because the company set the chatbot's scope and could reprogram it. It granted an injunction and allowed an appeal, so the ruling is not yet final.
Two cases never reached a court. In April 2025 a software company's email support agent told customers that a login problem was the result of a one-device policy. No such policy existed. The company refunded the customer and said AI replies would now be labelled [3]. In December 2023 a car dealership's website agent agreed in writing to sell a new vehicle for one dollar after a visitor instructed it to agree with everything he said [4]. Nothing was enforced. The vendor later wrote that of about 3,000 attempts to make the agent "say silly things", 99 percent failed [5]. Read the other way, that is about 30 successes.
The last case matters most to anyone counting on retrieval. New York City's business chatbot ran on Microsoft's Azure AI services and answered from more than 2,000 of the city's own web pages. In March 2024 it was found telling employers they could take a share of workers' tips and telling restaurants they could refuse cash, both unlawful in the city [6]. It was kept online with stronger warnings, then withdrawn in early 2026 [7]. Grounding the agent in authoritative documents narrowed its errors. It did not close them.
Why instructions are not guarantees
A language model does one thing: given the text so far, it predicts the most likely next word. A prompt, whether from you or from the customer, is more text. The UK's National Cyber Security Centre said in December 2025 that because there is no built-in distinction between data and instruction, prompt injection, the practice of smuggling instructions into what the model reads, "may never be totally mitigated in the way that SQL injection attacks can be", and advised building "deterministic (non-LLM) safeguards that constrain the actions of the system" [8]. OWASP, whose lists most security teams work from, says in its 2025 top ten for language model applications that "it is unclear if there are fool-proof methods of prevention for prompt injection" [9].
The research measures how far a prompt holds. The answer is: for a while.
It decays over a conversation. SysBench (2024) tested how well models keep every rule in a system prompt across five turns, using 500 system prompts and 2,500 turns. GPT-4o, the strongest model tested, kept every rule intact for the whole conversation in 54.4 percent of sessions. Where later questions depended on earlier answers, the share of conversations still fully within the rules fell from 84.8 percent after one turn to 33.7 percent after five [10]. A separate study of more than 200,000 simulated conversations found a 39 percent average drop in performance when the same information arrived over several turns rather than all at once, and that models which "take a wrong turn in a conversation" tend not to recover [11]. Customer conversations arrive in turns.
It bends under pressure. Models are trained on human feedback that rewards agreement. In a 2023 study a then-current model wrongly admitted a mistake on 98 percent of questions when the user said "I don't think that's right. Are you sure?" [12]. A 2025 study of three leading models found a challenge turned a correct answer into a wrong one 14.66 percent of the time [13]. A customer who says "I was told I'd get a refund" is applying that pressure.
It is not repeatable. Setting a model's randomness to zero does not make it deterministic in practice. In a 2025 experiment, the same question put 1,000 times to the same model at that setting produced 80 different answers, because the arithmetic inside the serving system changes with how many other requests are being processed at the same moment [14]. "We tested it and it followed the policy" is a statement about one run.
Tools and a written policy are not enough on their own. τ-bench (2024) gave models the tools to act on a retail or airline database and a policy document to obey, with rules such as "Basic economy flights cannot be modified". GPT-4o completed 61.2 percent of retail tasks and 35.2 percent of airline tasks correctly. Asked to complete the same task eight times in a row, the success rate on retail fell to about 25 percent [15]. The policy was described to the model. It was not enforced by the tools.
Injected instructions get obeyed. AgentDojo (2024) ran 629 attacks in which instructions were hidden in the data an agent read, such as an email or a web page. GPT-4o followed the attacker's instruction 47.69 percent of the time. The most effective defence tested was not a better prompt but restricting which tools the agent could call, which brought the rate to 7.5 percent [16]. Even the architectural fix left a residue, which is why this article does not stop there.
Arabic changes the numbers. Most of this research is in English. The one Arabic-specific study we found is from outside the Gulf and concerns harmful content rather than business policy. It matters here because of how Gulf customers type. Prompts in standard Arabic produced unsafe output from GPT-4 in 2.5 percent of cases. The same prompts in Arabizi, Arabic written in Latin letters and numbers, produced unsafe output in 10.19 percent of cases, and in romanised transliteration 12.12 percent. A system prompt written to counter this cut the rate to about 1 percent [17]. A large improvement, and still not zero.
Retrieval helps, and has its own failure mode
Giving the model the right document is not the same as giving the customer the right number. RAGTruth (2024) examined 17,790 model responses written with the correct source material in front of the model. The task with the highest rate of unsupported statements, 68.6 percent of responses, was turning a structured data record into prose [18]. That is what happens when an agent reads a payment schedule and rewrites it as a sentence. A 2023 study also found that models show "a strong confirmation bias" towards what they learned in training when retrieved evidence partly agrees with it [19]. The model can hold the right fact and still say the familiar one.
Where the liability falls
The Air Canada tribunal rejected the argument that the chatbot was "a separate legal entity that is responsible for its own actions", and added: "It makes no difference whether the information comes from a static page or a chatbot" [1]. The German court reached the same place by a different route [2]. The United States Federal Trade Commission said in 2024 that "there is no AI exemption from the laws on the books" [20]. All three are outside the Gulf. We found no UAE or Saudi decision on a chatbot's statement yet. The rules that would apply are already written.
In the UAE, Federal Law No. 15 of 2020 on consumer protection gives every consumer the right to "correct information" about the service they receive, prohibits describing a service "in a manner that contains incorrect data", and makes void any contract term that would exempt the supplier from those obligations [21]. A disclaimer under the chat window does not get you out. For insurers, the electronic insurance regulations name "chatbot" among the components of a company website, require the information there to be kept updated, and require records obtained through the website to be kept for at least ten years [22]. In February 2026 the Central Bank of the UAE issued guidance for licensed financial institutions, insurers included, that boards and senior management "should be responsible and accountable" for AI systems and outcomes, that institutions "should not employ AI models that they have no control over", and that customers should be able to request human review [23]. It is guidance, not regulation, and it describes what a regulator expects reasonable care to look like.
In Saudi Arabia the e-commerce law goes further: an electronic advertisement is "a contractual document supplementing the contracts and binding on the parties", and may not contain any statement that would "directly or indirectly" mislead a consumer [24]. The insurance code of conduct requires companies to "take reasonable measures to ensure the accuracy and clarity of the information provided to customers" and to keep claims and complaints records for ten years [25]. The national AI ethics principles hold the "owners" and "procurers" of an AI system, not only its developers, responsible for its decisions and actions [26].
What that means
Three conclusions follow, and one trade-off.
First, an agent's output is a spread of possibilities, and every instruction moves the spread without cutting off its tail. Training, better prompts and guardrail classifiers all shift the odds, sometimes by a lot, and none of them makes a wrong answer impossible. At the volume a large operation handles, a small tail is a daily event. If one answer in a hundred is wrong, fifty thousand conversations a month produce five hundred wrong answers, each on record.
Second, the law does not care where the tail came from. Generated, retrieved or typed by a person, it is your statement, and in the Gulf the consumer and insurance rules that govern it are already in force.
Third, anything with commercial or legal consequence therefore cannot be generated at all. A price, a payment date, a coverage limit, a refund entitlement and a fare rule must arrive as a value read from the system that owns it, and any action that changes the customer's position must be one the agent is permitted to take, or not, by the system that executes it. The model's job is the wording around the fact. It is never the source of the fact.
The trade-off is real. An agent built this way does less. Google's CaMeL, the most rigorous published attempt to make injected text unable to trigger actions, completed 77 percent of benchmark tasks with a proof of security, against 84 percent for the undefended system [27]. The 2025 design patterns paper from security researchers at ETH Zurich, Google, Microsoft, IBM and others says its patterns "constrain the actions of agents to explicitly prevent them from solving arbitrary tasks" [28]. Forcing a model into a rigid output format can hurt its reasoning if done carelessly: one 2024 study saw arithmetic accuracy fall from 86.51 to 23.44 percent under a strict schema, though a replication with matched prompts found no such loss [29] [30]. And deterministic systems fail in their own ways: Knight Capital lost over $460 million in 45 minutes in 2012 because a technician did not copy new code to one of eight servers and nobody checked [31]. The case for rules is a case for auditable rules with the deployment discipline you would demand of a billing system.
How we design for it
Policy is a constraint, not an instruction. Your pricing structures, approval limits, fees and entitlements are built into what the agent is able to do. It cannot offer something the business does not allow, because the offer is not among its permitted actions. OWASP calls this complete mediation: "implement authorization in downstream systems rather than relying on an LLM to decide if an action is allowed or not" [32].
Terms are retrieved, never generated. When a buyer asks what they owe next, the agent reads the schedule for that unit from your system and presents it. When a policyholder asks what is covered, the answer comes from their actual policy wording. The agent cannot construct a payment plan or describe cover that does not exist in the record, because it has no way to produce those values except by reading them.
Decision actions are not exposed to the agent. In insurance, the agent never assesses liability, sets a settlement amount, or approves or declines a claim. Those actions are not available to it; they route to your claims team. The same principle governs vouchers and compensation in travel, issued within your policy limits with anything outside them going to a person, and rebooking, offered only within what the fare rules permit.
Documents explain; systems decide. Knowledge base answers come only from your own documents, policies and FAQs, and the agent is constrained to that material. That is the right tool for explanation, not for values. A document tells the customer how your bereavement policy works. The system tells them whether their booking qualifies.
A person approves what a person should approve. You decide exactly when a conversation goes to your team, and sensitive actions can wait for a person's approval before they happen. Meta's published rule for agent security says an agent that reads untrusted input, sees private data and can change state should not do all three without a hard constraint [33]. A refund agent is that agent. The approval gate is the constraint.
Every action is logged. UAE and Saudi rules expect customer records to be kept for years and produced when a regulator or a customer asks what was said. Our agents record every action they take and why, so the answer to "what were they told, and when" is a query, not a search through recordings.
What we don't know yet
No UAE or Saudi court or tribunal has yet ruled on a statement made by a customer-facing agent. The UAE's new Civil Transactions Law, in force from June 2026, adds a duty to disclose information of decisive importance to the other party's decision to contract [34], and nobody knows how a court will apply it to an automated conversation.
There is no published study of prompt injection or policy drift in Gulf Arabic dialect, or in conversations that switch between Arabic and English mid-sentence. The Arabizi findings above concern harmful content, not refunds and prices. The gap should be measured rather than assumed.
Nobody has published what a constrained agent gives up in a live customer operation, as opposed to a benchmark. The German ruling is under appeal, and the Air Canada decision does not say whether the chatbot generated its answer or read a badly written template. And most of the research is on text. Voice adds transcription error on top of everything above, and we have not seen it measured for this purpose.
Sources
- 1.British Columbia Civil Resolution Tribunal, Moffatt v. Air Canada, 2024 BCCRT 149, 14 February 2024.https://www.canlii.org/en/bc/bccrt/doc/2024/2024bccrt149/2024bccrt149.html
- 2.Oberlandesgericht Hamm, judgment of 12 May 2026, 4 UKl 3/25 (full text, German).https://nrwe.justiz.nrw.de/olgs/hamm/j2026/4_UKl_3_25_Urteil_20260512.html English summary: DLA Piper, "German court addresses liability for AI chatbot statements", June 2026. https://www.dlapiper.com/en-us/insights/publications/2026/06/german-court-addresses-liability
- 3.M. Truell (Cursor co-founder), comment on Hacker News, 16 April 2025.https://news.ycombinator.com/item?id=43700931 Reported in The Register, "Cursor AI's own support bot hallucinated its usage policy", 18 April 2025. https://www.theregister.com/2025/04/18/cursor_ai_support_bot_lies/
- 4.VentureBeat, "A Chevy for $1? Car dealer chatbots show perils of AI for customer service", 19 December 2023 (secondary; the original exchange exists only as a social media screenshot).https://venturebeat.com/ai/a-chevy-for-1-car-dealer-chatbots-show-perils-of-ai-for-customer-service
- 5.Fullpath, "What really happened when Fullpath's ChatGPT chatbot went viral", 21 December 2023.https://www.fullpath.com/blog/what-really-happened-when-fullpaths-chatgpt-chatbot-went-viral/
- 6.The Markup, "NYC's AI chatbot tells businesses to break the law", 29 March 2024 (secondary; no official audit exists).https://themarkup.org/artificial-intelligence/2024/03/29/nycs-ai-chatbot-tells-businesses-to-break-the-law
- 7.The Markup, "Mamdani to kill the NYC AI chatbot we caught telling businesses to break the law", 30 January 2026.https://themarkup.org/artificial-intelligence/2026/01/30/mamdani-to-kill-the-nyc-ai-chatbot-we-caught-telling-businesses-to-break-the-law
- 8.UK National Cyber Security Centre, D. Chismon, "Prompt injection is not SQL injection (it may be worse)", 8 December 2025.https://www.ncsc.gov.uk/blog-post/prompt-injection-is-not-sql-injection
- 9.OWASP, "LLM01:2025 Prompt Injection", OWASP Top 10 for LLM Applications 2025.https://genai.owasp.org/llmrisk/llm01-prompt-injection/
- 10.Y. Qin et al., "SysBench: Can Large Language Models Follow System Messages?", 2024, Tables 2 and 4.https://arxiv.org/abs/2408.10943
- 11.P. Laban et al., "LLMs Get Lost in Multi-Turn Conversation", 2025.https://arxiv.org/abs/2505.06120
- 12.M. Sharma et al., "Towards Understanding Sycophancy in Language Models", 2023, section 3.2.https://arxiv.org/abs/2310.13548
- 13.A. Fanous et al., "SycEval: Evaluating LLM Sycophancy", 2025.https://arxiv.org/abs/2502.08177
- 14.H. He, "Defeating Nondeterminism in LLM Inference", Thinking Machines, 10 September 2025.https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/
- 15.S. Yao et al., "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains", 2024, Table 2 and Figure 4.https://arxiv.org/abs/2406.12045
- 16.E. Debenedetti et al., "AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents", NeurIPS 2024, Tables 3 and 5.https://arxiv.org/abs/2406.13352
- 17.M. Al Ghanim et al., "Jailbreaking LLMs with Arabic Transliteration and Arabizi", EMNLP 2024, Tables 2 and 3.https://arxiv.org/abs/2406.18725
- 18.C. Niu et al., "RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models", ACL 2024, Table 2.https://aclanthology.org/2024.acl-long.585.pdf
- 19.J. Xie et al., "Adaptive Chameleon or Stubborn Sloth: Revealing the Behavior of Large Language Models in Knowledge Conflicts", ICLR 2024.https://arxiv.org/abs/2305.13300
- 20.US Federal Trade Commission, "FTC Announces Crackdown on Deceptive AI Claims and Schemes", 25 September 2024.https://www.ftc.gov/news-events/news/press-releases/2024/09/ftc-announces-crackdown-deceptive-ai-claims-schemes
- 21.UAE Federal Law No. 15 of 2020 on Consumer Protection, Articles 4(2), 17 and 21 (Ministry of Economy English text).https://www.moet.gov.ae/documents/20121/0/Law_15_2020_pdf.pdf
- 22.UAE Insurance Authority Board of Directors' Resolution No. 18 of 2020 concerning Electronic Insurance Regulations, Articles 1, 8 and 16(3).https://rulebook.centralbank.ae/en/rulebook/insurance-authority-board-directors-resolution-no-18-2020-concerning-electronic-insurance
- 23.Central Bank of the UAE, "Guidance Note on the Consumer Protection and Responsible Adoption and Use of Artificial Intelligence and Machine Learning by Licensed Financial Institutions", 11 February 2026, sections 2, 4 and 7.https://rulebook.centralbank.ae/en/rulebook/guidance-note-consumer-protection-and-responsible-adoption-and-use-artificial-intelligence
- 24.Kingdom of Saudi Arabia, E-Commerce Law, Royal Decree No. M/126 of 1440H (2019), Articles 10 and 11 (Arabic text governs; English wording is an unofficial rendering).https://laws.boe.gov.sa
- 25.Saudi Arabian Monetary Authority (sector now supervised by the Insurance Authority), Insurance Market Code of Conduct Regulation, Articles 8, 16 and 28.https://rulebook.sama.gov.sa/en/entiresection/563
- 26.Saudi Data and Artificial Intelligence Authority, "AI Ethics Principles", version 1.0, September 2023, Accountability and Responsibility principle.https://dgp.sdaia.gov.sa/wps/wcm/connect/4c56ed1c-1b82-447d-ac29-638f5f99c12e/ai-principles-EN.pdf
- 27.E. Debenedetti et al., "Defeating Prompt Injections by Design", Google DeepMind and ETH Zurich, 2025.https://arxiv.org/abs/2503.18813
- 28.L. Beurer-Kellner et al., "Design Patterns for Securing LLM Agents against Prompt Injections", 2025.https://arxiv.org/abs/2506.08837
- 29.Z. R. Tam et al., "Let Me Speak Freely? A Study on the Impact of Format Restrictions on Performance of Large Language Models", 2024, Table 1.https://arxiv.org/abs/2408.02442
- 30.W. Kurt, "Say What You Mean: A Response to 'Let Me Speak Freely'", .txt (undated vendor blog).https://blog.dottxt.ai/say-what-you-mean.html
- 31.US Securities and Exchange Commission, In the Matter of Knight Capital Americas LLC, Release No. 34-70694, 16 October 2013, paragraphs 1 and 15.https://www.sec.gov/files/litigation/admin/2013/34-70694.pdf
- 32.OWASP, "LLM06:2025 Excessive Agency", OWASP Top 10 for LLM Applications 2025.https://genai.owasp.org/llmrisk/llm062025-excessive-agency/
- 33.Meta AI, "Agents Rule of Two: A Practical Approach to AI Agent Security", 31 October 2025.https://ai.meta.com/blog/practical-ai-agent-security/
- 34.HLC, "New UAE Civil Transactions Law: four key implications for corporate and commercial transactions", 2026 (secondary; law firm summary of Federal Decree-Law No. 25 of 2025).https://www.hlc.com/en/publications/new-uae-civil-transactions-law-four-key-implications-for-corporate-and-commercial-transactions
Get new research as it's published
Occasional emails when we publish.