Insight report · September 2026
All insight reportsWhether an AI-assisted result can be relied on is a question an organisation can answer, but not from a demonstration or a supplier's figures. It turns on whether the answer can be traced to its source, whether the system has been tested on the organisation's own work, and whether a named person with the authority to change the outcome checks it. This report sets out the evidence for each.
The short version
AI assistants get a material share of ordinary questions wrong, and they rarely decline to answer. In an audit by 22 public broadcasters, coordinated by the European Broadcasting Union and the BBC, 45 per cent of assistants' answers to news questions had at least one significant problem, most often with sourcing, and the assistants declined just 17 of 3,113 questions. Rates vary widely between tools and tasks, and they are improving, which is why a figure is only useful with its date.
Giving a system your own documents helps, but it does not remove the need to check. In a peer-reviewed study, specialist legal research tools that search their own databases still gave hallucinated answers to between 17 and 33 per cent of queries, and the errors the authors judged potentially more dangerous than an invented case were those that cited a real source that did not support the claim. The High Court has said that those who use AI for legal research in their professional work have a duty to check it against authoritative sources, and a public tracker records 69 UK decisions, as at 19 September 2026, in which a court or tribunal found, or implied, that a party had relied on hallucinated material.
Responsibility stays with people. UK data protection law, whose rules on automated decisions were rewritten with effect from February 2026, requires safeguards for significant decisions about people made with no meaningful human involvement. The regulator's draft guidance says a real check comes before the decision takes effect, is made by someone with the authority to change it, and covers every decision, not a sample. Accredited certification now exists for how an organisation governs AI, but that certificate does not show whether a particular system's answers are right. That still takes testing on the organisation's own work, and a named person who checks.
What the measurements show
The most useful measurements of AI error come from audits of realistic tasks run outside the AI companies, and one of the largest recent ones is unusually clear. In late May and early June 2025, journalists from 22 public service broadcasters in 18 countries put a shared set of news questions to four widely used AI assistants, in their free versions, asking them to use the broadcaster's own sources where possible, and then rated the answers.
Of 2,709 answers, 45 per cent had at least one significant issue, and 81 per cent had an issue of some kind. Sourcing was the biggest single cause, at 31 per cent, followed by accuracy at 20 per cent and missing context at 14 per cent. The share of answers with significant issues ranged from 30 to 76 per cent depending on the assistant. The broadcasters have their own interest in how AI assistants use their journalism, which is worth keeping in mind.
Two further findings from the same audit matter more for a business than the headline. The assistants declined to answer just 17 of 3,113 questions, a rate the report reads as a willingness to answer whether or not the assistant is capable of a high-quality answer. And the tools are improving: in the BBC's own comparison, the share of answers with significant issues fell from 51 to 37 per cent between December 2024 and mid-2025. Any error rate is a snapshot of particular tools on a particular date.
Other measurements point the same way on narrower tasks. When the Tow Center at Columbia University asked eight AI search tools to identify the source of quoted news excerpts in early 2025, they answered more than 60 per cent of queries incorrectly. A public leaderboard that scores models on summarising documents they have been given finds unsupported content in between 1.8 and 24.2 per cent of summaries, depending on the model, although it is scored by the operator's own commercial model.
Some model developers publish only relative improvements: one 2026 system card reports that its model's claims are more likely to be correct than its predecessor's, measured on conversations chosen because they were prone to error.
None of these figures describes a particular organisation's work. They show that error is material, that it varies widely by tool and task, and that it changes quickly. The only rate that matters for a decision is the rate on the organisation's own cases, which nobody else can measure for it.
Trace the information
The standard remedy for AI error is to ground the system in the organisation's own documents and make it cite them. The evidence supports doing so, with a caution. A preregistered, peer-reviewed study of specialist legal research tools, which retrieve from their publishers' own databases, found they hallucinated less than a general chatbot but still did so on between 17 and 33 per cent of queries.

The authors concluded that grounding "can reduce hallucinations" but that they "remain substantial, wide-ranging, and potentially insidious". The tools were tested in 2024 and have changed since. Studies in other fields, such as medicine, have also found that retrieval improves accuracy, but on different tasks, so the size of the benefit on an organisation's own documents has to be tested.
The subtler case is the one to design against. The study calls it misgrounding: an answer that cites a real source for a claim the source does not support, or cites a source that does not apply, and judges it potentially more dangerous than inventing a case outright, because it is harder to spot. A citation that exists is easy to mistake for a citation that has been checked.
The High Court described the same failures in 2025: AI tools "may cite sources that do not exist" and "may purport to quote passages from a genuine source that do not appear in that source". Independent tests of AI search tools have also found fabricated or broken links, and wrong answers given with confidence.
Tracing the information therefore means three practical things. Every claim in an AI-assisted output should carry its source, so a reviewer can open it. Someone should check that the source says what the answer claims, at least for anything that will be relied on. And the system should say plainly when its sources do not cover a question. That behaviour has to be designed and tested, because the evidence above suggests assistants will otherwise fill the gap.
Test the behaviour on your own work
A convincing demonstration shows how a system behaves on the examples chosen for it. It cannot show how often the system will be right on the organisation's own cases. The US National Institute of Standards and Technology puts it directly in its guidance on generative AI: avoid extrapolating performance "from narrow, non-systematic, and anecdotal assessments", and do not treat success on tests designed for humans, such as professional exams, as proof of validity or reliability. What a representative evaluation can establish is that a system is valid and reliable under the conditions it will be used in, with the limits of that finding written down.
Even good tests have limits. The International AI Safety Report 2026, written with guidance from more than 100 independent experts, including nominees from more than 30 countries and international organisations, found that "performance on pre-deployment tests does not reliably predict real-world utility or risk", and that it has become more common for models to distinguish test settings from real use. It notes that AI systems can be made more robust by layering several safeguards, an approach known as defence in depth.
Some risks cannot be tested away at all. The National Cyber Security Centre says prompt injection, where instructions hidden in a document or email hijack an AI system, may never be totally mitigated, so the aim is to reduce the risk and the impact by design.
In practice a test worth relying on is built from the organisation's own work: a sample of real cases, including the awkward and incomplete ones, with the right answers agreed in advance. It is run before launch and again whenever the model, the prompts or the documents change. The UK government's AI Playbook similarly advises testing fully before deployment and keeping regular checks on the live tool.
Free tools exist for running such evaluations, including Inspect, an open-source framework developed by the UK AI Security Institute and Meridian Labs. The demonstrations BespokeWorks builds are made to show how a workflow could function; the evaluation on real cases is a separate step, and it is the one that answers whether to rely on it.
Keep responsibility with a named person
Courts and regulators are consistent that using AI does not move responsibility onto the tool. In June 2025 the Divisional Court, in Ayinde v London Borough of Haringey and a linked case, said that freely available generative AI tools "are not capable of conducting reliable legal research", and that those who use AI for legal research have "a professional duty" to check it against authoritative sources before using it in their work. In one of the two cases, 18 of the 45 citations put before the court did not exist. The court referred the barrister and the supervising solicitor in the first case, and the solicitor in the second, to their regulators.
A public tracker maintained by the researcher Damien Charlotin records 2,045 decisions worldwide, 69 of them in the UK, in which a court or tribunal found, or implied, that a party had relied on hallucinated material.
Regulators say the same outside the courtroom. The Solicitors Regulation Authority's 2026 warning notice states that "AI has no separate legal personality", so solicitors remain accountable for their work however it was prepared. The Financial Conduct Authority has chosen not to write AI-specific rules and relies on existing ones, including the accountability of senior managers. The judiciary's own guidance, updated in October 2025, stresses the personal responsibility of judges for material produced in their name.
There is an older lesson in English law. In criminal proceedings, the courts presume that a computer was operating correctly unless there is evidence to the contrary, which the Ministry of Justice's own call for evidence summarised in January 2025 as "the computer is always right". The government said the Post Office Horizon scandal had shown the limits of that presumption starkly, and that any reform should cover evidence generated by software, "including Artificial Intelligence and algorithms".
The call for evidence closed in April 2025. We could find no published government response by September 2026, so the presumption still stands. For an organisation, the practical point is that an AI-assisted record may be believed by default, which makes checking it before it is relied on more important, not less.
Make the human check a real one
Where AI helps make decisions about people, UK law now sets out what follows when there is no real human check. The rules on automated decisions that the Data (Use and Access) Act 2025 inserted into the UK GDPR, Articles 22A to 22D, have applied in full since 5 February 2026.
A decision is solely automated when there is "no meaningful human involvement" in taking it, and a significant decision of that kind must come with safeguards: information about the decision, the chance to make representations, human intervention and the right to contest it. The law does not require human involvement as such. It attaches duties to decisions taken without it.
The Information Commissioner's Office set out its practical test in draft guidance in March 2026, which had not been finalised by September 2026.
The involvement has to come before the decision is applied, while it can still change. The person involved needs the ability to influence the outcome, the authority to alter it, and the training to understand the system's logic and limits. And it has to happen every time: the draft says ad hoc spot-checking is not sufficient, because some decisions would receive no check at all. A person who only enters data for the system to decide on does not count. The regulator's functions pass to a new Information Commission on 30 September 2026.
A check can also be weakened by the person doing it. In a peer-reviewed but observational study, doctors who had grown used to AI assistance in colonoscopy found fewer abnormalities when working without it, a detection rate that fell from 28.4 to 22.4 per cent.
A survey of 319 knowledge workers found that those with higher confidence in AI reported less critical thinking. And in a small trial in early 2025, experienced software developers believed AI had made them about 20 per cent faster when it had made them 19 per cent slower, although the same researchers said in February 2026 that developers were probably being sped up more by then, while cautioning that their newer data were unreliable. None of these proves that AI makes reviewers careless. Together they explain why oversight needs time, training and the authority to disagree.
Keep a record, and plan for the day it goes wrong
The guidance converges on a small set of records. The regulator's draft guidance says to keep a record of how the human reviewed each decision, to give the reviewer all the data and original facts the system used, and to monitor the outcomes of reviews over time. The National Cyber Security Centre advises logging enough to identify suspicious activity, potentially including the full input and output of the model and the tools it used, which itself needs retention rules where personal data is involved.
Central government departments, and the arm's-length bodies in scope, must publish records of their in-scope algorithmic tools under the Algorithmic Transparency Recording Standard; the wider public sector is encouraged to.
We found no single UK regime for reporting AI incidents. The practical guidance is to let users report a problem and have it prompt a human review, as the government's AI Playbook advises, to keep logs good enough to reconstruct what happened, and to be able to pause the tool. From December 2027 for uses such as hiring and education, and from August 2028 for AI in regulated products, EU deployers of high-risk systems must suspend a system that may present a risk and tell its provider and the market surveillance authority.
Certification helps with some of this and not the rest. Since January 2026, UKAS-accredited certification has been available in the UK for AI management systems against the international standard ISO/IEC 42001, with BSI the first body accredited. It attests to how an organisation governs AI. The government's own assurance roadmap says plainly that "certification of AI products is beyond the scope" of its work, so an ISO/IEC 42001 certificate does not say that a particular system's answers can be relied on.
A record worth keeping covers five things.
- The inputs and sources the system used for the output.
- The system and version, and the prompts or instructions in force.
- Who reviewed it, when, and what they changed.
- Any challenge or complaint, and how it was resolved.
- How a problem is reported, who can pause the tool, and who must be told.
What we could not establish
Independent benchmarks of how often models answer wrongly from their own memory now exist, including one published in November 2025, but they use set questions, not an organisation's own work, and developers report their own figures in different ways.
Grounding has been shown to reduce errors in law and to improve accuracy in medical question answering, but on tasks unlike most organisations' work. We could not establish how large the benefit is on an organisation's own documents.
The best experimental evidence of people failing to catch AI errors, as opposed to observational and survey evidence, was not established in our research. Nor could we find a UK-wide scheme for reporting AI incidents, a final version of the regulator's guidance on meaningful human involvement, or a government response on the presumption that computers are reliable.
How it was researched
The report draws on primary sources read in September 2026: legislation, court judgments, regulators' and government guidance, peer-reviewed studies and the audits' own reports. Every figure carries its date, because error rates move quickly. Measurements run by a company with an interest, such as a leaderboard scored by its operator's own model, are labelled as such, and no tool is ranked. Before publication, a separate reviewer checked every claim against its source and tried to disprove it. This is general guidance, not legal advice.
Questions for your team
- What evidence would justify relying on the proposed output?
- Which decisions need an accountable person to review them, and do they have the time and authority to?
- How would the organisation identify, record and respond to a failure?
Sources
Links open the primary source. Dates are publication dates, or the date read where a source changes daily.
Measurements of AI error
- European Broadcasting Union and BBC, News Integrity in AI Assistants, October 2025
- Tow Center for Digital Journalism, Columbia Journalism Review, We compared eight AI search engines, 6 March 2025
- Vectara, Hallucination Leaderboard, updated 11 May 2026
- Magesh and others, Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools, Journal of Empirical Legal Studies, 2025
- Xiong and others, Benchmarking Retrieval-Augmented Generation for Medicine, 2024
- OpenAI, GPT-5.5 System Card, 23 April 2026
- Jackson and others, AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models, 17 November 2025
Courts, regulators and responsibility
- R (Ayinde) v London Borough of Haringey; Al-Haroun v Qatar National Bank [2025] EWHC 1383 (Admin), 6 June 2025
- Damien Charlotin, AI Hallucination Cases database, read 19 September 2026
- Solicitors Regulation Authority, Misuse of AI: warning notice, 2026
- Financial Conduct Authority, AI and the FCA: our approach, updated 13 February 2026
- Courts and Tribunals Judiciary, Artificial Intelligence (AI) Judicial Guidance, October 2025
- Ministry of Justice, The use of evidence generated by software in criminal proceedings: call for evidence, January 2025
- GOV.UK, call for evidence page (open 21 January to 15 April 2025)
Testing and assurance
- National Institute of Standards and Technology, AI Risk Management Framework: Generative AI Profile (NIST AI 600-1), July 2024
- International AI Safety Report 2026, executive summary, 3 February 2026
- UK AI Security Institute, Inspect evaluation framework
- Government Digital Service, AI Playbook for the UK Government, February 2025
- National Cyber Security Centre, Prompt injection is not SQL injection (it may be worse), 8 December 2025
- Department for Science, Innovation and Technology, Trusted third-party AI assurance roadmap, 3 September 2025
- UKAS, UKAS grants first accreditation for ISO/IEC 42001, 15 January 2026
Oversight and records
- UK GDPR Articles 22A to 22D, legislation.gov.uk
- Information Commissioner's Office, draft guidance: What does the UK GDPR say about ADM?, March 2026
- Information Commissioner's Office, draft guidance: What are the ADM safeguards?, March 2026
- The Data (Use and Access) Act 2025 (Commencement No. 9 and Transitional and Saving Provisions) Regulations 2026, SI 2026/1015
- European Commission, AI Act: regulatory framework, updated 3 August 2026
- European Commission AI Act Service Desk, Article 14 (human oversight)
- European Commission AI Act Service Desk, Article 26 (obligations of deployers)
- Government Digital Service, Algorithmic Transparency Recording Standard hub
Over-reliance
- Budzyń and others, Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy, Lancet Gastroenterology and Hepatology, 2025
- Lee and others, The Impact of Generative AI on Critical Thinking, CHI 2025
- METR, Measuring the impact of early-2025 AI on experienced open-source developer productivity, 10 July 2025
- METR, We are changing our developer productivity experiment design, 24 February 2026

