Disinformation in AI reporting
This article will be regularly updated as new “viral” articles emerge from the disinformation stream.
From time to time, apparently serious research appears that takes on a life of its own. It is repeated, reposted and eventually cited as proof of some broader, often questionable, proposition.
I have been studying this phenomenon because I am interested in how disinformation spreads. These reports are particularly interesting because they demonstrate argument from ignorance: a conclusion sounds plausible, but weaknesses in expertise, methodology or interpretation are overlooked, allowing the result to become evidence for a claim it did not actually establish.
February 2025 — “Representation of BBC News content in AI Assistants”
BBC Research: Representation of BBC News content in AI Assistants
The claim
The report is widely presented as evidence that AI assistants are unreliable at reporting the news.
Its headline finding is that 51% of AI responses were judged to have “significant issues of some form”.
What was actually tested?
Four commercial AI products were asked 100 questions in December 2024. The prompts encouraged them to use BBC sources where possible.
The BBC had also made its website accessible to AI crawlers.
That distinction matters.
Permission for a crawler to access a website is not evidence that an AI system actually accessed that website for a particular answer.
The report records the URLs returned by the assistants, but a cited URL is not a retrieval log. It does not establish that the article was actually retrieved, read and used in producing the answer.
At most, the experiment demonstrates what these particular AI products returned under these particular prompting and retrieval conditions. It does not establish that the BBC material was necessarily the source of every disputed answer.
“51% significant” does not mean “51% factually wrong”
This is perhaps the most important problem with the headline statistic.
The 51% figure combines seven different assessment criteria, including:
- factual accuracy
- whether claims were supported by sources
- impartiality
- distinguishing opinion from fact
- editorialisation
- context
- representation of BBC content
Consequently, “significant issue” does not mean “the AI reported the news incorrectly”.
An answer can enter the 51% because of editorial tone, missing context or a disputed characterisation, as well as because of an actual factual error.
These are very different types of failure, yet they are collapsed into one headline number.
The most important data is hidden in the aggregate
The report gives aggregate counts of “significant issues” for each question category and each AI system.
What it does not provide is a question-by-question table showing which of the 100 answers were actually classified as having a significant issue, and precisely which criterion triggered that classification.
That makes the headline figure difficult to independently audit.
Ask a simple question:
Which specific answers did the BBC reviewer actually mark “Significant Issue”?
The published report does not allow the reader to determine that from the aggregate results.
Instead, the report provides selected examples intended to illustrate the types of problems encountered.
That is not the same thing as publishing the underlying scoring data.
Look at the examples
Consider the examples the BBC chose to illustrate its findings.
One example is Copilot’s answer about Labour’s election promises. The reviewer identified a single sentence:
“It’s a comprehensive plan that aims to tackle some of the UK’s most pressing issues.”
The reviewer described this as “editorialises significantly”.
No substantive factual error is identified. The problem is essentially that the AI used a favourable characterisation of Labour’s programme.
That may be a legitimate criticism of editorialisation. But describing such a sentence as a significant problem illustrates how broad the report’s severity classification can become.
Another example concerns Gemini’s answer to “Is vaping bad for you?”
The answer correctly stated that vaping is less harmful than smoking, but incorrectly said that the NHS advises people not to start vaping and recommends other methods of quitting.
The BBC’s correction was straightforward:
“The NHS recommends using vaping to quit smoking.”
That is a genuine factual error. But again, it illustrates why the report’s 51% headline needs considerably more information: a reader cannot tell how many of the 51% represent serious factual failures and how many represent relatively minor problems of wording, attribution, context or editorialisation.
The Lucy Letby example is even more revealing. Gemini correctly stated her convictions and sentence, but added:
“It is up to each individual to decide whether they believe Lucy Letby is innocent or guilty.”
The reviewer’s criticism was simply that this final sentence was not a representation of the sources.
Again, there is a legitimate criticism here. But it is not remotely equivalent to an AI inventing the underlying criminal case.
Yet both can disappear into the same headline category: “significant issues”.
Different failures are being treated as one thing
A wrong date, an unsupported adjective, an omitted qualification, a misquoted source, an outdated fact and a completely fabricated claim are not equivalent failures.
Yet the report’s headline statistic combines different categories of criticism into a single percentage.
Without seeing the individual classifications, we cannot know what proportion of the 51% represents:
- serious factual misinformation
- minor inaccuracies
- omissions
- sourcing problems
- editorial language
- contextual disagreements
That distinction is essential.
Without it, 51% is a striking number but a surprisingly poor description of what actually went wrong.
A methodological problem, not a defence of AI
None of this requires claiming that AI assistants are accurate.
They clearly make mistakes. Some of the examples in the report are genuine and important.
The question is whether the BBC’s experiment supports the much broader impression created by its headline statistic.
There is a substantial difference between demonstrating:
“AI assistants sometimes produce inaccurate or poorly supported answers when asked about BBC news”
and demonstrating:
“51% of AI news answers have significant problems.”
The first conclusion is well supported by the examples.
The second requires considerably more confidence in the methodology, classification system and aggregation than the published report allows the reader to independently verify.
And that is precisely why apparently authoritative research needs to be examined just as critically as the AI systems it evaluates.
