Google research finds factual errors are often failures of recall, not missing knowledge

New Google Research looks at why leading language models sometimes provide incorrect factual answers. Its researchers found that frontier models often appear to have encoded the correct information but fail to retrieve it reliably when asked a question.

The distinction matters because simply making models larger or giving them more training data may not solve the problem. Google found that allowing models to use additional reasoning can recover some information that could not initially be recalled, although it still does not eliminate factual errors.

Key facts

  • Google evaluated 13 language models using a new benchmark containing 2,150 facts.
  • For Gemini 3 Pro and GPT-5, researchers found that 95% to 98% of tested facts appeared to be encoded.
  • The models still failed to directly recall 26% to 34% of those facts, with additional reasoning reducing but not eliminating the gap.

Our take

This reinforces a basic rule for organisations using generative AI: fluent answers are not evidence that the underlying facts are correct.

For important work, asking a model to rely on authoritative source material remains preferable to relying on its internal knowledge alone. New Zealand organisations should design workflows around retrieval, source checking and human verification, particularly where an incorrect answer could affect a client, customer or important decision.

Sources

About the author

Campbell McKenzie is a Director at Incident Response Solutions, a New Zealand firm experienced in cyber incident response, digital forensics, investigations and technology risk. Through KiwiGen.AI, Campbell helps professional services firms adopt generative AI safely, with practical governance and controls.