What is happening?

systems find patterns in data and generate predictions or text.

Some systems recommend videos or recognise patterns in images; others answer questions and write text. An output may look polished and confident, but that does not guarantee accuracy. A system learns from data that may be incomplete, outdated or unsuitable for a new situation. Performing well on a familiar example does not mean it will work equally well everywhere.

How and why?

Output quality depends on data, task and checking. Errors and can hide behind fluent language.

Imagine asking a writing tool about a scientific paper. It might produce a fluent explanation yet mix up dates or name a source that does not exist. In systems that recommend or decide, errors may affect some groups more than others if the data do not represent everyone fairly. stresses that risk depends on the setting, the data and changes over time.

How do we know?

advises evaluating performance in context and communicating uncertainty clearly.

Reliability should be checked with tasks resembling real use, not just a striking demonstration. We need to measure types of errors, examine who may be harmed, document limitations and repeat tests when circumstances change. Where possible, check a claim against the original document or an expert institution. Communicating uncertainty clearly is part of responsible use, not a weakness.

Why does it matter?

For health, schoolwork or important decisions, check claims against reliable sources.

can save time when studying, translating or drafting. Yet decisions about health, money or people's rights should not rest on its answer alone. At school, useful questions include: where did this claim come from, can I open the source, and is there another explanation? Those habits build digital literacy even as the tools change.

Where do a model’s mistakes come from?

An system learns patterns from data, but a statistical pattern is not the same as a verified fact. A fluent model may confidently give a wrong answer, invent a source or miss an important condition in the question. Errors also arise when training data are biased, outdated or unrepresentative of the people and situations where the tool is used. A good score on one benchmark therefore cannot guarantee reliability in a new setting.

The level of checking should match the consequences of error. An idea for a headline may need editorial review; medical, financial or legal advice needs an authoritative source and qualified professional. For a scientific claim, open the original study or data, check its date, sample and method, and ask whether independent groups find something similar. A link is not enough: it must actually support the specific claim being made.

describes trustworthy through several distinct properties: validity, safety, robustness, explainability and fairness cannot be reduced to a single score. A system may be accurate on average but fail more often for one group; it may work well until circumstances change. Sensible use therefore requires a clear purpose, testing on realistic tasks, monitoring failures and a way for people to challenge and correct decisions.

A verification path: compare the model’s answer with the original source, its methods and independent evidence.
A verification path: compare the model’s answer with the original source, its methods and independent evidence.
Original NZM illustration · Sources: NIST

A further detail

A practical method is to break an answer into claims that can be checked. If says an institute conducted an experiment, verify the paper title, authors, date and actual result rather than judging the answer by its tone. Then look for study limitations: sample size, whether subjects were animals or people, and independent confirmation. Distinguish a summary from a direct quotation, because a model may change wording and meaning. If the source cannot be found, do not publish the claim as a fact. This is ordinary editorial checking, not delay for its own sake.

It is useful to record tasks the system fails. A collection of errors shows whether a problem affects a particular topic, language or group of users. The same tasks can be repeated with an updated version to test whether a claimed improvement really helped.

Key terms

— a systematic distortion that can affect a result.

Sources