Find out what the system could access
Ask which tools, files and online services were available during the demonstration. A system using a curated document collection is operating differently from one answering without it. The same applies to human help behind the scenes. A clear account should distinguish the model’s output from preparation, editing and the surrounding software.
Look beyond the successful example
One impressive result does not establish reliability. Were multiple attempts shown? Were failures included? Was the example selected after many trials? A good evaluation uses tasks that match the intended use and explains how performance was measured. Avoid comparing two systems if the conditions, prompts or available resources differ in ways the demonstration does not explain.
Match the claim to the evidence
Consider whether the demo supports a narrow capability or a much broader promise. Producing a plausible answer is different from checking it. Completing one task is different from handling an entire workflow. Ask what happens when information is missing, instructions conflict or the system is wrong. Finally, look for a reproducible example or documentation so that the result can be examined outside the promotional presentation.
Put these questions into practice: AI insiders disagree over the scale of existential risk





Join the conversation