LinkedIn analytics tracking pixel for AIDOLS AI consulting website performance measurement

Unit 4 · On your own work

Telling a result from a demo

12 min readFree · no sign-upTrack: Leadership

The principle

A demo is one run chosen by the person showing it; a result is many runs measured against a number that existed before the work started.

By the end: You can ask four evidence questions of any AI claim and tell whether a result exists behind it.

The situation

Forty minutes on a call. Someone from your own team walks you through a tool that drafts responses to incoming tenders. They paste a live request, it produces four pages, and the four pages are good. Better than the draft your team would have produced on a Thursday afternoon. Everyone on the call is pleased. You end it agreeing to "find the budget", walk to your next meeting, and cannot name one thing that changed. Nothing dishonest happened. You watched a demo and priced it like a result.

The principle

A demo is one run chosen by the person showing it; a result is many runs measured against a number that existed before the work started.

The distance between the two is a distribution. A demo is one draw from it, selected by somebody who wants your approval, judged by that same person. A result is the whole spread — including the runs they would not have shown you — held against a standard that predates the enthusiasm. Four questions close the distance, and you can ask all four in under ten minutes.

How many times has it run, and who chose the inputs? A number, and provenance. "Live tenders as the queue delivered them" is an answer. "We tried it on a few" is not.

What did this number look like before? The baseline has to predate the work. A baseline measured after the pilot started is not a baseline, it is a memory.

Show me the worst run. Anyone can show the best one. Somebody with a result has already gone looking for the failures and can describe them without flinching.

Who checks the output, and how long does that take? Review is a cost, and it belongs in the total. If the answer is "it does not really need checking", you are looking at a demo.

Worked example

Shown here in Microsoft 365 Copilot.

Two hours after the call your notes are a file in your work account. Before you write the follow-up email, sort the claims:

Job: Sort the claims in my notes file into checkable and not checkable.

Inputs: My notes from a 40-minute internal demo of a tool that drafts
tender responses. Use only that file. Do not use anything you know
about the category or any vendor.

Output: A two-column table. Left: every claim made in the notes,
quoted. Right: what somebody would have to show me for that claim to
be checkable — a count, a date, a named report, or a named system.

Standard: If no document could make a claim checkable, write "opinion"
in the right column. Do not judge whether any claim is true.

What comes back:

ClaimWhat would make it checkable
"Saves the team about a day a week"The report the day comes off — responses shipped per week, or hours logged, before and after a named start date
"Quality is as good as ours"The scoring standard used before the pilot, and the name of whoever applied it
"The team loves it"Opinion — no document makes this checkable

The third row is the one that sounded strongest in the room. Twenty minutes of sorting turns an impression into three specific requests, and you send those instead of a budget approval.

A good answer, when it comes back, sounds like this. The numbers are stand-ins for yours; the shape is the point:

It has run 140 times since March, on tenders the queue assigned — we did not pick them. Responses shipped per week sat at nine for the eight weeks before we started and eleven since; that is on the weekly pipeline report, not a spreadsheet we built. Roughly one in nine drafts gets rewritten from scratch, and here are three of them. Two people spot-check ten a day, about twenty minutes total.

Nobody producing that answer needs a demo.

In your tool

The work is reading a long, messy input — notes, a transcript, a thread — against a fixed question list.

ToolReading long notes or a transcript
ClaudePaste or attach; attach to a Project so later conversations can still read it
ChatGPTPaste or attach; a Project holds the notes and your standing question list together
GeminiPaste directly, or attach a file already stored in the same account
Microsoft 365 CopilotReads notes, meeting transcripts and files already captured in your work account, where your organisation records them

Product surfaces checked 2026-08-04.

Try it

Take the most recent AI claim you heard — from a vendor, a peer, or your own team. Write the four questions as a four-line email. Five minutes, no tool needed.

You are done when every line asks for a number, a date, a document or an artifact. If any line asks whether something is working well, rewrite it. That question has only one answer and you already know what it is.

Common failures

  • Grading the output instead of the process. The sample reads well, so the thing works. You are judging one run, chosen by the person who needs your signature, on an input they had time to pick.
  • Accepting time saved as reported by the people who saved it. Self-reported minutes are a feeling with a decimal point on it. Ask which number, on which report that already existed, moved.
  • Letting "it is early" close the question. Early is a perfectly good answer. But early means "no result yet, and here is the date we will have one" — not a story told with more conviction. Write the date down when you hear it.

You just read the whole lesson.

Guided builds, starter templates, instructor video, graded capstone, certificate, cohort office hours, org roster and completion reporting.

See what enrolment adds