Home

Services

SEO Services Social Media Website Development AI Automation About Reviews Portfolio Blog Contact Free First Look
AI Evaluation & Testing

“It Seemed Fine in the Demo” Is Not a Standard

Almost every AI build is signed off on a handful of examples chosen by the person who built it. Those examples work — they were picked because they work. It is the least informative test available, and it is the one nearly everybody runs. What tells you anything is a set of real cases from your business, including the ones that are hard, scored before the thing goes near live work.

No credit card required · 30-minute call · No obligation to sign up

12+ years in digital marketing Working with AI Evaluation businesses
✓ Built In Your Accounts, Not Ours ✓ Quoted & reported in GBP (£) ★★★★★ Trustpilot 5.0 ★★★★★ Google 4.9
An analyst reviewing scored test results on a monitor
Where the test cases come fromYour business
Accuracy percentages on this pageNone
The Numbers First

Six things that only show up on real cases

Each of these passes a demo comfortably. Each of them is how an AI build actually fails once it is doing real work.

  • It is confidently wrongOn the case nobody thought to try
  • It handles the tidy 80%And falls apart on the rest
  • It behaves differently on TuesdaySame input · different answer
  • The provider updated the modelNobody told you · behaviour moved
  • Somebody edited the instructionsAnd nothing was re-tested
  • It is wrong in one directionAlways too generous, or too strict

The fourth row is the one people find hardest to accept. The model underneath your build is somebody else’s product and it changes without asking you — so a system that was measured once was measured on a thing that no longer exists. That is not a reason to avoid AI; it is a reason to keep the test set and run it again, which takes minutes once it exists.

Straight Talk

Fifty real cases, with the answers agreed in advance

A test set is not complicated. It is a few dozen genuine examples from your own records — real enquiries, real documents, real messages — with the correct outcome for each one written down by somebody in your business before the system sees them. That last part is what makes it a test rather than a demonstration: if the right answer is decided after seeing the output, you have not measured anything.

The cases that matter most are the awkward ones, and they are the ones a demo never includes. The enquiry that is half a complaint. The one in two languages. The duplicate. The one that looks routine and is not. Building the set deliberately around those is uncomfortable because the score comes out lower, and a lower score you can trust is worth considerably more than a high one you cannot.

Then it gets a threshold, agreed with you, below which the thing does not go live. That is the sentence that makes evaluation real rather than ceremonial — a score with no consequence attached is a number in a report. And the direction of the errors matters as much as the count: a system that is consistently too generous with refunds is a different problem from one that is consistently too strict, even at the same score.

After launch the same set gets run again on a schedule, and whenever anything changes — the instructions, the model, the data it draws on. It takes minutes once it exists, and it is the only way anybody notices drift before a customer does. The test set is the deliverable that outlives the build, and you keep it whether or not you keep us.

  • Answers written before the system sees them — Otherwise it is a demonstration, not a test.
  • Build it around the awkward cases — The score comes out lower and it is worth more.
  • A threshold with a consequence — A score that cannot stop a launch is a number in a report.
  • Direction matters, not just count — Too generous and too strict are different problems.
  • Re-run it when anything changes — The model is somebody else’s product and it moves.
Test results and scoring laid out on a desk

If the right answer is decided after seeing the output, you have not measured anything.

.
Ghalib Ashrafi Founder & Digital Strategist · 12+ years across search, social & web

What an evaluation actually involves

Six things, and the deliverable outlives the build.

01

A Test Set From Your Records

A few dozen genuine cases, with the correct outcome written down by somebody in your business before the system sees any of them. That order is what makes it a test.

02

Weighted Towards the Awkward

Half-complaints, two languages, duplicates, the routine-looking one that is not. Deliberately uncomfortable, because a lower score you can trust beats a high one you cannot.

03

A Threshold With Teeth

A score agreed with you below which it does not go live. Without a consequence attached, an evaluation is ceremony — and everybody involved knows it.

04

Error Direction, Not Just Rate

Consistently too generous and consistently too strict are different business problems at the same score. Which way it leans is usually more actionable than how often it is wrong.

05

Re-Run on Every Change

New instructions, a new model version, new source data. Minutes once the set exists, and the only way drift is noticed by you rather than by a customer.

06

The Set Is Yours

Handed over in a plain format you can run against anything, including a different supplier’s build. It is the deliverable that outlives the project.

BeforeWhen the correct answers are written
AwkwardWhat the test set is weighted towards
YoursWho keeps the test set afterwards
0Accuracy percentages claimed here

What clients say

Real clients, quoted in their own words — published with their permission.

More of them, in full, on our reviews page.

— Our Proprietary Methodology —

The Visibility Framework™, applied in AI Evaluation

The method doesn’t change by market. What it’s pointed at does.

Step 01

Audit The Hours

Where time actually goes, task by task, scored on volume, repetition and the cost of getting it wrong. Ends in a ranked blueprint with estimated hours saved — yours to keep either way.

Step 02

Design The Guardrails

Before any building: what the agent may touch, where a human must approve, what happens when it is unsure, and which data is never allowed near a third-party model.

Step 03

Build & Evaluate

One workflow at a time, in your accounts, scored against real examples from your business before it touches live work. Shipped early so it meets reality while it is still cheap to change.

Step 04

Run & Improve

Monitored for cost, failures and quality drift. Models change, your business changes, and an automation nobody tends becomes a liability rather than an asset.

Honest, No-Nonsense Commitment

If the audit concludes that a task is not worth automating, we will tell you and refund the difference rather than build it anyway. And if a workflow we built does not hit the outcome we agreed in the blueprint, we keep working on it at no extra cost until it does or we take it out.

Investment

AI Automation pricing for AI Evaluation, in GBP

Included in every build we do rather than sold as an upgrade, because a build we cannot measure is one we would rather not ship. The prices below are for evaluating something we did not build — an existing system, or one another supplier is proposing.

One System

An existing build, scored honestly.

£1,200 – £2,800
  • Test set built from your own records
  • Scored, with error direction reported
  • The set handed over for you to keep
Get a Quote

Ongoing Monitoring

Drift caught before customers find it.

£600+ / month
  • Scheduled re-runs and change-triggered runs
  • Alerting when a score moves
  • Threshold reviewed as the business changes
  • Optional: combine all 4 services for full-funnel growth
Get a Quote

Every plan is scoped around your market — start with a free first look and we’ll recommend what fits, priced in GBP.

What you’re actually committing to

Most agencies keep this in a contract you only see after the sales call. We would rather you knew now, because it is the question everyone asks second — right after the price.

  • The audit is credited, not sunkPay for the audit, and the full amount comes off the build if you proceed. If you don’t, the blueprint is still yours to hand to anyone else.
  • A fixed build price after the auditQuoted once we know what we are building. If it takes longer than we estimated, that is our risk — the price only moves if you change the scope.
  • You own everythingAccounts, API keys, workflows, prompts, evaluation sets, logs and documentation — all in your name from day one, and still yours if we never work together again.
  • Running costs are yours and visibleAPI usage is billed by the provider directly to you. We never resell tokens or mark up usage, and you see the real number.
  • The retainer is month to month30 days’ notice, no exit fee. Stop it and your automations keep running — you are simply maintaining them yourself.
  • Human approval is the defaultAnything customer-facing or irreversible needs a person until the evaluation data justifies otherwise, and that decision is yours to make, not ours.

These are the terms as they appear in the agreement itself — nothing here is softened for the website. The full wording lives in our terms and conditions, and you get the agreement to read before anything is signed or invoiced.

AI Evaluation AI Automation questions, answered

Our supplier demoed it and it worked. Is that not enough? +
It is the least informative test available. Those examples were chosen because they work — that is why they were in the demo. What tells you something is a set of real cases from your own records, weighted towards the awkward ones, with the correct answers written down before the system sees them.
A few dozen genuine cases is usually enough to be useful, and far more useful than several hundred tidy ones. What matters is that they are real, that they include the difficult ones, and that somebody in your business decided the right answer in advance rather than after seeing the output.
Because the model underneath is somebody else’s product and it changes without asking you. A system measured once was measured on something that may no longer exist. Re-running takes minutes once the set exists, and it is the only way drift gets noticed by you rather than by a customer.
Then it does not go live, because the threshold was agreed before the score existed. That sentence is what separates evaluation from ceremony — a score that cannot stop a launch is a number in a report, and everybody in the room knows it.
Yes, and that is most of this work. An existing system, or one a supplier is proposing before you sign. The test set is yours afterwards in a plain format, so you can run it against whatever comes next — including a different supplier’s version.
Because a page about evidence standards cannot carry an unsourced number without contradicting itself. Any percentage we published would be from a build you cannot inspect, on a test set you did not see. Your own score, on your own cases, is the only figure that means anything here.
Keep Exploring

Related AI Automation pages

Same service, different angle — by market, by service and by industry.

Get A Free Evaluation Scope

Tell us what an AI system in your business is supposed to decide. We will tell you what a real test set for it would look like, and where we would expect it to fail.

Want a score you can actually trust?

Free 30-minute call · No obligation · Response within 24 hours

Get in touch with Ghalib Ashrafi HQ

Ghalib Ashrafi takes on work for brands in six markets — the UK, USA, UAE, Saudi Arabia, Australia and Pakistan. Every enquiry gets a reply within 24 hours.

Phone / WhatsApp: +92 343 2653224
Email: info@ghalibashrafi.com
Hours: Mon–Sat, 10am–7pm (PKT)
Still deciding? Get a free first look — no obligation. Get Free First Look