“It Seemed Fine in the Demo” Is Not a Standard
Almost every AI build is signed off on a handful of examples chosen by the person who built it. Those examples work — they were picked because they work. It is the least informative test available, and it is the one nearly everybody runs. What tells you anything is a set of real cases from your business, including the ones that are hard, scored before the thing goes near live work.
No credit card required · 30-minute call · No obligation to sign up
Six things that only show up on real cases
Each of these passes a demo comfortably. Each of them is how an AI build actually fails once it is doing real work.
- It is confidently wrongOn the case nobody thought to try
- It handles the tidy 80%And falls apart on the rest
- It behaves differently on TuesdaySame input · different answer
- The provider updated the modelNobody told you · behaviour moved
- Somebody edited the instructionsAnd nothing was re-tested
- It is wrong in one directionAlways too generous, or too strict
The fourth row is the one people find hardest to accept. The model underneath your build is somebody else’s product and it changes without asking you — so a system that was measured once was measured on a thing that no longer exists. That is not a reason to avoid AI; it is a reason to keep the test set and run it again, which takes minutes once it exists.
Fifty real cases, with the answers agreed in advance
A test set is not complicated. It is a few dozen genuine examples from your own records — real enquiries, real documents, real messages — with the correct outcome for each one written down by somebody in your business before the system sees them. That last part is what makes it a test rather than a demonstration: if the right answer is decided after seeing the output, you have not measured anything.
The cases that matter most are the awkward ones, and they are the ones a demo never includes. The enquiry that is half a complaint. The one in two languages. The duplicate. The one that looks routine and is not. Building the set deliberately around those is uncomfortable because the score comes out lower, and a lower score you can trust is worth considerably more than a high one you cannot.
Then it gets a threshold, agreed with you, below which the thing does not go live. That is the sentence that makes evaluation real rather than ceremonial — a score with no consequence attached is a number in a report. And the direction of the errors matters as much as the count: a system that is consistently too generous with refunds is a different problem from one that is consistently too strict, even at the same score.
After launch the same set gets run again on a schedule, and whenever anything changes — the instructions, the model, the data it draws on. It takes minutes once it exists, and it is the only way anybody notices drift before a customer does. The test set is the deliverable that outlives the build, and you keep it whether or not you keep us.
- ✓Answers written before the system sees them — Otherwise it is a demonstration, not a test.
- ✓Build it around the awkward cases — The score comes out lower and it is worth more.
- ✓A threshold with a consequence — A score that cannot stop a launch is a number in a report.
- ✓Direction matters, not just count — Too generous and too strict are different problems.
- ✓Re-run it when anything changes — The model is somebody else’s product and it moves.
If the right answer is decided after seeing the output, you have not measured anything.
What an evaluation actually involves
Six things, and the deliverable outlives the build.
A Test Set From Your Records
A few dozen genuine cases, with the correct outcome written down by somebody in your business before the system sees any of them. That order is what makes it a test.
Weighted Towards the Awkward
Half-complaints, two languages, duplicates, the routine-looking one that is not. Deliberately uncomfortable, because a lower score you can trust beats a high one you cannot.
A Threshold With Teeth
A score agreed with you below which it does not go live. Without a consequence attached, an evaluation is ceremony — and everybody involved knows it.
Error Direction, Not Just Rate
Consistently too generous and consistently too strict are different business problems at the same score. Which way it leans is usually more actionable than how often it is wrong.
Re-Run on Every Change
New instructions, a new model version, new source data. Minutes once the set exists, and the only way drift is noticed by you rather than by a customer.
The Set Is Yours
Handed over in a plain format you can run against anything, including a different supplier’s build. It is the deliverable that outlives the project.
What clients say
Real clients, quoted in their own words — published with their permission.
“It was a pleasure working with Ghalib and his team.”
“Ghalib’s clear guidance has helped improve our website’s SEO performance.”
“We’re very thankful to Ghalib and his team for building a fully-functioning website for our business.”
More of them, in full, on our reviews page.
— Our Proprietary Methodology —
The Visibility Framework™, applied in AI Evaluation
The method doesn’t change by market. What it’s pointed at does.
Audit The Hours
Where time actually goes, task by task, scored on volume, repetition and the cost of getting it wrong. Ends in a ranked blueprint with estimated hours saved — yours to keep either way.
Design The Guardrails
Before any building: what the agent may touch, where a human must approve, what happens when it is unsure, and which data is never allowed near a third-party model.
Build & Evaluate
One workflow at a time, in your accounts, scored against real examples from your business before it touches live work. Shipped early so it meets reality while it is still cheap to change.
Run & Improve
Monitored for cost, failures and quality drift. Models change, your business changes, and an automation nobody tends becomes a liability rather than an asset.
Honest, No-Nonsense Commitment
If the audit concludes that a task is not worth automating, we will tell you and refund the difference rather than build it anyway. And if a workflow we built does not hit the outcome we agreed in the blueprint, we keep working on it at no extra cost until it does or we take it out.
AI Automation pricing for AI Evaluation, in GBP
Included in every build we do rather than sold as an upgrade, because a build we cannot measure is one we would rather not ship. The prices below are for evaluating something we did not build — an existing system, or one another supplier is proposing.
One System
An existing build, scored honestly.
- Test set built from your own records
- Scored, with error direction reported
- The set handed over for you to keep
Pre-Purchase Review
Before you sign with anybody.
- Everything in One System
- The same set run against what is proposed
- A written view on whether it is ready
- Optional: bundle in SEO or Website Development
Ongoing Monitoring
Drift caught before customers find it.
- Scheduled re-runs and change-triggered runs
- Alerting when a score moves
- Threshold reviewed as the business changes
- Optional: combine all 4 services for full-funnel growth
Every plan is scoped around your market — start with a free first look and we’ll recommend what fits, priced in GBP.
What you’re actually committing to
Most agencies keep this in a contract you only see after the sales call. We would rather you knew now, because it is the question everyone asks second — right after the price.
- The audit is credited, not sunkPay for the audit, and the full amount comes off the build if you proceed. If you don’t, the blueprint is still yours to hand to anyone else.
- A fixed build price after the auditQuoted once we know what we are building. If it takes longer than we estimated, that is our risk — the price only moves if you change the scope.
- You own everythingAccounts, API keys, workflows, prompts, evaluation sets, logs and documentation — all in your name from day one, and still yours if we never work together again.
- Running costs are yours and visibleAPI usage is billed by the provider directly to you. We never resell tokens or mark up usage, and you see the real number.
- The retainer is month to month30 days’ notice, no exit fee. Stop it and your automations keep running — you are simply maintaining them yourself.
- Human approval is the defaultAnything customer-facing or irreversible needs a person until the evaluation data justifies otherwise, and that decision is yours to make, not ours.
These are the terms as they appear in the agreement itself — nothing here is softened for the website. The full wording lives in our terms and conditions, and you get the agreement to read before anything is signed or invoiced.
AI Evaluation AI Automation questions, answered
Related AI Automation pages
Same service, different angle — by market, by service and by industry.
Want a score you can actually trust?
Free 30-minute call · No obligation · Response within 24 hours
Get in touch with Ghalib Ashrafi HQ
Ghalib Ashrafi takes on work for brands in six markets — the UK, USA, UAE, Saudi Arabia, Australia and Pakistan. Every enquiry gets a reply within 24 hours.