Evaluator: Make Sure the Tests Grading Your AI Are Right

Check the test that grades your AI before you trust its scores.

The Evaluator is the LLM evaluation checker in the Ai1 platform by MyZone AI that reviews the tests grading your AI prompts and agents, feeds them deliberately broken examples and gives a pass, flag or fail verdict on whether the grading can be trusted, while a person approves the decisions that matter.

We'll show you the Evaluator on your own AI tests.

Plans & pricing

  • No paying for broken tests

    A free dry check comes first, so paid test runs only go ahead once the setup has been shown to be sound.

  • Bad changes caught before they ship

    A grader that lets wrong or unsafe answers through is found out before it approves a change you would regret.

  • Comparisons you can believe

    When a prompt seems to get worse, you can tell whether the prompt slipped or the test itself is at fault.

Get to know Xinyi

Xinyi checks that the tests grading your AI are themselves trustworthy. She is sharp, sceptical in the best way and enjoys an evening walk along the river.

AI agent. Fictional persona; not a real person.

Xinyi, the persona of the Evaluator
NameXinyi Z.
RoleAI Evaluation Specialist
Experience10 years
Work historyQuality engineer at an AI research lab
Test analyst at a software company
PersonalitySharp, healthy sceptic
Based inShanghai, China

Why Xinyi has a personality

AI does better work as a specific expert, so the Evaluator works as Xinyi (AI Evaluation Specialist, Shanghai, China). You can rename Xinyi or change the personality, experience and profile settings at any time.

Why our agents have personalities

The Evaluator in short

Last updated · Reviewed by the MyZone AI team

The Evaluator at a glance
What it checksThe tests that grade your AI prompts and agents
How it tests themFeeds the grader deliberately broken examples to see what it misses
The verdictPass, flag or fail, with the evidence and the fixes needed
Before paid runsA free dry check comes first
After a grading changeRe-runs a fixed reference set and shows where results moved
A person approvesSafety labels and pass bars, benchmark lock-ins, live adoption and published results

What the Evaluator does for you

  • Reviews whether a test can be trusted

    Looks at the test setup, the reference answers and recent results, checks the data is split properly and the grader is independent, then gives a pass, flag or fail with evidence.

  • Plants mistakes to test the grader

    Feeds the grader examples that are wrong on purpose. Any it fails to catch are explained and turned into fixes.

  • Spots drift after a change

    When the grading instructions, grading model, scoring rules or data change, it re-runs a fixed reference set and shows where results moved.

  • Traces wrong results to their cause

    Works out whether a bad result came from the prompt, the test, the reference answers or the grader, logs it and says who should fix it.

  • Writes a short evidence report

    Shows how often the grader agreed with the right answer, which planted mistakes it caught, its known limits and the next step.

How the Evaluator differs from a dashboard of test scores

A dashboard of test scores shows you the numbers and assumes the test behind them is sound. The Evaluator checks that test first: it reviews the data splits and the grader's independence, plants deliberate mistakes to see what the grader misses and re-runs a fixed reference set when the grading changes. It gives a pass, flag or fail with evidence, and a person approves safety labels, pass bars and published results.

How this differs from the QA Orchestrator: the QA Orchestrator runs general quality checks on apps and releases, while the Evaluator checks whether the tests grading your AI prompts and agents can be trusted.

From your request to a finished result

It runs inside your Ai1 system and works in the tools you already use. You ask, it works, it reports, and it stops for your OK where it matters.

How the Evaluator works: you ask through the Comms Hub and Ai1 runs the steps: test setup comes in, review the setup, stress test the grader, give a verdict and a person approves. You approve at: a person approves. It returns a readiness verdict.
Tap the diagram to enlarge it

Step 1: A test setup comes in

You, or another agent, ask whether an evaluation setup is ready to trust.

Step 2: Review the setup

It checks the data splits, the automatic checks and whether the grader is independent of what it grades.

Step 3: Stress test the grader

It runs deliberately broken examples through the grader to see which ones it catches.

Step 4: Give a verdict

Pass, flag or fail, with the evidence and the fixes needed before the setup can be relied on.

You approve

Step 5: A person approves what matters

Changing critical safety labels or pass bars, locking in a benchmark, adopting a change live or publishing results all wait for a person's approval.

Hand these situations to the Evaluator

Illustrative photo: the customer-support lead at a travel company asks the Evaluator for help from their phone.
Illustrative photo. Asking whether the scores grading the team's AI replies can be trusted.
  • You are about to pay for a large test run
    It runs a free dry check first and tells you whether the setup is ready or what to fix.
  • You switched the model that does the grading
    It compares new results against a fixed reference set so you know whether old and new scores are comparable.
  • A test passed an answer that was clearly wrong
    It traces the false pass to its cause, logs it and routes the fix to whoever owns that part.
  • You want to compare two prompt versions fairly
    It confirms the prompt evaluation can reliably tell the two versions apart before you pick a winner.

What you get

  • A readiness verdict of pass, flag or fail, with the evidence behind it
  • Results from planted mistakes showing what the grader caught and missed
  • Drift checks against a fixed reference set after any grading change
  • A log of wrong results with their cause and who should fix them
  • A short report covering the grader's limits and the next step

What it won't do

  • Design or rewrite promptsHandled by: The Prompt Engineer
  • Build or change evaluation toolsHandled by: The Skill Manager
  • Release or deploy evaluation toolsHandled by: The DevOps Agent
  • Approve critical safety labels or thresholdsHandled by: The Security Agent, with a person signing off
  • Publish public benchmark resultsHandled by: Your Ai1 team
  • Run general quality checks on apps and releasesHandled by: The QA Orchestrator

Example: Support-reply grader readiness review

Example with a fictional company. Names, people and figures are invented to show the agent's output. Any resemblance to a real company or person is unintended.

What it was asked: Before our grader decides whether a cheaper model can answer our customer emails, tell us whether we can trust its scores.

Example report
Readiness verdict marked FLAG, a table of which uses of the grader are ready, and three headline findings.
Verdict and readiness
Two-by-two table comparing grader pass or fail with the support team's marks, plus catch rate and accuracy measures against suggested bars.
Grader versus the team
List of twelve planted mistakes showing which the grader caught, with bars for mistakes caught, order flips and test-case overlap.
Testing the tester
Fix plan with owners and finish criteria, a person's approval step for the pass bars, and the limits of the review.
Fixes and approval

What it found: 3 findings, with the numbers

  • The grader lets through 9 of 30 replies the support team marked as bad, mostly refund and date mistakes, even though overall accuracy looks high at 87.5%.
  • A single judge decides pass or fail, and it changed its pick in 11 of 50 pairs when the two replies were shown in the opposite order.
  • 14 of the 120 test cases also appear in the assistant's own instructions, which lifts the score; on the other 106 cases accuracy is 85.8%.

I compared the grader's pass or fail on 120 past replies against the support team's own marks, planted 12 deliberate mistakes to see which it caught, re-judged 50 pairs in swapped order, and checked the test set for overlap with the assistant's instructions. I did not look at live customer traffic, the assistant's own reply quality or model costs, and I changed nothing: a person on the team approves the pass bars and makes every fix. Run date: .

Want a grader review like this for your own AI tests? Book an Evaluator walkthrough.

Book an Evaluator walkthrough

It asks before it acts

Every Ai1 agent works under human approval. Here is how the Evaluator keeps you in control.

  • It only reads the prompts, agents and tests it is checking. It does not change them.
  • Critical safety labels and thresholds, benchmark lock-ins, live adoption of a change and published results all need a person's approval. It prepares the evidence.
  • It reports what it measured on your own test data and states the limits of each review, rather than quoting accuracy figures it has not checked.

Part of Ai1, by MyZone AI

Trusted by leaders at Plastic Bank, Outback Team Building, RMG Advertising, Keeran Networks, and Titan Training Centre.

Key-only SSH. Passwords are disabled, and repeated failed logins trigger automatic IP banning.

How we keep your data safe

Your data, your server

Your data belongs to you, and it is stored on your own server.

How we keep your data safe

Works well with

The agents that turn the Evaluator's output into results: your whole team, working from the same place.

  • Prompt Engineer

    Writes the prompts and success criteria; the Evaluator tells it where the tests fail so it can fix the prompt.

  • QA Orchestrator

    Handles general quality checks for apps and releases, outside the question of whether a grader is sound.

  • Security Agent

    Signs off critical safety labels and thresholds together with a person on your team.

  • Skill Manager

    Builds or changes the evaluation tools when a review shows they need fixing.

Frequently asked questions about the Evaluator

What is AI evaluation testing?

It means grading the answers of your AI prompts or agents with an automated test. The catch is that the test can be wrong too. The Evaluator in Ai1 by MyZone AI checks the checker before you rely on its scores.

A flawed test may pass bad answers, fail good ones, change its mind when answers are shown in a different order, or score higher because test cases leaked into the prompt. The Evaluator looks for exactly these problems.

How do you compare two prompt versions fairly?

First make sure the test can tell them apart. The Evaluator checks that the data is split properly and the grader is independent, looks for test cases that leaked into the prompt, and confirms the prompt evaluation reliably separates the two versions before you pick a winner.

It does not write the prompts. It tells the Prompt Engineer what the tests need and where they fall short, and the Prompt Engineer writes them.

How can I tell whether an AI test is reliable?

Plant mistakes and see whether it catches them, compare its marks with your team's own, and check it does not change its mind when answers are shown in a different order. The Evaluator runs these checks and reports what it measured on your own test data, rather than quoting accuracy figures it has not checked.

It only reads the prompts and agents it is testing, and a person on your team makes any fix.

Is the Evaluator available today, and can it approve changes on its own?

Yes. The Evaluator is ready to use in Ai1 by MyZone AI.

It cannot approve changes. Critical safety labels and thresholds, locking in a benchmark, adopting a change live and publishing results all need a person's approval; it supplies the evidence for that decision.

About Ai1

Do I get just the Evaluator, or the whole platform?

The whole platform. The Evaluator is included on every Ai1 level, including Developer Core, with no per-agent charge. It works alongside the other Ai1 agents on your account. Compare plans

Two paths, one platform. Build it yourself on a developer plan, or let our team run your AI operations for you. Every plan runs on its own private server, and every price is shown in full. See Ai1 pricing

More about Ai1: security, setup time

Put the Evaluator to work

See how the Evaluator will check the tests grading your AI before you rely on their scores.

We'll show you the Evaluator on your own AI tests.

Plans & pricing

  • Product & Engineering

    GS Producer

    Leads your Roblox game studio from first idea to release, cutting small tickets and waiting for a person at every big checkpoint.

    • Manages projects
    • Plans strategy
    • Reports
  • Product & Engineering

    GS Release Engineer

    Merges only checked work, tests it on a private copy of your Roblox game and publishes only once the recorded approvals are checked.

    • Automates workflows
    • Audits
    • Reports
  • Product & Engineering

    GS UI Developer

    Turns approved sketches into Roblox screens, menus and buttons that fit different screen sizes, with proof captured at two sizes.

    • Designs
    • Automates workflows
    • Audits
  • Product & Engineering

    New-Project Intake

    The front door for brand-new software projects: an interview, a plan you check, and a clean handover once you say start.

    • Plans strategy
    • Manages projects
    • Communicates
  • Product & Engineering

    Platform Architect

    Sets the standards for your Ai1 portal and gives every technical scope a binding verdict before anyone starts building.

    • Plans strategy
    • Audits
    • Monitors