Prompt Engineer
Writes the prompts and success criteria; the Evaluator tells it where the tests fail so it can fix the prompt.
The Evaluator is the LLM evaluation checker in the Ai1 platform by MyZone AI that reviews the tests grading your AI prompts and agents, feeds them deliberately broken examples and gives a pass, flag or fail verdict on whether the grading can be trusted, while a person approves the decisions that matter.
We'll show you the Evaluator on your own AI tests.
No paying for broken tests
A free dry check comes first, so paid test runs only go ahead once the setup has been shown to be sound.
Bad changes caught before they ship
A grader that lets wrong or unsafe answers through is found out before it approves a change you would regret.
Comparisons you can believe
When a prompt seems to get worse, you can tell whether the prompt slipped or the test itself is at fault.
Meet your agent
Xinyi checks that the tests grading your AI are themselves trustworthy. She is sharp, sceptical in the best way and enjoys an evening walk along the river.
AI agent. Fictional persona; not a real person.
| Name | Xinyi Z. |
|---|---|
| Role | AI Evaluation Specialist |
| Experience | 10 years |
| Work history | Quality engineer at an AI research lab Test analyst at a software company |
| Personality | Sharp, healthy sceptic |
| Based in | Shanghai, China |
AI does better work as a specific expert, so the Evaluator works as Xinyi (AI Evaluation Specialist, Shanghai, China). You can rename Xinyi or change the personality, experience and profile settings at any time.
Why our agents have personalitiesAt a glance
Last updated · Reviewed by the MyZone AI team
| What it checks | The tests that grade your AI prompts and agents |
|---|---|
| How it tests them | Feeds the grader deliberately broken examples to see what it misses |
| The verdict | Pass, flag or fail, with the evidence and the fixes needed |
| Before paid runs | A free dry check comes first |
| After a grading change | Re-runs a fixed reference set and shows where results moved |
| A person approves | Safety labels and pass bars, benchmark lock-ins, live adoption and published results |
Capabilities
Looks at the test setup, the reference answers and recent results, checks the data is split properly and the grader is independent, then gives a pass, flag or fail with evidence.
Feeds the grader examples that are wrong on purpose. Any it fails to catch are explained and turned into fixes.
When the grading instructions, grading model, scoring rules or data change, it re-runs a fixed reference set and shows where results moved.
Works out whether a bad result came from the prompt, the test, the reference answers or the grader, logs it and says who should fix it.
Shows how often the grader agreed with the right answer, which planted mistakes it caught, its known limits and the next step.
A dashboard of test scores shows you the numbers and assumes the test behind them is sound. The Evaluator checks that test first: it reviews the data splits and the grader's independence, plants deliberate mistakes to see what the grader misses and re-runs a fixed reference set when the grading changes. It gives a pass, flag or fail with evidence, and a person approves safety labels, pass bars and published results.
How this differs from the QA Orchestrator: the QA Orchestrator runs general quality checks on apps and releases, while the Evaluator checks whether the tests grading your AI prompts and agents can be trusted.
How it works
It runs inside your Ai1 system and works in the tools you already use. You ask, it works, it reports, and it stops for your OK where it matters.
You, or another agent, ask whether an evaluation setup is ready to trust.
It checks the data splits, the automatic checks and whether the grader is independent of what it grades.
It runs deliberately broken examples through the grader to see which ones it catches.
Pass, flag or fail, with the evidence and the fixes needed before the setup can be relied on.
Changing critical safety labels or pass bars, locking in a benchmark, adopting a change live or publishing results all wait for a person's approval.
When to use it
Example grader review
Example with a fictional company. Names, people and figures are invented to show the agent's output. Any resemblance to a real company or person is unintended.
What it was asked: Before our grader decides whether a cheaper model can answer our customer emails, tell us whether we can trust its scores.
I compared the grader's pass or fail on 120 past replies against the support team's own marks, planted 12 deliberate mistakes to see which it caught, re-judged 50 pairs in swapped order, and checked the test set for overlap with the assistant's instructions. I did not look at live customer traffic, the assistant's own reply quality or model costs, and I changed nothing: a person on the team approves the pass bars and makes every fix. Run date: .
Want a grader review like this for your own AI tests? Book an Evaluator walkthrough.
Book an Evaluator walkthroughBuilt-in guardrails
Every Ai1 agent works under human approval. Here is how the Evaluator keeps you in control.
Why teams choose Ai1
Trusted by leaders at Plastic Bank, Outback Team Building, RMG Advertising, Keeran Networks, and Titan Training Centre.
Key-only SSH. Passwords are disabled, and repeated failed logins trigger automatic IP banning.
Your data belongs to you, and it is stored on your own server.
Better together
The agents that turn the Evaluator's output into results: your whole team, working from the same place.
Writes the prompts and success criteria; the Evaluator tells it where the tests fail so it can fix the prompt.
Handles general quality checks for apps and releases, outside the question of whether a grader is sound.
Signs off critical safety labels and thresholds together with a person on your team.
Builds or changes the evaluation tools when a review shows they need fixing.
Common questions
It means grading the answers of your AI prompts or agents with an automated test. The catch is that the test can be wrong too. The Evaluator in Ai1 by MyZone AI checks the checker before you rely on its scores.
A flawed test may pass bad answers, fail good ones, change its mind when answers are shown in a different order, or score higher because test cases leaked into the prompt. The Evaluator looks for exactly these problems.
First make sure the test can tell them apart. The Evaluator checks that the data is split properly and the grader is independent, looks for test cases that leaked into the prompt, and confirms the prompt evaluation reliably separates the two versions before you pick a winner.
It does not write the prompts. It tells the Prompt Engineer what the tests need and where they fall short, and the Prompt Engineer writes them.
Plant mistakes and see whether it catches them, compare its marks with your team's own, and check it does not change its mind when answers are shown in a different order. The Evaluator runs these checks and reports what it measured on your own test data, rather than quoting accuracy figures it has not checked.
It only reads the prompts and agents it is testing, and a person on your team makes any fix.
Yes. The Evaluator is ready to use in Ai1 by MyZone AI.
It cannot approve changes. Critical safety labels and thresholds, locking in a benchmark, adopting a change live and publishing results all need a person's approval; it supplies the evidence for that decision.
About Ai1
The whole platform. The Evaluator is included on every Ai1 level, including Developer Core, with no per-agent charge. It works alongside the other Ai1 agents on your account. Compare plans
Two paths, one platform. Build it yourself on a developer plan, or let our team run your AI operations for you. Every plan runs on its own private server, and every price is shown in full. See Ai1 pricing
See how the Evaluator will check the tests grading your AI before you rely on their scores.
We'll show you the Evaluator on your own AI tests.
Leads your Roblox game studio from first idea to release, cutting small tickets and waiting for a person at every big checkpoint.
Merges only checked work, tests it on a private copy of your Roblox game and publishes only once the recorded approvals are checked.
Turns approved sketches into Roblox screens, menus and buttons that fit different screen sizes, with proof captured at two sizes.
The front door for brand-new software projects: an interview, a plan you check, and a clean handover once you say start.
Sets the standards for your Ai1 portal and gives every technical scope a binding verdict before anyone starts building.