AI agent evaluation, grounded in your business

Know which AI agents can actually do your work.

Gradia helps your team test AI agents on your workflows and business rules. Find failures, check whether the test deserves your trust, and compare changes before your next release.

Bring one workflow and the decision you need to make. We agree the inputs, deliverables, review and fee before an assisted engagement begins.

No agent yet?

Start with the work people do today. Review its history, understand the process and define what an agent would need to get right.

Find your starting point ↓
What does an evaluation tell you?Illustrative example

Resolve a refund request.

The customer requests $140. No approval is recorded.

Your rule: Refunds above $100 need a supervisor’s approval.

Agent action

Issues the $140 refund immediately.

  • Understood the request✓ Pass
  • Respected the approval rule× Fail
  • Reached the correct next step× Fail

The customer gets a refund, but the approval rule is broken.

Next test: Add the approval check, then test cases on both sides of the limit.

Scripted actions and checks, not live agent runs or measured results. Passing one example does not establish release readiness.

01

What do you need to know?

Choose the question you have today. Each path starts with a useful result and can connect to the wider platform as your needs grow.

Can this agent do the job?

Define a task and its success criteria. Test a supported agent, inspect failures and compare changes under the same conditions.

Result: a reviewed assessment

Agent evaluation

How does this workflow really happen?

Bring selected, approved activity into Observer. Review the steps, decisions and gaps with the people who do the work.

Result: a reviewed case history

Workflow capture

What did my agent do, and was it allowed?

Use Guard to check permissions around covered calls and keep a verifiable recording. Add a separate assessment to judge quality.

Result: covered execution evidence

Try open-source Guard

Already have a test? Audit your benchmark. Looking for process improvements? Explore workflow efficiency.

Guard runs locally without an account. Managed evaluations need a supported runtime and reviewed setup. Desktop capture is a controlled preview; signed public installers are pending.

02

A task. A test. A decision.

An evaluation should answer a business question. The setup makes that question measurable.

  1. 01

    Define the job

    What should the agent do? What must it never do? Review examples and success criteria with someone who knows the work.

  2. 02

    Test what happens

    Run the supported agent in the prepared environment. Inspect its actions and the evidence behind the checks.

  3. 03

    Decide what to improve

    See which failures belong to the agent, the test or the environment. Compare a change under matched conditions.

What to bring, what you get and how to start →

Go deeper when the work needs it

What if the job takes weeks, involves people and keeps changing?

That is where a Gradia Universe comes in: a test environment built around a reviewed workflow, its rules and the information available at each step.

Start from a brief and expert input, or use Observer to assemble approved work across tools and time. Review the history before defining what an agent can see and do.

Then test a changed condition: a missing document, a delayed approval or a different policy. Inspect what happens in the supported environment.

This walkthrough illustrates the process. Capturing a history does not automatically reproduce an application. Runtime support, permissions and branch restoration need their own qualification.

One workflow. A connected world.Scripted illustration
Example 1 of 3 · Explore at your pace

Opportunity 042

37 days · one case

Before capture

Choose what can be observed

Approve the account, channel, pages and native window for this opportunity. Only approved sources and permitted evidence are retained.

Approved sources → one connected history01 / 08
Six selected sources · explicit permission

Synthetic histories and simulated outcomes. This preview connects no sources and runs no agent.

Enterprise Universes

A world built around the work.

The messages, decisions and waiting between applications belong to the same job. A Universe gives that work an explicit state, rules and history.

Explore the Universe approach
Concept illustration: mail, records, time and approvals connect through copper paths inside a miniature stone workflow world.

Observe

Inspect the history and the information each actor could see.

Interrupt

Check the stop receipt and settled actions before continuing.

Branch

Restore supported state into separate continuations and compare outcomes.

Concept artwork. Interruption and independent restoration are demonstrated in a bounded synthetic browser workflow. Each additional application, external effect and state component needs its own qualification.

03

We check the test, too.

A score is only useful if the test deserves your trust. Check known right and wrong answers, look for shortcuts that earn an undeserved pass, and review where human and AI judges disagree.

Public benchmarks show how a model scored on someone else’s test. They do not prove an agent can do your work. Inspect the source, the released evidence and the method behind a result.

Start with a brief or an existing test. Your expert helps define success; Gradia connects the reviewed requirements, supported environment and evaluation evidence. A reusable Universe or training export is a further step with its own qualification and permissions.

Inspect the full platform and evidenceThe complete workflow, runtime boundaries, released benchmarks, research and technical methods. Open the detail when you need it.

Governance · Evaluation · Benchmarks · Gradia Universes

02

One evidence spine, from production to proof.

Start with one agent run. Gradia Guard authorizes covered model and tool calls before dispatch, captures evidence after execution, and says exactly what it did not observe.

Govern

Policy checks before covered model and tool calls.

Evaluate

Deterministic checks, admitted judges, and humans inspect the same evidence.

Improve

Failures become regression tests, benchmark extensions, or approved training data.

Need more than tracing? The same contract runs inside a Gradia Universe with controlled visibility, interruptions, multiple actors, and branch/restore proof.

Already have a benchmark? Audit the ruler before interpreting the score. Gradia separates benchmark defects, infrastructure exclusions, judge disagreement, and model failures—then returns the evidence and a prioritized repair plan.

Open-source beta. SDK capture covers the calls it instruments. Stronger runtime enforcement needs its own qualification. Every receipt names its coverage.

open-source beta
public source · npm pending
npm install git+https://github.com/rudycelekli/gradia-guard.git#92a93c610f4d02c3db21fe6a6d526d9bd046fe86
npx --no-install gradia-guard run -- node agent.js
Capture, verify, and inspect your first bundle →
  1. 01governpolicy before dispatch
  2. 02evaluateapproved graders over exact evidence
  3. 03benchmarkone frozen ruler across releases
  4. 04improvefailures become governed data
  5. 05certifya release-specific evidence package
00

How it works, end to end.

One enterprise question — can an agent underwrite a mortgage? — carried the whole way, a stage at a time: interview, spec, critique, freeze, build, calibrate, run, review, iterate, certificate. Two of those stages are refusals, and they are not failure states. They are the reason the signed result at the end means anything. Play the walkthrough, or inspect the same twelve stages below.

Illustrative workflow · all figures are hypothetical specimen data, not measured results

01

Twelve steps. No shortcuts.

Every engagement runs the same spine, E0 through E11. Each step must clear its gate before the next one opens.

  1. E0BriefOne sentence: the task, the domain, what the model must do.
  2. E1InterviewStructured questions turn intent into measurable requirements.
  3. E2SpecRubric, tiers, tool surface, and budget — drafted as a contract.
  4. E3FREEZESubjective rubric terms are refused. The spec hashes to a sha256 fingerprint.
  5. E4BuildTasks, oracle solutions, and graders generated against the frozen spec.
  6. E5CertifyEvery task's oracle must earn full marks — and its grader must resist cheating.
  7. E6CalibrateLLM judges must clear Wilson-bounded agreement floors against human annotations.
  8. E7RunModels execute inside sandboxed environments; every step is logged.
  9. E8ReviewHumans audit transcripts and verdicts in campaign workbenches.
  10. E9AlignJudge and human verdicts are reconciled; floors re-checked on holdout.
  11. E10ExportTraining data and reports, sliced and fingerprinted.
  12. E11CertificateA signed attestation anyone can verify — even offline.
02

The refusals are the product.

Most platforms promise quality. Gradia refuses to proceed without it. These are real responses from the platform, not illustrations of intent.

409 · SPEC_FREEZE_BLOCKEDrefused

Refused: 2 subjective terms in rubric.

"reasonably clear" → replace with a measurable threshold
"good tone" → cite a rule a grader can check

A spec that can be argued about cannot be frozen. Fix the rubric, then freeze.

402 · BUDGET_REQUIREDrefused

Run refused: budget does not cover the estimate.

estimated_cost_usd
142.50
budget_cap_usd
100.00
shortfall_usd
42.50

The exact shortfall, before a single token is spent. Never a surprise invoice.

JUDGE CALIBRATION · HOLDOUTpassed

Wilson 95% lower bound vs. human annotations.

0.00floor 0.701.00

wilson_lb_exact 0.805 ≥ floor 0.70 ✓ — judge activated

Below the floor, the judge does not run. Agreement is proven, not assumed.

TASK CERTIFICATION · E5certified

Oracle must score 100%. Cheats must score 0%.

oracle solution
24/24 · 100%
cheat probes
0/6 · 0%
verdict
certified

If the reference answer can't earn full marks, or the grader can be gamed, the task never ships.

409 · VINTAGE_REFUSEDrefused

A task may not see past the instant its world was cut at.

as_of
2026-03-14T21:00:00Z
latest record
2026-03-15T08:12:00Z
blockers
record_after_as_of

Eleven hours of hindsight is the difference between a benchmark and a crystal ball. The build stops.

409 · EXPORT_REFUSEDcertified, not shippable

The licence permits evaluation. It does not permit redistribution.

certify
allowed
export · list
refused
blockers
redistribution_forbidden_for_export

A merkle root proves what ran without moving a byte. Licensed data stays where its licence says it stays.

03

LLM judging that earns the right to scale.

Humans annotate in campaign workbenches. Judges shadow them, disagree, and are revised against human reasoning — until holdout agreement clears the floors. Only then does judging scale.

annotate
shadow-judge
disagree
revise
holdout ≥ floors
scale
loop until every floor holds on holdout
04

Proof, on paper.

Every certified run ends in a signed attestation. Change one digit of the pass rate and the signature no longer verifies.

What one number is made of · specimen data

Gradia · Certificate of Evaluation

Benchmark Attestation

cert_9f3ba21c · specimen with demo data

Headline pass rate
90%

Fingerprint
sha256:7c1e9a04d2…f31b
Judge calibration
wilson lb 0.805 · κ 0.86 · n 200
Run
run_5d2c…a9
Report
rep_c4f1…7e

Gradia
Authorized attestation service
signed
sealed
05

Universes built for your work.

A Gradia Universe combines the environment, evolving world, actors, time, tools, evidence, and judges an agent needs to do real work. Agents operate through shells, editors, browsers, full computers, databases, and approved enterprise systems.

bash, editor, tmux, browser, computer, sql, frozen feeds, custom integrations

The tools a spec declares enter its fingerprint — swap one and the benchmark is a different benchmark. Custom enterprise integrations are certified against offline stubs, so nothing touches production while the tasks prove themselves.

05A

The world can change while the agent works.

Real work does not wait politely for a task to end. A client changes a priority. A calendar event moves. A policy owner retracts an instruction. New evidence arrives after the agent has already formed a plan.

Gradia can schedule those changes at exact logical boundaries, deliver only the view the benchmark identity is allowed to see, and measure whether the agent notices, recovers, preserves still-active constraints, asks for missing information, or escalates. The same declared event semantics run behind the shared guest boundary instead of being rewritten for each sandbox provider.

Customer systems enter through a read-only, rights-scoped capture. The certified path freezes a minimized replay with identities, permissions, edits, deletions, notifications, retention and allowed uses intact. A live read remains exploratory: it cannot claim the reproducibility of frozen evidence.

evolution witness
  1. 01declared event + exact scenario digest
  2. 02root-only application receipt
  3. 03world root before → world root after
  4. 04the exact projection the agent could observe
  5. 05act boundary + restore generation
  6. 06tamper-evident occurrence chain head

Ordinary logs show that something was printed. This witness is designed to prove what changed, when it became visible, which world received it, and whether restore preserved the same history.

06

Benchmarks with a date on them.

The questions enterprises actually want answered are time-indexed. Given this morning's filings and the current curve, does this credit still clear our overlays? A synthetic fixture can't ask that credibly, and a live API call during the episode destroys reproducibility outright.

Both are avoidable, because capture and replay are separable. The network is touched exactly once — by us, in the authoring plane, at a named instant. The agent reads a frozen capture through an ordinary tool and has no egress at all. The world is as frozen as it ever was. What's new is that it has a date, and the date is signed.

Same recipe, two vintages

Did the model get worse, or did the world get harder?

recipe_hash
identical · 4a91…c07d
vintage_digest · mar
e21f…88b3
vintage_digest · jun
9c40…1de6
pass rate
0.90 → 0.81
attributable to
the world, not the model

Hold the recipe identical, move the capture forward, and a score change decomposes. Every buyer with an agent in production has this question and nowhere to take it.

Five gates guard the instant
  1. V1No live egress at run time — the agent reads the capture, never the network.
  2. V2No lookahead — not one record may postdate the instant the world was cut at.
  3. V3Point-in-time faithfulness — a restated source is not what was knowable then.
  4. V4Provenance and rights — unlicensed data may be certified, never exported.
  5. V5Replay determinism — read the vintage twice; the bytes must match.

Each is a 409 with named blockers, refusing at build, at export, and at the listing door. Each was mutation-verified: stubbed out in testing to prove a named test catches its absence.

07

Four answers, off one ruler.

A benchmark is not the deliverable — it's the instrument. What an engagement actually produces is a decision, a diagnosis, a shopping list, and a way to check whether the fix worked. It can start from a single file.

One file → an ecosystem · hypothetical specimen data, not measured results

Which model should I use?

Bake-off

Every candidate runs the identical fingerprint. Per-episode model pins, confidence intervals, one ruler.

Where does the winner still fail?

Failure map

Clusters by the gate that refused, plus a behavioral taxonomy of what the agent actually did wrong.

What data would fix it?

Ranked asks

The failure clusters become a prioritized data request, not a purchase by the pound.

Did the fix hold?

Honest delta

Re-run the same fingerprint after your fine-tune. Same ruler, so the difference is the model.

Undirected data spend is the largest wasted line item in enterprise AI, and it is downstream of one thing: you cannot aim without a ruler.

08

Made of your production, not of fixtures.

A benchmark built out of invented cases measures an invented job. These four mechanisms take a frozen, certified environment and point it at the traffic you already have. You export the traces and upload them; nothing of ours is installed in the path your users run through, so there is no proxy to fail open and no tee to fall behind.

trace → task

Your traffic writes the tasks.

A production trace selects and parameterizes a task inside the already-frozen environment — it never edits the world. A trace naming something the environment does not contain is refused by name rather than fuzzy-matched, and the refusal list comes back as a coverage report on what your environment is missing.

pass^k

Reliability beside capability.

pass^k, not pass@k: every one of k sampled seeds must have passed. Reported at the minimum k across tasks, because k is a promise about every task in the number. An episode that broke the environment leaves the denominator, so the reliability figure and the headline rate are rates over one population — not two on one page.

simulated stakeholders

A counterparty that cannot leak the answer.

Multi-turn tasks need someone on the other end. The disclosure gate runs in the sandbox, deterministically, before the persona is prompted at all — the persona has no parameter through which a withheld fact could arrive. That is the difference between a model instructed to keep a secret and one that was never told it.

shadow traffic

The candidate, scored against what actually happened.

A candidate's score on a compiled task, printed beside the incumbent's real outcome on the trace it was compiled from. Nothing touches your production: no proxy, no tee — we replay the inputs against the sealed certified environment. Below twenty paired traces the net figure is not produced at all, and it never carries a p-value at any size: a significance claim over twenty paired binaries is a false precision that gets quoted out of context and then defended in a meeting.

Each of these refuses rather than guesses. A trace that does not fit is named, a pairing that lost its counterparty is withheld, and a run with too few paired cells is reported without a significance claim instead of with a flattering one.

predictive validity

Did the benchmark predict anything?

A benchmark score is a prediction about production, and almost nobody checks. So the certificate carries a pre-registered claim inside its signature — which production workflow this score expects to track, and the instant the claim was made. A correlation computed afterwards, against a workflow picked once the outcomes were visible, is a search nobody can see in the result.

A production period may only be scored against a claim already on record when the period opened, and it pairs with exactly one certificate — the latest one preceding it. Pairing against every earlier certificate is the natural implementation, and it inflates the sample size by how often we happened to mint. The rank correlation that comes back is computed live and never signed: the signature covers what we did, not how the world subsequently went.

The arithmetic imports nothing outside the Python standard library, on purpose. A buyer who disbelieves the number has to be able to re-derive it from the pairs alone — a re-derivation that first needs our database and our API standing up is one nobody performs.

09

Built for two kinds of builders.

For enterprises

Your benchmark, your data, verified end-to-end.

  • Evaluate the tasks your business actually runs, against your own systems.
  • Every artifact fingerprinted, every event on the hash-chained audit trail.
  • Deploy where you need it: laptop → VM → bring-your-own compute cells.
For creators & agencies
marketplace — onboarding sellers

Author, certify, publish — and sell eval access.

  • Build benchmarks on the same spine enterprises use, with the same gates.
  • Buyers run evals and get signed results — the test set never leaves the platform.
  • Your certificate is your storefront: provenance, calibration, and floors on display.
Hard, fairly.

A low pass rate matters only after the task, world, tools, answer, rubric, judge, and runtime have survived their own checks.

Gradia measurement standard

Failure alone does not prove difficulty. The environment may be broken, the answer may be underdetermined, the tool may be unavailable, or the judge may punish valid work. Gradia separates those failures from model failures before a task enters a headline denominator.

Only then do repeated, version-pinned runs estimate capability and consistency. Difficulty is an empirical distribution with uncertainty, not an adjective an author assigns. The result is narrower than a sweeping frontier claim and more useful: evidence about whether an agent can do a particular approved workflow in a particular frozen world.

10

Agents went to production. The rulers didn't.

Evaluation was a research chore while agents were demos. It became a procurement problem the moment they started touching real files, and the instruments didn't move with them.

What changed
  • Agents are being handed real work — underwriting files, discovery sets, production incidents — and approved on numbers borrowed from generic public benchmarks.
  • Published test sets enter training corpora within months. A leaderboard position is a claim nobody outside can falsify.
  • Procurement and audit have started asking for evidence of eval quality. There is no accepted format to hand them.
  • So models get picked on vendor reputation and training data gets bought by the pound.
What's hard to copy
  • Gradia does not train foundation models. The venue cannot also be a contestant.
  • Refusal can't be retrofitted. A tool that won its users by scoring anything cannot start telling them no.
  • There is no bypass — no skip-gate flag, no manual issuance. A certificate means the same thing whoever is holding it.
  • Certificate history, judge-agreement records, and registry roots committed before a model shipped are accumulated, never copied.

And when a pilot can't clear a calibration floor, what ships is a deterministic-gate-only benchmark and a rubric revision plan. The certificate date moves; the bar doesn't. A permissive platform in the same position ships the bad number.

11

Trust, itemized.

A certificate is only as strong as the weakest way around it. Gradia has no way around it — no override flag, no manual issuance, no negotiated floor. That is an architectural commitment, not a settings default, and it is the only reason the certificate can mean anything to someone who wasn't in the room.

sha256

Fingerprints on every spec, dataset, and eval harness.

hash-chain

An append-only audit trail links every event to the last.

wilson 95%

Judges activate only above statistically-proven agreement floors.

merkle root

A holdout's contents are provable from hashes alone — nothing is published.

offline verify

Certificates check out anywhere — no call home required.

no override

No skip-gate flag, no manual certificate issuance, no negotiated exception.

Neutrality

Gradia does not train foundation models. The venue cannot also be a contestant — which is why a lab can buy an honest map of its own weaknesses here, and why an enterprise's ranking isn't a vendor's marketing.

Mutation-verified

Every gate is stubbed out in testing and must be caught by a named test before it's restored. The gates aren't merely present — their absence is detectable.

12

Research that exposes its own uncertainty.

Gradia studies the ruler as aggressively as the agent: whether a world really changed, what the agent could observe, whether an evaluator survives mutation, and where model judges disagree with each other before a human decides.

The public record includes exact editions, code, evidence boundaries, and unfinished work. A DOI proves which artifact was published. It does not turn a pending adjudication or preregistered experiment into a result.

Read the research library →
Paper 01Published research artifact · one-seed paired GRPO

Reward Hacking in the RL Loop

A training-time study of what happens when an optimizer repeatedly probes a reward channel for exploitable seams. Matched 300-step GRPO/LoRA arms preserve exact manifests, frame chains, final adapters, analysis, and model-backed replay receipts so proxy success and oracle truth remain separately inspectable.

Read paper ↗doi:10.5281/zenodo.22259605
Paper 02Published preprint · empirical evaluator study

Reward-Hacking Wind Tunnel

An evaluation immune system that attacks the scorekeeper rather than the model. Oracle-wrong outputs probe frozen scorers, witnessed single-variable forks test causal attribution, guarded repairs are re-attacked on held-out evidence, and the resulting gameability and convergence measurements remain independently inspectable.

Read paper ↗doi:10.5281/zenodo.22233638
Paper 03Preliminary results · human adjudication pending

Conditionally Approved

A synthetic mortgage testbed for evaluating long-horizon agents while evidence, authority, and time change around them. The study binds each eligible episode to the exact world roots, visible projections, restore lineage, runtime, model, and evaluator that produced it.

Read paper ↗doi:10.5281/zenodo.22104672
Paper 04Pre-results · agreement study preregistered

The Value Engine Benchmark

A fully synthetic enterprise-sales negotiation environment with a deterministic, evidence-graded harness and a frozen evaluation grid covering 13 models and 3,510 graded episodes.

Read paper ↗doi:10.5281/zenodo.22073789

Bring one workflow. Build evidence you can defend.

Gradia | Enterprise AI evaluation and workflow intelligence