Can this agent do the job?
Define a task and its success criteria. Test a supported agent, inspect failures and compare changes under the same conditions.
Result: a reviewed assessment
Agent evaluation →Gradia helps your team test AI agents on your workflows and business rules. Find failures, check whether the test deserves your trust, and compare changes before your next release.
Bring one workflow and the decision you need to make. We agree the inputs, deliverables, review and fee before an assisted engagement begins.
No agent yet?
Start with the work people do today. Review its history, understand the process and define what an agent would need to get right.
Find your starting point ↓The customer requests $140. No approval is recorded.
Your rule: Refunds above $100 need a supervisor’s approval.
Agent action
Issues the $140 refund immediately.
The customer gets a refund, but the approval rule is broken.
Next test: Add the approval check, then test cases on both sides of the limit.
Scripted actions and checks, not live agent runs or measured results. Passing one example does not establish release readiness.
Choose the question you have today. Each path starts with a useful result and can connect to the wider platform as your needs grow.
Define a task and its success criteria. Test a supported agent, inspect failures and compare changes under the same conditions.
Result: a reviewed assessment
Agent evaluation →Bring selected, approved activity into Observer. Review the steps, decisions and gaps with the people who do the work.
Result: a reviewed case history
Workflow capture →Use Guard to check permissions around covered calls and keep a verifiable recording. Add a separate assessment to judge quality.
Result: covered execution evidence
Try open-source Guard →Already have a test? Audit your benchmark. Looking for process improvements? Explore workflow efficiency.
Guard runs locally without an account. Managed evaluations need a supported runtime and reviewed setup. Desktop capture is a controlled preview; signed public installers are pending.
An evaluation should answer a business question. The setup makes that question measurable.
01
What should the agent do? What must it never do? Review examples and success criteria with someone who knows the work.
02
Run the supported agent in the prepared environment. Inspect its actions and the evidence behind the checks.
03
See which failures belong to the agent, the test or the environment. Compare a change under matched conditions.
Go deeper when the work needs it
That is where a Gradia Universe comes in: a test environment built around a reviewed workflow, its rules and the information available at each step.
Start from a brief and expert input, or use Observer to assemble approved work across tools and time. Review the history before defining what an agent can see and do.
Then test a changed condition: a missing document, a delayed approval or a different policy. Inspect what happens in the supported environment.
This walkthrough illustrates the process. Capturing a history does not automatically reproduce an application. Runtime support, permissions and branch restoration need their own qualification.
37 days · one case
Approve the account, channel, pages and native window for this opportunity. Only approved sources and permitted evidence are retained.
Synthetic histories and simulated outcomes. This preview connects no sources and runs no agent.
Enterprise Universes
The messages, decisions and waiting between applications belong to the same job. A Universe gives that work an explicit state, rules and history.
Explore the Universe approach →
Inspect the history and the information each actor could see.
Check the stop receipt and settled actions before continuing.
Restore supported state into separate continuations and compare outcomes.
Concept artwork. Interruption and independent restoration are demonstrated in a bounded synthetic browser workflow. Each additional application, external effect and state component needs its own qualification.
A score is only useful if the test deserves your trust. Check known right and wrong answers, look for shortcuts that earn an undeserved pass, and review where human and AI judges disagree.
Public benchmarks show how a model scored on someone else’s test. They do not prove an agent can do your work. Inspect the source, the released evidence and the method behind a result.
Start with a brief or an existing test. Your expert helps define success; Gradia connects the reviewed requirements, supported environment and evaluation evidence. A reusable Universe or training export is a further step with its own qualification and permissions.
Governance · Evaluation · Benchmarks · Gradia Universes
Start with one agent run. Gradia Guard authorizes covered model and tool calls before dispatch, captures evidence after execution, and says exactly what it did not observe.
Policy checks before covered model and tool calls.
Deterministic checks, admitted judges, and humans inspect the same evidence.
Failures become regression tests, benchmark extensions, or approved training data.
Need more than tracing? The same contract runs inside a Gradia Universe with controlled visibility, interruptions, multiple actors, and branch/restore proof.
Already have a benchmark? Audit the ruler before interpreting the score. Gradia separates benchmark defects, infrastructure exclusions, judge disagreement, and model failures—then returns the evidence and a prioritized repair plan.
Open-source beta. SDK capture covers the calls it instruments. Stronger runtime enforcement needs its own qualification. Every receipt names its coverage.
npm install git+https://github.com/rudycelekli/gradia-guard.git#92a93c610f4d02c3db21fe6a6d526d9bd046fe86
npx --no-install gradia-guard run -- node agent.jsCapture, verify, and inspect your first bundle →One enterprise question — can an agent underwrite a mortgage? — carried the whole way, a stage at a time: interview, spec, critique, freeze, build, calibrate, run, review, iterate, certificate. Two of those stages are refusals, and they are not failure states. They are the reason the signed result at the end means anything. Play the walkthrough, or inspect the same twelve stages below.
Illustrative workflow · all figures are hypothetical specimen data, not measured results
Every engagement runs the same spine, E0 through E11. Each step must clear its gate before the next one opens.
Most platforms promise quality. Gradia refuses to proceed without it. These are real responses from the platform, not illustrations of intent.
Refused: 2 subjective terms in rubric.
A spec that can be argued about cannot be frozen. Fix the rubric, then freeze.
Run refused: budget does not cover the estimate.
The exact shortfall, before a single token is spent. Never a surprise invoice.
Wilson 95% lower bound vs. human annotations.
wilson_lb_exact 0.805 ≥ floor 0.70 ✓ — judge activated
Below the floor, the judge does not run. Agreement is proven, not assumed.
Oracle must score 100%. Cheats must score 0%.
If the reference answer can't earn full marks, or the grader can be gamed, the task never ships.
A task may not see past the instant its world was cut at.
Eleven hours of hindsight is the difference between a benchmark and a crystal ball. The build stops.
The licence permits evaluation. It does not permit redistribution.
A merkle root proves what ran without moving a byte. Licensed data stays where its licence says it stays.
Humans annotate in campaign workbenches. Judges shadow them, disagree, and are revised against human reasoning — until holdout agreement clears the floors. Only then does judging scale.
Every certified run ends in a signed attestation. Change one digit of the pass rate and the signature no longer verifies.
What one number is made of · specimen data
cert_9f3ba21c · specimen with demo data
Candidate scope: incident response, CI/CD pipelines, and infrastructure work.
Candidate scope: evidence-grounded analysis, credit memos, and portfolio work.
Candidate scope: contract review, discovery response, and compliance workflows.
A Gradia Universe combines the environment, evolving world, actors, time, tools, evidence, and judges an agent needs to do real work. Agents operate through shells, editors, browsers, full computers, databases, and approved enterprise systems.
The tools a spec declares enter its fingerprint — swap one and the benchmark is a different benchmark. Custom enterprise integrations are certified against offline stubs, so nothing touches production while the tasks prove themselves.
Real work does not wait politely for a task to end. A client changes a priority. A calendar event moves. A policy owner retracts an instruction. New evidence arrives after the agent has already formed a plan.
Gradia can schedule those changes at exact logical boundaries, deliver only the view the benchmark identity is allowed to see, and measure whether the agent notices, recovers, preserves still-active constraints, asks for missing information, or escalates. The same declared event semantics run behind the shared guest boundary instead of being rewritten for each sandbox provider.
Customer systems enter through a read-only, rights-scoped capture. The certified path freezes a minimized replay with identities, permissions, edits, deletions, notifications, retention and allowed uses intact. A live read remains exploratory: it cannot claim the reproducibility of frozen evidence.
Ordinary logs show that something was printed. This witness is designed to prove what changed, when it became visible, which world received it, and whether restore preserved the same history.
The questions enterprises actually want answered are time-indexed. Given this morning's filings and the current curve, does this credit still clear our overlays? A synthetic fixture can't ask that credibly, and a live API call during the episode destroys reproducibility outright.
Both are avoidable, because capture and replay are separable. The network is touched exactly once — by us, in the authoring plane, at a named instant. The agent reads a frozen capture through an ordinary tool and has no egress at all. The world is as frozen as it ever was. What's new is that it has a date, and the date is signed.
Did the model get worse, or did the world get harder?
Hold the recipe identical, move the capture forward, and a score change decomposes. Every buyer with an agent in production has this question and nowhere to take it.
Each is a 409 with named blockers, refusing at build, at export, and at the listing door. Each was mutation-verified: stubbed out in testing to prove a named test catches its absence.
A benchmark is not the deliverable — it's the instrument. What an engagement actually produces is a decision, a diagnosis, a shopping list, and a way to check whether the fix worked. It can start from a single file.
One file → an ecosystem · hypothetical specimen data, not measured results
Every candidate runs the identical fingerprint. Per-episode model pins, confidence intervals, one ruler.
Clusters by the gate that refused, plus a behavioral taxonomy of what the agent actually did wrong.
The failure clusters become a prioritized data request, not a purchase by the pound.
Re-run the same fingerprint after your fine-tune. Same ruler, so the difference is the model.
Undirected data spend is the largest wasted line item in enterprise AI, and it is downstream of one thing: you cannot aim without a ruler.
A benchmark built out of invented cases measures an invented job. These four mechanisms take a frozen, certified environment and point it at the traffic you already have. You export the traces and upload them; nothing of ours is installed in the path your users run through, so there is no proxy to fail open and no tee to fall behind.
A production trace selects and parameterizes a task inside the already-frozen environment — it never edits the world. A trace naming something the environment does not contain is refused by name rather than fuzzy-matched, and the refusal list comes back as a coverage report on what your environment is missing.
pass^k, not pass@k: every one of k sampled seeds must have passed. Reported at the minimum k across tasks, because k is a promise about every task in the number. An episode that broke the environment leaves the denominator, so the reliability figure and the headline rate are rates over one population — not two on one page.
Multi-turn tasks need someone on the other end. The disclosure gate runs in the sandbox, deterministically, before the persona is prompted at all — the persona has no parameter through which a withheld fact could arrive. That is the difference between a model instructed to keep a secret and one that was never told it.
A candidate's score on a compiled task, printed beside the incumbent's real outcome on the trace it was compiled from. Nothing touches your production: no proxy, no tee — we replay the inputs against the sealed certified environment. Below twenty paired traces the net figure is not produced at all, and it never carries a p-value at any size: a significance claim over twenty paired binaries is a false precision that gets quoted out of context and then defended in a meeting.
Each of these refuses rather than guesses. A trace that does not fit is named, a pairing that lost its counterparty is withheld, and a run with too few paired cells is reported without a significance claim instead of with a flattering one.
A benchmark score is a prediction about production, and almost nobody checks. So the certificate carries a pre-registered claim inside its signature — which production workflow this score expects to track, and the instant the claim was made. A correlation computed afterwards, against a workflow picked once the outcomes were visible, is a search nobody can see in the result.
A production period may only be scored against a claim already on record when the period opened, and it pairs with exactly one certificate — the latest one preceding it. Pairing against every earlier certificate is the natural implementation, and it inflates the sample size by how often we happened to mint. The rank correlation that comes back is computed live and never signed: the signature covers what we did, not how the world subsequently went.
The arithmetic imports nothing outside the Python standard library, on purpose. A buyer who disbelieves the number has to be able to re-derive it from the pairs alone — a re-derivation that first needs our database and our API standing up is one nobody performs.
A low pass rate matters only after the task, world, tools, answer, rubric, judge, and runtime have survived their own checks.
Gradia measurement standard
Failure alone does not prove difficulty. The environment may be broken, the answer may be underdetermined, the tool may be unavailable, or the judge may punish valid work. Gradia separates those failures from model failures before a task enters a headline denominator.
Only then do repeated, version-pinned runs estimate capability and consistency. Difficulty is an empirical distribution with uncertainty, not an adjective an author assigns. The result is narrower than a sweeping frontier claim and more useful: evidence about whether an agent can do a particular approved workflow in a particular frozen world.
Evaluation was a research chore while agents were demos. It became a procurement problem the moment they started touching real files, and the instruments didn't move with them.
And when a pilot can't clear a calibration floor, what ships is a deterministic-gate-only benchmark and a rubric revision plan. The certificate date moves; the bar doesn't. A permissive platform in the same position ships the bad number.
A certificate is only as strong as the weakest way around it. Gradia has no way around it — no override flag, no manual issuance, no negotiated floor. That is an architectural commitment, not a settings default, and it is the only reason the certificate can mean anything to someone who wasn't in the room.
Fingerprints on every spec, dataset, and eval harness.
An append-only audit trail links every event to the last.
Judges activate only above statistically-proven agreement floors.
A holdout's contents are provable from hashes alone — nothing is published.
Certificates check out anywhere — no call home required.
No skip-gate flag, no manual certificate issuance, no negotiated exception.
Gradia does not train foundation models. The venue cannot also be a contestant — which is why a lab can buy an honest map of its own weaknesses here, and why an enterprise's ranking isn't a vendor's marketing.
Every gate is stubbed out in testing and must be caught by a named test before it's restored. The gates aren't merely present — their absence is detectable.
Gradia studies the ruler as aggressively as the agent: whether a world really changed, what the agent could observe, whether an evaluator survives mutation, and where model judges disagree with each other before a human decides.
The public record includes exact editions, code, evidence boundaries, and unfinished work. A DOI proves which artifact was published. It does not turn a pending adjudication or preregistered experiment into a result.
Read the research library →A training-time study of what happens when an optimizer repeatedly probes a reward channel for exploitable seams. Matched 300-step GRPO/LoRA arms preserve exact manifests, frame chains, final adapters, analysis, and model-backed replay receipts so proxy success and oracle truth remain separately inspectable.
An evaluation immune system that attacks the scorekeeper rather than the model. Oracle-wrong outputs probe frozen scorers, witnessed single-variable forks test causal attribution, guarded repairs are re-attacked on held-out evidence, and the resulting gameability and convergence measurements remain independently inspectable.
A synthetic mortgage testbed for evaluating long-horizon agents while evidence, authority, and time change around them. The study binds each eligible episode to the exact world roots, visible projections, restore lineage, runtime, model, and evaluator that produced it.
A fully synthetic enterprise-sales negotiation environment with a deterministic, evidence-graded harness and a frozen evaluation grid covering 13 models and 3,510 graded episodes.