Agentic acceptance testing

Revisore

compliance and traceability

An agentic harness for acceptance testing, with evidence for the EU AI Act technical file.

tracedone verdict per requirementTicketand discussionRequirementsfrom the ticketTestsfrom requirementsRunon the real productResultsone per testReportbugs, evidenceVerdictper requirementgate: code checks every steptraced: requirement to test to result Ticketand its discussionRequirementswritten from the ticketTestswritten from the requirementsRunon the real productResultsone per testReportbugs and evidenceVerdictone per requirementgate: code checks every steptraced: requirement to test to result

What it is

An agentic harness for acceptance testing

Writes

Requirements from the ticket, then tests from the requirements.

Runs

On the real product, built from the branch.

Traces

Every requirement to its tests, every test to its results.

Proves

A dated, signed report with the evidence.

Why Revisore

What sets it apart

30

Trained on 30 projects

Automotive, dating, storage, MedTech, entertainment. Since 2010.

Code

The gate is code, not a model

Every step checked in 3.8 s median. No model grades its own work.

300×

A map instead of grep

Agents read 300× less repository context per turn.

1:1

One verdict per requirement

Traced to its tests and results. Dated evidence for the AI Act file.

Who it's for

You don't write the tests. You evaluate them.

Test engineer

Reviews tests instead of writing them.

Technical product manager

Judges the verdict per requirement.

Compliance officer

Collects testing evidence for the AI Act file.

Anyone who does acceptance testing and can make a reasoned judgement on how well the testing went.

How it works

How the harness is built

Architecture. In: the ticket, its discussion and the code branch. Inside an isolated Docker environment per run: a test container where agents run the flow, with checks between stages, tests and a repository map, next to the system under test built from the ticket's branch. Out: a test report, traceability and a test-code pull request. Observability and the watchdog sit outside the run.

Watchdog and observability

The run sends every step to a watchdog that sits outside it. The watchdog has its own clock and a limits file that is re-read during a run. It measures each step against limits on time, turns and money, feeds a live dashboard and a record of every call, check and verdict, and can warn, retry or stop the run.

Test code map for agents

where-are-we reads test code, product code and history in one tree walk of about ten seconds, offline and without a model, and writes a map to disk. The agent's prompt carries a pointer of about 212 tokens instead of 64,000; the agent asks a question and gets only the matching rows.
The map tool is open source: github.com/ngavrish/where-are-we

Test flows

Five test flows, one harness

FLOW 1 OF 5

Functional tester

A ticket reaches test and needs full coverage.

  • Requirements
  • Bugs
  • Tests
  • A report
  • A run on the branch
  • A test-code PR
A ticket becomes requirements, then a run, then a report; a failed test becomes a bug.

FLOW 2 OF 5

Happy-path tester

You want every acceptance criterion checked, fast.

  • One scenario per criterion
  • A person approves the design
Four acceptance criteria, each with one scenario, and a person who approves.

FLOW 3 OF 5

Data tester

The ticket changes a metric, a report or a pipeline.

  • Every row traced, source to destination
  • Lost or changed rows named
Rows followed from a source table to a destination table: most arrive the same, one is changed, one is lost.

FLOW 4 OF 5

Analytic

The ticket is still in design; nothing to run yet.

  • The full requirement set
  • A test plan
  • Both before the code exists
A design turned into requirements and a test plan, with the code not written yet.

FLOW 5 OF 5

Ad-hoc check

You have one question and no ticket.

  • One answer about the running product
  • The evidence behind it
A question goes to the running product and comes back as an answer with evidence.

The flow

Four chapters every test flow inherits

  1. Ticket requirements, discussion, the branch
  2. 1 Introduction code map · plan · environment
  3. gate
  4. 2 Requirements from the ticket and thread
  5. gate
  6. 3 Testing written and run on the product
  7. gate
  8. 4 Reporting bugs · report · test-code PR
  9. gate
  10. Verdict one per requirement, for you to evaluate

Outside the run the watchdog sees every step and enforces limits on time, turns and money.

Every step passes a gate

Agentdoes the work Rule gatecode checks the result Next steponly when green Repair agentfixes what the gate found green red checked again Agentdoes the work Rule gatecode checks the result Repair agentfixes what it found Next steponly when green red again green

Real numbers

Most first attempts don't pass the gate

63%red on the first check297 of 474 steps
110red first, green at the endof those 297 steps
3.8 smedian time for one gatecode, not a model

745 rule-gate checks in 131 runs, 21 August to 28 September 2026. Round n is the n-th check of the same step.

The output

What every run produces

A test report marked FAIL with three passed and two failed scenarios, and one defect showing the expected value, the actual value and its Jira link.
The test report.
Requirements traceability: each requirement with the tests that prove it, a failing requirement in red, and one requirement no test covers yet.
Requirements traceability.

Verdict

One per requirement

pass, fail or untested

Defects

Expected against actual

each linked to its Jira bug

Evidence

Dated and kept

for the technical file

What it is

EU law for AI, by risk

Regulation (EU) 2024/1689. The higher the risk, the more you must prove: tested, documented, logged.

Who it applies to

AI used in the EU

Providers who build it and deployers who use it, EU-based or not. Heaviest duties for high-risk AI.

Where Revisore's evidence fits

WHAT THE ACT ASKS OF THE PROVIDER

WHAT REVISORE CONTRIBUTES

Testing against metrics set in advanceArt. 9(8)
Pass criteria fixed before tests run
Test and validation proceduresArt. 17(1)(d)
The same flow on every ticket
Test logs and reports, dated and signedAnnex IV 2(g)
Test report and traceability, per run
Documentation kept for 10 yearsArt. 18
Every call, check and verdict kept

Applies only if your product is a high-risk AI system. Revisore supplies evidence; compliance stays with the provider.

When it applies

  1. New deadlines in forceReg. (EU) 2026/1744
  2. High-risk AI systemsHiring, credit, education, essential services
  3. AI in regulated productsMachinery, medical devices, vehicles

Assessment route varies

Self-check or notified body, Art. 43

Kept 10 years

Reports stay in the technical file

On request

Shown to an authority when asked

Tech stack

What it runs on

Models
Dynamic
chosen per phase
Orchestrator
Go · gRPC
polls Jira, dispatches runs, streams every step
UI
Go + plain JS
observability dashboard: runs, cost, limits, verdicts
Tests
In a container
run next to the product, isolated per run
Search
Keyword and meaning search, re-ranked
local, on CPU; no vector database
Watchdog
Go · SQLite
limits, budgets, the record of every call and verdict
Configs
JSON flows · Markdown agents · YAML limits
rules as Markdown, in git
Infrastructure
Docker Compose
one isolated environment per run

Summary

Four things to take home

  1. Observability firstEvery run's cost, time and limits, watched from outside.
  2. Fire is useful in a forgeRevisore holds the agents to YOUR requirements.
  3. Give agents a map, not grep300× less repository context per turn.
  4. The record is evidenceReports and traceability for the AI Act technical file.

Contact

Get in touch

Nikita Gavrish · Evaluation Engineer · Chiron Systems