Evaluation method

Measure usefulness, not just agreement.

This is a proposed evaluation protocol. It is not a claim that an independent benchmark, fairness audit, or compliance assessment has been completed.

Proposed protocol

Agree on the reference standard first.

  • Define the intended task, eligibility boundaries, and job-related criteria.
  • Select a representative, authorised sample before reviewing results, keeping a held-out evaluation set where appropriate.
  • Have qualified reviewers adjudicate evidence with a documented disagreement-resolution process. Blind model and source labels where feasible.
  • Compare the existing workflow or defined baseline against the proposed process on the same cases.
  • Report denominators, failed runs, exclusions, uncertainty, and observed review effort, not a selected set of success stories.
  • Document limits. A small pilot does not establish broad fairness, reliability, or regulatory compliance.
Measures to agree before a pilot
MeasureDefinitionInterpretation limit
Useful issues identifiedCount issues confirmed by qualified reviewers against the agreed job-related evidence.Disagreement alone is not an error.
False alarmsCount escalations independently judged not to require correction or additional information.A high escalation rate can add work without improving review.
Review timeRecord the actual reviewer time per case and the distribution, not just an average.Keep setup and training time separate from steady-state estimates.
Cost per caseInclude model and API cost, infrastructure, and any included human or support effort.Do not confuse a platform fee with all-in cost.
Record completenessCheck required inputs, responses, failures, timestamps, reviewer rationale, and outcome.Completeness is not proof that the outcome is correct.
Baseline comparisonUse the same cases and a pre-specified single-model or existing-workflow baseline.Avoid choosing only examples where the panel looks better.

Separate model calls do not establish statistically independent errors. Treat model-provided confidence as self-reported, not calibrated accuracy.

Published evidence

Results belong beside their method.

No evaluated performance results are published here. That does not mean no private testing exists; it means no results are being claimed on this page.