Evaluation method
Measure usefulness, not just agreement.
This is a proposed evaluation protocol. It is not a claim that an independent benchmark, fairness audit, or compliance assessment has been completed.
Proposed protocol
Agree on the reference standard first.
- Define the intended task, eligibility boundaries, and job-related criteria.
- Select a representative, authorised sample before reviewing results, keeping a held-out evaluation set where appropriate.
- Have qualified reviewers adjudicate evidence with a documented disagreement-resolution process. Blind model and source labels where feasible.
- Compare the existing workflow or defined baseline against the proposed process on the same cases.
- Report denominators, failed runs, exclusions, uncertainty, and observed review effort, not a selected set of success stories.
- Document limits. A small pilot does not establish broad fairness, reliability, or regulatory compliance.
| Measure | Definition | Interpretation limit |
|---|---|---|
| Useful issues identified | Count issues confirmed by qualified reviewers against the agreed job-related evidence. | Disagreement alone is not an error. |
| False alarms | Count escalations independently judged not to require correction or additional information. | A high escalation rate can add work without improving review. |
| Review time | Record the actual reviewer time per case and the distribution, not just an average. | Keep setup and training time separate from steady-state estimates. |
| Cost per case | Include model and API cost, infrastructure, and any included human or support effort. | Do not confuse a platform fee with all-in cost. |
| Record completeness | Check required inputs, responses, failures, timestamps, reviewer rationale, and outcome. | Completeness is not proof that the outcome is correct. |
| Baseline comparison | Use the same cases and a pre-specified single-model or existing-workflow baseline. | Avoid choosing only examples where the panel looks better. |
Separate model calls do not establish statistically independent errors. Treat model-provided confidence as self-reported, not calibrated accuracy.
Published evidence
Results belong beside their method.
No evaluated performance results are published here. That does not mean no private testing exists; it means no results are being claimed on this page.