← Back to home

Evidence

What this is

A structured evaluation comparing AI-generated tax workings against outputs from qualified accountants. Same inputs, same jurisdictions, blind review.

Methodology

Each test case consists of a realistic taxpayer scenario with source documents: bank feeds, receipts, invoices, prior returns. Two arms produce workings independently.

Arm A: AI system using Open Accountants jurisdiction skills, Claude, and the OA connector. No human in the loop during computation.

Arm B: Qualified accountant in the relevant jurisdiction, working from the same source documents, using their normal tools and professional judgment.

A third-party reviewer scores both arms against the correct filing position. Discrepancies are classified by type: arithmetic, legislative interpretation, missing deduction, overclaim, procedural.

Results

Evaluation data will be published here as test cases are completed. Check the Open Accountants repository for the latest.

Known weaknesses

This section will document the specific categories where AI workings consistently underperform, the failure modes we've identified, and what we're doing about each one. No system works everywhere. The point is knowing where it doesn't.