AI audit reviewer
A draft financial-statement package goes in; a categorised exception report and a PDF come back in about 90 seconds, at a measured $0.24–0.28 per review.
- Role
- Sole engineer
- Status
- Delivered
- Client
- CPA firm
- Python
- FastAPI
- Anthropic Claude
- Next.js
- TypeScript
- Supabase
- python-docx
- openpyxl
- pdfplumber
- reportlab
- LibreOffice
- Docker
- Railway
- Vercel

The problem
- Context
- A reviewer at a CPA firm checks every outgoing financial-statement package by hand before it reaches a client: the arithmetic, the tie-outs between the statements, missing disclosures and the firm's house formatting. The tool takes the draft Word report and its Excel workpapers and returns a categorised exception report and a ready-to-send PDF, for the firm's reviewers and administrators.
- Constraints
- Formatting rules could not be hard-coded: they had to come from the firm's own style guide, so the tool carries no conventions of its own. Uploads are large, the whole request had to be bounded before anything is parsed, and the documents under review must never be stored. It had to run the same on a laptop and in a container.
- What was at stake
- A model that states a wrong number confidently is worse than no check: a reviewer who trusts it could sign off a defective statement. Enforcing a rule the firm never wrote creates busywork and erodes trust, and one failed model call must not sink the whole review.
What I built
Accuracy: numbers belong to code
A deterministic tie-out engine
Code refoots every total, balances the balance sheet, rolls equity and cash forward, ties ending cash and net income across the statements, and reconciles working-capital and related-party movements, within a one-unit tolerance. It even handles totals separated from their last line by a blank row, so layout quirks do not raise false alarms.
The model's numbers never overrule the arithmetic
The facts the code has verified are passed to the model, and any model finding that contradicts them is dropped, with the report saying so. The model's job is judgement on disclosures and wording; the arithmetic belongs to the engine.
Every finding anchored to a page
Each model finding has to include a short verbatim quote, which is matched back to the document to give it a section and a page number, so a reviewer can go straight to it.
Rules from the firm's own guide
A fixed schema, nothing invented
About forty formatting rules (fonts, sizes, capitalisation, underlines, currency signs, column widths, contents pages) are extracted from the firm's uploaded style guide into a fixed schema. A rule the guide does not state stays empty and its check is skipped.
Extract first, replace second
A new style guide is read and its rules extracted and saved before the old file is replaced, so a failed extraction leaves the previous guide working.
Workbooks read with intent
Only the statement tabs are parsed, and only up to the columns that hold the statements, so scratch calculations off to the side are never mistaken for figures.
Reliability: failure is designed
A failed model call is never fatal
The model gets at most two attempts, and only for empty or unreadable replies; refusals, rate limits and length limits surface immediately. If the review cannot run, the arithmetic, formatting and date findings still ship, with a warning naming the cause.
Requests bounded before they are read
Anything over about 51 MB is refused before the server spools the upload to disk, which it would otherwise do before authentication even runs.
Nothing reviewed is kept
Each review runs in a temporary directory that is deleted afterwards, and responses are marked never to be cached.
Security: fails closed
Tokens verified locally
Sign-in tokens are verified on the server against the identity provider's published keys, cached and refreshed when keys rotate, with expiry and audience always checked.
Roles read live, outages refused
Each request reads the user's role and blocked status fresh, so demoting or blocking someone takes effect immediately. If the identity provider cannot be reached, the request is refused rather than admitted.
Database rules as a second wall
Row-level security limits what each user can read, and only the server can change roles, block users or touch the firm's settings.
Cost, measured
Usage logged on every call
Input and output tokens are logged for each model call, which is how the $0.24–0.28 per review figure was measured on sample documents rather than estimated.
How it works
Tell me what you’re building and where it’s stuck.
I’ll tell you the cleanest path forward, including if it’s “don’t build that.”
Or write tocontact@alihassan.dev
