How to Turn Scanned Invoices into Structured Data with an Agent
Everyone who's tried to automate invoice intake knows the wall: OCR reads the scan fine, but tabular layouts come back scrambled — descriptions divorced from their amounts, dates split across lines. The regex parser that worked on Monday's invoice breaks on Tuesday's.
That wall is exactly where a language model shines. With Apdf connected over MCP, the OCR tools deliver the characters and the agent does what regexes can't: reassemble meaning from messy text, then hand you clean JSON.
Connect the agent to Apdf
claude mcp add --transport http apdf https://apdf.io/mcp/main
Ask for records, not text
https://files.northlight.example/inbox/invoice-scan-1.pdf
https://files.northlight.example/inbox/invoice-scan-2.pdf
Look at what the agent actually had to work with — the raw OCR return for the first invoice:
{
"pages_total": 1,
"characters_total": 215,
"pages": [
{
"page": 1,
"characters": 215,
"content": "Brenner\nINVOICE\n\nBuerotechnik\n\n2026-0714\n\n—\n\nissued\n\nJuly\n\nDescription\nToner\n\ncartridges\n\nPrinter\n\nTOTAL\n\nPayment\n\n2026\n\nAmount\nHP\n\n207X\n\nmaintenance\n\nJuly\n\nDUE\n\n2,\n\nGmbH\n\n(4x)\n\n312.00\n\nEUR\n\n145.00\n\nEUR\n\n457.00\n\nEUR"
}
]
}
A regex parser dies here. The agent reassembles it:
[
{
"vendor": "Brenner Buerotechnik GmbH",
"invoice_no": "2026-0714",
"date": "2026-07-02",
"lines": [
{ "description": "Toner cartridges HP 207X (4x)", "amount": 312.00 },
{ "description": "Printer maintenance July", "amount": 145.00 }
],
"total": 457.00,
"currency": "EUR",
"sum_check": "312.00 + 145.00 = 457.00 ✓"
},
{
"vendor": "Skyline Catering UG",
"invoice_no": "R-88231",
"date": "2026-07-09",
"lines": [
{ "description": "Team lunch July 8 (14 pax)", "amount": 406.00 },
{ "description": "Delivery fee", "amount": 18.50 }
],
"total": 424.50,
"currency": "EUR",
"sum_check": "406.00 + 18.50 = 424.50 ✓"
}
]
Make the sum check non-negotiable
The sum_check field isn't decoration —
it's the guard that makes agentic extraction accountable. OCR's classic failure
mode is a mangled digit, and a line-items-versus-total cross-check catches it
cold: if the sum doesn't match the stated total, the agent flags the invoice for
human eyes instead of feeding a wrong number into your books.
From here, scale is a sentence — “do the same for everything in the inbox folder and give me a CSV” — and the agent loops the same two tools over the batch.
Where to go from here
The same connection carries every PDF operation — and the engagement layer behind them.
Related tutorials
OCR Scanned PDFs by Asking Claude
That folder of scans — old leases, signed contracts, photographed paperwork — becomes searchable conversationally: Claude checks the text layer, runs the OCR job through the Apdf MCP tools, and answers questions from the recovered text with page-level receipts.
Extract Text from Scanned Invoices for Automated Data Entry
Automate data entry from scanned invoices and receipts using Node.js and the Apdf OCR Read API. Extract vendor names, invoice numbers, and totals from image-based PDFs and feed them directly into your accounting system.
Convert Scanned PDFs to Searchable Documents with Python
Transform image-based scanned PDFs into fully searchable documents using Python and the Apdf OCR API. Perfect for digitizing paper archives, enabling text selection, and making legacy documents ready for modern search systems.
Your code made the PDF.
Then it went dark.
Opened, read, re-read, dropped on page 4 — you never see any of it. Share the PDFs you generate through Apdf recipient links, and every signal becomes something you can act on: ping Slack, update the CRM, let an agent follow up. Same account, same API token, one more call.
API · Webhook · MCP