In development — no upload/extraction demo is live yet

Structured data out of Vietnamese
PDFs, scans, and photos

An AI-native pipeline (Gemini for a fast vision pass, escalating to Claude when confidence is low) for turning messy, real-world Vietnamese documents into tables, CSV, or JSON — with a confidence score attached to every field, not just a wall of text. That pipeline itself is still being built.

What already exists and is fully verifiable: a complete case study applying this approach's underlying discipline — source-traceable extraction, measured accuracy, honest reporting of what didn't work — to 83 real Vietnamese government files.

What's real today, and what isn't yet

Built and verified

  • · Full source-to-database traceability for every extracted record
  • · Measured, published accuracy and completeness — not claimed, measured
  • · A real-world case: 35 projects, 4,343 verified transactions, 96.8% field accuracy after fixes, from 83 government PDF/Excel files
  • · A honest, itemized "what's still wrong" list, not a hidden one

Not built yet

  • · A general-purpose upload-any-document converter (Phase 1, in progress)
  • · Output format selection (JSON/CSV/XLSX/text) for arbitrary documents
  • · A versioned accuracy benchmark across document types (Phase 2)
  • · Fine-tuned models — deliberately later-stage, see the methodology page

Running a business that deals with Vietnamese paperwork?

Invoices, receipts, supplier documents — see the business-focused page.

For business