MLQU 006 · Live
Document transcription
Scanned pages to formatted Word, with stamps and signatures kept in place.
The problem
Ordinary OCR reads the text on a scanned document and throws away everything else. For certified and legal work the stamp and the signature are the point, and losing them makes the output useless.
How it’s built
Read each page
One Claude Vision call per page, run concurrently under a configurable ceiling, so a hundred-page file does not crawl through serially.
Find the non-text
OpenCV and scikit-image locate stamps, seals and signatures and record where on the page they sat.
Two output modes
Placeholder replaces each element with a text marker. Preserve extracts the image and re-embeds it at the detected position.
Async by default
Submit a job, poll its status, download through a presigned link. Storage is S3 or MinIO, with Alembic managing the schema.
What changed
- Stamps and signatures survive the conversion instead of vanishing
- Long documents process page-parallel rather than one at a time
- Output is a formatted DOCX, ready to edit
Want one of these?
Two lines is enough. We reply within a day.