Add on-device OCR to Proton Drive document scanning for searchable PDFs
Proton Drive offers document scanning, but the resulting PDFs are image-only without embedded text. This makes scanned documents (leases, medical records, IDs, receipts) unsearchable and inaccessible via text selection or Proton Drive's search functionality.
Why This Matters
Users need to find specific content across their document archive without manual review of each file. Without OCR, workflows like locating a lease clause, extracting an address from an ID, or finding a prescription date require opening and reading each scanned document individually.
Proposed Solution
Add on-device OCR capability to Proton Drive's document scanner that:
1. Processes scanned pages locally using zero-access encryption
2. Embeds searchable text layers in the resulting PDF (searchable PDF format)
3. Allows text selection, copying, and full-text search within Proton Drive
4. Optionally downloads language packs for offline OCR support
Precedent & Feasibility
- FairScan, an open-source Android scanner, includes local OCR with text-searchable PDF output [FairScan blog]