Self-hosted PDF OCR that never uploads your documents
Most OCR services process your documents on their servers. SnapOtter runs OCR locally, so contracts and records stay on infrastructure you control.
Local AI, one-time bundle install
What most teams do today
You upload a scanned PDF to an online OCR tool and download searchable text or a tagged PDF. The document leaves your network to get there.
Why that's a problem for sensitive files
Scanned documents are exactly the sensitive kind: contracts, medical records, statements. Online OCR services process and often retain uploads under terms you do not set, which is hard to square with GDPR or HIPAA and impossible to audit.
How SnapOtter does it privately
SnapOtter runs the OCR model inside your instance. Install the OCR bundle once, then it works offline, air-gapped included. The PDF never leaves your network, and there is no per-page charge.
Self-host in one command
Process a file over the REST API
Last reviewed July 10, 2026. Commands match the current single-container image and REST API.
Frequently asked questions
- Does it work air-gapped?
- Yes. After the one-time OCR bundle install, OCR runs entirely offline, so it works in air-gapped and classified networks.
- Are my documents retained anywhere?
- No. The PDF is processed inside your instance. There is no auto-save; you choose whether to keep a result.
- Which languages are supported?
- The OCR bundle covers a range of languages; pass the target language in the tool settings.
More self-hosted workflows
Deploying this across a team?
Audit trails, per-tool permissions, SSO, and air-gapped deployment for regulated environments. Talk to us about running SnapOtter as shared infrastructure.