Self-hosted PDF OCR that never uploads your documents

Most OCR services process your documents on their servers. SnapOtter runs OCR locally, so contracts and records stay on infrastructure you control.

Local AI, one-time bundle install

What most teams do today

You upload a scanned PDF to an online OCR tool and download searchable text or a tagged PDF. The document leaves your network to get there.

Why that's a problem for sensitive files

Scanned documents are exactly the sensitive kind: contracts, medical records, statements. Online OCR services process and often retain uploads under terms you do not set, which is hard to square with GDPR or HIPAA and impossible to audit.

How SnapOtter does it privately

SnapOtter runs the OCR model inside your instance. Install the OCR bundle once, then it works offline, air-gapped included. The PDF never leaves your network, and there is no per-page charge.

Self-host in one command

Process a file over the REST API

Last reviewed July 10, 2026. Commands match the current single-container image and REST API.

Frequently asked questions

Does it work air-gapped?
Yes. After the one-time OCR bundle install, OCR runs entirely offline, so it works in air-gapped and classified networks.
Are my documents retained anywhere?
No. The PDF is processed inside your instance. There is no auto-save; you choose whether to keep a result.
Which languages are supported?
The OCR bundle covers a range of languages; pass the target language in the tool settings.

Deploying this across a team?

Audit trails, per-tool permissions, SSO, and air-gapped deployment for regulated environments. Talk to us about running SnapOtter as shared infrastructure.