PDF Translation Quality Methodology
PDFTranslate evaluates whether a document can be processed, whether the promised output files are delivered, and whether important page structure remains usable. These checks establish product behavior; they do not establish perfect translation accuracy for every language, subject, scan, or layout.
What does the production service currently expose?
The public production capability contract is the source of truth for limits exposed to the application.
| Capability | Current production value |
|---|---|
| Input | Native and scanned PDF documents |
| File limit | 50 MB and 100 source pages |
| Languages | 13 source languages and 19 target languages |
| OCR | Available for image-only pages |
| Outputs | Translated-only PDF and bilingual PDF from one job |
| Quote | Source-page count and required credits shown before processing |
Capabilities can change. The product interface and capability endpoint take precedence over this page if a controlled rollout changes a limit.
How is a release checked?
Release checks cover four separate layers:
- Contract checks verify file limits, language allowlists, page accounting, job states, output kinds, and failure behavior.
- Document checks use native-text PDFs, image-only scans, structured pages, tables, links, and multilingual text.
- Artifact checks require both output files to be valid PDFs with the expected page count, readable text where applicable, and no development watermark.
- Runtime checks request the deployed site, initial server-rendered HTML, the public capability endpoint, indexability controls, and key routes after release.
Automated tests are necessary but not sufficient. Layout and translation quality also need visual and language-aware review.
What first-party evidence is available today?
The following checks have been completed during controlled release verification:
- A one-page native-text PDF completed the layout-preserving translation path and produced a one-page output with extractable translated text.
- A one-page image-only English scan completed OCR, translation, and reconstruction; translated-only and bilingual outputs contained visible, extractable Simplified Chinese without the original scanned glyphs covering the translation.
- Output handling is contract-tested for translated-only and bilingual artifacts, page-based credit accounting, failed-job credit release, and owner-scoped downloads.
These narrow checks establish that the paths work for the tested fixtures. They are not a general accuracy percentage and do not prove equivalent performance for handwriting, noisy scans, unusual fonts, complex tables, every language pair, or every PDF producer.
Which limitations should users expect?
- Machine translation can mistranslate, omit, or alter meaning.
- OCR quality depends on resolution, contrast, rotation, handwriting, and background complexity.
- Longer translated text can change line wrapping and spacing.
- Complex tables, forms, annotations, embedded media, and unusual fonts may not reproduce exactly.
- Certified, sworn, legal, medical, financial, immigration, and other high-stakes uses require qualified human review.
We do not publish an accuracy percentage until a versioned, representative corpus, scoring method, reviewer process, and reproducible results are publicly available.
Which components inform the method?
The document path uses published open-source tooling for OCR and PDF reconstruction. Relevant primary technical sources include the OCRmyPDF documentation, the Tesseract user manual, and the BabelDOC project. These sources describe the tools; the product claims above come from PDFTranslate's own contracts and controlled checks.
Questions about a result can be sent through the Contact page. Do not attach confidential PDFs to email.