Skip to main content
Zimei TechnologyEnterprise AI · Development and delivery
EnglishEN
Discuss a project
Back to articles

Which AI tool can translate a PDF? Handling tables, scans, and long documents

Ordinary text PDFs can be translated in their entirety; scanned documents must be OCRed first; tables must be checked for rows and columns; long documents must be unified in terminology and then processed by chapters. Tools should be selected according to the file type.

This article was generated and organized by AI. Translation functions, file restrictions and supported languages ​​may change, please refer to the current official instructions; important documents should be reviewed by professionals.

Two documents of different colors are connected by a short arrow, with a table square in the middle, indicating that text and layout are processed at the same time when translating the PDF.

First determine whether the PDF is text or pictures

Let’s talk about the answer first: you can try ordinary text PDF, Google Translate and DeepL directly; if you want to select some pages and continue processing in the Acrobat workflow, you can look at Acrobat’s PDF translation; for scanned documents, you must first confirm whether the text can be recognized; for tables and long documents, you can’t just check whether the translation is smooth, but also check the format, terminology and missing pages.

Before selecting a tool, drag the text with the mouse. The ability to select text and the whole page being just a picture are two completely different PDFs.

Many so-called "PDF translation failures" are not because the translation model cannot understand the language, but because the tool does not read the text correctly at all, or when the translation becomes too long, the original tables and pages cannot fit in it.

Normal text PDF: Translate the entire document first

If the text can be copied and the page is dominated by consecutive paragraphs, the easiest way is to hand over the entire document to a document translation tool. Google Translate currently supports uploading PDF, DOCX, PPTX, and XLSX, and downloading translated documents; its public limitations are that files cannot exceed 10 MB, PDFs cannot exceed 300 pages, and document translation does not support small screens and mobile phones.

DeepL also provides PDF file translation, which can be used on both web and desktop applications. Different plans have different file times and download formats. Some paid plans can download the translated PDF as DOCX for continued modification. DeepL officials also recommend that if the PDF results are not satisfactory, try to upload the original DOCX or PPTX.

Therefore, when there are only a few pages of ordinary text, it is enough to first try a tool that can directly download the entire translation. Don't copy each page to the chat window first; you'll lose title hierarchy, page numbers, and paragraph relationships.

Scan image PDF: recognize text first, then translate

A scanned PDF may look like it has words, but it may actually be a set of pictures. Google Translate clearly states that text from images and scanned pages will be found in the output file, but the text will not be translated. Adobe also explains that PDFs that are scanned, encrypted, too large, or have complex structures may skip translation.

DeepL can process scanned documents, but officials remind that the quality of the scan will affect the quality of the translation. Skews, blurs, low contrast, and handwritten annotations can all cause errors in text recognition and then bring errors into translation.

  1. First zoom in and check: whether the edges of the text are clear and whether the page is skewed or shadowed.
  2. First do OCR to turn the image into text that can be searched and copied.
  3. Randomly check the names, numbers, units and titles to confirm that there are no wrong lines in the identification.
  4. Translate again and keep the original scan and compare it page by page with the translation.

If OCR recognizes a "0" as an "O," the translation software won't know that the original was written incorrectly. The first thing to test on scanned documents is recognition, not language style.

Table PDF: It’s not enough for the translation to be correct, the rows and columns must also be wrong

The difficulty with tables is not to translate each word into another language, but to maintain the row-column relationship. The translation is often longer than the original text, cells wrap, column widths change, and headers may squeeze onto the next page. Financial tables, parameter tables, and schedules cannot yet tolerate numbers running into the next row.

This type of file will give priority to original Excel, Word or typesetting files. DeepL officials also recommend uploading the original document when the PDF quality is not satisfactory. When there is only PDF, you can first translate the entire translation and keep the general layout, and then check the table headers, rows and columns, numbers, units, footnotes and cross-page content table by table.

  • Randomly select three lines and check horizontally whether the original text and the translated text are still on the same line.
  • Search for the amount, percent sign, date and model number to confirm that the quantity has not increased or decreased.
  • Check whether cross-page headers, merged cells, and footnotes are missing.
  • Export once before official delivery to check for font substitutions, truncation and blank pages.

Long documents: unify terminology first, then process by chapters

Long documents cannot be judged solely by "can they be uploaded". Within 300 pages there may also be inconsistencies in terminology, missing chapters, and misplaced figure descriptions. It is more useful to compile product names, organization names, professional terms and abbreviations that should not be translated into a one-page glossary instead of repeatedly asking for "professional translation".

If the tool exceeds file size or page limits, split by chapter rather than mechanically cutting every ten pages. Chapter headings, tables, and footnotes should be kept intact; use the same glossary for each section, and finally consolidate the table of contents, numbering, and cross-references.

General-purpose tools such as ChatGPT are more suitable for explaining a certain paragraph after translation, checking terminology, or comparing modifications, but are not suitable for formatting restoration of the entire complex PDF by default. OpenAI also currently states that visual reading of images and charts in PDF is only available under specific corporate accounts and upload methods. If you cannot see "Support PDF upload", it is assumed that all accounts will read charts.

Select directly by file type

  • Normal, Chinese text-optional PDF: Try Google Translate or DeepL for full document translation first.
  • You need to select the page, preview it, and continue to modify it in the Adobe environment: see the translation process between Acrobat and Adobe Express.
  • Scanned documents: OCR first, spot check the recognition results, and then translate.
  • Complex tables: If you can find the original Word or Excel, don't use the PDF; if you can't find it, check the rows and columns table by table.
  • Long document: Make a glossary first, split it into complete chapters, and finally unify the numbering and table of contents.

There is one more thing that cannot be omitted: before uploading contracts, customer information, unpublished reports or documents with personal information, first confirm whether the company allows uploading to the service and the data processing rules of the corresponding account. The convenience of translation does not mean that documents can be handed over to external platforms at will.

official information

The documentation capabilities and limitations of this article are derived from the following official sources, verified as of August 17, 2026:

Discuss a project