How to OCR Chinese, Arabic, Thai and Other Languages

Extract text from scanned documents in 29 languages — Chinese, Japanese, Korean, Arabic, Hindi, Thai and more — with tips for accurate non-Latin OCR.

3 min read ·

Free OCR tools often assume English. PDF OCR reads 29 languages, including Chinese (Simplified and Traditional), Japanese, Korean, Arabic, Hebrew, Hindi, Thai and Vietnamese. Add the scanned PDF or photo, pick the document's language from the list, and click Start OCR. The key to good results is choosing the right language — here's why, and how to handle the trickier scripts.

Supported languages

Latin script: English, Spanish, French, German, Italian, Portuguese, Dutch, Polish, Czech, Romanian, Hungarian, Swedish, Danish, Norwegian, Finnish, Turkish, Vietnamese, Indonesian.

Other scripts: Greek, Russian, Ukrainian, Arabic, Hebrew, Hindi, Thai, Japanese, Korean, Chinese (Simplified), Chinese (Traditional).

Step by step

  1. Open PDF OCR and add the scanned PDF, or photos of the pages.
  2. Choose the language the document is written in.
  3. Click Start OCR. The language data downloads the first time you use a language — Chinese and Japanese data are larger, so allow a little extra time.
  4. Copy the text or download it as a .txt file.

Why the language setting matters so much

OCR doesn't just match shapes; it uses knowledge of each language — its alphabet, accents and common words — to decide between similar-looking characters. With the wrong language selected:

  • Accented letters lose their accents (é becomes e) or turn into symbols.
  • Non-Latin scripts come out as nonsense Latin letters.
  • Chinese characters may be read as Japanese ones, or vice versa.

So if a result looks like gibberish, check the language first. The setting also decides what is downloaded: only the chosen language's data is fetched, which keeps the first run quick for most languages.

Tips by script

Chinese, Japanese and Korean

  • Choose Simplified for mainland China and Singapore documents, Traditional for Taiwan, Hong Kong and Macau.
  • Horizontal text reads best. Vertical columns, common in Japanese books and older Chinese texts, are read much less reliably.
  • Characters have fine strokes, so sharp scans matter: 300 dpi or a well-focused phone photo.
  • OCR output may contain spaces between characters; remove them with find-and-replace if needed.

Arabic and Hebrew

  • These read right to left. The extracted text is in logical order, so it looks right when pasted into an app that supports right-to-left text, such as Word, Google Docs or a notes app.
  • Clean printed fonts work well; decorative calligraphy and handwriting do not.

Thai, Hindi and other Indic scripts

  • Thai has no spaces between words, and the OCR output follows the same convention.
  • Hindi's connected headline (shirorekha) needs a crisp scan; blurred photos lose the small marks above and below the line.

Documents with two languages

A bilingual form or a textbook with English terms in a Chinese text mixes scripts. Choose the language of most of the text first. If an important part comes out badly, run OCR again with the other language and keep the better result for that part.

What next: translate the text

Once you have the text, you can translate it: paste it into the AI Translator, which handles long documents and keeps paragraphs in place. For a single photo of a sign, menu or label, Image to Text is a quicker start.

Frequently asked questions

Can it read handwriting?

Not reliably in any language. Neat printed letters work; joined-up writing usually doesn't.

My PDF has selectable text already — do I need OCR?

No. Pages with real text are read directly, which is faster and exact, and PDF to Text does that for the whole document.

Is the document uploaded?

No. The engine and language data are downloaded to your browser and the recognition runs on your device.

Tools mentioned in this guide

← All guides