Why Arabic Text Looks Broken When You Copy It from a PDF
The three-part problem
When you copy Arabic out of a PDF — or convert it to Word — you often get one of three symptoms: disconnected letters (each ح standing alone), reversed words that read backwards, or swapped diacritics floating before their letters. None of this means the PDF is "corrupted". It means the PDF stores Arabic in a form designed for printing, not for text processing.
1. Shaped glyphs instead of letters
Arabic letters change shape depending on their position. A PDF does not store "ب"; it stores the exact drawn shape (an initial, medial, final or isolated glyph) from the font. Text extraction therefore returns presentation forms — the Unicode block made for those shaped variants — instead of plain letters.
2. Visual order instead of logical order
The printing engine draws glyphs strictly left-to-right across the page, because that is how the file is rendered. So the "first" character in the stored stream is often the last letter of the first word. Copying in stream order gives you backwards text.
3. Lam-alef and diacritics
Ligatures like لا are stored as a single glyph, and harakat are separate marks positioned above the letters. Without normalization, Word receives the wrong building blocks and cannot re-form the words.
How a proper converter fixes it
A converter worth using performs three repairs in order: detect shaped presentation forms and fold them back to base letters (Unicode NFKC normalization handles ligatures too), detect lines whose reading order is right-to-left and rebuild them logically, and re-attach combining marks to their base letters. Word's own text engine then shapes and connects the Arabic exactly as it should — editable, searchable, copy-paste friendly.
Try it with the PDF to Word tool: the editable mode applies exactly this pipeline locally in your browser, and the exact-layout mode keeps the page appearance as high-resolution images when the design matters more than the text.
Tips for Arabic PDFs
- If the PDF is text-based, prefer the editable mode — the repaired text is searchable and reusable.
- Numbers and Latin brand names inside Arabic lines stay left-to-right; a good converter keeps their internal order while fixing the surrounding Arabic.
- Heavily diacritized texts (Quran, poetry) are the hardest: check a sample page before converting a large document.
- Scanned Arabic PDFs contain images of text — use the exact-layout mode, or run dedicated OCR software.
Once repaired, Arabic text in the exported Word file is fully editable: you can reflow it, change the font to any Arabic typeface installed on your device, and search inside it like any other document.
Last updated: 2026-10-05