Arabic PDF text extraction that doesn't quietly corrupt your corpus: logical word order, base letters not presentation forms, ligatures intact and a per-page verdict telling you which pages genuinely need OCR. Rust + Python, no OCR.

58 stars 5 forks 58 watchers Rust GNU General Public License v3.0
1 Open Issue Need Help Last updated: Sep 18, 2026

Open Issues Need Help

View All on GitHub
bug good first issue approved needs-triage

Arabic PDF text extraction that doesn't quietly corrupt your corpus: logical word order, base letters not presentation forms, ligatures intact and a per-page verdict telling you which pages genuinely need OCR. Rust + Python, no OCR.

Rust