qalam
Arabic PDF text extraction that doesn't quietly corrupt your corpus: logical word order, base letters not presentation forms, ligatures intact and a per-page verdict telling you which pages genuinely need OCR. Rust + Python, no OCR.
File Explorer
- release.yml
- extract.rs
- main.rs
- Cargo.toml
- arabic.rs
- bidi.rs
- blocks.rs
- cff.rs
- content.rs
- detect.rs
- document.rs
- encoding.rs
- error.rs
- font.rs
- graphics.rs
- html.rs
- images.rs
- layout.rs
- lib.rs
- parser.rs
- structure.rs
- tables.rs
- types.rs
- golden.rs
- tagged.rs
- Cargo.toml
- lib.rs
- Cargo.toml
- pyproject.toml
- README.md
- style.scss
- columns-xy-cut.png.jpg
- hero-side-by-side.png
- inspect-verdict.png
- table-to-csv.png
- _config.yml
- index.md
- compare.py
- .gitignore
- Cargo.lock
- Cargo.toml
- CONTRIBUTING.md
- LICENSE
- PLAN.md
- README.md
- rust-toolchain.toml
# Use via CDN
jsDelivrjsDelivr serves any public GitHub repository as a CDN with zero setup. Pick a version and a file to get a ready-to-paste link and snippet.
Link
Example
// repository documentation
Was this content helpful?
(0 ratings)
