arabic-img2md
Arabic Img2MD
Introduced 2024-11-19
Click to add a brief description of the dataset (Markdown and LaTe# Arabic Img2MD
Dataset Summary
The arabic-img2md dataset consists of 15,000 examples of PDF pages paired with their Markdown counterparts. The dataset is split into:
- Train: 13,700 examples
- Test: 1,520 examples
This dataset was created as part of the open-source research project Arabic Nougat to enable OCR and Markdown extraction from Arabic documents. It contains mostly Arabic text but also includes examples with English text. X enabled).
Provide:
- a high-level explanation of the dataset characteristics
- explain motivations and summary of its content
- potential use cases of the dataset