arabic-img2md

Arabic Img2MD

Introduced 2024-11-19

Click to add a brief description of the dataset (Markdown and LaTe# Arabic Img2MD

Dataset Summary

The arabic-img2md dataset consists of 15,000 examples of PDF pages paired with their Markdown counterparts. The dataset is split into:

  • Train: 13,700 examples
  • Test: 1,520 examples

This dataset was created as part of the open-source research project Arabic Nougat to enable OCR and Markdown extraction from Arabic documents. It contains mostly Arabic text but also includes examples with English text. X enabled).

Provide:

  • a high-level explanation of the dataset characteristics
  • explain motivations and summary of its content
  • potential use cases of the dataset