Data Science Roundup

  • In “OpenITI MAKHZAN: An Open Annotated Dataset of Arabic, Persian, Ottoman Turkish, and Urdu Print and Manuscript Data” (Journal of Open Humanities Data), Jonathan Parkes Allen (University of Maryland) and others introduce OpenITI MAKHZAN, a “large aggregation of Arabic-script ground truth and evaluation data drawn from a wide variety of Persian, Arabic, Ottoman Turkish, and Urdu scribal print and handwritten (manuscript) documents.” Their article “explains the different types of data in this large dataset and how this data was compiled and verified and suggests potential use cases for it, such as the training and evaluation of new print and handwritten transcription models.”

Leave a Reply