Digital Humanities Projects
-
al-Kashshāf (upcoming)
Al-Kashshāf is an open-source web and desktop application for advanced search of pre-modern Arabic texts. It searches across a meta-corpus of 7,199 Arabic texts (up to 1348 AH/1930 CE), comprising over 5.7 million pages and nearly 1 billion tokens, drawn from al-Maktaba al-Shamela, OpenITI, KITAB, and nuṣūṣ. Its search capabilities include morphological analysis, root, lemma, and surface queries, Boolean and proximity search, and specialized name search. The primary goal of the project is to increase access to digitized texts that are not represented in the major searchable corpora or are not available to the non-technical user. Initial funding, under the project's earlier name mutūn, was provided by NYU's Faculty DH Seed Grant program. Al-Kashshāf also incorporates the name search functionality previously found on tabādīl. As of fall 2026, al-Kashshāf is undergoing private testing, with a planned release in late winter 2026 or early 2027.
-
tabaṣṣur
Tabaṣṣur is a simple web application for viewing short manuscripts or manuscript excerpts that I've transcribed. All texts here are unpublished unless otherwise noted. I hope it will serve not only to bring attention to unpublished works, but also as an aid to those studying Arabic paleography.
-
tabādīl
Tabādīl is a search tool that generates and searches all possible permutations of a name from a given kunya + nasab + nisba combination. It shares its backend with al-Kashshāf (above) and searches the same meta-corpus of 7,199 Arabic texts and nearly 1 billion tokens. The app is designed to assist with prosopography and general biographical research. NB: the functionality of tabādīl has been incorporated into al-Kashshāf.
-
nuṣūṣ
Nuṣūṣ is a corpus of digitized Arabic texts designed to fill gaps in existing digital corpora. Originally a collection of early Sufi and Sufi-adjacent texts, nuṣūṣ has since expanded to include early works on kalām, falsafa, and Christian theology. Through the website, users can browse text metadata, including author biographies; read the works online; and, most importantly, search the contents of the corpus. The digitized texts can also be downloaded from nuṣūṣ for computational textual analysis or other purposes. All of the project's data is available on GitHub.
-
ishtiqāq
Many searchable Arabic dictionaries are already available online, and providing another is not the focus of this tool, even though ishtiqāq shares much of that functionality. At its core, ishtiqāq is an aid for reading manuscripts when you come across a rasm (consonantal skeleton) that is unclear or whose letters are ambiguous. With ishtiqāq, you can mark the unclear letter(s) and, with a single search, retrieve results for every possible spelling of the word. To learn more about its functionality, see the How To page.
This tool was inspired by LexiQamus and uses indices from ejtaal.net, a flat-file database of Arabic words and their roots compiled by Abdalaziz Alsaydi, and a noun list made by Taha Zerrouki. English dictionary data was scraped from an OCR'd version of Hans Wehr and from the Living Arabic Project, and the classical lexica were scraped from al-Maktaba al-Shamela and lesanarab.com.
-
sham-scrap
Sham-scrap is a metadata browser for the texts of al-Maktaba al-Shamela, scraped on July 2, 2024. The texts themselves can also be browsed on the Shamela website and through its application. I am in the process of cleaning the scraped texts and making them available to all. Individual texts can currently be downloaded, and the entire corpus will soon be available on GitHub. While many of these texts have already been scraped and made available through the Open Islamicate Texts Initiative and the KITAB project's meta-corpus, a significant portion has not. Furthermore, the texts in the OpenITI/KITAB corpus are encoded in its custom mARkdown schema. While this has its benefits for token- and phrase-level tagging, it is somewhat cumbersome, and many edge cases make cleaning the texts difficult. The metadata for the OpenITI/KITAB files is also not standardized, with varying field names, missing items, and so on, so I have standardized both the metadata and the texts in .json format.