← All work Next case study: Bidaya Marketing Communications →
Documents & NLP
Abjjad
10× editorial productivity for an EBRD Star Venture publisher.
Lead NLP consultantVisit website ↗
The challenge
Raw source files (HTML, CSS, Adobe IDML, OCR output) were often corrupt, and every book needed manual repair, chaptering and tagging.
What we built
- Automated ingestion to parse and repair source files, clean OCR, insert chapter breaks and auto-tag rich metadata.
- AI-assisted review tools replaced most manual QA.
Results
10× editorial productivity
5,000+ new e-books released
Mobile app refresh doubled organic downloads
200,000+ new readers during COVID-19 lockdown without paid advertising
Turnaround from days to minutes
Stack
PythonFastAPIPostgreSQLReactNLPOCR post-processingHTML/CSS/IDML parsing