Adalwise
“Rescuing classical Islamic discourse and Iqbalian philosophy from the entropy of social media feeds through strict bibliographic schema pipelines and client-side bilingual search.”
A bilingual scholarly archive and study platform engineered with Next.js 16, sub-millisecond in-memory MiniSearch, and automated YouTube ingestion pipelines.
- Role
- Lead Platform Architect & Engineer
- Context
- 8 Weeks (Active Production)
- Team
- 2 Members (Co-founded with Dr. Hafiz Haseeb)
- Core Stack
- Next.js 16, React 19, TypeScript, MiniSearch, Node.js Automation Scripts, Tailwind CSS

Fig 1.0 — Architecture execution snapshot (Adalwise)
The Friction
Why engineer a custom digital platform instead of relying on YouTube playlists or Substack?
Over hundreds of hours of recorded seminars on Lisan ul Quran (classical Arabic grammar), the philosophical reconstruction of Allama Iqbal, and domestic political critiques (Twasi al-Haq), serious scholarship was getting buried under YouTube's algorithmic churn, opaque search ranking, and fragmented WhatsApp study circles.
Generic publishing tools (WordPress, Ghost, Substack) fail completely when handling bilingual Perso-Arabic and Latin scholarship. They lack support for diacritic-insensitive search (where Arabic A'raab and Tashkeel break basic string queries), Quranic notation resolvers (e.g., mapping '2:255' to Surah Al-Baqarah and Ayat ul Kursi), and structured curriculum hierarchies for multi-part lecture series.
We engineered Adalwise to treat spoken and written scholarship with bibliographic discipline: an automated headless ingestion engine that syncs with external video feeds, coupled with a zero-latency client-side inverted index and typography harmonized across Nastaliq, classical Arabic, and Latin serifs.
Deliberate Constraints
The system architecture was not chosen in an unconstrained vacuum. Each structural decision emerged directly from four non-negotiable technical boundaries.
Urdu and Arabic search queries break when user input lacks short vowels (A'raab / Tashkeel) or uses different unicode glyph variants (Alef Maksura vs. Yeh, Heh Goal vs. Heh Do-Chashmi).
Built a custom regex-based normalizer and tokenizer that strips Tashkeel, normalizes orthographic letter variants, cleans zero-width joiners, and tokenizes across both Latin and Perso-Arabic punctuation marks.
Relying on external hosted search engines (Algolia, Meilisearch) introduces network latency, subscription costs, and vendor lock-in.
Implemented an in-memory client-side MiniSearch inverted index with custom term weighting, prefix matching, and fuzzy search that indexes thousands of lectures, notes, and articles in under 5ms directly in the browser.
Manual data entry for 200+ lectures with video durations, thumbnails, and descriptions is unsustainable for a small scholarly team.
Engineered automated Node.js ingestion scripts (sync-youtube.mjs, audit-catalogue.mjs, fetch-video-stats.mjs) that poll the YouTube Data API, validate category schemas, audit missing metadata, and commit normalized JSON files.
Pairing Urdu Nastaliq (Noto Nastaliq Urdu), classical Quranic Arabic (Amiri), and English literary serif (EB Garamond) causes jarring baseline shifts and optical size mismatch.
Fine-tuned font metric overrides, optical letter-spacing, and line-height multipliers across responsive breakpoints to create a tranquil, book-like reading environment.
System Architecture & Data Pipeline
A decoupled, static-first architecture. Offline Node.js pipelines harvest, normalize, and audit external lecture catalogs into typed JSON files. At runtime, Next.js 16 App Router renders static course matrix views while an in-memory MiniSearch provider delivers sub-millisecond search across English, Urdu, and Quranic citations.
Alto, Cultus, Corolla. Standard tiered rental base.
Audi A6, BMW 7, Land Cruiser. Chauffeur insurance rate.
Sportage, Tucson, Fortuner. All-terrain security deposit.
Bolan, Hiace, Coaster. High-capacity commercial rate.
Dynamic Polymorphism at Runtime: The orchestrator holds a single container std::vector<Vehicle*> fleet. When executing reservations or computing quotes, method calls to v->calculateCost(days) dynamically dispatch to the concrete subclass implementation through each instance's vtable pointer.
Subsystem Decomposition
Headless Synchronization & Audit Pipeline
scripts/sync-youtube.mjs & audit-catalogue.mjsExtracts video IDs, durations, view telemetry, and descriptions; validates metadata against strict schema contracts.
Bilingual Search & Normalization Engine
src/lib/search/minisearch-provider.tsBuilds and caches an in-memory inverted index supporting fuzzy search, category filters, and cross-lingual synonym expansion.
Curriculum & Matrix Navigator
src/features/lectures/ & majlis/Structures open-access course pathways: Lisan ul Quran grammar tracks, Surah matrix navigators, and Majlis chronological archives.
Scholarly Fellowship Intake
src/features/fellowship/components/IntakeForm.tsxManages student applications and community access for offline Majlis gatherings and intensive study cohorts.
The Hard Part: Bilingual Perso-Arabic Orthography & Sub-Millisecond Search Normalization
How invisible zero-width joiners, letter variants, and diacritics quietly broke string matching across scripts.
In testing search for lectures titled in both English and Urdu (e.g., 'سورۃ کہف کی تفہیم' vs 'Surah Al-Kahf'), searches for 'کہف' failed if the query contained an Alef with Madd (آ) or a different Yeh glyph (ی vs. ي). Even worse, users searching by Quranic reference like '2:255' or 'Para 30' received zero results because raw text indices don't understand scriptural notation.
Urdu and Arabic fonts utilize multiple unicode points for visually identical letters (e.g., U+064A Arabic Yeh vs U+06CC Farsi/Urdu Yeh). Furthermore, vowel diacritics (A'raab: Fatha, Damma, Kasra) alter raw byte representations. A standard substring search (.includes()) or generic Latin tokenizer fails completely on these character boundaries.
const ARABIC_DIACRITICS_REGEX = /[\u064B-\u065F\u0670\u06D6-\u06ED]/g;
const ZERO_WIDTH_REGEX = /[\u200B-\u200D\uFEFF]/g;
const ALEF_VARIANTS_REGEX = /[\u0622\u0623\u0625\u0671]/g;
const HEH_VARIANTS_REGEX = /[\u0629\u06C2\u06C3\u06BE]/g;
const YEH_VARIANTS_REGEX = /[\u064A\u0649\u0626]/g;
const KAF_VARIANT_REGEX = /\u0643/g;
export function normalizeUrduArabic(text: string | undefined | null): string {
if (!text) return "";
return text
.replace(ARABIC_DIACRITICS_REGEX, "") // Strip short vowels/A'raab
.replace(ALEF_VARIANTS_REGEX, "\u0627") // Normalize all Alef forms to ا
.replace(HEH_VARIANTS_REGEX, "\u06C1") // Normalize Heh variants to ہ
.replace(YEH_VARIANTS_REGEX, "\u06CC") // Normalize Yeh forms to ی
.replace(KAF_VARIANT_REGEX, "\u06A9") // Normalize Arabic Kaf to Urdu ک
.replace(ZERO_WIDTH_REGEX, "") // Strip invisible joiners
.toLowerCase()
.trim();
}
export function tokenizeBilingual(text: string | undefined | null): string[] {
if (!text) return [];
return text
.split(/[\s,./\\;:'"[\]{}|!@#$%^&*()_+=\-–—؟،۔«»‹›"“”'‘’`~]+/u)
.map((token) => token.trim())
.filter((token) => token.length > 0);
}
Fig 2.0 — Live MiniSearch query ('musa') resolving bilingual English/Urdu titles and Arabic Surah tags in sub-5ms.
We authored a specialized normalization pipeline coupled with an O(1) Quranic notation mapper (mapping chapter numbers, Latin transliterations, and traditional Arabic names to common root keys). When a user searches '2:255' or 'Baqarah', the query expander augments the search tokens automatically.
Building search for non-Latin languages teaches you that text is never just characters—it is human culture codified into unicode. You cannot import a Western library and expect it to respect the orthographic reality of Urdu or Arabic.
Catalog Audit & Ingestion Test Suite
Running the automated catalogue audit script to verify 100% metadata compliance across lecture series and schema integrity.
System Interface & Pedagogical Architecture
High-resolution captures of the live production interface across core learning tracks.



Engineering Reflection
“The modern internet is optimized for immediacy and dopamine; classical thought demands stillness, patience, and structure.”
When building software for classical study, the hardest challenge is not the code—it is building an interface that encourages deep contemplation rather than frantic skimming. If your platform looks or behaves like an engagement-driven social network, you have failed the content before the reader has even finished the first paragraph.
By deliberately removing algorithmic recommendations, clickbait thumbnails, and distracting widgets, we built a digital sanctuary. The user is greeted with calm typography, structured curricula, and a search tool that works silently and instantly.
Adalwise taught me that architecture is an act of stewardship: preserving serious ideas with the technical craftsmanship they deserve.
Interested in discussing this architecture?
I'm always open to technical dialogue, code reviews, and exploring system constraints.