<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0"><channel><title><![CDATA[Kirnu Dev Notes]]></title><description><![CDATA[Kirnu Dev Notes]]></description><link>https://kirnu.hashnode.dev</link><image><url>https://cdn.hashnode.com/uploads/logos/6ac77090181984e5c1889abd/8a1d4737-d41a-4302-96cc-ae56f5be565f.png</url><title>Kirnu Dev Notes</title><link>https://kirnu.hashnode.dev</link></image><generator>RSS for Node</generator><lastBuildDate>Sat, 10 Oct 2026 16:07:41 GMT</lastBuildDate><atom:link href="https://kirnu.hashnode.dev/rss.xml" rel="self" type="application/rss+xml"/><language><![CDATA[en]]></language><ttl>60</ttl><item><title><![CDATA[Searching Arabic text in JavaScript: why "احمد" doesn't find "أحمد" (and how to fix it)]]></title><description><![CDATA[A user types «احمد» (Ahmad) into your search box and gets nothing, even though the page says «أحمد». Or they search for «مدرسه» (school) and the text has «مدرسة». Or they search «محمد» and the text is]]></description><link>https://kirnu.hashnode.dev/searching-arabic-text-in-javascript-why-doesn-t-find-and-how-to-fix-it</link><guid isPermaLink="true">https://kirnu.hashnode.dev/searching-arabic-text-in-javascript-why-doesn-t-find-and-how-to-fix-it</guid><category><![CDATA[JavaScript]]></category><category><![CDATA[unicode]]></category><category><![CDATA[Arabic ]]></category><category><![CDATA[i18n]]></category><dc:creator><![CDATA[Kirnu (كرنو)]]></dc:creator><pubDate>Fri, 09 Oct 2026 09:00:20 GMT</pubDate><content:encoded><![CDATA[<p>A user types «احمد» (Ahmad) into your search box and gets nothing, even though the page says «أحمد». Or they search for «مدرسه» (school) and the text has «مدرسة». Or they search «محمد» and the text is vowelled, «مُحَمَّد», or stretched with tatweel, «مـحـمـد». To the user it's the same word every time. JavaScript's string matching doesn't treat these spellings as equivalent.</p>
<p>This post covers four traps when searching Arabic text in JavaScript, then builds a search-normalization function, a function that finds match positions so you can highlight them, and finally when you should <em>not</em> normalize.</p>
<p><em>Disclosure: I work on an open-source Arabic text library that has normalization functions; I mention it briefly at the end. Everything before that is plain JavaScript, no libraries.</em></p>
<h2>Trap 1: includes() matches characters exactly</h2>
<pre><code class="language-js">'محمد أحمد'.includes('احمد');  // false
'مُحَمَّد'.includes('محمد');      // false
'مـحـمـد'.includes('محمد');     // false
</code></pre>
<p><code>includes</code>, <code>indexOf</code> and regular expressions do exact matching, with no Arabic-specific normalization. «أ» (alef with hamza, U+0623) is not «ا» (bare alef, U+0627); the vowel marks (harakat) are separate characters between the letters, and so is the tatweel «ـ» (U+0640). People also spell hamza forms, teh marbuta «ة» and alef maksura «ى» inconsistently, so literal search fails a lot.</p>
<h2>Trap 2: two strings that look identical but aren't</h2>
<pre><code class="language-js">const a = 'أحمد';
const b = 'ا\u0654حمد';   // alef + separate hamza above (U+0654)
a === b;                       // false
a === b.normalize('NFC');      // true
</code></pre>
<p>«أ» can arrive as one code point, or as two: a bare alef followed by a combining hamza. They render the same, but they don't compare equal. This shows up in text copied from some programs and in file names from some systems; <code>normalize('NFC')</code> unifies these canonically equivalent forms.</p>
<p>There's a second case NFC doesn't cover. Text copied from some PDFs arrives in Arabic "presentation forms": dedicated code points for certain letter shapes and ligatures.</p>
<pre><code class="language-js">const fromPdf = '\uFEE3\uFEA4\uFEE4\uFEAA';   // «محمد» in presentation forms
fromPdf === 'محمد';                   // false
fromPdf.normalize('NFC') === 'محمد';  // false
fromPdf.normalize('NFKC') === 'محمد'; // true
</code></pre>
<p><code>NFKC</code> maps them back to the base letters, but it also changes other characters (ﷺ, for example, expands into a whole phrase), so we'll use it only inside the search key.</p>
<h2>Trap 3: Intl.Collator compares, it doesn't search</h2>
<p><code>Intl.Collator</code> looks like the answer, since it compares strings using language rules. On Node 24.21 / ICU 78.3:</p>
<pre><code class="language-js">const collator = new Intl.Collator('ar', { sensitivity: 'base' });
collator.compare('أحمد', 'احمد');     // 0
collator.compare('مُحَمَّد', 'محمد');   // 0
collator.compare('مـحـمـد', 'محمد');  // 0
collator.compare('على', 'علي');       // 0
collator.compare('مدرسة', 'مدرسه');   // -1
</code></pre>
<p><code>0</code> means equal. It's great for sorting and for comparing two whole words, but it has three problems for search:</p>
<ul>
<li>It compares a whole string to a whole string. It doesn't find a word inside a text, and JavaScript has no search API built on it.</li>
<li>Its rules aren't your rules: here it treats «على» (the preposition "on") and «علي» (the name Ali) as equal, yet keeps «مدرسة» and «مدرسه» apart.</li>
<li>Results depend on the ICU version in the browser or server, so they can differ between environments.</li>
</ul>
<h2>Trap 4: characters you can't see</h2>
<p>Text copied from web pages or chat apps can carry invisible characters inside words, such as direction marks (U+200F, U+061C) or a zero-width space (U+200B):</p>
<pre><code class="language-js">'محمد\u200F'.includes('محمد');  // true
'مح\u200Bمد'.includes('محمد');  // false
</code></pre>
<p>The first works because the mark sits at the end of the word; the second fails because the invisible character is in the middle. The user can't see it and won't understand why the result is missing.</p>
<h2>The fix: a normalized search key</h2>
<p>Don't search the text as-is. Compute a normalized "search key" for each text, compute the same key for the query, and compare keys. Keep the original text for display.</p>
<p>The rules below are a search policy that suits most general Arabic text, not an exhaustive list; adjust them for your app:</p>
<pre><code class="language-js">// Removed from the key: Arabic combining marks (including the separate hamza U+0654/U+0655), tatweel,
// and selected invisible and direction-control characters
const DROP = /[\u064B-\u065F\u0670\u0640\u200B-\u200F\u061C\u202A-\u202E\u2066-\u2069]/;
// Unified: alef forms, alef maksura, teh marbuta, the Persian keheh (ک) and yeh (ی)
const MAP = { 'أ': 'ا', 'إ': 'ا', 'آ': 'ا', 'ٱ': 'ا', 'ى': 'ي', 'ة': 'ه', 'ک': 'ك', 'ی': 'ي' };

// Key for one character: NFKC maps presentation forms to letters, then drop and unify
function keyOf(ch) {
  let out = '';
  for (const c of ch.normalize('NFKC')) {
    if (!DROP.test(c)) out += MAP[c] ?? c;
  }
  return out;
}

function searchKey(text) {
  let key = '';
  for (const ch of text.normalize('NFC')) key += keyOf(ch);
  return key;
}

searchKey('أحمد');           // 'احمد'
searchKey('مُحَمَّد');         // 'محمد'
searchKey('مـحـمـد');        // 'محمد'
searchKey('مدرسة');          // 'مدرسه'
searchKey('مستشفى');         // 'مستشفي'
searchKey('ا\u0654حمد');     // 'احمد'
searchKey('مح\u200Bمد');     // 'محمد'
searchKey(fromPdf);          // 'محمد'

searchKey('ذهب محمد أحمد إلى المدرسة').includes(searchKey('احمد'));  // true
</code></pre>
<p>Notes:</p>
<ul>
<li>Apply it to both sides, the text and the query. Normalize only one and they won't match.</li>
<li>Don't store it instead of the original: it loses information. «مدرسه» in the key could be «مدرسة» or «مدرسه» in the source.</li>
<li>It leaves «ؤ», «ئ» and the standalone hamza «ء» alone, because stripping the hamza from them changes the word more than it helps search. If you need that, add them to <code>MAP</code> deliberately.</li>
<li>This is substring search, so «علي» also matches inside «عليكم». For whole words, tokenize the text with a word-boundary strategy that suits your app, or check the boundaries around each match.</li>
<li><code>NFKC</code> inside the key also folds other characters, such as fullwidth digits and Latin letters. We apply it per character to keep offsets; that's enough for Arabic presentation forms, but it isn't exactly the same as running NFKC on the whole string.</li>
</ul>
<h2>Highlighting matches</h2>
<p>The key differs from the text in length and positions (harakat and tatweel are gone, and NFKC can make it longer), so a match position in the key isn't its position in the text. The fix: while building the key, record the start and end of the character each key unit came from. We walk the text by code point with <code>for...of</code>, and record offsets in UTF-16 code units, because that's what <code>slice</code> and <code>indexOf</code> use:</p>
<pre><code class="language-js">// Returns match ranges [start, end) in the NFC text, including consecutive dropped characters (e.g. harakat) after the last match
function findMatches(text, query) {
  text = text.normalize('NFC');
  let key = '';
  const starts = [];
  const ends = [];
  let i = 0;
  for (const ch of text) {
    const k = keyOf(ch);
    key += k;
    for (let u = 0; u &lt; k.length; u++) {
      starts.push(i);
      ends.push(i + ch.length);
    }
    i += ch.length;
  }
  const q = searchKey(query);
  const matches = [];
  if (!q) return matches;
  for (let at = key.indexOf(q); at !== -1; at = key.indexOf(q, at + q.length)) {
    const start = starts[at];
    let end = ends[at + q.length - 1];
    while (end &lt; text.length &amp;&amp; DROP.test(text[end])) end++; // keep a haraka with its letter
    const prev = matches[matches.length - 1];
    if (prev &amp;&amp; start &lt; prev[1]) prev[1] = Math.max(prev[1], end); // two matches in one character (like ﷺ): merge them
    else matches.push([start, end]);
  }
  return matches;
}

const escapeHtml = (s) =&gt; s.replace(/[&amp;&lt;&gt;"']/g, (c) =&gt; `&amp;#${c.charCodeAt(0)};`);

function highlight(text, query) {
  text = text.normalize('NFC');
  let html = '';
  let last = 0;
  for (const [start, end] of findMatches(text, query)) {
    html += escapeHtml(text.slice(last, start)) + '&lt;mark&gt;' + escapeHtml(text.slice(start, end)) + '&lt;/mark&gt;';
    last = end;
  }
  return html + escapeHtml(text.slice(last));
}

highlight('قال مُحَمَّد: أهلاً يا محمد', 'محمد');  // 'قال &lt;mark&gt;مُحَمَّد&lt;/mark&gt;: أهلاً يا &lt;mark&gt;محمد&lt;/mark&gt;'
highlight('ذهبت إلى المـدرسـة', 'مدرسه');       // 'ذهبت إلى ال&lt;mark&gt;مـدرسـة&lt;/mark&gt;'
highlight('😀 مُحَمَّد', 'محمد');                 // '😀 &lt;mark&gt;مُحَمَّد&lt;/mark&gt;'
</code></pre>
<p>The output shows the NFC text with its harakat and tatweel, even though the search ignored them (NFC doesn't change how the text looks). And the function escapes HTML before inserting <code>&lt;mark&gt;</code>, so user text can't inject markup here. This escaping is enough for text inside an element, as in this example, not for attributes or URLs.</p>
<h2>In the database</h2>
<p>Don't compute the key for every row on every search. Store it in its own column when you save, and search that:</p>
<ul>
<li>A <code>name</code> column for display and a <code>name_search</code> column = <code>searchKey(name)</code>, updated whenever the name changes.</li>
<li>Normalize the query with the same function before querying, on the server, not only in the browser. Pass it as a parameter in a prepared statement, and with <code>LIKE</code>, escape <code>%</code> and <code>_</code> in the query so they're treated as literal characters.</li>
<li>An ordinary index helps exact matches and prefix searches, but usually not <code>LIKE '%term%'</code>. For word search in long text, consider a full-text index; arbitrary substring search usually needs an n-gram or trigram index, depending on your database.</li>
<li>If you change the <code>searchKey</code> rules later, recompute the column for every row, or the keys won't match.</li>
</ul>
<h2>When not to normalize</h2>
<p>Normalization widens results, and some of them aren't what the user wants:</p>
<ul>
<li>«على» (preposition) and «علي» (the name Ali) get the same key.</li>
<li>«حماة» (the city of Hama) and «حماه» get the same key.</li>
<li>In Quranic text or vowelled poetry, the reader may be searching for the vowels themselves.</li>
</ul>
<p>So:</p>
<ul>
<li>Use normalization for search only; keep the original for display and storage.</li>
<li>Rank results: exact matches first, then normalized matches. If the user searched «على», show «على» before «علي».</li>
<li>Offer an exact-search option where vowels or hamza forms carry meaning.</li>
</ul>
<h2>Summary</h2>
<ul>
<li><code>includes</code> matches characters exactly, and <code>Intl.Collator</code> compares whole strings rather than searching inside them.</li>
<li>Unify forms with NFC, and presentation forms with NFKC inside the search key, then apply your policy: drop marks, tatweel and selected invisible characters; unify alef forms, teh marbuta and alef maksura.</li>
<li>Apply the key to both the text and the query, and keep offsets so you can highlight matches.</li>
<li>Store the key in its own column; never replace the original with it.</li>
</ul>
<p>Disclosure: the open-source <a href="https://github.com/getkirnu/arabic-core">@kirnu/arabic-core</a> library has <code>normalizeArabic</code> (alef forms, yeh, teh marbuta, the Persian ک and ی, and tatweel, each rule a separate option) and <code>removeTashkeel</code> (vowel marks), with tests; their rules are close to this article's policy but not identical. If you want to try normalization on a text without writing code, there's a free <a href="https://getkirnu.com/ar/text/arabic-normalizer/?utm_source=hashnode">Arabic text normalizer</a> (the page is in Arabic).</p>
<p>What rules do you use for Arabic search in your apps? Do you unify teh marbuta and alef maksura, or leave them?</p>
]]></content:encoded></item><item><title><![CDATA[Arabic digits in web forms - why a valid phone number gets rejected, and how to fix it in JavaScript]]></title><description><![CDATA[A user in Riyadh types their mobile number with Arabic digits: ٠٥٠١٢٣٤٥٦٧. Your form rejects it, even though it's a perfectly valid number. Or worse, it accepts it and stores it as typed, so your data]]></description><link>https://kirnu.hashnode.dev/arabic-digits-in-web-forms-why-a-valid-phone-number-gets-rejected-and-how-to-fix-it-in-javascript</link><guid isPermaLink="true">https://kirnu.hashnode.dev/arabic-digits-in-web-forms-why-a-valid-phone-number-gets-rejected-and-how-to-fix-it-in-javascript</guid><category><![CDATA[JavaScript]]></category><category><![CDATA[Web Development]]></category><category><![CDATA[i18n]]></category><category><![CDATA[Regex]]></category><dc:creator><![CDATA[Kirnu (كرنو)]]></dc:creator><pubDate>Thu, 08 Oct 2026 14:05:58 GMT</pubDate><content:encoded><![CDATA[<p>A user in Riyadh types their mobile number with Arabic digits: <strong>٠٥٠١٢٣٤٥٦٧</strong>. Your form rejects it, even though it's a perfectly valid number. Or worse, it accepts it and stores it as typed, so your database now holds two "different" numbers for the same person: <code>0501234567</code> and <code>٠٥٠١٢٣٤٥٦٧</code>. Arabic-Indic digits represent the same decimal values as 0–9, but they are different Unicode characters.</p>
<p>This post covers where the problem comes from, four traps in JavaScript, and then how to normalize, validate and store one canonical form.</p>
<p><em>Disclosure: I maintain an open-source Arabic text library that has a digit-conversion function; it's mentioned briefly at the end. Every example before that is plain JavaScript.</em></p>
<h2>Three sets of digits</h2>
<ul>
<li><strong>Western digits</strong> 0–9: U+0030 to U+0039.</li>
<li><strong>Arabic-Indic digits</strong> ٠–٩: U+0660 to U+0669. They may be entered by Arabic keyboard layouts or Arabic locale settings.</li>
<li><strong>Extended Arabic-Indic (Persian/Urdu) digits</strong> ۰–۹: U+06F0 to U+06F9. Most look like the Arabic ones, but some are shaped differently, such as ۴, ۵ and ۶ (and ۷ in Urdu), and all ten are separate code points.</li>
</ul>
<p>Number formatting also has its own marks: the Arabic decimal separator ٫ (U+066B), the Arabic thousands separator ٬ (U+066C) and the Arabic percent sign ٪ (U+066A).</p>
<h2>Trap 1: <code>\d</code> doesn't match Arabic digits</h2>
<pre><code class="language-js">/^\d+$/.test('٠٥٠١٢٣٤٥٦٧');  // false
/^\d+$/u.test('٠٥٠١٢٣٤٥٦٧'); // false
</code></pre>
<p>In JavaScript, <code>\d</code> means 0–9 only, even with the <code>u</code> flag. Any validation built on <code>\d</code> rejects a number typed with Arabic digits.</p>
<h2>Trap 2: <code>Number()</code> and <code>parseInt()</code> return NaN</h2>
<pre><code class="language-js">Number('١٢٣');      // NaN
parseInt('١٢٣');    // NaN
parseFloat('١٢٫٥'); // NaN
Number('12٣');      // NaN
</code></pre>
<p>JavaScript doesn't convert Arabic-Indic digits to numbers, and a single Arabic digit inside a Western number is enough to make the whole conversion fail.</p>
<h2>Trap 3: <code>\p{Nd}</code> accepts more than you want</h2>
<p><code>\p{Nd}</code> (any Unicode decimal digit) looks like the fix:</p>
<pre><code class="language-js">/^\p{Nd}+$/u.test('٠٥٠١٢٣٤٥٦٧'); // true
/^\p{Nd}+$/u.test('৫৫৫');        // true (Bengali digits)
</code></pre>
<p>It does accept Arabic digits, but also the digits of many other scripts: 770 characters in 77 digit sets in Node 24. And it converts nothing: <code>Number()</code> still returns NaN afterwards. Validation alone isn't enough: <strong>normalize first</strong>, then validate.</p>
<h2>Trap 4: invisible direction marks</h2>
<p>Copy a number from an Arabic page, or from <code>Intl</code> output, and invisible bidi control characters can come along. On Node 24.21 / ICU 78.3:</p>
<pre><code class="language-js">const date = new Intl.DateTimeFormat('ar-SA', { timeZone: 'UTC' }).format(new Date(Date.UTC(2024, 2, 11)));
date;                    // '١١‏/٣‏/٢٠٢٤'
date.includes('‏'); // true: a RIGHT-TO-LEFT MARK (RLM) after the day and the month

const percent = new Intl.NumberFormat('ar-EG', { style: 'percent' }).format(0.25);
percent.includes('؜'); // true: an ARABIC LETTER MARK (ALM) after the percent sign
</code></pre>
<p>These can break a strict <code>^...$</code> check that expects digits only, even though the text looks fine on screen. In numeric fields you can strip them before validating, or reject them explicitly, depending on your input policy. Don't strip them from general Arabic text: there they control the order of mixed Arabic/Latin words.</p>
<h2>The fix: normalize, validate, then store one form</h2>
<p>Normalization and validation are different steps: normalization changes the representation; validation decides whether the resulting value is allowed in your application.</p>
<p>Step one is a function that only normalizes: it converts Arabic-Indic and Persian digits to 0–9 and strips bidi marks (for numeric fields only):</p>
<pre><code class="language-js">// For numeric fields: strips bidi marks and converts ٠–٩ and ۰–۹ to 0–9; touches nothing else.
function normalizeDigits(input) {
  return input
    .replace(/[‎‏؜‪-‮⁦-⁩]/g, '') // bidi marks
    .replace(/[٠-٩]/g, (d) =&gt; String(d.charCodeAt(0) - 0x0660))     // ٠–٩ → 0–9
    .replace(/[۰-۹]/g, (d) =&gt; String(d.charCodeAt(0) - 0x06f0));    // ۰–۹ → 0–9
}

normalizeDigits('٠٥٠١٢٣٤٥٦٧'); // '0501234567'
normalizeDigits('٠50١٢٣٤٥٦٧'); // '0501234567' (mixed digits)
normalizeDigits('۰۵۰');        // '050'
</code></pre>
<p>Don't drop separators without validating, or ١٬٢ silently becomes 12. For amounts and decimals, check the format first:</p>
<pre><code class="language-js">// An amount like ١٬٢٥٠٫٥ or 1250.5: thousands separators only in valid positions; null otherwise.
function parseAmount(raw) {
  const s = normalizeDigits(raw).replace(/٫/g, '.').replace(/٬/g, ',');
  if (!/^(?:\d{1,3}(?:,\d{3})+|\d+)(?:\.\d+)?$/.test(s)) return null;
  return Number(s.replace(/,/g, ''));
}

parseAmount('١٬٢٥٠٫٥'); // 1250.5
parseAmount('١٢٫٥');    // 12.5
parseAmount('١٬٢');     // null
</code></pre>
<p>For phone numbers, normalizing digits isn't enough: <code>0501234567</code>, <code>+966501234567</code> and <code>00966501234567</code> are one number in three formats. Convert to one international form (E.164) and store that:</p>
<pre><code class="language-js">// Saudi mobile in local or international form → '+9665XXXXXXXX', or null if the format doesn't match.
// Spaces and hyphens are allowed because they're only formatting.
function toSaudiE164(raw) {
  const s = normalizeDigits(raw).replace(/[\s-]/g, '');
  if (/^05\d{8}$/.test(s)) return '+966' + s.slice(1);
  if (/^(?:\+|00)9665\d{8}$/.test(s)) return '+966' + s.replace(/^(?:\+|00)966/, '');
  return null;
}

toSaudiE164('٠٥٠١٢٣٤٥٦٧');        // '+966501234567'
toSaudiE164('+٩٦٦ ٥٠ ١٢٣ ٤٥٦٧');   // '+966501234567'
toSaudiE164('00966501234567');    // '+966501234567'
toSaudiE164('0501234567‏');  // '+966501234567'
toSaudiE164('٠٥٠١٢٣٤٥');          // null
</code></pre>
<p>This is a simplified example that checks the format only, not whether the number is assigned or active.</p>
<p>Three practical rules:</p>
<ul>
<li><strong>Store the canonical form</strong>, not what the user typed: Western digits for numbers, E.164 for phone numbers, so the same value never exists in two forms and searches don't miss it.</li>
<li><strong>Normalize and validate on the server too.</strong> Browser validation can be bypassed, and requests may come from other clients.</li>
<li><strong>Don't rely on <code>inputmode</code> alone.</strong> <code>inputmode="numeric"</code> asks for a number keypad; it doesn't guarantee which digits you receive.</li>
</ul>
<h2>Displaying numbers: pick the numbering system explicitly</h2>
<p>Display matters too. Don't rely on the default when you want Arabic or Western digits on screen. On Node 24.21 / ICU 78.3, <code>new Intl.NumberFormat(locale).format(1234567.89)</code> gives:</p>
<ul>
<li><code>ar</code>: 1,234,567.89</li>
<li><code>ar-SA</code> and <code>ar-EG</code>: ١٬٢٣٤٬٥٦٧٫٨٩</li>
<li><code>ar-AE</code>: 1,234,567.89</li>
<li><code>ar-MA</code>: 1.234.567,89 (dot for thousands, comma for decimals)</li>
</ul>
<p>Defaults differ by country and can change between ICU versions, so put the numbering system in the locale:</p>
<pre><code class="language-js">new Intl.NumberFormat('ar-SA-u-nu-latn').format(1234567.89); // '1,234,567.89'
new Intl.NumberFormat('ar-u-nu-arab').format(1234567.89);    // '١٬٢٣٤٬٥٦٧٫٨٩'
</code></pre>
<p>Use this for display only: never parse formatted text back into a number; keep the original numeric value.</p>
<h2>Summary</h2>
<ul>
<li><code>\d</code> and <code>Number()</code> don't understand Arabic-Indic digits; <code>\p{Nd}</code> accepts too much and converts nothing.</li>
<li>Normalize digits first, then validate the format, then store one canonical form (E.164 for phones), in the browser and on the server.</li>
<li>Don't drop separators without validating, and strip bidi marks only from numeric fields.</li>
<li>When displaying, set <code>-nu-latn</code> or <code>-nu-arab</code> explicitly.</li>
</ul>
<p>Disclosure: my open-source library <a href="https://github.com/getkirnu/arabic-core">@kirnu/arabic-core</a> has a tested <code>toWesternDigits</code> function that converts Arabic-Indic and Persian digits (and separators between digits). For a non-technical explanation of Arabic vs Western digits, there's a <a href="https://getkirnu.com/ar/guides/arabic-english-numbers/?utm_source=hashnode">short guide on Kirnu</a> (Arabic).</p>
<p>Have you hit this in your apps? Where do you normalize: in the browser, on the server, or in the database?</p>
]]></content:encoded></item><item><title><![CDATA[Hijri dates in JavaScript - what Intl can do, what it can't, and four traps]]></title><description><![CDATA[Saudi Arabia's official Hijri calendar is Umm al-Qura, and dates like "1 Ramadan 1448" show up in contracts, government forms, school calendars and HR systems across the Gulf. JavaScript can format Hi]]></description><link>https://kirnu.hashnode.dev/hijri-dates-in-javascript-what-intl-can-do-what-it-can-t-and-four-traps</link><guid isPermaLink="true">https://kirnu.hashnode.dev/hijri-dates-in-javascript-what-intl-can-do-what-it-can-t-and-four-traps</guid><category><![CDATA[JavaScript]]></category><category><![CDATA[TypeScript]]></category><category><![CDATA[i18n]]></category><category><![CDATA[Web Development]]></category><dc:creator><![CDATA[Kirnu (كرنو)]]></dc:creator><pubDate>Thu, 08 Oct 2026 10:55:52 GMT</pubDate><content:encoded><![CDATA[<p>Saudi Arabia's official Hijri calendar is <strong>Umm al-Qura</strong>, and dates like "1 Ramadan 1448" show up in contracts, government forms, school calendars and HR systems across the Gulf. JavaScript can format Hijri dates without a library thanks to <code>Intl</code>, but it has no direct API for the reverse (Hijri → Gregorian), and there are a few traps that cause classic off-by-one-day bugs.</p>
<p><em>Disclosure: I maintain an open-source Hijri/Arabic library, mentioned at the end. Everything before that uses only built-in</em> <code>Intl</code><em>.</em></p>
<h2>Gregorian → Hijri with Intl</h2>
<pre><code class="language-ts">const fmt = new Intl.DateTimeFormat('en-u-ca-islamic-umalqura', {
  timeZone: 'UTC', year: 'numeric', month: 'numeric', day: 'numeric',
});

function toHijri(date: Date) {
  const p = Object.fromEntries(fmt.formatToParts(date).map((x) =&gt; [x.type, x.value]));
  return { year: parseInt(p.year), month: Number(p.month), day: Number(p.day) };
}

toHijri(new Date(Date.UTC(2024, 2, 11))); // { year: 1445, month: 9, day: 1 }  (1 Ramadan 1445)
</code></pre>
<p>Use <code>formatToParts</code> rather than parsing the formatted string: the string's layout changes with the locale (and with locales like <code>ar</code>, so do the digits), while the parts are stable.</p>
<p>That's the easy direction. Here are the four things that go wrong.</p>
<h2>Trap 1: "islamic" isn't one calendar</h2>
<p>ICU ships several Hijri variants, and they don't always agree. For 11 March 2024:</p>
<table>
<thead>
<tr>
<th>Calendar</th>
<th>Result</th>
</tr>
</thead>
<tbody><tr>
<td><code>islamic-umalqura</code></td>
<td>1 Ramadan 1445</td>
</tr>
<tr>
<td><code>islamic</code></td>
<td>1 Ramadan 1445</td>
</tr>
<tr>
<td><code>islamic-civil</code></td>
<td>1 Ramadan 1445</td>
</tr>
<tr>
<td><code>islamic-tbla</code></td>
<td><strong>2</strong> Ramadan 1445</td>
</tr>
</tbody></table>
<p><code>islamic-civil</code> and <code>islamic-tbla</code> are arithmetic calendars (fixed rules for month lengths); <code>islamic-umalqura</code> uses ICU's data for the Umm al-Qura calendar published in Saudi Arabia. If your users are in Saudi Arabia (or you're matching government documents), ask for <code>islamic-umalqura</code> explicitly.</p>
<h2>Trap 2: don't rely on the locale's default calendar</h2>
<p>It's tempting to write <code>new Intl.DateTimeFormat('ar-SA')</code> and expect Hijri output. In Node 24 (ICU 78), that formats 11 March 2024 as <strong>١١‏/٣‏/٢٠٢٤</strong>, a Gregorian date, and <code>resolvedOptions().calendar</code> is <code>"gregory"</code>. Locale defaults are data, and data changes between runtime versions. Always put the calendar in the locale string (<code>-u-ca-islamic-umalqura</code>) or pass <code>calendar: 'islamic-umalqura'</code>.</p>
<h2>Trap 3: time zones</h2>
<p>A <code>Date</code> is an instant, not a calendar day. Take midnight on 11 March 2024 in Riyadh:</p>
<pre><code class="language-ts">const d = new Date('2024-03-11T00:00:00+03:00');
// formatted with timeZone: 'UTC'          → 29 Sha'ban 1445  (it's still 10 March in UTC)
// formatted with timeZone: 'Asia/Riyadh'  → 1 Ramadan 1445
</code></pre>
<p>Pick one convention and stick to it. Either build dates with <code>Date.UTC(...)</code> and always format in UTC (what the snippets here do), or always pass the user's time zone. Mixing the two is where "the date is off by one" reports come from.</p>
<h2>Trap 4: the official calendar isn't the announced date</h2>
<p>Umm al-Qura is a calendar computed in advance. The start of Ramadan and the Eids, however, is usually decided by an official announcement based on moon sighting, so the observed date can differ from the calendar by a day, and can differ between countries. Your conversion can be correct and still not match the announced date, so say so in the UI wherever it matters (countdowns, Ramadan timetables).</p>
<h2>Hijri → Gregorian: the direction Intl can't do</h2>
<p><code>Intl</code> only <strong>formats</strong>. There's no API that takes "1 Ramadan 1448" and gives you a <code>Date</code>. (The answer you'll often find, "use <code>Intl.DateTimeFormat</code> with an Islamic calendar", only covers the other direction.)</p>
<p>You can still build it on top of <code>Intl</code>: estimate the date, format it with the Hijri calendar, and step until the formatter agrees. The function below is my implementation of that search; <code>Intl</code> does the calendar work, the loop just finds the right day.</p>
<pre><code class="language-ts">// Hijri → Gregorian with Intl alone: estimate, then step until the formatter agrees.
function hijriToGregorianIntl(year: number, month: number, day: number): Date {
  // A rough starting point near the start of the Hijri calendar; the mean Islamic year is ≈ 354.367 days.
  // It's only a starting guess: the loop corrects it using Intl.
  const estimate = Date.UTC(622, 6, 16) + ((year - 1) * 354.367 + (month - 1) * 29.53 + (day - 1)) * 86_400_000;
  let t = Math.floor(estimate / 86_400_000) * 86_400_000; // UTC midnight, so the result is the start of the day
  for (let i = 0; i &lt; 60; i++) {
    const h = toHijri(new Date(t));
    const monthDiff = (year * 12 + month) - (h.year * 12 + h.month); // continuous month index
    const diff = monthDiff * 29.5 + (day - h.day);
    if (h.year === year &amp;&amp; h.month === month &amp;&amp; h.day === day) return new Date(t);
    t += Math.sign(diff) * Math.max(1, Math.round(Math.abs(diff))) * 86_400_000;
  }
  throw new RangeError('No such Hijri date in this calendar');
}

hijriToGregorianIntl(1448, 9, 1).toISOString(); // '2027-02-08T00:00:00.000Z'
hijriToGregorianIntl(1446, 9, 30); // throws: that Ramadan had 29 days
</code></pre>
<p>I checked this function against a table-based implementation on every valid Hijri date from 1356 to 1500 AH (51,383 dates, including every year boundary): all matched, using on average about 2 formatter calls per conversion and never more than 3. (Node 24.21 / ICU 78.3.)</p>
<p>Note the last line: Hijri months have 29 or 30 days, and which one varies by year (Ramadan was 30 days in 1445, 29 in 1446 and 30 in 1447). So <strong>validate Hijri input</strong> rather than assuming 30-day months.</p>
<h2>Runtime support</h2>
<p>The <code>islamic-umalqura</code> calendar comes from the runtime's ICU data. Node has shipped full ICU by default since version 13, and current browsers support the Islamic calendars through <code>Intl</code>, but older or stripped-down runtimes may not. Check <code>new Intl.DateTimeFormat('en-u-ca-islamic-umalqura').resolvedOptions().calendar</code>: if it doesn't say <code>islamic-umalqura</code>, you've silently fallen back to another calendar.</p>
<h2>When a table beats Intl</h2>
<p>The search above works, but a lookup table is better when:</p>
<ul>
<li><p><strong>You convert a lot.</strong> In a rough benchmark (Node 24.21 on a Windows laptop, the 51,383 dates above converted in a loop, averaged), the <code>Intl</code> search took about 15 µs per date and a table lookup about 0.3 µs. Your numbers will differ; the gap is the point.</p>
</li>
<li><p><strong>You need the same answer everywhere.</strong> A table gives identical results in every browser, Node version and runtime, independent of the ICU data each one ships.</p>
</li>
<li><p><strong>You need month lengths, validation or date differences</strong> without probing the formatter.</p>
</li>
</ul>
<p>I maintain <a href="https://github.com/getkirnu/arabic-core">@kirnu/arabic-core</a>, a zero-dependency TypeScript library that includes the Umm al-Qura table for 1343–1500 AH (2 August 1924 to 16 November 2077). It agrees with <code>Intl</code>'s <code>islamic-umalqura</code> on every day from 1937 to 2077 that I checked:</p>
<pre><code class="language-ts">import { gregorianToHijri, hijriToGregorian, hijriMonthLength, HIJRI_RANGE } from '@kirnu/arabic-core';

gregorianToHijri({ year: 2024, month: 3, day: 11 }); // { year: 1445, month: 9, day: 1 }
hijriToGregorian({ year: 1448, month: 9, day: 1 });  // { year: 2027, month: 2, day: 8 }
hijriMonthLength(1446, 9);                            // 29
// Dates outside HIJRI_RANGE throw a HijriRangeError instead of guessing.
</code></pre>
<p>It also does Hijri ages and date differences, and (the reason it exists) Arabic number-to-words. If you only need a quick answer rather than code, there's a <a href="https://getkirnu.com/ar/dates/hijri-to-gregorian/?utm_source=hashnode">free online Hijri ↔ Gregorian converter</a> built on the same table (Arabic interface). The library's source and tests are at <a href="https://github.com/getkirnu/arabic-core">github.com/getkirnu/arabic-core</a>.</p>
<p>If you've hit a Hijri-date bug that isn't covered here, I'd like to hear about it in the comments.</p>
]]></content:encoded></item><item><title><![CDATA[Converting numbers to Arabic words in JavaScript is harder than you think]]></title><description><![CDATA[In English, turning a number into words is close to a lookup table: 15 is fifteen, and 15 dollars is fifteen dollars. In Arabic, the same number is written differently depending on the noun that follo]]></description><link>https://kirnu.hashnode.dev/converting-numbers-to-arabic-words-in-javascript-is-harder-than-you-think</link><guid isPermaLink="true">https://kirnu.hashnode.dev/converting-numbers-to-arabic-words-in-javascript-is-harder-than-you-think</guid><category><![CDATA[JavaScript]]></category><category><![CDATA[TypeScript]]></category><category><![CDATA[i18n]]></category><category><![CDATA[Open Source]]></category><category><![CDATA[Arabic ]]></category><dc:creator><![CDATA[Kirnu (كرنو)]]></dc:creator><pubDate>Thu, 08 Oct 2026 10:53:44 GMT</pubDate><content:encoded><![CDATA[<p>In English, turning a number into words is close to a lookup table: 15 is <em>fifteen</em>, and 15 dollars is <em>fifteen dollars</em>. In Arabic, the same number is written differently depending on the noun that follows it and on its role in the sentence. Fifteen Saudi riyals is <strong>خمسة عشر ريالاً</strong> (<em>khamsata ʿashara riyālan</em>), while fifteen halalas (the Saudi riyal's fractional unit) is <strong>خمس عشرة هللة</strong> (<em>khamsa ʿashrata halalatan</em>): both words of "fifteen" change. That's why most quick number-to-words implementations for Arabic produce text a native speaker immediately spots as wrong.</p>
<p>Writing amounts in words isn't a niche need in the Arab world: cheques, invoices and contracts in Saudi Arabia, the UAE, Egypt and elsewhere write the amount in words, and on a cheque the words usually take precedence over the digits. The practice even has its own name, <strong>tafgeet</strong> (تفقيط).</p>
<p>This post walks through the rules an Arabic number-to-words engine has to model and what each one means for your code. All examples were generated by <a href="https://github.com/getkirnu/arabic-core">@kirnu/arabic-core</a>, an open-source TypeScript library I work on that powers the tools on <a href="https://getkirnu.com/ar/?utm_source=hashnode">Kirnu</a>, a free Arabic tools site. The Arabic examples use Modern Standard Arabic; regional spoken Arabic can differ.</p>
<h2>Why a lookup table can't work</h2>
<p>The correct output depends on information that isn't in the number:</p>
<ul>
<li><strong>The gender of the counted noun.</strong> "Three books" and "three pages" use different words for <em>three</em>.</li>
<li><strong>Grammatical case.</strong> "Twelve" changes form when it's the object of a verb or follows a preposition.</li>
<li><strong>The number's range.</strong> 3–10, 11–19, the tens and the hundreds each follow different rules.</li>
<li><strong>The form of the noun.</strong> After 3–10 the noun is plural; after 11–99 it's <em>singular</em>; after 100 it's singular again but in a different case.</li>
<li><strong>The currency.</strong> Each currency has its own main and fractional unit, each with its own gender, and they don't all have 100 subunits.</li>
<li><strong>Fractions</strong> and, when needed, <strong>cheque format</strong>.</li>
</ul>
<p>So an Arabic <code>numberToWords(n)</code> needs more inputs than <code>n</code> (at least gender and case), and converting <em>amounts</em> needs data about each currency. The library used here covers seven: the Saudi, Qatari and Omani riyal, the UAE dirham, the Egyptian pound, and the Kuwaiti and Bahraini dinar.</p>
<p>Let's go through the rules one at a time.</p>
<h2>1. Numbers 3–10 take the <em>opposite</em> gender</h2>
<p>The rule people forget most: from three to ten, the number takes the opposite gender of the noun. With a masculine noun it gets the feminine-looking ending <em>-a</em> (ة); with a feminine noun it drops it.</p>
<table>
<thead>
<tr>
<th>Number</th>
<th>With a masculine noun (book)</th>
<th>With a feminine noun (page)</th>
</tr>
</thead>
<tbody><tr>
<td>3</td>
<td>ثلاثة <em>thalātha</em></td>
<td>ثلاث <em>thalāth</em></td>
</tr>
<tr>
<td>5</td>
<td>خمسة <em>khamsa</em></td>
<td>خمس <em>khams</em></td>
</tr>
<tr>
<td>8</td>
<td>ثمانية <em>thamāniya</em></td>
<td>ثماني <em>thamānī</em></td>
</tr>
</tbody></table>
<p>One and two do the opposite and <em>agree</em> with the noun: one book is كتاب واحد, one page is ورقة واحدة. <strong>Consequence for the API: your function needs to know the noun's gender</strong>, and a sensible default is masculine.</p>
<h2>2. 11–19: two words, two rules</h2>
<p>Compound numbers combine both behaviours. In <em>fifteen</em>, the "five" part follows the 3–10 rule (opposite gender) while the "ten" part agrees with the noun:</p>
<ul>
<li>masculine: خمسة عشر <em>khamsata ʿashara</em></li>
<li>feminine: خمس عشرة <em>khamsa ʿashrata</em></li>
</ul>
<p>So a single gender flag flips the two words in opposite directions, which means you can't generate each part independently. Eleven and twelve have their own forms: أحد عشر / إحدى عشرة for 11, and اثنا عشر / اثنتا عشرة for 12.</p>
<h2>3. Grammatical case changes the word itself</h2>
<p>"Twelve" is اثنا عشر in the nominative but اثني عشر in the accusative and genitive. The same goes for "two thousand" (ألفان → ألفين) and "twenty-two" (اثنان وعشرون → اثنين وعشرين). In "I received twelve orders", the nominative form would be a mistake. <strong>That's the second input your function needs: case</strong>, with nominative as the safe default when the number stands alone.</p>
<p>So far we've only produced the number. Real-world uses like invoices and cheques have a noun after it, and the noun has rules of its own.</p>
<h2>4. The noun changes with the number</h2>
<p>Even with the number right, the noun (say, <em>Saudi riyal</em>) changes across four ranges:</p>
<table>
<thead>
<tr>
<th>Range</th>
<th>Noun form</th>
<th>Example</th>
</tr>
</thead>
<tbody><tr>
<td>1 and 2</td>
<td>singular / dual, number after it or omitted</td>
<td>ريال سعودي واحد, ريالان سعوديان</td>
</tr>
<tr>
<td>3–10</td>
<td>plural</td>
<td>ثلاثة ريالات سعودية</td>
</tr>
<tr>
<td>11–99</td>
<td>singular, accusative</td>
<td>خمسة عشر ريالاً سعودياً</td>
</tr>
<tr>
<td>100+</td>
<td>singular, genitive</td>
<td>مئة ريال سعودي, ألفا ريال سعودي</td>
</tr>
</tbody></table>
<p>The adjective ("Saudi") follows the noun's form too. And the range is decided by the <strong>last two digits</strong>, not the whole number: 1250 ends in fifty, so the noun is singular accusative, giving ألف ومئتان وخمسون ريالاً سعودياً. Your code has to split the number into groups before it can pick the noun form.</p>
<h2>5. Hundreds, thousands and spelling</h2>
<ul>
<li>200 is مئتان (a dual), 300 is ثلاثمئة, written as one word in modern spelling.</li>
<li>Thousands follow the counted-noun rules: ألفان (2,000), ثلاثة آلاف (3,000), أحد عشر ألفاً (11,000).</li>
<li>"Hundred" has a modern spelling (مئة) and a classic one (مائة). Both are in use, so make it an option: 1250 is ألف ومئتان وخمسون or ألف ومائتان وخمسون.</li>
<li>Parts are joined with <em>wa</em> ("and"): 1,001,011 is مليون وألف وأحد عشر.</li>
</ul>
<h2>6. Currencies and fractions</h2>
<p>For amounts, every rule above applies twice: once to the main unit and once to the subunit, including the subunit's gender.</p>
<ul>
<li><strong>Saudi riyal:</strong> 100 halalas, and <em>halala</em> is feminine, so 0.25 is خمس وعشرون هللة, not the masculine خمسة وعشرون.</li>
<li><strong>Kuwaiti dinar:</strong> 1,000 fils (three decimal places), so 1.250 is دينار كويتي واحد ومئتان وخمسون فلساً.</li>
<li><strong>Egyptian pound:</strong> 100 piastres, so 3.03 is ثلاثة جنيهات مصرية وثلاثة قروش.</li>
</ul>
<p>Plain fractions are read after the word <em>fāṣila</em> ("point"): 12.5 is اثنا عشر فاصلة خمسة. And cheques wrap the amount in فقط … لا غير ("only … nothing more") so nothing can be added before or after it:</p>
<blockquote>
<p>فقط ألف ومئتان وخمسون درهماً إماراتياً وخمسون فلساً لا غير</p>
</blockquote>
<h2>7. Write the tests before the code</h2>
<p>With this many interacting rules, a fix for one case easily breaks ten others. What worked for us: a table of hand-written cases reviewed for grammar (number, gender, case, expected output) that runs on every change, plus dedicated tests for the edge cases: 11 and 12, numbers ending in 01 and 02, dual thousands, zero, and Eastern Arabic digits (١٢٣) alongside Western ones (123).</p>
<h2>Using @kirnu/arabic-core in JavaScript and TypeScript</h2>
<p>All of the above is packaged as an MIT-licensed TypeScript library with no dependencies. It runs in the browser and in Node:</p>
<pre><code class="language-bash">npm install @kirnu/arabic-core
</code></pre>
<pre><code class="language-ts">import { tafgeet, currencyToWords } from '@kirnu/arabic-core';

// masculine noun (the default)
tafgeet(15);                                   // خمسة عشر

// feminine noun
tafgeet(15, { gender: 'feminine' });           // خمس عشرة

// grammatical case
tafgeet(12, { case: 'accusative' });           // اثني عشر

// Eastern Arabic digits with a decimal separator
tafgeet('١٢٫٥');                               // اثنا عشر فاصلة خمسة

// an amount in cheque format
currencyToWords('1250.50', 'SAR', { cheque: true });
// فقط ألف ومئتان وخمسون ريالاً سعودياً وخمسون هللة لا غير
</code></pre>
<p>Besides the seven currencies, it also removes Arabic diacritics (tashkeel), normalises Arabic text, and converts between Hijri (Umm al-Qura) and Gregorian dates.</p>
<ul>
<li>Source: <a href="https://github.com/getkirnu/arabic-core">https://github.com/getkirnu/arabic-core</a></li>
<li>npm: <a href="https://www.npmjs.com/package/@kirnu/arabic-core">https://www.npmjs.com/package/@kirnu/arabic-core</a></li>
<li>Try it without code: <a href="https://getkirnu.com/ar/text/tafgeet/?utm_source=hashnode">Kirnu's number-to-words tool</a> (Arabic interface)</li>
</ul>
<p>If you find a case it gets wrong, or need another currency, please open an issue. Corrections from Arabic speakers are just as welcome.</p>
]]></content:encoded></item></channel></rss>