All posts

#software bugs#unicode#java

The Turkish I Problem: How One Dotted Letter Breaks Software Worldwide

Turkish's i-İ and ı-I distinction is one of the most famous sources of software bugs. Why does your app break on a Turkish user's computer, why does "title" become "TİTLE", and how do you protect yourself?

The Turkish I Problem: How One Dotted Letter Breaks Software Worldwide
Contents 7

There's a famous class of bugs known as "the Turkish I problem", important enough that companies from Microsoft to Oracle warn about it in their documentation. Its source is the Turkish alphabet itself: Turkish has four different i letters, while most languages have two.

An app works perfectly on a developer's machine in the US. On a computer whose system language is Turkish, the same app can't log in, can't read its config file, or crashes outright. All because of one letter.

In short:

  • In English, i uppercases to I. In Turkish, i uppercases to İ and ı uppercases to I.
  • In many languages case conversion follows the system language, so on a Turkish system "title" uppercases to "TİTLE" and no longer matches "TITLE".
  • In Python, "İ".lower() returns not one letter but a two-character string.
  • The fix: locale-independent conversion for machine-facing text (keywords, file names, protocols), explicit Turkish conversion for human-facing text.

Four i's: the root of the problem

Most languages using the Latin alphabet have one pair: lowercase i and uppercase I. When Turkish switched to the Latin alphabet in 1928, it defined two separate pairs for two different sounds:

Lowercase Uppercase Unicode
i (dotted) İ (dotted) U+0069 → U+0130
ı (dotless) I (dotless) U+0131 → U+0049

Unicode handled this by leaving English i and I as they were and adding two new characters for Turkish İ and ı. So the "i" in a Turkish text and the "i" in an English text are the same character. What differs is its uppercase form.

That's why case conversion depends on the answer to "according to which language?":

  • With English rules: i → I, I → i
  • With Turkish rules: i → İ, I → ı

How the bug shows up

In many programming languages, case conversion follows the operating system's language by default. A developer, unaware of that, writes something like:

String command = userInput.toUpperCase();
if (command.equals("QUIT")) {
    exit();
}

On an English computer, a user typing "quit" can exit. On a Turkish computer, "quit".toUpperCase() returns "QUİT". It doesn't match "QUIT", and the program doesn't exit. Java's own official documentation gives this exact example: in a Turkish locale, "TITLE".toLowerCase() returns not "title" but "tıtle" with a dotless ı.

This bug can appear anywhere you compare HTTP headers, file extensions (.HTML or .html?), SQL keywords, config keys or email addresses. Code that lowercases an "ID" field to compare with "id" gets "ıd" on a Turkish system and misses the match.

Language by language

Java and Kotlin

In Java, toUpperCase() and toLowerCase() without arguments use the default locale. Pass Locale.ROOT for every machine-facing conversion:

"quit".toUpperCase(Locale.ROOT);                    // "QUIT" (on every system)
"istanbul".toUpperCase(Locale.forLanguageTag("tr")); // "İSTANBUL"

Kotlin's uppercase() and lowercase(), introduced in 1.5, are locale-independent by default; for Turkish conversion you pass the locale explicitly.

C# / .NET

In .NET, ToUpper() and ToLower() use the current culture. Microsoft's docs call out this problem by name. The fix is ToUpperInvariant() and, for comparisons, StringComparison.OrdinalIgnoreCase:

string.Equals(command, "QUIT", StringComparison.OrdinalIgnoreCase);

JavaScript

In JavaScript, toUpperCase() and toLowerCase() are locale-independent, so they're safe for machine text. The opposite problem appears instead: to uppercase Turkish text correctly you need toLocaleUpperCase('tr-TR'):

"istanbul".toUpperCase();              // "ISTANBUL"  (wrong for Turkish)
"istanbul".toLocaleUpperCase("tr-TR"); // "İSTANBUL"

If you uppercase headings with CSS (text-transform: uppercase), don't forget lang="tr" in your HTML. Browsers convert according to that attribute; without it, "İstanbul" shows up as "ISTANBUL".

Python

In Python, str.upper() and str.lower() are locale-independent. But Python has a different surprise:

>>> "İ".lower()
'i̇'
>>> len("İ".lower())
2

Under Unicode's general rules, lowercasing the dotted capital İ adds a combining dot (U+0307) to "i" to preserve the dot. It looks like "i" on screen, but it's two characters. That's why "İstanbul".lower() == "istanbul" returns False. Java shows the same behavior with Locale.ROOT.

Databases and search

The problem doesn't stop at application code. Case-insensitive search in databases also depends on the chosen collation. With a non-Turkish collation, a search for "ışık" may not find "IŞIK"; or, the other way round, English rules treat "i" and "ı" as different letters and produce unexpected results.

In a system where users search in Turkish, pick a Turkish collation for the search field or search against a normalized column.

The golden rule: ask who the text belongs to

One question clears up this mess: does this text belong to a human or a machine?

Machine text (commands, keywords, file extensions, HTTP headers, JSON field names, email domains):

  • Use locale-independent conversion: Locale.ROOT in Java, Invariant and Ordinal in .NET.
  • Better yet, don't convert at all; compare case-insensitively directly.

Human text (names, cities, headings, product names):

  • Convert explicitly with the right language's rules.
  • Don't leave the user's language to the system's language; set it deliberately.

A simple testing tip: set your test environment's default language to Turkish once and run the whole test suite. In Java that's as easy as starting the JVM with -Duser.language=tr -Duser.country=TR. Most Turkish I bugs show up on the first run.

Frequently asked questions

Is this only a Turkish problem?

Turkish and Azerbaijani, which share the same alphabet rule, are the best-known cases. But language-specific casing rules exist elsewhere too, such as in Lithuanian, or the uppercase form of German ß. Turkish gives its name to the problem because it's the most common and bug-prone example.

Why didn't Unicode define a single "i"?

Unicode had to stay compatible with earlier character sets. English i and I already existed in the oldest ASCII set. Adding Turkish İ and ı as separate characters was the only way not to break billions of existing texts.

Does this affect me if my app isn't used in Turkey?

Yes. A single user or server anywhere in the world with Turkish as the system language is enough. That's why testing with a Turkish locale is considered good practice in international software projects.

How do I avoid this on my website?

Set lang correctly in your HTML, do uppercase display with CSS text-transform, and use toLocaleUpperCase('tr-TR') in JavaScript for Turkish text. When generating URLs and slugs, convert Turkish characters with a deliberate mapping table (ı → i, İ → i and so on).

That a single dot can cause so many bugs shows how much small details matter in software. If you want web and mobile apps that work flawlessly across languages, reach us through our web application development page. For more expensive bugs, see this post.

Sources

ShareLinkedInXWhatsApp
Need help with this?

If you would like to apply what this post covers to your own project, let’s look at it together.

Write to us
YE

Founder of EngerekTech. Builds web, mobile and enterprise software for businesses with Angular, Spring Boot and Flutter, and made the KPSS Düello and Kelime Kavanozu apps. On the blog he covers AI tools and software development as he uses them in his own projects.