Java String Length: Code Units vs Code Points
Understand how Java's String.length() counts UTF-16 code units, why that differs from visible characters, and when to use codePointCount() for Unicode-aware measurement.
In Java, String.length() returns the number of UTF-16 code units in the string, not necessarily the number of characters a user sees. This distinction matters when your application handles emoji, accented characters, or other text outside the Basic Multilingual Plane. This article explains what length() returns, when to use codePointCount() instead, and why the difference affects validation, indexing, and storage.
How to Get the Length of a String in Java
The simplest way to get the length of a string is to call the length() method:
String message = "Hello, world!"; int len = message.length(); System.out.println(len); // 13
length() returns an int representing the number of UTF-16 code units in the string. For most strings composed of ASCII or basic Latin characters, this equals the number of characters. The method runs in constant time because it reads the stored backing-array length without scanning the content.
What length() Actually Returns: UTF-16 Code Units
Java's String API exposes strings as sequences of UTF-16 code units, represented by char values. Characters from the Basic Multilingual Plane (BMP), which includes most common scripts, are represented by a single code unit. Characters outside the BMP, such as many emoji and some rare CJK characters, are represented by a pair of code units called a surrogate pair.
Consider this example:
String emoji = "😀"; // U+1F600 System.out.println(emoji.length()); // 2
Even though the string contains one visible character, length() returns 2 because a surrogate pair occupies two char positions. This behavior is often surprising and leads to bugs when code assumes length() gives the number of characters for display or validation.
Counting Unicode Code Points with codePointCount()
To count the actual Unicode code points—the values that represent each character—use codePointCount(int beginIndex, int endIndex):
String emoji = "😀"; int codePoints = emoji.codePointCount(0, emoji.length()); System.out.println(codePoints); // 1
This method scans the string and treats surrogate pairs as a single code point. For most standalone characters, that matches the visible character count. However, codePointCount() counts code points, not grapheme clusters. Combining sequences and emoji ZWJ sequences can contain multiple code points while displaying as one character.
The codePointAt(int index) method returns the code point at a specific index, but it also expects an index measured in code units, not code points. If you need to iterate over code points, use codePoints() (Java 8+) to get an IntStream:
String text = "A😀B"; text.codePoints().forEach(cp -> System.out.println(Character.toChars(cp)));
When the Difference Between Code Units and Code Points Matters
The distinction becomes critical in several real-world scenarios:
- Input validation: If you enforce a visible-character limit with
length(), strings with emoji can be miscounted. For example, a 10-emoji username haslength() == 20, so a limit of 10 written aslength() <= 10will reject it even though it has 10 code points. - Database column sizing: Database Unicode limits can count code points, code units, or bytes depending on the database and column type. Check that definition before using
length()to predict whether a value will fit. - Text truncation: Truncating a string by
length()can split a surrogate pair, producing an invalid sequence. UsecodePointCountandoffsetByCodePointsto cut safely. - Substring operations:
substring(int beginIndex, int endIndex)also operates on code unit indices. Passing a code point index will produce incorrect results.
Performance and Memory Characteristics of length()
String.length() is an O(1) operation because it reads the stored backing-array length without traversing the characters. This makes it safe to call repeatedly in loops or validation logic without performance concerns.
codePointCount() is O(n) because it must scan the string to identify surrogate pairs. For strings that are predominantly BMP characters, this overhead is small, but it is not free. If you only need to know whether a string is empty, prefer isEmpty() over length() == 0 for clarity, though both are constant time.
At the storage level, Java's UTF-16 code units are two bytes each in the classic char[] representation. Java 9+ compact strings can store Latin-1 text as one byte per character, but strings that contain emoji still use two bytes per code unit. A string with 10 emoji characters contains 20 UTF-16 code units, so its character data takes 40 bytes (plus object overhead), even though it has only 10 code points.
Common Mistakes and Edge Cases
- Assuming
length()equals character count: Always consider surrogate pairs when the string may contain non-BMP characters. - Using
length()for array indexing: If you need to access individual characters, remember thatcharAt(index)uses code unit indices. A surrogate pair occupies two indices. - Mixing code point and code unit indices: Methods like
substring()andcodePointAt()expect code unit indices. Convert usingoffsetByCodePoints()when needed. - Ignoring the null character: A
Stringcan contain the null character U+0000, andlength()counts it normally. This is rarely an issue but can confuse debugging.
Related Methods: isEmpty() and Array Length
isEmpty() returns true if length() is zero. It is a clearer way to check for an empty string.
String s = ""; if (s.isEmpty()) { // preferred // handle empty }
For arrays, the length is a public field, not a method: int[] arr = new int[10]; int n = arr.length;. This is a common source of confusion for developers new to Java, but it is unrelated to String.length().
When you need a count of Unicode code points instead of UTF-16 code units, use codePointCount(0, s.length()). For most business logic that operates on code units, length() is sufficient. If you need a true user-perceived character count for display, use a grapheme-cluster-aware API such as BreakIterator.getCharacterInstance() or a Unicode library.