Text

SMS Character Counter: Segments, GSM-7, and Unicode Check

SMS character counter: see segments, GSM-7 or Unicode encoding, and the exact characters that raise the segment count. Runs in your browser, nothing uploaded.

Use the SMS Character Counter: Segments, GSM-7, and Unicode Check

Counted in your browser. Nothing is uploaded or stored.

Segments0nothing to send
Encoding-up to 160 per segment
Characters0
Septets used0

Everything runs in your browser. What you enter is never uploaded or stored.

A text message looks like one message, but the network sends it in segments. A message of up to 160 characters in the basic GSM alphabet is one segment. A message with a single character outside that alphabet, even a curly apostrophe pasted from a word processor, is sent in Unicode, where a segment holds only 70 characters. Longer messages are cut into pieces that hold a little less, because each piece carries a header. Senders who pay per segment often find that a 130-character message was billed as two, or a 300-character message as five.

This page counts the message the way the network does. It tells you the encoding, the number of segments, and how much room is left, and when the message is Unicode it shows the exact characters that caused it and offers plain replacements. It also shows each segment, including the places where a character that takes two units had to move to the next piece. Everything is counted in your browser. Your message is not uploaded or stored.

How to use the SMS Character Counter: Segments, GSM-7, and Unicode Check

  1. Type or paste the messageType in the box, or paste the text from your template or marketing tool. The count updates as you type. Nothing is sent anywhere.
  2. Read the countThe tiles show the number of segments, the encoding, the characters, and the units used. The line under them says what limit applies and how many units are left in the last segment.
  3. Look at the flagged charactersIf the message is Unicode, the characters that caused it are marked in your text and listed with their codes. Each says whether a plain character can replace it.
  4. Make it GSM-7Use the changed text to replace curly quotes, long dashes, ellipses, and special spaces. You can also drop accents, which saves segments but can change spelling. Emoji and other alphabets stay.
  5. Check merge fields and recipientsIf the message has fields such as {{first_name}}, enter a typical value for each to count the message as it will be sent. Enter the number of recipients to see the total number of segments.

Two alphabets, two limits

A text message holds 140 bytes of data. When the text is written with the GSM 7-bit default alphabet, each character takes 7 bits, so 160 characters fit. The alphabet is defined in 3GPP TS 23.038. It has 127 characters you can type: the English letters and digits, the usual punctuation, a few currency signs, some accented letters such as é, è, ü, ñ, and ß, and the capital Greek letters Δ, Φ, Γ, Λ, Ω, Π, Ψ, Σ, Θ, and Ξ. It does not have the lowercase ç, or á, í, ó, ú, â, ê, or ô. French, Spanish, and Portuguese text therefore often falls out of the alphabet.

If any character is not in the alphabet, the message is sent in Unicode, called UCS-2, where each character takes 16 bits. Only 70 fit in 140 bytes. The whole message changes, not just the character: one emoji in a 100-character message makes it a Unicode message of 100 characters, which needs two segments instead of one. Twilio's documentation gives smart quotation marks as a common cause, and says that adding such a character reduces the limit from 160 to 70.

Why long messages hold fewer

A message that does not fit in one segment is cut into pieces. Each piece starts with a user data header that tells the phone how to put the pieces back in order. Twilio's documentation says the header takes seven characters' worth of space, leaving 153 characters of data per GSM-7 segment and 67 per Unicode segment. Its example is a 161-character GSM-7 message, which is sent as two segments: one with 153 characters and one with 8.

This is why the second segment is a bigger step than it looks. A 160-character message is one segment. A 161-character message is two, and it stays two up to 306 characters. Three segments reach 459. In Unicode the same steps are 70, 134, and 201. Twilio notes an exception for US and Canadian toll-free numbers, which allow 152 and 66 in a long message, and the page has a setting for it.

Characters that count twice

The GSM alphabet has a second table, the extension table, for ten more characters: the euro sign, the square brackets, the curly braces, the backslash, the caret, the tilde, the vertical bar, and the form feed. Each is sent as an escape code followed by the character, so it takes two septets. Twilio's documentation says the escape characters count as two characters in a GSM-7 message. A template with four {{fields}} has sixteen curly braces, so it is sixteen units longer than it looks.

The standard says a phone that receives an escape it does not understand shows it as a space, so the extension characters are the first to go wrong on an old handset. The page shows which characters in your text count twice.

Where a segment ends

A character that takes two units cannot be cut in half. The standard says a receiving device decodes each segment of a long message on its own, so an escape code at the end of a segment would arrive without its character. A well-made sender moves the whole character to the next segment, and the previous segment is left one unit short. The same is true for an emoji in a Unicode message, which is stored as two 16-bit units. This page splits the message this way and tells you when it happens, so the segments you see are the ones that would be sent.

Many counters add up the units and divide by the segment size, which gives the right number most of the time and a wrong one when a two-unit character sits on a border. The page counts piece by piece.

Look-alike characters

Most Unicode messages are not meant to be Unicode. A word processor turns a straight apostrophe into a curly one, a hyphen into a long dash, and three dots into a single ellipsis character. A web page or a PDF adds a non-breaking space or an invisible zero-width space. They look right and cost a lot. The page lists each such character with its code, such as U+2019 for the curly apostrophe, and offers the plain character.

Some providers do the replacement for you. Telnyx documents a smart encoding feature that substitutes more than 200 characters, such as curly quotes with straight ones and em dashes with hyphens, and notes that it cannot convert emoji or non-Latin scripts. It also notes that an ellipsis becomes three periods, which adds two characters. This page uses a shorter list of the same kind and shows every change before you accept it. Accents are not replaced unless you ask, because dropping one can change a word.

SMS and MMS

An MMS is not cut into segments. AWS End User Messaging documents that the text body of an MMS can hold 1,600 characters from any character set, and that an MMS is not broken into multiple parts. The limit differs by provider and by country, and an MMS is usually priced differently, so this page counts only SMS segments. Twilio supports up to 1,600 characters in one SMS as well, and the page warns when the text is longer.

How the counter was tested

The alphabet was compared with the Python gsm0338 codec for every character in the Basic Multilingual Plane. Both agree on the 127 single-unit characters and the 10 two-unit characters. The segmenting was compared with the Python smsutil library on 758 texts, each in the standard setting and the toll-free setting. The texts were built to land on and around every segment border, with a euro sign, a brace, a backslash, or an emoji on the last place of a segment. The page gave the same encoding, the same unit count, and the same text in every segment on all of them. Both programs follow the same rule for a two-unit character, so the agreement shows the page counts correctly under that rule and not that every carrier follows it.

Limits and accuracy

  • The page counts the message as the network would split it, but your provider decides what you are charged. It does not know any prices, so it shows segments and not money.
  • Only the GSM 7-bit default alphabet and its extension table are used. The standard also defines national language tables, which allow more accented letters in some languages. Most sending platforms do not use them, so this page does not either.
  • An emoji is counted as two 16-bit units, the way UTF-16 stores it. A provider or carrier may count differently, so a message with emoji can be off by a segment on some routes.
  • The page does not know the country, the route, or the carrier. Some carriers filter or change messages, and long messages can arrive in pieces on very old phones.
  • Merge fields are counted with the sample values you enter. The real values vary, so a message close to a border can be one segment longer for a customer with a long name.
  • RCS and MMS are not counted. They are not sent in SMS segments.
  • The replacements are a short list of look-alikes. They are not a translation, and they cannot turn an emoji or another alphabet into GSM-7.

Frequently asked questions

How many characters are in one SMS?

A single segment holds 160 characters in the GSM 7-bit alphabet, or 70 in Unicode. A longer message is split into segments of 153 characters (GSM-7) or 67 (Unicode), because each piece carries a header that joins them.

Why does my message count as Unicode?

At least one character is outside the GSM 7-bit alphabet. Curly quotes, long dashes, the ellipsis character, special spaces, accents such as á or ç, emoji, and other alphabets all do it. The page marks each one and shows what could replace it.

Does an emoji cost more than a letter?

It does two things. It switches the whole message to Unicode, where a segment holds 70 characters instead of 160, and it takes two of the 16-bit units itself. A short message with an emoji is often still one segment, but a long one takes many more.

Why do the euro sign and the curly braces count twice?

They are in the extension table of the GSM alphabet, so each is sent as an escape code plus the character. That takes two septets. The other characters that count twice are the square brackets, backslash, caret, tilde, vertical bar, and form feed.

What is a segment?

It is one 140-byte piece of a text message. A short message is a single segment. A long message is sent as several, and the phone joins them. Providers usually charge for each segment, so the count matters more than the number of messages.

What is the limit for MMS?

An MMS is not split into segments. AWS End User Messaging documents a text body of up to 1,600 characters from any character set, and other providers set their own limits. This page counts SMS segments only.

Is my message uploaded or stored?

No. The message is counted in your browser, and the code behind the page makes no network requests. It is not saved, so paste it again if you reload the page.

Research and references

This page was written and checked against the sources below.

  1. 3GPP TS 23.038: Alphabets and language-specific information
  2. Twilio: How long can a message be?
  3. Twilio: What is GSM-7 character encoding?
  4. Telnyx: Smart Encoding
  5. AWS End User Messaging: MMS file types, size and character limits