
English is the label on 52.4% of the X accounts that carry a language label, measured across a random sample of 96,030 accounts drawn from our 19,363,964-account X directory on 2026-09-02. Japanese is second at 10.5%, Spanish third at 9.8%. The number that matters more than any of those: 28.0% of sampled accounts carry no language label at all, and for 99.5% of the catalog that label is inferred from profile text rather than declared by the account or supplied by the platform. X publishes no language field on a user object. Every "Twitter users by language" figure you have read, this one included, is somebody's classifier output.
Geography is in worse shape. In the same sample, 61.0% of accounts had typed something into the free-text location box and 36.5% resolved to a country through our own place-name table. The country column on our account rows is not a platform field either. It is derived entirely from that free text, and in an independent 48,342-account replication sample it was populated on 4 accounts that had left location blank. Two X accounts in three cannot be placed in a country at all, and no amount of follower analysis fixes that.
Both facts land on one commercial point. If you are buying an account to reach a market, audience language decides whether that audience can read you at all, and it is the single variable that neither the platform, nor our listings, nor any competing marketplace actually states.
What language are X accounts actually in?
Shares below are of the 69,159 sampled accounts that carry a label, not of all 96,030, because a share of the whole would fold the 26,871 unlabeled accounts into a phantom language. Median follower counts come from the same sample, in which 66 distinct language codes appeared.
| Language | Accounts in sample | Share of labeled | Median followers | Median tweets posted |
|---|---|---|---|---|
| English | 36,225 | 52.4% | 795 | 2,391 |
| Japanese | 7,279 | 10.5% | 851 | 3,575 |
| Spanish | 6,811 | 9.8% | 639 | 3,252 |
| Arabic | 2,956 | 4.3% | 864 | 2,910 |
| Portuguese | 2,462 | 3.6% | 644 | 3,592 |
| French | 1,524 | 2.2% | 756 | 2,078 |
| Turkish | 1,334 | 1.9% | 935 | 2,739 |
| Chinese | 1,018 | 1.5% | 2,644 | 193 |
| German | 987 | 1.4% | 690 | 2,320 |
| Indonesian | 888 | 1.3% | 604 | 2,696 |
| Thai | 662 | 1.0% | 1,520 | 816 |
| No label | 26,871 | not applicable | 193 | 810 |
The unlabeled row is the interesting one. Accounts with no language label hold a median of 193 followers against 795 for English-labeled accounts, a 4.1x gap. They are not dead: their median tweet count is 810, so they post. What they lack is enough text in name and bio for a classifier to read, so abstention concentrates exactly where accounts are cheapest.
Where does the language label come from?
Our X engine assigns language through a five-stage cascade, and the order is the only way to read that table honestly. First choice is X's own per-tweet language label, aggregated per account and applied where we have captured at least three labeled tweets and one language holds 60% of them. That is the only ground truth in the set. Everything below it is inference.
Second is script: Hangul, kana, Devanagari, Arabic and Cyrillic settle the language themselves, and a script hit vetoes a detector answer from outside that family. Third is stopwords and language-exclusive characters, for short bios where n-grams fail but a Turkish dotless i does not. Fourth is an n-gram detector over the bio alone, abstaining below 0.90 confidence on bios of eight words or more. Fifth is the resolved country, reached only when the text said nothing.
The size of that ground-truth layer is the figure nobody publishes. We measured 104,724 distinct accounts carrying at least one language-labeled tweet in our archive of 1,818,088 captured posts, which is 0.54% of the catalog. Narrowing to accounts with three or more labeled tweets and one language holding 60% of them leaves 95,948 accounts, or 0.50%. For the other 99.5%, the label is a guess made from a bio.
How accurate is an inferred language label?
The obvious test compares the stored label against per-tweet ground truth where both exist. We sampled 25,000 of the 95,948 ground-truth accounts and joined them back to the catalog. The stored label matched the dominant tweet language on 23,766 of them, 95.1%, with 413 stored null and 821 disagreeing.
That 95.1% is not published here as an accuracy figure, because it is circular. A backfill pass inside our own engine overwrites the language field from exactly that signal at exactly those thresholds, so comparing the field to the source it was copied from measures a copy operation rather than a classifier. Killed.
A distribution comparison survives instead. Ground truth gives an account-level mix for 95,948 accounts; the sample gives an inferred mix for 69,159. Unbiased inference would make the two agree. Where they diverge by a factor, inference is inventing or erasing a language.
| Language | Inferred share (n=69,159) | Ground-truth share (n=95,948) | Inferred divided by truth |
|---|---|---|---|
| English | 52.38% | 51.26% | 1.02x |
| Japanese | 10.53% | 10.47% | 1.01x |
| Russian | 0.36% | 0.36% | 0.98x |
| Spanish | 9.85% | 11.04% | 0.89x |
| Turkish | 1.93% | 2.88% | 0.67x |
| Arabic | 4.27% | 6.61% | 0.65x |
| Thai | 0.96% | 1.64% | 0.58x |
| Hindi | 0.43% | 0.99% | 0.43x |
| Dutch | 0.54% | 0.19% | 2.84x |
| German | 1.43% | 0.46% | 3.08x |
| Catalan | 0.85% | 0.13% | 6.7x |
| Danish | 0.35% | 0.03% | 13.3x |
| Norwegian | 0.43% | 0.03% | 13.4x |
| Estonian | 0.32% | 0.01% | 33.6x |
English holds up. So does Japanese, and so does Russian, which we expected to be the failure case and which matched its ground-truth share to within 2%. Russian is genuinely rare in this directory at 0.36% of labeled accounts, and that is a fact about how our crawl grew rather than a classification error. Anyone setting these shares against a Twitter-era language breakdown should note the difference in basis: ours describes accounts our crawler saw between 2026-06-09 and 2026-09-02, not a fixed historical panel.
Which language labels are inventions?
Two are outright fiction. Afrikaans sits on 206 sampled accounts and Somali on 161, a combined 0.53% of labeled accounts, and both appear on zero of the 95,948 accounts where X's own tweet labels give a dominant language. Estonian sits on 218 sampled accounts against 9 in ground truth. Catalan, Norwegian and Danish are over-assigned by 6.7x, 13.4x and 13.3x. Each is a Latin-script language that a short bio can be mistaken for.
The pattern is diagnostic. Non-Latin-script languages come out equal to or below their ground-truth share, which is what you expect when script detection is reliable and errors live where script gives no help. Every language inflated more than 2x uses the Latin alphabet and has a small real population on X. The residue turns up where it cannot be: six Estonian-labeled accounts resolve to Japan, and US-located accounts run Norwegian 43 and Danish 33 against Spanish 121.
A rule follows for anyone reading a language facet on the X account directory or anywhere else. Trust the top of the list and the non-Latin scripts, and treat a small Latin-script European language as unproven until the posts confirm it.
How many X accounts even say where they are?
Location is user-typed free text, optional, and never validated by the platform, which means it can say anything at all. We measured its coverage twice, on two independent random samples drawn with different seeds, to confirm the figure was stable rather than an artifact of one draw.
| Field | Sample A (n=96,030) | Share | Sample B (n=48,342) | Share |
|---|---|---|---|---|
| Language label present | 69,159 | 72.0% | 34,959 | 72.3% |
| Location text present | 58,543 | 61.0% | 29,406 | 60.8% |
| Country resolved | 35,009 | 36.5% | 17,709 | 36.6% |
| Follower count present | 96,030 | 100% | 48,342 | 100% |
The samples agree to within 0.3 points on every field, so these are real coverage rates rather than a sampling accident. Of accounts that do declare a location, 60.2% resolve to a country and 39.8% do not, across 129 distinct countries. Anyone quoting an accounts-per-country figure for X is working from this same 36.5% base or from something weaker still.
Resolution accuracy is a different question from coverage, and it is the half that holds up. We re-derived country independently in SQL, matching location segments against our 1,414-row place-name table, then compared that against the engine's stored value. On the 31,936 accounts where both methods returned exactly one country they agreed on 31,922, or 99.96%, disagreeing on 14. Where a location matched two countries, the engine's pick was among the candidates 350 times out of 355. X geography is a coverage problem, not an accuracy problem.
What do people actually type in the location field?
Two very different failures sit inside the unresolved 39.8%, and conflating them would be a mistake. The first is users putting something in the box that is not a place. Across seven such strings our sample holds 586 occurrences and the resolver returned a country for none, which is correct: a wrong country is worse than a null one, because it files an account on a page it does not belong to.
| Location text | Occurrences in sample | Resolved to a country | Failure type |
|---|---|---|---|
| earth | 151 | 0 | Not a place |
| she/her | 125 | 0 | Bio overflow |
| worldwide | 108 | 0 | Not a place |
| global | 84 | 0 | Not a place |
| everywhere | 51 | 0 | Not a place |
| somewhere | 37 | 0 | Not a place |
| web3 | 30 | 0 | Not a place |
| Nashville, Pittsburgh, Minneapolis, Baltimore | 194 | 0 | Lookup table gap |
The second failure is ours. Those four American cities take 194 occurrences in the sample and resolve to nothing, because a 1,414-row place table does not carry every mid-size US city and the resolver will not infer a country from a bare state abbreviation. The unresolved 39.8% is part user noise and part table size.
The "she/her" row, at 125 occurrences, is the tell for what the field has become: a meaningful slice of X users treat the location box as spare bio space. That is why the history checks that actually verify an X account lean on account age, posting record and follower composition instead.
Does a resolved country tell you the audience language?
No, and this is the finding that should change how you price a purchase. We cross-tabulated resolved country against inferred language for every sampled account carrying both. Country is resolved on 36.5% of accounts and language on 72.0%, so this table describes accounts with fuller-than-average profiles rather than the whole directory. Among the eleven most common resolved countries, the share writing in something other than the country's main language runs from 5.9% to 41.1%.
| Resolved country | Accounts with both fields | Most common language | Its share | English share |
|---|---|---|---|---|
| United States | 8,071 | English | 90.4% | 90.4% |
| United Kingdom | 2,468 | English | 94.1% | 94.1% |
| Brazil | 1,258 | Portuguese | 74.6% | 12.2% |
| India | 1,063 | English | 75.4% | 75.4% |
| Spain | 994 | Spanish | 69.4% | 17.5% |
| Mexico | 879 | Spanish | 79.6% | 11.9% |
| Canada | 831 | English | 88.0% | 88.0% |
| Japan | 793 | Japanese | 81.8% | 12.9% |
| France | 699 | French | 58.9% | 33.5% |
| Turkey | 692 | Turkish | 82.1% | 10.8% |
| Saudi Arabia | 582 | Arabic | 83.7% | 14.3% |
France is the extreme case: 33.5% of French-located accounts in our sample write in English, so a third of the audience you would buy for a French campaign is not writing French. Spain runs 17.5% English, Saudi Arabia 14.3%, Japan 12.9%, Brazil 12.2%, Mexico 11.9%, Turkey 10.8%. Read the other way, 9.6% of US-located accounts carry a label other than English.
Does language change how hard an audience engages?
It changes the number a great deal. We measured engagement per view on 1,286,000 captured tweets posted between 2024-01-01 and 2026-09-02 with more than 100 views, grouped by X's own per-tweet label, which is ground truth here rather than inference. These are each account's top-performing captured posts, so absolute levels are selection-inflated and only the cross-language ratios are usable.
| Tweet language | Tweets measured | Median views | Median engagements | Median engagement per view |
|---|---|---|---|---|
| Urdu | 5,726 | 9,704 | 280 | 3.80% |
| Hindi | 17,079 | 14,342 | 438 | 3.15% |
| Portuguese | 51,189 | 16,625 | 455 | 2.31% |
| Spanish | 131,345 | 10,710 | 173 | 1.68% |
| Japanese | 162,351 | 25,146 | 364 | 1.55% |
| English | 647,150 | 23,312 | 345 | 1.51% |
| Thai | 18,817 | 30,473 | 438 | 1.46% |
| Turkish | 41,371 | 19,107 | 254 | 1.35% |
| French | 27,164 | 15,298 | 212 | 1.29% |
| Arabic | 84,832 | 11,892 | 89 | 0.78% |
The spread between Urdu at 3.80% and Arabic at 0.78% is 4.9x on the same platform, the same metric and the same window. Reach and engagement do not move together either: Thai posts take a median 30,473 views, 31% more than English, while converting fewer of them. Benchmarking against a 1.5% engagement rate quoted for X means benchmarking against the English figure, and applying it to an Arabic-language account makes a healthy account look broken.
Is the language label better on bigger accounts?
Yes for coverage, no for content. We bucketed the same sample by follower count, at n=96,023 for this cut because the live table moved by seven rows between queries. The share carrying a language label climbs from 48.8% under 100 followers to 97.4% above a million, because bigger accounts write longer bios and the classifier has more to read.
| Follower band | Accounts | Location present | Country resolved | Language labeled | English share of labeled |
|---|---|---|---|---|---|
| Under 100 | 19,911 | 42.3% | 27.0% | 48.8% | 49.3% |
| 100 to 999 | 39,247 | 63.0% | 37.9% | 72.3% | 53.2% |
| 1,000 to 9,999 | 29,370 | 69.3% | 40.2% | 83.4% | 53.3% |
| 10,000 to 99,999 | 6,622 | 67.8% | 39.0% | 87.1% | 49.2% |
| 100,000 to 999,999 | 796 | 66.1% | 43.6% | 91.1% | 54.8% |
| 1,000,000 and above | 77 | 68.8% | 51.9% | 97.4% | 43 of 75 labeled |
A tempting claim died here. Raw English share does rise with size, from 24.1% of all sub-100-follower accounts to 55.8% above a million, which reads like English dominating the head of the platform. Taken over labeled accounts only it is flat: 49.3%, 53.2%, 53.3%, 49.2%, 54.8%. The rise was entirely the label-coverage effect, and the top band holds only 75 labeled accounts. Location declaration plateaus near 67% above a thousand followers.
What should a buyer conclude before paying for an account?
Start with the constraint. We checked our own listings table for any column holding language, country, region, locale or audience geography and found none, across 691 listings and 141 completed sales. No marketplace we know of carries that field, ours included, so the verification is yours to do.
Do it from the posts, not the profile. A recent timeline carries X's own per-tweet language label, the one signal here that is not inference, and twenty posts settle in a minute what a bio-derived label cannot. Check the follower side too, because an account can post in one language and hold an audience in another. Our X audience overlap tool shows whose followers an account shares, and the audience finder works the same signal in reverse.
Then price against the right benchmark. That 1.5% engagement rate is the English number: Arabic runs 0.78% and Hindi 3.15% on our measurement, so an identical raw figure means different things by language. Rank the account inside its own tier with the X follower rank tool, and read what an X account should actually cost in 2026 for the ladder itself. Where the audience is in the wrong language for your market, the correct discount is large.
Sellers face the mirror image. An account with an identifiable language audience should say so when you list an account for sale, with evidence from its own posts, because no buyer can filter for it and the ones who need it will pay for certainty. Shopping the live account marketplace, assume any language claim without post-level evidence is unverified, exactly as our guide to checking whether an X account has real followers treats follower claims.
What we measured and what we threw away
Two independent random samples of xdir_accounts were drawn with Postgres system sampling at 0.5% and 0.25%, giving 96,030 and 48,342 rows from a catalog of 19,363,964. A full scan of that table reads 28 GB and we do not run one. Ground truth came from 1,818,088 captured tweets carrying X's per-tweet label. First-seen timestamps run 2026-06-09 to 2026-09-02.
Four measurements were taken and discarded. The 95.1% agreement between our stored language and per-tweet ground truth was rejected as circular. The rise in English share with account size was rejected as a label-coverage artifact. Apparent unresolvability of accented location strings, including the native spellings of Mexico and Turkiye, was an artifact of our replication SQL rather than the engine, which folds accents. And the country column, listed in our internal notes as effectively empty, measured 36.5% populated on both samples.
One caveat cannot be designed away. The ground-truth set is the 0.50% of accounts whose tweets we captured, which skews toward more active accounts, so a modest over or under-representation in the comparison table could be composition rather than classifier error. A 33.6x gap or a count of zero cannot be. That is why the Afrikaans and Somali result is stated as a finding while the Hindi and Thai gaps are not.
Questions buyers ask about X account language and geography
The six questions below come up in nearly every conversation about buying an X account for a non-English market, and most are still asked under the platform's old Twitter name rather than the current one. Each answer here is measured, drawn from the samples described above.
What percentage of X users are English speakers?
English is the label on 52.4% of the 69,159 language-labeled accounts in our 96,030-account random sample, measured 2026-09-02. On the 95,948 accounts where X's own per-tweet labels give a dominant language, English is 51.3%. Both describe our directory of 19,363,964 accounts rather than X as a whole, and neither is platform-published.
Does X tell you what language an account is in?
Not on the account. X labels each individual tweet with the language it detected at post time, but the user object carries no language field. Any account-level language you see, on our directory or anywhere else, is either aggregated from those per-tweet labels or inferred from the bio. That route covers 104,724 accounts in our catalog, 0.54% of the total.
Can you filter X accounts by country?
Only for about a third of them. In our 96,030-account sample, 61.0% had typed something into the location field and 36.5% resolved to a country across 129 distinct countries. The remaining 63.5% cannot be placed. Any country filter, ours included, runs on that resolved minority, and reflects a volunteered self-description rather than a verified location.
Is a foreign-language X account worth less?
It is worth less to you if you cannot address the audience, which is a different claim from being worth less in general. Median follower counts by language sit in a narrow band, from 598 for Russian to 935 for Turkish, with English at 795. What changes materially is engagement per view: 0.78% for Arabic against 3.80% for Urdu, a 4.9x spread on top tweets from 2024-01-01 to 2026-09-02.
How do I verify the audience language before buying?
Read the account's own recent posts, because each carries X's per-tweet language label and that is the only non-inferred signal available. Check the follower side separately, since an account can post in one language and hold an audience in another. Do not rely on the location field: 39.0% of accounts leave it blank, and 125 occurrences of "she/her" turned up in our sample as location text.
Which language labels should I not trust?
Small Latin-script European ones. Afrikaans and Somali were assigned to 367 sampled accounts and appear on zero of the 95,948 ground-truth accounts. Estonian is over-assigned 33.6x, Norwegian 13.4x, Danish 13.3x and Catalan 6.7x. English, Japanese, Russian and every non-Latin-script language came within roughly 2x of their ground-truth shares.
Related Articles
Continue learning with these related guides.

Does Twitch Account Age Still Matter? We Measured 432,192
Across 432,192 Twitch channels read on 2026-09-08, account age is almost pure accumulation, and at a matched follower count the newer channel wins.

X Bookmark Rate Benchmarks From 94,318 Scored Accounts
Half the X accounts we scored on ten or more posts bookmark under 0.0273% of their views, and a third of original posts get no bookmarks at all.
