background image blur
background image
  • Blog
    >
  • News
    >
  • The One-Language Internet Is a Form of Censorship

The One-Language Internet Is a Form of Censorship

Image of author
By Tech Writer and VPN Researcher Gintarė Mažonaitė
clock icon
Last updated: 30 September, 2026
A woman in India using her phone

Key Takeaways

  • Most of the world's languages are functionally absent online, and the effect on speakers resembles censorship even though nobody issued a ban.
  • AI tools perform well in languages with large training data sets and considerably worse elsewhere, while sounding equally confident in both.
  • The Gates Foundation is convening 60 organizations to build more representative language data, aiming to reach more than 3 billion people over five years.
  • How that data gets collected matters as much as whether it exists. Communities are asking who controls translations of their languages and whether consent was given.

Today, September 30th, is International Translation Day, a UN observance recognizing the work of translators and the role translation plays in connecting people across languages. The connection to internet freedom is more direct than it sounds.

Censorship is usually understood as something done to content after it exists: a block, a takedown, a ban. But if a language is barely present online in the first place, its speakers are excluded from most of the internet without anyone deciding to exclude them. The outcome resembles censorship. The mechanism gets no attention because there's nothing to point at.

Exclusion That Nobody Ordered

When a government blocks a website, there's a decision, an order, and usually a public argument about it.

When a language has almost no digital presence, there's none of that. There's an absence. Search returns little, tools work poorly, and the material people need either doesn't exist in their language or was never translated into it.

We've written before about how content moderation systems fail speakers of languages platforms don't staff for and how Indigenous languages are disappearing from digital spaces. AI has made the problem harder to see rather than easier. A model trained mostly on English produces confident output in every language it covers, and the confidence doesn't drop when the accuracy does.

The consequences aren't abstract. A mistranslation in a medical context can render guidance meaningless or dangerous, and neither the system nor the person relying on it would necessarily notice.

Something Is Actually Being Funded

The Gates Foundation is convening 60 organizations, including Anthropic, Google, and the OpenAI Foundation, to build more representative language datasets for AI tools. The stated goal is to reach more than 3 billion people over five years, and it follows a commitment of $1 billion toward AI work in health, education, and agriculture for underserved communities.

What makes it more interesting than a funding headline is the approach. Google's Project Vaani is collecting over 150,000 hours of audio across every district in India, working at dialect level rather than language level, because dialects within a language can vary enough that a model trained on one performs badly on another.

That kind of work requires local partners, fieldwork, and years. It can't be scraped.

We covered the announcement in more detail when the coalition was first reported. The short version is that it treats language coverage as infrastructure rather than as a localization step applied after a product already works, which is a genuinely different starting position.

Who Controls the Translation

The collection method is doing as much work here as the funding.

The Mozilla Data Collective, part of the coalition, has described training data scraped from the internet as the original sin of AI language representation, and is instead working with communities to contribute linguistic and cultural material on their own terms.

That distinction matters more than it might appear. As AI systems expand coverage, under-served communities are being asked to contribute language data, or finding their material used without being asked.

A community that has spent decades protecting a language has reasonable grounds to ask who benefits from its digitization, who holds rights over the result, and whether the version that ends up in a model reflects the language as speakers actually use it. UNESCO's work on moral and material rights in Indigenous translation raises exactly these questions.

Inclusion imposed without consent is a different thing from inclusion people asked for.

What the Day Is Worth Marking For

I think translation belongs in the internet freedom conversation more squarely than it usually sits.

The right to seek and receive information doesn't specify a language. In practice, it's exercised in whichever ones the infrastructure supports, and everyone else gets a narrower internet. Nobody had to ban anything for that to happen.

Whether a coalition announced in September turns into the fieldwork, funding, and governance it would take to close the gap is a question September can't answer. Sixty organizations agreeing on a problem is not the same as solving it.

But International Translation Day recognizes translators as the people who make cross-language communication possible, and the digital version of that work has mostly been carried out by volunteers, academics, and underfunded projects for languages spoken by hundreds of millions. An effort to treat it as infrastructure instead is worth acknowledging when it appears.


Share on
Facebook share Twitter share Reddit share Linkedin share

Be part of the resistance, quietly.

Get Mysterium VPN Arrow icon
awareness campaign banner img
Image of author
Gintarė Mažonaitė
Tech Writer and VPN Researcher

Gintarė is a cybersecurity writer at Mysterium VPN, where she explores online privacy, VPN technology, and the latest digital threats in editorial pieces. With hands-on experience researching and writing about data protection and digital freedom, Gintarė makes complex security topics accessible and actionable.

Read our editorial policy here.

Read more by this author
© Copyright 2026 UAB "MN Intelligence"