Anglicism Checker Methodology

Full disclosure of how we compute the clarity score: formula, weights, dictionary sources and tokenization decisions. No black boxes — every number is transparent and reproducible.

📐 Scoring formula

The final score is a number from 0 to 100. It's based on a smoothed density of weighted matches: each weight depends on severity, the necessity hierarchy, the domain, and context (for homonyms).

weight = SEV[level] × NECESSITY[tier] × domain_factor × context × strictness weighted = Σ weight context: exact=1.0, possible=0.3, native_sense=0 density = weighted / ( word_count + 60 ) ← smoothing for short texts penalty = 100 × density / ( density + 0.04 ) score = max( 0 , round( 100 − penalty ) )

Each match counts once; multi-word phrases count as one. Homonyms are judged by context: not penalized in their English sense, gently in ambiguous cases. Smoothing and a gradual curve keep the score stable across text lengths and comparable between texts.

⚖️ Weights and factors

A match's weight combines severity, the necessity hierarchy (VDS model), domain, and strictness. All values are baked into the code and identical across languages.

Severity level

LevelWeightMeaning
low0.5Naturalised or neutral borrowing ("manager").
medium1.0OK in some contexts, redundant in others ("content").
high2.0Fully replaceable by a native word with no loss ("deadline").

Strictness mode

ModeFactorMeaning
lax×0.6Lax (×0.6). Ignores borrowings that are acceptable within the domain.
balanced×1.0Balanced (×1.0, default). Counts everything, but gently.
strict×1.6Strict (×1.6). Holds you to a higher bar.
purist×2.4Purist (×2.4). No leniency, even for naturalized words.

Necessity hierarchy (tier) — VDS Anglizismenindex model

ModifierFactorWhen applied
Needed (×0.15)×0.15No exact English equivalent — terminology. Barely penalized.
Nuance (×0.5)×0.5Adds a shade of meaning or has caught on. Penalized at half weight.
Avoidable (×1.0)×1.0A solid English replacement exists. Full weight.

Modifiers

ModifierFactorWhen applied
Domain match (×0.4)×0.4The word is acceptable in the chosen domain (e.g., an IT term in IT writing).
Homonym context×1.0 / ×0.3 / 0For homonyms: exact sense ×1.0, ambiguous ×0.3, native English sense — not counted.

🎯 Letter grades

Numeric score is mapped to a letter for quick orientation. Boundaries are hard-coded:

A
90–100
Excellent
B
75–89
Good
C
60–74
Mixed
D
40–59
Poor
F
0–39
Heavily borrowed

🧠 Tokenization & noise filtering

Before matching, we mask fragments that look like words but aren't. Masking preserves source offsets so highlighting stays accurate.

  • URLs (https://…) — masked entirely, so we don't flag "site" inside "https://site.com".
  • Email addresses — same.
  • @mentions and #hashtags — skipped as social markup.
  • Inline `code` in backticks — ignored.
  • camelCase and snake_case identifiers — ignored as code.
  • Multi-word phrases (e.g. "user experience") are matched before single words and consume their tokens, so they don't count twice.
  • Known brands and proper nouns (Apple, Yandex, ChatGPT) are skipped. The list extends per language via `meta.brands`.

📚 Per-language dictionaries

All dictionaries are freely available as JSON. Use them in your projects, cite in research, send corrections.

Frequently-cited entries carry their own source (`src`) — visible in the per-word popover. As curation progresses we are seeding every entry with a per-entry citation.

Russian 2805 · v2026-07-23-r5 · CC0

Sources
  • RAN reference dictionaries
  • Gramota.ru
  • VDS Anglizismenindex (cross-ref)
  • НКРЯ (Национальный корпус русского языка)
  • Словарь молодёжного сленга 2020-25

German 1187 · v2026-07-16-de2 · CC0

Sources
  • VDS Anglizismenindex (cross-ref)
  • Duden Fremdwörterbuch (descriptive)
  • VDS Anglizismenindex
  • Duden
  • Wahrig Fremdwörterlexikon
  • Goethe-Institut Wortschatz

French 2460 · v2026-07-16-fr1 · CC0

Sources
  • Académie française — Dire, ne pas dire
  • OQLF Vitrine linguistique
  • TERMIUM Plus
  • Larousse
  • FranceTerme (CULTURE.FR)
  • Office québécois de la langue française (OQLF)
  • Le Robert

Spanish 968 · v2026-07-16-es1 · CC0

Sources
  • RAE Diccionario panhispánico de dudas
  • Fundéu RAE
  • Diccionario de la lengua española (DRAE)
  • RAE — Diccionario de la lengua española
  • RAE — Diccionario panhispánico de dudas
  • Fundéu BBVA
  • CORPES XXI
  • Banco de neologismos del CSIC

Portuguese 1030 · v2026-07-16-pt1 · CC0

Sources
  • Academia Brasileira de Letras (ABL)
  • Ciberdúvidas da Língua Portuguesa
  • Vocabulário Ortográfico da Língua Portuguesa (VOLP)
  • Priberam
  • Houaiss
  • Aulete
  • Acordo Ortográfico de 1990

Korean 767 · v2026-07-16-ko1 · CC0

Sources
  • 국립국어원 외래어 표기 용례
  • 국립국어원 우리말 다듬기
  • 표준국어대사전
  • 국립국어원 (NIKL)
  • 우리말샘
  • 외래어 표기법
  • 한국어 어문 규범

Indonesian 745 · v2026-07-16-id1 · CC0

Sources
  • KBBI (Kamus Besar Bahasa Indonesia)
  • Badan Bahasa — Pedoman Padanan Istilah Asing
  • Pusat Bahasa
  • Badan Pengembangan dan Pembinaan Bahasa
  • Tata Bahasa Baku Bahasa Indonesia

Arabic 663 · v2026-07-16-ar1 · CC0

Sources
  • مجمع اللغة العربية بالقاهرة
  • مجمع اللغة العربية الأردني
  • معجم الوسيط
  • المعجم العربي الأساسي
  • مجمع اللغة العربية في دمشق
  • معجم اللغة العربية المعاصرة

🤝 Honest about limitations

We aim to be a reference — so we're upfront about what the tool does and doesn't do.

What we don't do yet: Lemmatization is approximate (regex over root and endings). Proper nouns are not fully separated — "Site Smith" might be flagged. Contextual grammar is not considered. These improvements are on the roadmap; until they ship, the most reliable mode is Balanced strictness with naturalised enabled.

📜 License

All eight dictionaries are released under CC0 — public domain. Use without restriction or attribution. The tool itself is closed source, but the data is yours.