Lista słów polskich - pochodzenie i licencja Polish word list - origin and licence ============================================ Source: SJP.PL spelling dictionary for Hunspell/MySpell (pl_PL), https://sjp.pl/sl/ort/ file https://sjp.pl/sl/ort/sjp-myspell-pl-20260901.zip SHA-256 47a5ab02cf3def374994b1b03efd9fed2de868349b8ddcebac281b484353d157 containing pl_PL.zip -> pl_PL.dic (ISO-8859-2) SHA-256 8dce9daca500da62af829958206e7a50ecd96827ca52a1582b5d79c4a0d3600d and pl_PL.aff SHA-256 82973651651aa930335c865b339b98db376ca3dbf3a661b70b9eeb71fdf41dca original notice: README_pl_PL-original.txt (version generated 2026-09-01) Retrieved: 2026-09-28 Licence: "licensed under GPL 2, LGPL 2.1, MPL (Mozilla Public License) 1.1, Apache 2.0 and Creative Commons Attribution 4.0 International" (README_pl_PL.txt; licence of your choice, sjp.pl/sl/ort: "Udostępniane na licencjach (do wyboru)"). Used here under Apache-2.0 (LICENSE-Apache-2.0.txt), which is compatible with GPLv3. Processing (scripts/build-wordlist.mjs, function plFormen, run: node scripts/build-wordlist.mjs pl): - every entry of pl_PL.dic plus all suffix forms from pl_PL.aff (SFX rules: declension, conjugation, comparison); the prefix rule "nie-" (PFX b) is not applied - lower-case words only (a-z and ą ć ę ł ń ó ś ź ż): proper names, abbreviations in capitals and multi-word entries are dropped; 2-20 letters, at least one vowel - a few slurs removed together with all their forms (PL_RAUS, hand-checked) - only forms that occur in wordfreq (pl) with Zipf >= 1.3 (PL_MIN_FINDER; see README_wordfreq.txt): all attested forms (307,000) would make 9.txt/10.txt about 480 KB, the target is <= 400 KB per file Result: 244,000 word forms in .txt (largest file 9.txt, ~380 KB). Wordle: wordle.txt = forms with exactly 5 letters of the Polish alphabet (without q, v, x; ą ć ę ł ń ó ś ź ż are letters of their own, as in the Polish Wordle "Literalnie") and wordfreq Zipf >= 3 -> 3,945 words. Licence of the lists in this folder (.txt, wordle.txt): GPL-3.0-or-later (COPYING-GPL-3.0.txt), because they combine SJP.PL (Apache-2.0) with wordfreq data (CC BY-SA 4.0).