Daftar kata bahasa Indonesia - sumber dan lisensi Indonesian word list - origin and licence ========================================= Source: hunspell-id (Indonesian spelling dictionary, version 2022.09.21), hunspell-id Authors (Ammar Shadiq, Andika Triwidada, Arno Brevoort, Benitius Brevoort, Kurniadi, Viko Adi Rahmawan, Volker Mueller, Yurizal Susanto; https://github.com/shuLhan/hunspell-id), as shipped in the LibreOffice dictionaries, commit 32b006a2c22a4ac7e8ed3f03346f7b3d85a970a4: https://raw.githubusercontent.com/LibreOffice/dictionaries/32b006a2c22a4ac7e8ed3f03346f7b3d85a970a4/id/id_ID.dic SHA-256 775ff17a52801f3d2ee80120952fb44cc9bac6c3cf61740de775e439851a5803 .../id/id_ID.aff SHA-256 c625d5b237a489c452cf1f9c666600103e8093667dff89e7030a24217995dc79 .../id/README-dict.adoc (here: README-hunspell-id-original.adoc), .../id/LICENSE-dict (GNU LGPL version 3 text, here unchanged: LICENSE-dict) Retrieved: 2026-09-28 Licence: "GNU Lesser General Public License as published by the Free Software Foundation, either version 3 of the License, or (at your option) any later version" (README-dict.adoc, section "Lisensi", Copyright (C) 2004-2022 hunspell-id Authors). The LGPLv3 is the GPLv3 plus additional permissions, so the data may be passed on under the GPLv3. Processing (scripts/build-wordlist.mjs, function idFormen, run: node scripts/build-wordlist.mjs id): - id_ID.aff (FLAG long, ISO-8859-1) is expanded with its prefix and suffix rules: prefixes (me-, di-, ber-, ter-, pe-, per-, se-, ke- ...), suffixes (-an, -i, -kan) and circumfixes (ke-...-an, pe-...-an, me-...-kan, se-...-nya, ke-...-nya; flag A1 = CIRCUMFIX: prefix and suffix only together), NEEDAFFIX A2 (the stem alone is not a word), cross products of prefix and suffix, prefixes licensed by a suffix continuation class - without clitics: the particles -lah/-kah and the pronouns -ku/-mu/-nya (classes l0, n0, nl, o0 and the variants -annya/-kanlah ... of a0/i0/k0) are not counted as words of their own (rumahku, bukunya, makanlah), except a hand-checked list of lexicalised adverbs (ID_NYA: akhirnya, biasanya, misalnya, rupanya, tampaknya ...) - lower-case words a-z only (proper names and acronyms such as Jakarta or ABRI are written with capitals in id_ID.dic; reduplications with hyphen such as anak-anak are left out), 2-20 letters, at least one vowel (no abbreviations such as km, dll) - only forms that occur in wordfreq (id) - wordfreq 3.1.1 has only the "small" Indonesian list, so all forms have Zipf >= about 3 (see README_wordfreq.txt) - proper names that id_ID.dic also lists in lower case removed: country names from CLDR (Intl.DisplayNames 'id') and a hand-checked list (ID_NAMEN: indonesia, jawa, islam, amerika ...; derived forms such as keislaman stay). id_ID.dic also lists many person, place and brand names in lower case (henry, linda, angga, glodok, tiktok); these were found by reviewing the stems of the finished list that are not in the LibreOffice Indonesian thesaurus (th_id_ID_v2, used only for this review, not in the build) and added to ID_NAMEN - slurs and the coarsest vulgar words removed together with their forms (ID_RAUS) Result: 15,288 word forms in .txt. Wordle: wordle.txt = forms with exactly 5 letters a-z (like Katla, github.com/pveyes/katla: keyboard a-z, "kata valid 5 huruf sesuai KBBI") -> 2,316 words. Limits: wordfreq does not know upper and lower case; KBBI words that are also common names (haris, salim) stay. Derived forms missing in id_ID.dic (e.g. diproduksi, digelar) are not generated. Licence of the lists in this folder (.txt, wordle.txt): GPL-3.0-or-later (COPYING-GPL-3.0.txt), because they combine hunspell-id (LGPL-3.0-or-later) with wordfreq data (CC BY-SA 4.0).