Wiktionary:Beer parlour
Wiktionary > Discussion rooms > Beer parlour
| Information desk start a new discussion | this month | archives Newcomers’ questions, minor problems, specific requests for information or assistance. |
Tea room start a new discussion | this month | archives Questions and discussions about specific words. |
Etymology scriptorium start a new discussion | this month | archives Questions and discussions about etymology—the historical development of words. |
Beer parlour start a new discussion | this month | archives General policy discussions and proposals, requests for permissions and major announcements. |
Grease pit start a new discussion | this month | archives Technical questions, requests and discussions. |
| All Wiktionary: namespace discussions 1 2 3 4 5 – All discussion pages 1 2 3 4 5 |

Welcome to the Beer Parlour! This is the place where many a historic decision has been made, and where important discussions are being held daily. If you have a question about fundamental aspects of Wiktionary—that is, about policies, proposals and other community-wide features—please place it at the bottom of the list below (click on Start a new discussion), and it will be considered. Please keep in mind the rules of discussion: remain civil, don’t make personal attacks, don’t change other people’s posts, and sign your comments with four tildes (~~~~), which produces your name with timestamp. Also keep in mind the purpose of this page and consider before posting here whether one of our other discussion rooms may be a more appropriate venue for your questions or concerns.
Sometimes discussions started here are moved to other pages for further development. In particular, changes to a major policy or guideline may be discussed on the corresponding talk page and “simple votes” (as opposed to drawn-out discussions) can be conducted on our votes page.
Questions and answers typically remain visible on this page for one to two months, but they can always be found in the appropriate monthly archive (based on the date discussion was initiated). While we make a point to preserve all discussions that were started here, talk that is clearly not appropriate for this page may be deleted. Enjoy the Beer parlour!
| Beer parlour archives edit | ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| ||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||
|
June 2026
Object animacy in translations
[edit]Adding translations of English terms is really easy thanks to the translation tool, but the one problem I have with the resource is that there seems to be no way to describe the animacy of a noun without an unassisted edit. This particularly annoys me as a Polish speaker, however some more widely spoken languages like Hindi and Arabic could also be affected by this issue. While I can't speak for those two, I know that nouns like "członek" (member) or "miś" (bear) change their animacy depending on which context they are used in, so a translation of English member in the sense of person who belongs to a group and in the sense of body part have different animacies.
My suggestion is to add an optional parameter that lets the user select animate, inanimate and animal as a noun's animacy (or deselect it, if they clicked the option by accident) to the translation template. To te umarłe słowa (talk) 10:42, 2 June 2026 (UTC)
- Since this is a technical request, it belongs at the Grease Pit. If you move this conversation there, you are liable to get better discussion and more involved editors seeing the request. ―Justin (koavf)❤T☮C☺M☯ 11:08, 2 June 2026 (UTC)
- Oh sorry for for that. Do I delete the thread and repost it there or is there another way to do it? To te umarłe słowa (talk) 13:35, 2 June 2026 (UTC)
- No worries. You can just Ctrl+X here and Ctrl+V there. ―Justin (koavf)❤T☮C☺M☯ 17:39, 2 June 2026 (UTC)
- Oh sorry for for that. Do I delete the thread and repost it there or is there another way to do it? To te umarłe słowa (talk) 13:35, 2 June 2026 (UTC)
Proto-Romance entries are messy
[edit]@Catonif, Nicodene Adjectives and masculine nouns are indiscriminately lemmatized with either classical endings (-us, -a, -um / -us, -∅, etc.) or reconstructed ones (-um, -em), compare *capillūtus and *carnūtum, *gemellicum/*melem and *plōppus. Are there guidelines on that subject? Saumache (talk) 20:15, 3 June 2026 (UTC)
- Not messy, just inconsistent. Our editing practices varied over the years. The latest one is to reconstruct in the accusative. Latest practice would also be not to have a separate entry for *melem or *plōppus at all and only mentioning them in the descendant section of the attested spelling. We have WT:AROA but it's more of a linguistic essay rather than editorial guidelines. Catonif (talk) 13:31, 4 June 2026 (UTC)
- I think many of these have been moved by Nicodene. The page Wiktionary:Proto-Romance entry guidelines (which I think is not particularly perfected yet, and so should maybe not be given much weight) seems to call for using the nominative singular: "The name of the entry should reflect, generally, the spelling of classical Latin (but not the pronunciation). [...] the 1st conjugation of verbs has -are, 2nd declension of nouns has -us, and so on, as usual." Advantages of using the nominative (-us, etc.): it matches the convention for attested Latin terms; the Latin nominative ending -s survived into Early Old French, and so can plausibly be reconstructed for Proto-Romance in the broadest, earliest sense; it makes the gender of masculine nouns clearer (even though we list gender anyways, "gemellicum" looks like a second-declension neuter). Advantages of using the accusative: for certain words, the nominative may be uncertain (e.g. third declension; second declension -er/-(e)rus); the accusative singular is almost always the phonetic source of the present-day Romance singular form (hence, this form can be solidly reconstructed for Proto-Romance of any time or place). Disadvantage of using the accusative: it creates an apparent inconsistency between the first declension and others: *alba, *domnicella are not at *albam, *domnicellam.--Urszag (talk) 21:59, 8 June 2026 (UTC)
Levantine Arabic
[edit]Isn't it time to correct our South Levantine Arabic Wiktionary entries to just Levantine Arabic in line with ISO's definition? Example entry with the older title: ملح. They used to be defined as two dialects, but there was not enough difference to set them apart. --Esperfulmo (talk) 00:48, 4 June 2026 (UTC)
- Wiktionary:Language treatment requests#Levantine Arabic, Wiktionary:Language treatment requests#Levantine merger. - -sche (discuss) 00:54, 4 June 2026 (UTC)
(Notifying workgroup: Eirikr, Atitarev, Fish bowl, Poketalker, Cnilep, Marlin Setia1, Kiril kovachev, 荒巻モロゾフ, Shen233, Cpt.Guapo, Sartma, Lugria, LittleWhole, Chuterix, Mcph2, Theknightwho, MedK1, Horse Battery, Ookap, Emanuele6):
Seemingly small change. It has always bugged me, and it also came up in an unrelated discussion around here.
For example, the page 頭
Instead of currently:
It should be:
Just trying to gather any support, caveats or disagreements? Shlyst (talk) 03:12, 4 June 2026 (UTC)
- Definitely agreed. I never noticed it before, but, now that I see it, it does look quite strange as it is now. I'd have assumed it were the other way if asked and think that way makes more sense. Ookap (talk) 03:15, 4 June 2026 (UTC)
- @Shlyst I agree. We should also include the accent in the IPA. At the moment all Japanese IPA transcriptions are incomplete. Accent in Japanese is phonemic, it can't be omitted. [a̠ta̠ma̠] should be [a̠ta̠ma̠ꜜ]. — Sartma 【𒁾𒁉 ● 𒊭 𒌑𒊑𒀉𒁲】 09:21, 4 June 2026 (UTC)
- It may be a good idea to represent pitch/accent in the IPA for Japanese, but I would advocate against using "ꜜ" for this purpose, since unfortunately, it is not the correct usage of the IPA symbol (see discussion on Wikipedia: w:Talk:Japanese_phonology#What's_with_the_downward_arrow_being_used_to_mark_accent?).--Urszag (talk) 13:21, 4 June 2026 (UTC)
- @Urszag Oh, yeah, I hate that too, I was just managing resistance there. If I had it my way it would indeed be the acute accent all the way. We've been discussing how to show the accent in this thread too (even though it was about transliterations, not IPA). With the acute accent it would look like this, then:
- [a̠ta̠má̠] or [a̠tá̠ma̠] — Sartma 【𒁾𒁉 ● 𒊭 𒌑𒊑𒀉𒁲】 13:54, 4 June 2026 (UTC)
- The IPA string is intended to represent the consonant and vowel values for the Japanese term, which are (generally) commonly used by all dialects of Japanese. (Note that this does not include the Ryukyuan languages, treated here at the EN Wikt as separate languages, despite often being described as "dialects" in Japanese-language materials.)
- The pitch accent, however, is specific to dialect.
- Due to this detail, in past discussions, we had settled on deliberately not including any pitch accent information in the IPA line. I continue to think that, if we are to include a separate "IPA" line, this line should not include any pitch markers.
- As an alternative notion, we could dispense with the separate "IPA" line, and just include IPA in each listed dialect line. However, in such a case, we should carefully consider:
- Are we going to give a strict phonetic description, in [brackets]? Or a looser phonemic description, in /slashes/?
- The output of
{{ja-pron}}currently -- and mistakenly -- uses [brackets] on each dialect-specific line to output what is actually a looser phonemic description.
- Strict or loose, if we include pitch markers in the output string, is the output visually clear?
- The strict notation especially frequently includes various additional diacritics, and in past discussions, we were worried that diacritic overload too often made the combined "strict IPA + pitch marker" string too visually messy to be useful.
- If we discontinue the use of the downward-arrow marker ꜜ to show the pitch drop, I would prefer it if we replaced that symbol with something rather than just removed it -- it's much more visually clear that "something special happens phonologically at this point" to use a glyph here, rather than just a combining diacritic accent on the vowels -- not least as the acute, grave, and circumflex accents are overloaded and mean different things for different languages, greatly increasing the chance for confusion on a site like Wiktionary that has entries for (ostensibly) all languages.
- ‑‑ Eiríkr Útlendi │Tala við mig 22:59, 4 June 2026 (UTC)
- @Eirikr I'm really not sure about not showing the accent in IPA transcriptions. Accent is not optional in Japanese. We don't do this for any other language where words can be pronounced in different ways. If we keep "Tokyo", "Kansai", "Kyoto", &c. in front of each transcription, we're being clear enough. I don't see how giving a half-baked transcription would help. We already give transliterations at the headword level with no indication of accent. Those should suffice to cover all dialects. I really wouldn't want to compromise on giving precise information on the basis of regional/dialectal pronunciations. If anyone would care to add those at some point, they do have a place to do so (the very pronunciation section).
- If it was up to me, I would only give the phoneMic transcription (the one you call "loose", between slashes
//), deleting most of those diacritics that are not needed to represent phonemes (like the line under the a̠, the + under the ɯ̟, &c. They just complicate things unnecessarily (especially for a language with such a simple phonetic set like Japanese). - If we get rid of all unnecessary diacritics, the acute accent on the accented syllable shouldn't be a problem at all.
- My opinion on this all thing is that no-one will ever learn how to correctly pronounce the Japanese pitch accent just by looking at its representation on paper (whatever the representation), so I'd prefer to use something that's clear, precise, and pertinent to the level of description we're giving, over any other creative or wrong use of IPA symbols. This is especially true in this case, since we've been using the acute (=high pitch!) accent to write Classical Greek's pitch accent for at least 2 millennia in this part of the world, so it would not only be a correct use of the IPA diacritic, but one in line with what we've been already doing elsewhere to show pitch accent since pretty much forever in Western history. — Sartma 【𒁾𒁉 ● 𒊭 𒌑𒊑𒀉𒁲】 09:14, 8 June 2026 (UTC)
- I am open to dispensing with the independent line currently labeled "IPA", so long as the lines for each dialect are clearly marked for that dialect, and themselves have correct notation. My intent in my post just above was to explain why this separate "IPA" line exists in its current form, without pitch markers. ‑‑ Eiríkr Útlendi │Tala við mig 20:54, 8 June 2026 (UTC)
- @Eirikr@-sche This is how I'd like the entirety of the pronunciation section to be:
- That's all you need to know, really. No Atamadaka/nakadata/&c., no numbers, no kana with lines going up and down (it's all redundant information, the same piece of information repeated in different forms). Japanese pronunciation is incredibly simple. This fact should be reflected in the pronunciation section.
- (I should add that Kyoto (Kansai) accent has two phonemic relevant elements, not just one as in standard Japanese. In /átama/ it's not relevant because the accent is on the first syllable, but if the tonic mora is not initial on top of the tonic mora we also need to show whether the word starts high or low. I'm not an expert of Kansai accent and never thought of a way of showing it in IPA, but it should be done somehow). — Sartma 【𒁾𒁉 ● 𒊭 𒌑𒊑𒀉𒁲】 09:30, 9 June 2026 (UTC)
I worry that this notation is too minimal, requiring that readers must already be familiar with Japanese.
IPA high-low pitch-accent markers should be applied consistently to all morae for which pitch is important. This makes it trivial to indicate whether a Kansai-dialect word starts lòw (grave accent) or hígh (acute accent), just by using the correct accent marker on that first mora. This also means that we do in fact need some marker to indicate where the pitch drops, since pitch-drop can happen word-finally, affecting any following particle, and thus cannot be shown clearly using just the IPA high-low pitch-accent markers. This is where the downward arrow ꜜ and numbering notation makes things clear in a way that accent diacritics cannot.
In addition, we must keep aware of the fact that we are not operating in isolation. Many of our readers may refer to various articles on the EN Wiki to read about the Japanese language, and so far, all of the ones I've seen that reference Japanese pitch appear to use the downward arrow ꜜ to show where the pitch drops. See, for instance, w:Japanese_phonology#Pitch_accent or the extensive examples at w:Japanese pitch accent. Diverging from that to come up with our own independent and novel notation system is problematic -- even before we get into issues of applicability (ensuring that whatever we come up with can accurately express patterns for more than just Tokyo-standard Japanese), clarity (accounting for word-final pitch drops), etc.
‑‑ Eiríkr Útlendi │Tala við mig 17:22, 9 June 2026 (UTC)- @Eirikr Well, any notation under any pronunciation section in any language presupposes that the reader is already familiar with that language. It usually also presupposes that the reader is familiar with IPA transcription. The system I favour (and to be clear, I did not invent it; other dictionaries already use it) is “minimal”, but since when has that been a disadvantage?
- Japanese dictionaries do the same when they give you the number of the accented mora, and so does the newest NHK pronunciation and accent dictionary, with the slash after the accented mora; likewise the latest 新選国語辞典 (10th edition), which marks the accented mora instead of giving a number, and the 三省堂国語辞典 (8th edition), which uses a hook (┐), &c. All Japanese dictionaries that mark accent are being clear in giving the only phonemically relevant piece of information, i.e. the accented mora. There is no need to indicate the pitch of any other mora in a word, because it can all be determined automatically once the accented one is given.
- We are currently doing what the Greeks also did millennia ago when they indicated the pitch of every single vowel, writing forms such as Θὲόδὼρὸς (Thèódṑròs), until they realised it was, to quote Allen’s Vox Graeca, "uneconomical and inelegant", and switched to indicating only the acute accent (Θεόδωρος (Theódōros)).
- Marking the pitch of each and every vowel is utterly redundant: everything before the accent is high, everything after is low. It is as simple as that. (The initial lowering of pitch on the first mora is a phrasal phenomenon and should not be included in a word’s phonetic representation, just as we do not mark an accent on the final syllable of every French word.)
- A phoneMic transcription like the one I am proposing (between slashes,
//) only needs to show phoneMes, not phones, so there is no need to mark anything that can be naturally inferred from the phonetics of the language. A single accent mark is both sufficient and desirable. This is what we already do in all other languages: we mark only the accented syllable and leave the unaccented ones unmarked. - As for the idea that we should not work in isolation: I see what you mean, but we already do many things differently from Wikipedia in a number of other respects, so this argument seems rather weak to me. I also believe in copying or maintaining a practice if it is good or correct. Unfortunately, I do not believe this is the case here, so I have no qualms about improving things on our side and letting Wikipedia catch up when it is ready. — Sartma 【𒁾𒁉 ● 𒊭 𒌑𒊑𒀉𒁲】 13:33, 10 June 2026 (UTC)
- > "so there is no need to mark anything that can be naturally inferred from the phonetics of the language"
- I think this may be (at least part of) where you and I diverge on this topic.
- One key difference is that we purport to describe Japanese, not just the NHK broadcast standard, which is based on Tokyo dialect and is the focus of the NHK pronunciation dictionary and the pitch-accent details provided in other monolingual general-purpose Japanese dictionaries. Since our entries are about Japanese, not just Tokyo dialect, whatever notation standard we use must be flexible enough to unambiguously account for more than just Tokyo dialect. What we have used so far appears to cover these more varied bases, so I am loath to jump off of that ship onto some other as-yet-untested conveyance.
- What Greeks did millenia ago is, well, Greek to me, and we cannot assume that it is any more relevant or well-known to other readers of our EN Wikt entries for Japanese terms. 😉
- The Japanese I am most familiar with is Kantō, basically Tokyo dialect. From what I've read about Kansai, observed in my limited intake of Japanese-language video, and what I've heard from friends native to that region, words can have high pitch throughout (high pitch, no accent), or high pitch until a pitch drop (high pitch, accented mora just before the drop), or low pitch until the last mora (low pitch, no accent), or low pitch until a high-pitched mora followed by a pitch drop (low pitch, accented mora followed by the drop). This seems to align with the descriptions at w:Japanese_pitch_accent#Keihan_(Kyoto–Osaka_type).
- → This difference in placement of high and low pitches would seem to mean that the proposed Ancient Greek analog, where only a single high-pitch mora is marked, cannot work for Kansai -- and thus cannot work for Japanese as a whole. ‑‑ Eiríkr Útlendi │Tala við mig 22:28, 10 June 2026 (UTC)
- Based on the description, I believe we could indicate the phonemic distinctions of a Kyoto–Osaka type accentuation system using IPA accent marks like this:
- initial accent: /híga/, /híkara/
- high initial + no accent: /kiga/, /kikara/
- high initial + non-initial accent: /atáma/
- low initial + no accent: /kìga/, /kìkara/
- low initial + non-initial accent: /hàrú/, /hàrúɡa/
- (Using marks for both states of a binary opposition is redundant, and not marking high initial tone in unaccented words is consistent with the statement that "Low tone is considered to be marked".) I don't feel the status quo is necessarily intolerably bad or needs to be changed in any hurry, but I think there's a real problem with using IPA symbols the wrong way, even if it's become common in practice in certain spheres for Japanese accent notation. My original comment was because I would see the change of adding ꜜ to the IPA transcriptions as worse than just leaving it out (as at present) and using it only in the non-IPA individual dialect transcriptions that follow.--Urszag (talk) 23:19, 10 June 2026 (UTC)
- Sorry, I'm a bit confused -- if "Low tone is considered to be marked", shouldn't we mark it? The /ma/ in Kansai /atama/ is lower than the preceding /ta/, for instance. ‑‑ Eiríkr Útlendi │Tala við mig 23:57, 10 June 2026 (UTC)
- In the context of that article, I believe "Low tone is considered to be marked" is supposed to refer to the binary distinction of starting with a low tone vs. starting with a high tone, not saying that low tone is phonologically marked on each mora in which it phonetically occurs. For example, the article does not write /ataꜜ˩ma/: it only uses /˩/ at the start of words in the provided phonemic transcriptions.--Urszag (talk) 00:15, 11 June 2026 (UTC)
- Thank you for clarifying.
- Looking again at the Kansai examples here, I note that "unmarked" is ambiguous:
- In "initial accent", unmarked = low.
- In "high initial + no accent", unmarked = high.
- In "high initial + non-initial accent", unmarked = both high and low.
- In "low initial + no accent", unmarked = both high and low.
- In "low initial + non-initial accent", unmarked = low.
- I find this all very confusing, requiring a good bit of cognitive overhead to puzzle through. I am concerned that this will be clear as mud to other readers less well-versed in Japanese phonotactics. ‑‑ Eiríkr Útlendi │Tala við mig 08:08, 11 June 2026 (UTC)
- @Eirikr It's not confusing, you're making it confusing. The only difference with the Tokyo accent is that a word can be high or low, on top of having an accent or not. It's as simple as that. Since apparently low words are the marked ones, that's what we need to show in an eventual phonemic transcription. We can use the acute for the accent like we do for the Tokyo accent, and something else for "low pitch word". That's what @Urszag did with his examples. If the grave accents confuse you, you can try with a low macron or something (I use Romanizasions here because they are easier to type):
- high - no accent: 気 ki, 風 kaze, 止める yameru
- high - accent: 日 hí, 川 káwa, 白い shíroi, 頭 atáma
- low - no accent: 木 ˍki, 糸 ˍito, 起きる ˍokiru
- low - accent: 春 ˍharú, 薬 ˍkusúri, マッチ ˍmacchí,
- (1) and (2) are basically the same as Tokyo, but in Kansai they don't lower the first mora at phrasal level. Words in (4) work exactly like Classical Greek: all moræ are low apart from the accented one.
- Low no-accent words (3) do have a rise in pitch on the last mora, but that's also at the phrasal level (if you attach a particle, the high pitch moves to the end of the particle), so showing a high mora at the end of low atonic word would also be a misrepresentation. Using the kana + line system you would be misrepresenting how the accent system in Kansai works.
- I know you don't like to hear this, but unfortunately one needs to have prior knowledge of Japanese accent (standard or regional) in any case, if they want to be able to pronounce words and phrases correctly. The kana + line system is not superior nor helps more. In fact, I have been arguing and will continue arguing that it's just a misleading misrepresentation of what Japanese accent actually is, and in my opinion it creates an obstacle to learning it. We're misleading our readers when we give the kana + line representation, actively preventing them from forming a more accurate understanding of what Japanese accent actually is about. — Sartma 【𒁾𒁉 ● 𒊭 𒌑𒊑𒀉𒁲】 12:53, 11 June 2026 (UTC)
- There is no functional difference between the system I described above (/híga/, /kiga/, /atáma/, /kìga/, /hàrú/) and a system that uses the downwards arrow (ꜜ) to mark the position of a pitch drop (/hiꜜga/, /kiga/, /ataꜜma/, /kìga/, /hàruꜜ/). My opinion is that neither is obvious to someone without knowledge of Japanese pitch accent patterns, but at least the former doesn't misuse IPA symbols. I don't care a ton about this, but I agree with Sartma that there's a fundamental conflict between trying to provide phonemic transcriptions, and trying to spell out non-phonemic details for readers who don't know the basic rules of how pitch accent works in the accent in question, and that we should prioritize the former since the phonemic pattern is the lexeme-specific information.
- I agree with Sartma that you have phrased things in a way that increases confusion rather than reducing it. Vowels are left unmarked when their pitch is predictable, and the way to predict it isn't to puzzle through an exhaustive list of combinations like "In "initial accent", unmarked = low. In "high initial + no accent", unmarked = high. In "high initial + non-initial accent", unmarked = both high and low. In "low initial + no accent", unmarked = both high and low. In "low initial + non-initial accent", unmarked = low." There are clearly more parsimonious ways to interpret the Kyoto–Osaka type accentuation system: I'll give one here that uses just three rules for the Osaka dialect system given by Wikipedia. 1) If the first mora has a low tone (notated with the grave accent mark), pitch remains low until the accented mora or until the word's final mora, whichever comes first. 2) If the first mora does not have a low tone (it is unmarked or notated with an acute accent mark), pitch starts high and remains high up through the accented mora or the word's final mora, whichever comes first. 3) If there is an accented mora (marked with the acute accent), pitch is high on the accented mora itself, but low on every mora following it. These rules are not shown by the phonemic transcription, because they are part of the general set of phonological rules for the language, like rules for consonant aspiration in English. If it's considered useful, it is certainly possible here just as it is elsewhere to display a phonemic transcription followed by a more detailed phonetic transcription.--Urszag (talk) 16:57, 11 June 2026 (UTC)
- @Sartma, @Urszag, I think you are both assuming too much foreknowledge on the part of readers.
- This is intended to be informational for people who are not already experts.
- That is one of the reasons for why we decided years ago to give the pitch information in multiple different ways, to match what different resources provide: some sources give the overbar notation, some give the accents + downstep arrow, some give the mora number, and some give the atamadaka etc. descriptors. Our reasoning in past discussions was that it is better to provide more information than less -- WT:NOTPAPER, so there is no driving need to be parsimonious.
- Beyond that, and setting aside the actual main issue of this thread (the position of the non-pitch-marked "IPA" line), I don't see any need to radically rework the notation that we already use. What problem do these various proposals solve? As best I can tell, a large motivator here appears to be subjective opinions on aesthetics. ‑‑ Eiríkr Útlendi │Tala við mig 17:39, 11 June 2026 (UTC)
- Chiming in again. I had only posted the topic here to discuss the one small issue, the output order of
{{ja-pron}}. It seems that this other topic resurfaced again here too. - I still agree with Eiríkr that we should keep the current system for those stated reasons, mainly the sheer practicality and versatility. I don't exactly see how the proposed overhaul that abandons this could be considered an upgrade either, all things considered. It seems like a stretch to say that the current output is some kind of gross misrepresentation of pitch accent. Shlyst (talk) 19:25, 11 June 2026 (UTC)
- Right. To be clear, I support your suggestion of this format:
- To summarize my later comments: I don't agree with Sartma's suggestion of adding "ꜜ" to Japanese IPA transcriptions ("[a̠ta̠ma̠ꜜ]"). If we do desire to mark the accent in IPA, I would support using the IPA acute/grave accents for that, and as far as I see they are sufficient for that purpose. But I'm fine with leaving accent marking only as a feature of the non-IPA dialect transcriptions.--Urszag (talk) 19:43, 11 June 2026 (UTC)
- @Shlyst Sorry for going off-topic. I do still believe that a phonetic transcription without accent notation is incorrect, though. I guess that can be a different thread.
- By the way, I never said the current system is a “gross” misrepresentation. I simply said it is a misrepresentation and that it is misleading, and I explained why.
- If you ever find yourself teaching Japanese with that system and you see students starting every single word with a low pitch or being completely overwhelmed by the redundant information, thinking they can't possibly remember the pitch of each mora and deciding to give up even trying, you'll probably start to see my point. — Sartma 【𒁾𒁉 ● 𒊭 𒌑𒊑𒀉𒁲】 08:07, 12 June 2026 (UTC)
- For what it's worth, I think that there could also be some merit to that theory, and it would influence my stance on this. The concern being: does the kana pitch line system confuse any learners? Japanese pitch accent isn't rocket science. You mention that there can be fundamental misunderstandings of a phonemic system that is otherwise very intuitive and practical. What you're calling "redundant information" I only see as useful information that complements and enhances learners' understanding. It makes it obvious that you're not meant to memorize the intonation of every single mora.
- The argument is similar to arguing that the Chinese Zhuyin tone markers are imprecise because they do not veraciously represent phonetic tone as opposed to the phonemic visuals. For example, 感覺 (ㄍㄢˇ ㄐㄩㄝˊ), the Mandarin third tone (ˇ) is always described and taught as "dip and rise," although in reality, it's just more like a low flat tone.
- Something like a Japanese learner starting every single word with a low pitch is just another example of rigid phonemic exaggeration. Would that alone warrant abolishing the entire system? And well, I'm no Japanese teacher, so I can't attest to how pitch accent is taught to that effect, or if pitch accent even gets taught at all for that matter. Shlyst (talk) 10:48, 12 June 2026 (UTC)
- @Shlyst Unfortunately, I don’t think I have any energy left to argue my case. You will probably take years to realize that I’m right, and nothing I say here now will change your mind. I just feel sorry that many people will be put off learning Japanese pitch accent because we (and other dictionaries using this system) are making it seem unnecessarily difficult and confusing. The overline system is useful for explaining how pitch accent works in Japanese, but not as a method for memorizing the accent of new words. Japanese is like any other language with a phonemic accent system. The only information a learner should need to retain is which syllable is accented, if any. Anything beyond that is just irrelevant cognitive load that discourages learning or leads students to speak unnaturally. — Sartma 【𒁾𒁉 ● 𒊭 𒌑𒊑𒀉𒁲】 — Sartma 【𒁾𒁉 ● 𒊭 𒌑𒊑𒀉𒁲】 11:31, 15 June 2026 (UTC)
- You might take years yourself to realize the practical implications of what you're arguing for. The reason you feel bad for beginners seems to be that, well, they're beginners. It's no more confusing than Chinese tones, or even IPA itself, if a layman who is unfamiliar with IPA uses this website. You cannot expect a beginner learner to have perfect pronunciation, no matter the system. Not to mention that, even granted that we decide to switch to your accent-overhauled Hepburn, we also must accommodate the other dialects of Japanese with differing pitch accent systems under your own system. I'm calling it yours because, although you say you didn't invent it yourself, it would appear that you did here and even documented a whole dossier for it: User:Sartma/Japanese#Sartma Transliteration. Just looking at this, and it is ironically even more convoluted than the pitch overlines.
- We're erring on the side of caution which is why we want to keep using a practical established system, so when you come along and say that your new system is better, that's why we ask the questions:
- Is that really true?
- What about other Japanese dialects?
- Does it clearly mark words with flat pitch?
- What about compound words, and what about words with multiple アクセント句?
- Is it worth making the change?
- Then once you're met with these inevitable questions, you say that you "don't have the energy" to elaborate further, which just leaves nothing to chew on. Shlyst (talk) 15:09, 15 June 2026 (UTC)
- @Shlyst I definitely don’t have the time or energy to respond to your strawman arguments. You already have all your answers in this thread and this one. — Sartma 【𒁾𒁉 ● 𒊭 𒌑𒊑𒀉𒁲】 09:34, 16 June 2026 (UTC)
- All right, but this is frustrating for all of us too. I wasn't even entirely opposed to your idea, but you tapping out at the crux without clearing up what seems to be that you eventually want to monopolize the whole site with your own personal "Sartma transliteration" which I don't think many will agree on especially while you avoid the implications of doing so. Not sure if you forgot or something, but the thread that you've just linked to is one that I was a part of (and what's even funnier is that I created it). Shlyst (talk) 11:46, 16 June 2026 (UTC)
- @Shlyst I don’t see how I can have a constructive discussion with someone who is clearly not reading what I write, consistently misinterpreting it, and repeatedly misrepresenting my position. For example, when have I ever proposed using my own transliteration system in this thread or in the one about theは/わ issue? Never. I have always referred to the system used by Samuel E. Martin in his grammar and dictionaries, and even that was off topic.
- Regarding the discussion on phonetic transcription, I suggested using the acute accent because it is what the IPA uses to mark high pitch. My point was that we only need to indicate the accented syllable, not the high/low pitch of every vowel in a word, as that is what's normally done for other languages.
- Believe me, this is far more frustrating for me. You are simply defending the status quo and seem completely unwilling to consider any argument that does not align with your existing views. I have already spent too much time on this discussion. Unfortunately, that's on me: I really should have know better after previous "discussions" with you. So, goodbye.
- PS: I'm not writing just for you. Others are reading this and might not be aware of the other thread. — Sartma 【𒁾𒁉 ● 𒊭 𒌑𒊑𒀉𒁲】 13:07, 16 June 2026 (UTC)
- > when have I ever proposed using my own transliteration system in this thread or in the one about the は/わ issue? Never.
- You literally used it there. To quote:
- ふくろう【梟】㋗㋺ = fukúrō or fukurô
- Tell me what the heck is a "ô"? Shlyst (talk) 15:51, 16 June 2026 (UTC)
- @Shlyst Again: when have I ever proposed using my own transliteration system? You have some serious reading comprehension problems. — Sartma 【𒁾𒁉 ● 𒊭 𒌑𒊑𒀉𒁲】 15:57, 16 June 2026 (UTC)
- I explicitly asked you how you would like your proposal to display exactly on pages, and that's what you replied. I guess I have serious reading comprehension problems. Shlyst (talk) 16:08, 16 June 2026 (UTC)
- @Shlyst No. What I replied is:
- > [...] my personal preference would be to give an accented Romanisation followed by its IPA transcription, so something like:
- > (Tokyo) fukúrō [ɸɯ̟̊kɯ̟↓ɾo̞o̞]
- > (Tokyo) fukurô [ɸɯ̟̊kɯ̟ɾo̞↓o̞]
- > And if that is too "unusual", then the IPA transcription alone would suffice.
- I told you what my preference was, i.e. a general accented Romanisation followed by its IPA transcription. That's why I said "something like", and not "exactly this". And I even said that if that was too weird the IPA transcription would be enough. I use the circumflex where Samuel E. Martin would use the macron with an acute accent ⟨ō´⟩ (because it's easier to type, more precise, and more elegant) but I never proposed anywhere "to use my own transliteration system" (which is much more than just a circumflex). I clearly said "my personal preference" would be that, but the whole discussion was a brainstorming session. The discussion page of a word is no place for proposals.
- The part you quoted, I wrote when giving examples of what other dictionaries do:
- > Some dictionary now tells you the accented mora insted of a number, so something like:
- > ふくろう【梟】㋗㋺ = fukúrō or fukurô
- Now, can you please stop miquoting, misreading, and misrepresenting what I write? Thanks. — Sartma 【𒁾𒁉 ● 𒊭 𒌑𒊑𒀉𒁲】 16:37, 16 June 2026 (UTC)
- I explicitly asked you how you would like your proposal to display exactly on pages, and that's what you replied. I guess I have serious reading comprehension problems. Shlyst (talk) 16:08, 16 June 2026 (UTC)
- @Shlyst Again: when have I ever proposed using my own transliteration system? You have some serious reading comprehension problems. — Sartma 【𒁾𒁉 ● 𒊭 𒌑𒊑𒀉𒁲】 15:57, 16 June 2026 (UTC)
- All right, but this is frustrating for all of us too. I wasn't even entirely opposed to your idea, but you tapping out at the crux without clearing up what seems to be that you eventually want to monopolize the whole site with your own personal "Sartma transliteration" which I don't think many will agree on especially while you avoid the implications of doing so. Not sure if you forgot or something, but the thread that you've just linked to is one that I was a part of (and what's even funnier is that I created it). Shlyst (talk) 11:46, 16 June 2026 (UTC)
- @Shlyst I definitely don’t have the time or energy to respond to your strawman arguments. You already have all your answers in this thread and this one. — Sartma 【𒁾𒁉 ● 𒊭 𒌑𒊑𒀉𒁲】 09:34, 16 June 2026 (UTC)
- @Shlyst Unfortunately, I don’t think I have any energy left to argue my case. You will probably take years to realize that I’m right, and nothing I say here now will change your mind. I just feel sorry that many people will be put off learning Japanese pitch accent because we (and other dictionaries using this system) are making it seem unnecessarily difficult and confusing. The overline system is useful for explaining how pitch accent works in Japanese, but not as a method for memorizing the accent of new words. Japanese is like any other language with a phonemic accent system. The only information a learner should need to retain is which syllable is accented, if any. Anything beyond that is just irrelevant cognitive load that discourages learning or leads students to speak unnaturally. — Sartma 【𒁾𒁉 ● 𒊭 𒌑𒊑𒀉𒁲】 — Sartma 【𒁾𒁉 ● 𒊭 𒌑𒊑𒀉𒁲】 11:31, 15 June 2026 (UTC)
- Chiming in again. I had only posted the topic here to discuss the one small issue, the output order of
- @Eirikr It's not confusing, you're making it confusing. The only difference with the Tokyo accent is that a word can be high or low, on top of having an accent or not. It's as simple as that. Since apparently low words are the marked ones, that's what we need to show in an eventual phonemic transcription. We can use the acute for the accent like we do for the Tokyo accent, and something else for "low pitch word". That's what @Urszag did with his examples. If the grave accents confuse you, you can try with a low macron or something (I use Romanizasions here because they are easier to type):
- In the context of that article, I believe "Low tone is considered to be marked" is supposed to refer to the binary distinction of starting with a low tone vs. starting with a high tone, not saying that low tone is phonologically marked on each mora in which it phonetically occurs. For example, the article does not write /ataꜜ˩ma/: it only uses /˩/ at the start of words in the provided phonemic transcriptions.--Urszag (talk) 00:15, 11 June 2026 (UTC)
- Sorry, I'm a bit confused -- if "Low tone is considered to be marked", shouldn't we mark it? The /ma/ in Kansai /atama/ is lower than the preceding /ta/, for instance. ‑‑ Eiríkr Útlendi │Tala við mig 23:57, 10 June 2026 (UTC)
- Based on the description, I believe we could indicate the phonemic distinctions of a Kyoto–Osaka type accentuation system using IPA accent marks like this:
- If it was up to me, I would only give the phoneMic transcription (the one you call "loose", between slashes
- @Eirikr I'm really not sure about not showing the accent in IPA transcriptions. Accent is not optional in Japanese. We don't do this for any other language where words can be pronounced in different ways. If we keep "Tokyo", "Kansai", "Kyoto", &c. in front of each transcription, we're being clear enough. I don't see how giving a half-baked transcription would help. We already give transliterations at the headword level with no indication of accent. Those should suffice to cover all dialects. I really wouldn't want to compromise on giving precise information on the basis of regional/dialectal pronunciations. If anyone would care to add those at some point, they do have a place to do so (the very pronunciation section).
- @Urszag Oh, yeah, I hate that too, I was just managing resistance there. If I had it my way it would indeed be the acute accent all the way. We've been discussing how to show the accent in this thread too (even though it was about transliterations, not IPA). With the acute accent it would look like this, then:
- It may be a good idea to represent pitch/accent in the IPA for Japanese, but I would advocate against using "ꜜ" for this purpose, since unfortunately, it is not the correct usage of the IPA symbol (see discussion on Wikipedia: w:Talk:Japanese_phonology#What's_with_the_downward_arrow_being_used_to_mark_accent?).--Urszag (talk) 13:21, 4 June 2026 (UTC)
Support. All other languages (except one) already display IPA first and then any additional information, plus indeed it scrambles the order when there's dialectal pitch information. lattermint (talk) 12:19, 4 June 2026 (UTC)
- For this reason especially I agree. Kiril kovachev (talk・contribs) 13:03, 4 June 2026 (UTC)
- I don't edit Japanese, but FWIW I too prefer putting IPA first. If accent is added to the IPA as suggested above, would all IPA be given before all other lines, like this, or would things be nested (the way some entries nest e.g. British audio under the British IPA line and American audio under the American IPA line), like this? - -sche (discuss) 21:50, 4 June 2026 (UTC)
Moving WT:AXXX shortcuts to WT:EG:xxx
[edit]Requested by @Juwan on my talk page. I figure this is noncontroversial but best to ask. I'll run the job in a week if there are no objections. Saph (talk) 06:32, 6 June 2026 (UTC)
- Still keep the shorter shortcuts, make brand new shortcuts if you wish Hftf (talk) 07:07, 6 June 2026 (UTC)
- Agree with Hftf; perhaps way later we could discuss deprecating the older shortcut. We must be mindful of veteran editors. TranqyPoo [💬 | ✏️] 18:02, 6 June 2026 (UTC)
- Not to mention old discussions in forums and talk pages. While we wouldn't want to keep every obsolete link blue, this was a very basic part of the architecture of the site for many years, there's no potential for confusion, and redirects are cheap. Chuck Entz (talk) 19:08, 6 June 2026 (UTC)
- May I suggest that the new shortcuts be
WT:EG:xxx, rather thanWT:EG:XXX? This would be in keeping with the variousWT:RE:xxxshortcuts for requested-entries pages. 0DF (talk) 19:21, 6 June 2026 (UTC)
- This highlights another difference between the old shortcuts and the new: AXXX was all one block of letters with "A" for "About" and XXX for the language name "AEN" being an abbreviation of "About English". EG:XXX, on the other hand, rearranges things: "EG:EN" is not an abbreviation for "EG:English" or even "Entry guidelines:English". It may not be desirable to have abbreviations where the first part is variable in length ("ETEG", "ETTEG", etc., not to mention those based on XXX-XXX codes), so I'm not sure we could avoid it- but it does make things less straightforward. The separation between the "EG" and the "XXX" does make it less important to have them match as far as case, though. Chuck Entz (talk) 20:24, 6 June 2026 (UTC)
- I think that is the intended request, as Juwan's post indicates such, and likely a typing mistake of Saph. TranqyPoo [💬 | ✏️] 20:35, 6 June 2026 (UTC)
- Yes, that's what I meant, sorry; in the offsite request Juwan brought up consistency with shortcuts like AP:BIB:pt too. Saph (talk) 01:46, 7 June 2026 (UTC)
- How on earth can be it "uncontroversial" to remove a set of shortcuts that has been used for years? — SURJECTION / T / C / L / 21:15, 6 June 2026 (UTC)
- supporting as requester. for the older shortcuts, I say to keep them for backwards compatibility, as redirects are cheap. these will work, just not be displayed at the main page. Juwan 🕊️🌈 06:53, 7 June 2026 (UTC)
- I support this as I was first to propose it in 2025. About WT:AXX shortcuts, I oppose their deletion for backwards compatibility, but I believe they should be removed from the
{{shortcut}}calls on the policies themselves and should not be newly created for new policy pages. Catonif (talk) 21:51, 8 June 2026 (UTC) - calling on @Saph to go forth with the proposal. as my read on the discussion is generally positive: create the new redirects and replace them in the shortcut box, all while keeping the old ones, just not creating more of them. Juwan 🕊️🌈 18:23, 1 July 2026 (UTC)
Literally vs Sum of parts (glossary)
[edit]What is the difference between literally (like the first meaning of go through) and SoP? JMGN (talk) 12:49, 7 June 2026 (UTC)
- Not much AFAICT. Same as
{{&lit|en|go|through}}(Used other than figuratively or idiomatically: see go, through.). DCDuring (talk) 19:20, 7 June 2026 (UTC)
- SoP parsings are often literal, but not all literal uses are a sum of any parts. For example, if I say that her feet are gnarly, I may mean that they are weirdly cool or that they are physically twisted/deformed. Thus SoP is a subset (subclass) of the things [in the universe] that can be literal. Additionally, there are terms that are viewable as SoP even though they involve some figurativeness (but no idiomaticness, though, in the sense of etymonically undiscoverable meaning). For example, the compound noun "cognitive flexibility" is viewable as SoP in the respect that it is merely "cognitive" (adj) plus "flexibility" (n), although "flexibility" is being used in a nonliteral sense within that compound noun (i.e., its abstract sense not its physical sense). Quercus solaris (talk) 22:33, 8 June 2026 (UTC)
- I think I agree. DCDuring (talk) 23:50, 8 June 2026 (UTC)
- well-said. considering adding a note like this at WT:IDIOM. Juwan 🕊️🌈 18:25, 1 July 2026 (UTC)
- SoP parsings are often literal, but not all literal uses are a sum of any parts. For example, if I say that her feet are gnarly, I may mean that they are weirdly cool or that they are physically twisted/deformed. Thus SoP is a subset (subclass) of the things [in the universe] that can be literal. Additionally, there are terms that are viewable as SoP even though they involve some figurativeness (but no idiomaticness, though, in the sense of etymonically undiscoverable meaning). For example, the compound noun "cognitive flexibility" is viewable as SoP in the respect that it is merely "cognitive" (adj) plus "flexibility" (n), although "flexibility" is being used in a nonliteral sense within that compound noun (i.e., its abstract sense not its physical sense). Quercus solaris (talk) 22:33, 8 June 2026 (UTC)
Plural lemma forms in Belarusian (and potentially all East Slavic languages)
[edit]Could we allow plural as a lemma form for Belarusian (and potentially all East Slavic languages), for words that are mostly used in plural?
Most other dictionaries do this, but we seem not to. E.g. we list чараві́к (čaravík, “shoe”) under singular form, but most Belarusian dictionaries would list it under its plural form, чараві́кі, and give singular form after it.
It looks even stranger for words like суро́к (surók, “evil eye”), ве́да (vjéda, “(piece of) knowledge”), прысма́к (prysmák, “delicacy”), which are overwhelmingly used in plural. In fact, they are listed as pluralia tantum in most Belarusian dictionaries. Singular forms are attestable per WT:CFI (I have 3 examples for singular forms), so probably should be added — but many people would consider them wrong (e.g. my Belarusian teacher was very unhappy when she heard a singular form of прысмакі on TV). It feels very strange to add such rare proscribed forms as the main form, instead of the plural form.
I suggest allowing plural forms as a lemma form in such cases.
The current headword is like this:
- суро́к • (surók) m inan (genitive суро́ка or суро́ку, nominative plural суро́кі, genitive plural суро́каў)
I propose having a headword like this:
- суро́кі • (suróki) m inan pl (genitive plural суро́каў, nominative singular суро́к, genitive singular суро́ка or суро́ку)
(This proposal is only for East Slavic languages, to bring Wiktionary in line with how other East Slavic dictionaries handle such cases.
It will, necessarily, bring it further away from other Slavic languages, which do use singular forms in such cases, e.g. Czech bota, Polish but — these languages seem not to have a similar tradition.)
WT:Lemmas allows plural lemmas for Welsh, so I guess allowing it for East Slavic languages could be OK, too. Would you be OK with this?
Note that this would require changes to templates like {{be-noun|...}}, which currently don't have an option for non-pluralia tantum plurals.
What do you think?
(Pinging Belarusian editors: @Ssvb, @Kanstancin9, @Plaga med.) Хтосьці (talk) 17:29, 7 June 2026 (UTC)
- We already have some plural-only headwords, such as людзі or содні. --Ssvb (talk) 19:17, 7 June 2026 (UTC)
- These cases are a bit different from what I’m talking about:
- со́дні (sódni) is pluralia tantum, which has no singular form at all.
- лю́дзі (ljúdzi) and чалаве́к (čalavjék) have semantic difference: лю́дзі (ljúdzi, “others”) is effectively a pluralia tantum (чалаве́к (čalavjék) cannot mean “other”), while in other meanings it’s honest plural of чалаве́к (čalavjék, “person”). These forms have non-overlapping meanings so can be treated as two different words.
What I’m talking about are cases where (a) both singular and plural exist, but plural vastly outweights singular, (b) there is no significant difference in meaning. This includes:- words for things that normally come in pairs or many portions:
- words for “almost pluralia tantum”, that are almost always plural (but nevertheless have their singular forms attested with little difference in meaning):
- ве́ды (vjédy, “knowledge”) is marked as pluralia tantum in TSBM («адз. няма.»), but it has attestable singular forms per WT:CFI:
- 1923, Якуб Колас, “Што яны страцілі”, in Збор твораў, Minsk: Мастацкая літаратура, published 1972:
- Толькі тая веда робіцца нашым сталым здабыткам, калі мы даходзім да яе, здабываем яе самі.
- Tólʹki taja vjeda róbicca našym stalym zdabytkam, kali my daxódzim da jaje, zdabyvajem jaje sami.
- That knowledge becomes our permanent possession only when we reach it, get it ourselves.
- 1924, Кузьма Чорны, Быльнікавы межы:
- Важна тое, што ў віхрастай галаве ёсць веда: «На свеце існуе тое, чалавекам створанае, чым можна заглушыць воўчыя галасы ў быльнікавых межах і саламяны крык вёскі пад ударамі буйнага ветру».
- Važna tóje, što w vixrastaj halavje joscʹ vjeda: “Na svjecje isnuje tóje, čalavjekam stvóranaje, čym móžna zahlušycʹ vówčyja halasy w bylʹnikavyx mježax i salamjany kryk vjóski pad udarami bujnaha vjetru”.
- What’s important is that there is knowledge in its tufted head: “In the world, there is something human-created that can both silence the wolves’ voices in the mugwort boundaries and the straw cry of the village under the blows of the strong wind.”
- 2023, Ганна Янкута, Час пустазелля, Warsaw: Выдавецтва Янушкевіч, →ISBN, page 176:
- Усё не так, як падказвае штодзённая веда, і толькі ўяўленне дапамагае з гэтым асвоіцца.
- Usjó nje tak, jak padkazvaje štodzjonnaja vjeda, i tólʹki wjawljennje dapamahaje z hetym asvóicca.
- Everything is different from what our day-to-day knowledge suggests us, and only imagination helps us to get used to this.
- ве́ды (vjédy, “knowledge”) is marked as pluralia tantum in TSBM («адз. няма.»), but it has attestable singular forms per WT:CFI:
- Хтосьці (talk) 21:16, 7 June 2026 (UTC)
- These cases are a bit different from what I’m talking about:
- Generally, I support the proposition only if it means, that we follow the tradition of most academical dictionaries. I see, that it is true for the TSBMs, TSBLM and BRS 2012. At the same time, I see, that all bilingual dictionaries don't follow this tradition, which should be also taken into consideration. Also, we need always to have the source indicated, showing the need for plural form as the main lemma form. This way there will be no unnecessary debates in the community, for which words we should apply this rule, and for which not. By default, we should apply singular for all lemmas, I think. Plaga med (talk) 21:39, 7 June 2026 (UTC)
- I strongly oppose the idea of copying existing dictionaries word-for-word. I think that instead we need to formulate general principles as a rules, like it’s done for Welsh in WT:Lemma, and then follow this general principle. Хтосьці (talk) 22:03, 7 June 2026 (UTC)
- There's Belarusian GrammarDB dictionary under the Wiktionary-compatible Creative Commons Attribution/Share-Alike 4.0 license. It's maintained by the Belarusian official government linguists and its latest release is 2026.01 at the moment. A web interface for this dictionary is available at https://bnkorpus.info/grammar.be.html and alternatively also at https://verbum.by/grammardb. And, for example, here's the "сурокі" entry: https://verbum.by/grammardb/сурокі
- I think it makes no sense not to use it, since we have no license-related of copyright-related obstacles. And picking the same lemmas would reduce our workload. That said, the "сурокі" entry is plural-only in GrammarDB (the singular form is basically proscribed by the official academics), yet you found citations for the singular form in the works of the reputable Belarusian authors, spanning over a century. We have no obligation to obey the academics and we have an agency of our own, so the singular form can be added to Wiktionary. But I agree that treating the plural form as the lemma is reasonable, since it's much more common.
- As for the
{{be-noun}}template, I can implement an alternative way of adding declension tables, something like{{be-grammardb-202601-noun}}or even{{be-grammardb-202601}}(with the automatic part of speech identification) via importing the whole GrammarDB data into Wiktionary as a Lua module. In the past I implemented Module:User:Ssvb/ru-accentdict/documentation (rejected by the Russian editors), which was relying on the Bloom filter data structure. Now I believe that the DAFSA data structure is superior (such as the CFSA2 format from Morfologick), because it provides better compression and can allow all lemma's wordforms retrieval. What do you think? --Ssvb (talk) 16:28, 8 June 2026 (UTC)- I'm not a fan about using GrammarDB directly. I think it's important that Wiktionary is editable, so that people can improve it.
- be-noun is easier to edit and improve than a hypothetical be-grammardb-202601-noun, so I think we should keep be-noun. I think having a black-box non-customizable template is not a good idea.
- I think a better way forward would be some mechanism to convert GrammarDB data to be-noun, and storing be-noun in actual articles (maybe via a JS widget, maybe via some {{subst:...}} magic). Хтосьці (talk) 21:25, 8 June 2026 (UTC)
- The
{{be-noun}}template is not easier to edit, as it has a rather steep learning curve. What I propose is the possibility to quickly get the declension table in a Wiktionary article in the same way how it looks in GrammarDB, without having to provide anything other than just the lemma string as a template parameter. GrammarDB itself is a publication in an electronic format, not very much different from a paper dictionary. It's maintained by the government linguists, the same people who publish paper dictionaries. A Wiktionary editor can check the GrammarDB's declension table, and if it looks correct, then just use it directly. The technical challenge is how to ensure fast lookups and low memory footprint, but I think that it's going to be fine (and it's a topic for Grease pit). - And, of course, the existing
{{be-noun}}is not going to be removed. It's necessary for the cases when we are not happy about the information from GrammarDB and want to correct something. - A JavaScript widget could possibly generate
{{be-ndecl-manual}}, but generating correct{{be-noun}}parameters in all cases may be tricky. --Ssvb (talk) 02:36, 9 June 2026 (UTC)- "is not easier to edit, as it has a rather steep learning curve" — yes, but if you want to edit some error from GrammarDB, you need to learn it anyway, and on top of that, you need to first convert GrammarDB to be-noun.
- I think the end result of this will be: new editors won't even touch such headwords, and won't try to fix them. And I think it's an undesirable end result. Хтосьці (talk) 06:43, 9 June 2026 (UTC)
- "you need to first convert GrammarDB to be-noun" — to be-ndecl-manual actually, and this is easily doable.
- "new editors won't even touch such headwords, and won't try to fix them" — they can ask questions about how to fix such headwords.---Ssvb (talk) 03:23, 12 June 2026 (UTC)
- I see Wiktionary as a source of its own, on par with GrammarDB, TSBM, etc. If you embed GrammarDB into Wiktionary, it’s no longer an independent source, is it?
- I would very much prefer it to stay independent, so that any data we add is separately curated. For that goal, having a separate data format (like used in
{{be-noun}}) is desirable: it’s like a filter. Right now, you can’t just copy data from GrammarDB, you need to read it before copying, and I think that’s a good thing: it forces you to verify data before adding it here. - I can't prevent you from adding these
{{be-grammardb-202601-noun}}templates, but I would very much prefer them not to exist. Then people won’t verify GrammarDB data, they will just copy it. - Current solution:
- copying data is harder (requires an extra step: rewriting in a different format),
- verifying and editing data is easier.
- Your proposed solution:
- copying data is easy,
- verifying and editing data is harder (requires an extra step: rewriting in a different format).
- I think the current solution is better! I don't want people to uncritically copy any other source into Wiktionary, so making this easier is not a good thing. I want people to think about what they’re coping, double-check and fix errors, so adding an extra step there is a bad thing.
- Like, sure, it’s doable. But you’re adding an extra step in wrong place, IMHO. We shouldn’t encourage uncritical copying, IMHO. We should encourage editing and making the entries our own, based on our research.
- But anyway, I'm not sure this is the best place to discuss such a thing. Хтосьці (talk) 07:17, 12 June 2026 (UTC)
- The
- I support limited use of grammatdb, but take a note, that while it is pretty usefull tool, it is not an official prescription, it is mainly a draft for future dictionaries. If you'll check the word "чаравік", you'll see, that grammardb does not answers the proposed issue, it is always using the singular form for words, that have it in the DB. The fact that "сурокі" is only plural only means, that creators of the DB did not added the singular form (yet?). Plaga med (talk) 21:39, 8 June 2026 (UTC)
- But it is de-facto the official prescription of the government linguists ("Інстытут мовазнаўства імя Якуба Коласа" Цэнтра даследаванняў беларускай культуры, мовы і літаратуры НАН Беларусі). GrammarDB incorporates data from a number of paper dictionaries, published earlier by pretty much the same institution (or its soviet predecessors). The https://verbum.by/grammardb/чаравік entry lists sources as
krapivabr2012,nazounik2008,sbm2012,tsblm1996,tsbm1984. Each of these identifiers refers to an actual published paper dictionary. - The initial 2021's release of GrammarDB had an ISBN number assigned to it, and it was probably published as a booklet with a CD disk (that's just my guess, I don't have a real copy of it): http://www.iml.basnet.by/publikacyi/hramatycnaja-baza-bielaruskaj-movy - "Граматычная база беларускай мовы / уклад. Уладзімір Кошчанка, Алесь Булойчык. - Электронныя тэкставыя даныя і праграма (350 Мб). - Мінск: Тэхналогія, 2021. - 1 электронны аптычны дыск (CDR). Сістэмныя патрабаванні: HTML5-сумяшчальны браўзер. ISBN 978-985-458-321-1., DOI: 10.13140/RG.2.2.17868.33929"
- Yes, GrammarDB treats the singular form of "чаравік" as the lemma (just like the academic paper dictionaries). And my opinion is that Wiktionary has no good reason to diverge. We can follow the GrammarDB's choice of lemmas. --Ssvb (talk) 12:58, 9 June 2026 (UTC)
- "We can follow the GrammarDB's choice of lemmas" — GrammarDB has some pretty questionable choices like splitting чарніцы 'bilberries' and чарніца 'single belberry', I don't think it's a good idea to follow it uncritically.
- I think it should be treated as one of many sources to consider, not as something to copy directly. I'm in general very sceptical about copying anything directly from one source. Хтосьці (talk) 16:39, 9 June 2026 (UTC)
- "(just like the academic paper dictionaries)" — this is not true, BTW. GrammarDB diverges from e.g. TSBM in this. Хтосьці (talk) 19:32, 9 June 2026 (UTC)
- What's not true? The "Grammatical dictionary of the noun"
- Граматычны слоўнік назоўніка / Нац. акад. навук Беларусі, Цэнтр даслед.
- беларус. культуры, мовы і літ., філ. «Ін-т мовы і літ. імя Якуба Коласа і Янкі Ку-
- палы»; уклад. Г. У. Арашонкава [і інш.] ; навук. рэд. В. П. Русак. – 2-е выд.,
- дапрац. – Мінск : Беларус. навука, 2013. – 1245 с.
- ISBN 978-985-08-1559-0
- diverges from TSBM too. It treats the singular form of чаравік as the lemma. Moreover, if shoes are a pair, then e.g. human legs form a pair too. This may quickly become ridiculous if classification exists for the sake of classification. My take on it is that the lemma choice isn't probably super important. If only the plural form exists for some word, then the non-existing singular form obviously can't be the lemma. In all other cases the singular form is likely preferable. But if the singular form is exceptionally uncommon and difficult to attest, then the plural form can be the lemma, at least provisionally. No big deal in my opinion. --Ssvb (talk) 06:31, 10 June 2026 (UTC)
- My take is that grammatical dictionaries (GrammarDB, Граматычны слоўнік назоўніка) have different goals from explanatory and bilingual ones, and are not very relevant for our case. Russian Zaliznyak's dictionary is similarly quite divergent from mainstream Russian dictionaries.
- Чаравікі is maybe not the best example, туфлі would be a better one. My understanding is: because plural is so overwhelming, many people would have trouble getting a correct singular (that it's туфель and not туфля), so listing the more common plural form makes it easier to find the word. I think this is the rationale behind listing footwear in plural. Legs don't have this problem.
- While I agree than in the grand scheme of things it's not important. Maybe we can not add anything to the rules, and just follow the de facto example of Ukrainian. If Ukrainian editors do this without changing WT:Lemmas explicitly, so can we. Хтосьці (talk) 07:04, 10 June 2026 (UTC)
- WT:What_Wiktionary_is_not (not a paper dictionary). There's no need to first guess the term's lemma and then flip pages, navigating through the alphabetically sorted headwords. Any word form can be searched directly, and this search will land on a hit in the term's declension table. This works at least for the Cyrillic spelling. Łacinka is a bit more problematic (currently we only provide red link lemmas in the "Alternative forms" section to serve as search anchors). --Ssvb (talk) 12:36, 10 June 2026 (UTC)
- That's not an argument for the singular, because similarly plural headword can be easily opened if someone uses singular form. Хтосьці (talk) 18:30, 10 June 2026 (UTC)
- Ultimately it comes down to whether we choose more uniformity (singular) or Belarusian lexicographic tradition (plural). I don't think we'll come to an agreement here, we just prioritise different things.
- Anyway, I think Ukrainian precedent is enough for me to add things like сурокі, веды in the plural. I hope that one is uncontroversial enough.
- As for footwear etc., that decision can be postponed, I guess. Хтосьці (talk) 20:08, 10 June 2026 (UTC)
- WT:What_Wiktionary_is_not (not a paper dictionary). There's no need to first guess the term's lemma and then flip pages, navigating through the alphabetically sorted headwords. Any word form can be searched directly, and this search will land on a hit in the term's declension table. This works at least for the Cyrillic spelling. Łacinka is a bit more problematic (currently we only provide red link lemmas in the "Alternative forms" section to serve as search anchors). --Ssvb (talk) 12:36, 10 June 2026 (UTC)
- What's not true? The "Grammatical dictionary of the noun"
- But it is de-facto the official prescription of the government linguists ("Інстытут мовазнаўства імя Якуба Коласа" Цэнтра даследаванняў беларускай культуры, мовы і літаратуры НАН Беларусі). GrammarDB incorporates data from a number of paper dictionaries, published earlier by pretty much the same institution (or its soviet predecessors). The https://verbum.by/grammardb/чаравік entry lists sources as
- I strongly oppose the idea of copying existing dictionaries word-for-word. I think that instead we need to formulate general principles as a rules, like it’s done for Welsh in WT:Lemma, and then follow this general principle. Хтосьці (talk) 22:03, 7 June 2026 (UTC)
- Here are some examples in Ukrainian:
- консе́рва (konsérva) (etymology 2), nominative singular non-lemma form of консе́рви (konsérvy) (etymology 1).
- за́сіб ма́сової інформа́ції (zásib másovoji informáciji), nominative/accusative singular non-lemma form of за́соби ма́сової інформа́ції (zásoby másovoji informáciji).
- негара́зд (neharázd) (noun), nominative/accusative singular non-lemma form of негара́зди (neharázdy).
- Voltaigne (talk) 22:00, 7 June 2026 (UTC)
- Oh, nice! Thanks for these example. So, de facto this practice already exists for Ukrainian — all the more reasons to expand it to Belarusian. Хтосьці (talk) 22:04, 7 June 2026 (UTC)
- I propose adding this formulation to WT:Lemmas#Nouns:
- Belarusian: usually singular, but can be plural for words that are chiefly used in plural, when such description follows the established lexicographical tradition (e.g. footwear, some portioned dishes)
- Would you be OK with that? Хтосьці (talk) 23:07, 7 June 2026 (UTC)
- I agree, but when you say "established lexicographical tradition" we need to rely on some source of truth. This is what I wrote about above in my initial reaction to the proposal. But from your reply above I do not understand your position on how we will show the established lexicographical tradition. Maybe I don't understand something. Could you please clarify your position on that? Plaga med (talk) 21:42, 8 June 2026 (UTC)
- I've tried to incorporate your feedback, that's why I've added "established tradition". I personally would prefer not to reference it at all; I've only added this condition based on your comment.
- As for showing it, we do it using the common consensus mechanisms, as with everything in Wiktionary. We discuss and eventually come to conclusion, like we are doing now. This proposal doesn't include any changes to this. Хтосьці (talk) 22:07, 8 June 2026 (UTC)
- (To be clear: my preferred way of formulating this would be to list categories directly: paired items, portioned food, berries, abstract concept primarily used in plural.) Хтосьці (talk) 22:10, 8 June 2026 (UTC)
- "Could you please clarify your position on that" — expanding on the answer above, I think such usage of plural headwords would be obvious and noncontroversial.
- And if it's not — then we can discuss it on case-by-case basis, and evaluate arguments as we go. (Although I don't expect this will be needed.) Хтосьці (talk) 22:17, 8 June 2026 (UTC)
- I agree, but when you say "established lexicographical tradition" we need to rely on some source of truth. This is what I wrote about above in my initial reaction to the proposal. But from your reply above I do not understand your position on how we will show the established lexicographical tradition. Maybe I don't understand something. Could you please clarify your position on that? Plaga med (talk) 21:42, 8 June 2026 (UTC)
- @Ssvb @Plaga med Please have a look at what I’ve done in сурокі/сурок (and in
{{be-noun}},{{uk-noun}}, Module:uk-be-headword). Feel free to undo my changes if you don’t agree with them. Хтосьці (talk) 21:09, 10 June 2026 (UTC)- Generally, I agree to follow the Ukrainian case. I would keep it as you proposed. Thank you for opening this discussion and making the contribution! I think, that if we will see in practice, that we need to use more concrete rules with clear classification of lemmas, or we will meet any other difficulties and unclear moments, we can come back to this discussion later. Plaga med (talk) 13:42, 11 June 2026 (UTC)
- This is fine. This also happens to match the lemma choice of GrammarDB, so there's no controversy. The added singular form is attested by quotations and that's also fine.
- The Wiktionary:Lemmas#Nouns page is careful enough to use the words "usually" and "normally", it does not strictly prohibit such special unusual entries in Belarusian or Ukrainian per se. Would you rather add something to WT:ABE? Which is supposed to be the most important information about editing Belarusian entries, consolidated in one place. --Ssvb (talk) 03:54, 12 June 2026 (UTC)
- Thanks!
- I don't think we have reached a consensus that is worth adding to WT:ABE, have we?
- To recap, here's our positions how I understand them:
- Ssvb: it's better to follow GrammarDB. We need a good reason to deviate from it.
- Plaga med: it's better to follow a curated set of normative dictionaries. We need a good reason to deviate from them.
- Хтосьці: it's better to formulate rules based on most pre-existing dictionaries (ignoring outliers like GrammarDB) and then follow them.
- There's little overlap, is there? So, I guess, this needs to remain an open question.
- Maybe other Belarusian editors will decide. Maybe we'll have another discussion when we have more arguments. Idk. But for now I don't see much overlap between our positions. Хтосьці (talk) 06:55, 12 June 2026 (UTC)
- I agree. Overall, I am not oposing completely to the idea of making some rules for Wiktionary in future if there will be consensus and some strong argumentation. And also I hope, that we will not sink in some discussions in future. This is also why I think that it is better to be carefull with making such rules, and why overall I think that it is better to stick to the dictionaries by default and try to not invent airplanes. Plaga med (talk) 22:24, 13 June 2026 (UTC)
- @Хтосьці If I understand it correctly, the two of you questioned the credibility of GrammarDB (due to its electronic format or whatnot), and only respect paper editions of the dictionaries. I guess, it's a topic for a separate discussion.
- As for formulating the rules, the choice of lemma (singular vs. plural) differs for some terms in different academic paper dictionaries. You mentioned potentially easier search as an argument in favor of some plural lemmas (e.g. footwear), but I don't think that this advantage is applicable to Wiktionary due to its electronic format and advanced search capabilities. Do you see any other advantages?
- I suggest the following guideline: "If at least one Belarusian academic paper dictionary uses a valid attestable singular form as the lemma, then use this singular form as the lemma in Wiktionary as well." Are there any objections? --Ssvb (talk) 20:53, 15 June 2026 (UTC)
- Yes, I strongly object to this.
- My problem with GrammarDB is not that it's digital. My problem with it is that it's an outlier (in this question): it makes choices that are different from the choices of most other dictionaries.
- This is the same reason why I object to your formulation: it will allow using any single outler dictionary to force the choice of the lemma form.
- We should not be beholden to whims of any single editor or edition, we should look at the broad picture. I think we need to check how most dictionaries do this: not follow any single one, but follow the tradition shared by most dictionaries. Хтосьці (talk) 22:48, 15 June 2026 (UTC)
- agree here + as, I've pointed out before, the fact that it is developped by the Institute and was published in CD-format does not make it the single source of truth, if we're speaking about the standarts in lexicography. This source have proven to be very usefull in many cases, but still it was made as an auxillary tool for spellcheckers, specialists, and creators of dictionaries. GrammarDB is not ready to replace the academical grammar dictionaries yet. We can check our assumptions with it, but we cannot just follow every statement there blindly, also considering that this source is not ideal and have some minor issues in it. Plaga med (talk) 18:38, 16 June 2026 (UTC)
- "You mentioned potentially easier search as an argument in favor of some plural lemmas" — the truth is, I expected this to be an obvious question. Like, "obviously everyone would expect to use plural lemma for footwear" is what I was thinking. For me it's like using nominative case or infinitives for lemmas — it's just something I take for granted.
- The reasons for search are my attempt at rationalising it. The real reason is that for me, plural lemma is just an obvious and unmarked default. It's how I conceptialise this in my head, based on using dictionaries in the past. It's common sense for me.
- During the conversation we've found out that you see it differently. And, like, OK. I've learned that my common sense is not that common. I still don't think we should codify the reverse option. I still very much prefer the plural. Хтосьці (talk) 23:04, 15 June 2026 (UTC)
- @Ssvb @Plaga med I’ve just thought of something... I wouldn’t use «*два сурокі» 'two jinxes', «*дзве веды» 'two pieces of knowledge'. That sounds wrong to me. I also can’t find examples of such usage.
- This would mean that суро́к (surók), ве́да (vjéda) are not singular forms of суро́кі (suróki), ве́ды (vjédy). They are synonyms / alternative forms: we have plural form, and we have uncountable singular form — but they're not part of the same paradigm! You can say «гэтыя сурокі», you can say «гэты сурок», but you can't count сурокі as if it's a countable noun!
- I’m planning to make this change:
- make суро́к (surók), ве́да (vjéda) uncountable singular nouns,
- remove ability to specify plural headwords, as it was before this discussion (since суро́кі (suróki), ве́ды (vjédy) was the only case on which we had consensus, if we change that case to alternative forms / synonyms, basically we have no consensus left).
- Objections? Хтосьці (talk) 07:56, 17 June 2026 (UTC)
- BTW, some information can be found in the old textbooks ("суроцы" is mentioned there): https://be.wikisource.org/wiki/Старонка:Сынтакс_беларускае_мовы_(1926).pdf/104 And the word урок (urók) might be an alternative form of "сурок" / "сурокі" / "суроцы" too.
- Another observation is that "капцы" (end, death) was supposed to be plural too, but today the singular form капец appears to be predominantly in use (possibly mimicking the Russian language). We observe the ongoing process of Russification of the Belarusian language, erasing or changing its original features. Maybe this can be explained in the usage notes of the relevant entries? --Ssvb (talk) 10:21, 17 June 2026 (UTC)
- Yeah, суро́цы (surócy) is worth adding, although I'd label it "dialectal".
- I also have a hunch that singular суро́к (surók) is influence of Russian сглаз (sglaz), but I'd prefer not to rely on hunches. Although I guess we could hedge it with "probably" and add that information.
- But I don't think treating уро́кі (uróki) as alternative form of суро́кі (suróki) is a good idea. It's a prefixed form, they often have similar meaning. There are a lot of similar words that only differ in prefix. If we treat forms different as prefix as alternative forms, where do we stop? I suggest not treating prefixed forms as alternative forms. Хтосьці (talk) 11:00, 17 June 2026 (UTC)
- I guess, such words should be used with such numerals, like "двое", "трое" etc. Plaga med (talk) 13:34, 17 June 2026 (UTC)
Proposed methodology for documenting Orok (Uilta) language entries
[edit]- Discussion moved from Wiktionary talk:Beer parlour.
I would like to formalize and open a discussion regarding the methodology for adding and updating entries for the Orok (Uilta) language (ISO 639-3: oaa), a critically endangered Southern Tungusic language. Given the extreme scarcity of active speakers and the lack of a standardized, legally binding orthography, documenting this language requires cross-referencing modern copyrighted materials and historical public-domain texts.
To strictly comply with Wiktionary's policies on copyright, original research, and attestation, the following systematic workflow is proposed:
1. Orthography and Phonetic Transcription
- Issue: Extant dictionaries and textbooks (e.g. Ikegami 1997; Ozolinia 2003, Ikegami 2008) employ distinct, non-standardized Cyrillic or modified Latin scripts, which are subject to individual editorial copyrights.
- Solution: No direct replication of any single modern editor's customized orthographic system will be performed. Instead, lexicographical data will be processed through systematic phonetic conversion rules into standard International Phonetic Alphabet (IPA) and the established, transparent Latin transliteration conventions adopted by the linguistic consensus for Tungusic languages.
2. Definition and Cross-Referencing
- Issue: Copying or directly translating definitions or semantic notes from modern Japanese or Russian Orok dictionaries constitutes a copyright violation.
- Solution: Lexical meanings will be treated strictly as objective linguistic facts. Definitions will be independently drafted in English based on core semantic values. To ensure accuracy and academic rigor, terms will be cross-referenced and verified against cognates in closely related Nanaic branch languages (specifically Nanai and Ulch), where documentation is mutually illustrative.
3. Attestation and Avoidance of Modern Neologisms
- Issue: Past administrative actions have highlighted the strict boundary against unendorsed modern coinages or non-attested compounds.
- Solution: This project restricts itself exclusively to documented historical and traditional Orok vocabulary. No modern neologisms, unrecognized compound phrases, or speculative constructions will be introduced. Public-domain historical corpus materials (such as pre-1945 phonetic records in Katakana and early exploration vocabulary lists) will be prioritized as primary historical baselines to satisfy attestation requirements.
4. Exclusion of Copyrighted Contextual Data
- Issue: Example sentences and specialized academic commentaries in modern dictionaries possess independent creative copyright.
- Solution: No example sentences or descriptive analytical texts from modern copyrighted sources will be extracted or reformatted. Entries will focus strictly on the structural presentation of the lemma, part of speech, cognate data, and verified core meanings.
Feedback, suggestions, or technical guidance from administrators and experienced editors regarding the optimization of this template before proceeding with large-scale data input are highly welcomed. Please feel free to let me know if there are any further points that require clarification, supplementation, or refinement. I look forward to your valuable input. --MiiCii (talk) 15:56, 7 June 2026 (UTC)
- I will ignore the fact that this comment was seemingly written completely by AI, which is a major red flag, and focus on the content right now:
- I would not recommend using IPA - members of the Orok community largely use Cyrillic, I think it would be fair to them to also use this writing system.
- I don't see why we should avoid neologisms if they are attestable.
- Other than that what you are proposing is just common sense already deployed by other language communities. However, mass-addition to the dictionary using AI is a very, very bad idea (please do not do this). Thadh (talk) 18:46, 7 June 2026 (UTC)
- Thank you for reminding not to use AI for mass additions, I will rely on manual editing for future contributions from now on.
- I have noticed that the Cyrillic orthography systems used in the Orok-Russian Dictionary and Russian-Orok Dictionary [Орокско-русский и русско-орокский словарь] (2003) by L.V. Ozolinya and I.Ya. Fedyaeva and in Uiltadairisu [Уилтадаирису] (2008) by Jiro Ikegami et al. do not always agree on the spelling of the same words. For instance:
- Ⅰ. The cardinal number "4" appears as "ӡ̌ӣн" on page 50 of the 2003 dictionary, but as "ӡӣн" on page 94 of the 2008 book.
- Ⅱ. The word for "song" appears as "я̄я" on page 186 of the 2003 dictionary, but as "ја̄ја" on page 95 of the 2008 book.
- While the 2008 book distinguishes between Southern and Northern dialect pronunciations using the markers (ю) and (с); for instance, the word for "winter" is listed on page 103 as "тувэни" for the Southern dialect and "тувэ" for the Northern dialect, whereas the 2003 dictionary does not specify the dialect for the entries it lists.
- We can found an Orok-language version of the Universal Declaration of Human Rights online, with the title written as "Чипа̄линне̄сал деклара̄сијачи нари доролбони". The word "деклара̄сијачи" consists of the Russian loanword stem "деклара̄сија" (meaning "declaration") and the third-person plural (3PL) personal marker "-чи." As neither the root nor this specific derivative form (typical of agglutinative languages) appears as an independent entry in the 2003 dictionary or the 2008 book. Thus, just in my opinion, a cautious approach should be taken regarding modern neologisms.
- How should we handle these spelling discrepancies, and how can we avoid copyright infringement? --MiiCii (talk) 17:20, 8 June 2026 (UTC)
- The fact that a word means something is a fact. And under US law, facts cannot be copyrighted. This means that it is not copyright infringement to upload basic definitions, like "я̄я" meaning "song". I would say that you could get away with translating all the definitions into English, as the goal here is to inform readers of the Uilta language.
- As for the orthography, count me in as another support of the Cyrillic script. From my reading, it seems that the preferred script of this language is Cyrillic. As for spelling differences, I would say use the spelling the community uses as the main one, and the ones that they would likely use as a backup.
- I have no clue as to why (attestable!) neologisms shouldn't be included alongside the normal words. They're just as part of the language as any other word, and adding them shows that the tongue is still in use in today. CitationsFreak (talk) 21:45, 8 June 2026 (UTC)
- I think the fact you used AI to write a response to a comment asking you not to use AI is a major red flag.
- Other than that - spelling variation is not a problem. We just have to settle on one consistent orthography for lemmatisation, and list the others as alternative forms. And I side with CitationsFreak on the fact that neologisms are not a problem Thadh (talk) 10:46, 9 June 2026 (UTC)
- After reading our discussion, I have changed ideas on neologisms, and agree with CitationsFreak and Thadh.
- When it comes to settling on one consistent spelling, which is mainly used in the Uilta community, just in my opinion, the 2008 version (Ikegami et al.) is recommended. At the same time, we can find that this version of spelling is adopted in English Wiktionary (We can search the contents of "See also" in the page of Ԩ#Orok).
- Page 7 of the 2003 dictionary (Ozolinya and Fedyaeva) listed a Uilta spelling plan including a gamma-like Cyrillic alphabet after "Гг", which is hardly found in the Cyrillic Unicode. By the way, "Ԩԩ" in the 2008 version is corresponded with "Н'н'" in the 2003 one. MiiCii (talk) 15:54, 9 June 2026 (UTC)
PIE *dʰR in Proto-Italic
[edit]@Graearms, Urszag, Catonif, Mellohi!, who reconstructs PIE *dʰR as yielding Proto-Italic *ðR? De Vaan, Lipp, Stuart-Smith, and Untermann all reconstruct *fR, while Sen and Weiss reconstruct *bR. See sources on RC:Proto-Italic/staðlom. --{{victar|talk}} 19:32, 9 June 2026 (UTC)
- De Vaan actually provides a reconstruction *staþlo- for *staðlom. In the beginning of his dictionary, he also writes that "I will note voiceless fricatives in my Plt. reconstructions, but it seems likely that they were voiced word-internally."[1] Therefore, De Vaan presumably really reconstructs *staðlo-, with a voiced dental fricative, as opposed to *staθlo- with a voiceless dental fricative. Meiser also reconstructs *δ as an intermediary between the Latin development of *dʰ to b.[2] Graearms (talk) 21:19, 9 June 2026 (UTC) Graearms (talk) 21:19, 9 June 2026 (UTC)
- Also, Victar, I doubt that *stablom is really an alternative reconstruction of *staðlom. I assume both Weiss and Sen were refferring to a Proto-Latin stage prior to the anaptyxis. Graearms (talk) 21:21, 9 June 2026 (UTC)
- @victar I don't think I know more than what I said previously in the following discussion: Wiktionary_talk:Proto-Italic_entry_guidelines#Does_the_change_of_word-internal_/θr/_to_/fr/_predate_Proto-Italic? As Graearms said, -ðl-, -ðl- etc were presumably real intermediate stages, whereas -b- only seems supportable at the pre-Latin stage.--Urszag (talk) 22:31, 9 June 2026 (UTC)
- Thanks for correcting me on de Vaan. Even so, he appears to be the only one who reconstructs *þR /ẟR/. To clarify, I agree that Proto-Italic underwent *dʰ > *ẟ { #_, _R }, but then, as in word-initially, *ẟ before liquids subsequently became *f. This is the view supported by most scholars, and the one I believe we should adopt. --
{{victar|talk}}04:37, 10 June 2026 (UTC)- I'm not sure that the majority of scholars reconstruct *f
- Weiss reconstructs *đ.[3]
- Bakkum reconstructs *ð[4]
- Sihler reconstructs *θ.[5]
- Schrijver appears to reconstruct *ð. He was talking about Italo-Celtic, but he doesn't mention any change word-medially in Italic even when describing changes made to this system in Italic.[6]
- Fortson does seem to accept a development of *dʰ > *f word-medially.[7] But, he also does have a passage where he seems to accept a development of *dʰ > *ð.[8]
- Graearms (talk) 17:52, 10 June 2026 (UTC)
- Sihler then continues on the following page to stating that "where intervocalic *f and *θ fall together only when the *θ is adjacent to r, l, or a ū̆". As for Weiss, he only discusses medial *dʰ on the page you cite. If you see Weiss (2009:164), he explicitly reconstructs *b before liquids.[9] I do not have access to Bakkum's book, so I cannot comment on it. To be clear, no one his disputing the development *-dʰ- > *-ð-; only whether *ð subsequently became *f before liquids. --
{{victar|talk}}18:55, 10 June 2026 (UTC)- Sihler was referring to a specifically Latin development. On note 1 of the same page, he writes "If medial "θ became *f adjacent to liquids and high back vocoids as early as PItal, their developments in L would be unchanged from the usual view; but in that case Sab. requires a subsequent merger of the remaining "0 with *f. This is not impossible, only over-elaborate; in any case, the communis opinio is that the partial merger of medial *θ and *f was a specifically Latin-Faliscan development." Also, I'm still not convinced that Weiss was referring to a PIt development; he could easily just be describing a pre-Latin sound change. Graearms (talk) 19:06, 10 June 2026 (UTC)
- Sihler then continues on the following page to stating that "where intervocalic *f and *θ fall together only when the *θ is adjacent to r, l, or a ū̆". As for Weiss, he only discusses medial *dʰ on the page you cite. If you see Weiss (2009:164), he explicitly reconstructs *b before liquids.[9] I do not have access to Bakkum's book, so I cannot comment on it. To be clear, no one his disputing the development *-dʰ- > *-ð-; only whether *ð subsequently became *f before liquids. --
- I'm not sure that the majority of scholars reconstruct *f
- Thanks for correcting me on de Vaan. Even so, he appears to be the only one who reconstructs *þR /ẟR/. To clarify, I agree that Proto-Italic underwent *dʰ > *ẟ { #_, _R }, but then, as in word-initially, *ẟ before liquids subsequently became *f. This is the view supported by most scholars, and the one I believe we should adopt. --
I just noticed a data point that seems relevant: the (currently unsourced) etymology at acerbus derives it from Proto-Italic *akriðos. Assuming this etymology is correct, that would require a change after Proto-Italic (postdating the change of *ri to [r̩]) from *-rð- to -rb-, and while it could be a repeat, it seems simpler to assume that happened only once, which would place it after Proto-Italic. That's a lot of assumptions, though, so I'd appreciate further thoughts on this. We don't see such a change in sacerdōs, although I expect one could appeal to leveling with other terms with final or medial -dō. Overall, my feeling at this point is that even though only conditional remnants survive of the distinction of /θ/ from /f/, they probably remained distinct in all positions in Proto-Italic. It seems doubtful to me that *ð survived for long as a distinct phoneme after the loss of initial /θ/, so it seems simpler to date the Latin conditional changes to a post-Proto-Italic date, with the Latin-specific change of *ð > -d- soon following, rather than supposing the conditioned changes of *ðr, ðl, rð > *βr, βl, rβ and unconditioned change of [θ] > [f] were an early feature of the entire proto-language.--Urszag (talk) 20:14, 7 July 2026 (UTC)
- Most sources appear to treat the change of *ðr, ðl, rð > *βr, βl, rβ as a post-Proto-Italic development. De Vaan, for instance, reconstructs *werþo- instead of *werβom, which was previously reconstructed by Wiktionary before I moved the page. Graearms (talk) 19:41, 8 July 2026 (UTC)
References
[edit]- ^ De Vaan, Michiel (2008), Etymological Dictionary of Latin and the other Italic Languages (Leiden Indo-European Etymological Dictionary Series; 7), Leiden, Boston: Brill, →ISBN, page 7
- ^ Meiser, Gerhard (2017), “Chapter VIII: Italic”, in Klein, Jared S., Joseph, Brian D., Fritz, Matthias, editors, Handbook of Comparative and Historical Indo-European Linguistics: An International Handbook (Handbücher zur Sprach- und Kommunikationswissenschaft [Handbooks of Linguistics and Communication Science]; 41.2), volume 2, Berlin; Boston: De Gruyter Mouton, →ISBN, § The phonology of Italic, page 744
- ^ Olander, T., editor (2022), The Indo-European Language Family: A Phylogenetic Perspective, Cambridge: Cambridge University Press, , →ISBN, page 124
- ^ Bakkum, Gabriël C.L.M. (2009), The Latin dialect of the Ager Faliscus: 150 years of scholarship[1], Amsterdam University Press, →ISBN, page 67
- ^ Sihler, Andrew L. (1995), New Comparative Grammar of Greek and Latin, Oxford, New York: Oxford University Press, →ISBN, page 139
- ^ https://www.google.com/books/edition/The_Reconstruction_of_Indo_European_Stop/zb29EQAAQBAJ?hl=en&gbpv=1&dq=%22Proto-Italic%22+%22aspirates%22&pg=PA272&printsec=frontcover, pages 275-276
- ^ Fortson IV, Benjamin W. (2017), “Chapter VIII: Italic”, in Klein, Jared S., Joseph, Brian D., Fritz, Matthias, editors, Handbook of Comparative and Historical Indo-European Linguistics: An International Handbook (Handbücher zur Sprach- und Kommunikationswissenschaft [Handbooks of Linguistics and Communication Science]; 41.2), volume 2, Berlin; Boston: De Gruyter Mouton, →ISBN, § The dialectology of Italic, page 836: “Labialization of PIE *dʰ and *gʷʰ to f word-initially and f (often showing up as v or β) word-internally”
- ^ Fortson IV, Benjamin W. (2017), “Chapter VIII: Italic”, in Klein, Jared S., Joseph, Brian D., Fritz, Matthias, editors, Handbook of Comparative and Historical Indo-European Linguistics: An International Handbook (Handbücher zur Sprach- und Kommunikationswissenschaft [Handbooks of Linguistics and Communication Science]; 41.2), volume 2, Berlin; Boston: De Gruyter Mouton, →ISBN, § The dialectology of Italic, page 852: “It is striking that *-δ- (from *-dh- and in some cases *-s-) became labialized to -β- in both Sabellic and Faliscan across the board, but only some of the time in Latin.”
- ^ Weiss, Michael L. (2009), Outline of the Historical and Comparative Grammar of Latin, Ann Arbor: Beech Stave Press, →ISBN, page 164
Transcribing quotations for audiovisual media
[edit]the documentations and informal consensus for {{quote-av}} may be incomplete with what to do with transcribing dialogues and the added glossing that it needs. following the suggestions given in documentation, this type of formatting is what I expect. in the entry for English fluffle:
* {{quote-av|en|year=2025|writer=w:Jared Bush|directors=Jared Bush;w:Byron Howard|actor=w:Ginnifer Goodwin;w:Jason Bateman|role=w:Judy Hopps;w:Nick Wilde|title=w:Zootopia 2|format=film|location=United States|publisher=w:Walt Disney Studios|at=<!-- timestamp needed --> |passage=Judy (Ginnifer Goodwin): {{q-g|sobbing; to Nick}} I should never have left you. And I do need a herd of therapy animals, and I should have told you that you're the only partner I would ever want because… you're my '''fluffle'''.{{pb}}<!-- -->Nick (Jason Bateman): {{q-g|perplexed; to Judy}} Hm... Um...?{{pb}}<!-- -->Judy: {{q-g|sobbing; to Nick}} That's a bunch of rabbits!}}
2025, Jared Bush, directed by Jared Bush and Byron Howard, Zootopia 2 (film), spoken by Judy Hopps and Nick Wilde (Ginnifer Goodwin and Jason Bateman), United States: Walt Disney Studios:
Judy (Ginnifer Goodwin): [sobbing; to Nick] I should never have left you. And I do need a herd of therapy animals, and I should have told you that you're the only partner I would ever want because… you're my fluffle.Nick (Jason Bateman): [perplexed; to Judy] Hm... Um...?Judy: [sobbing; to Nick] That's a bunch of rabbits!
@Randomcanarian (in an offwiki discussion) asked how to cite a TikTok video for Alemannic German Uufgab (see initial revision). using the format above, the best that I could come up with is below:
* {{quote-web|gsw|date=2026-04-20|author=luusbueb [lucischnucki]|title=geheimagent odr so#kopfhörer #schule #prüfung #schweiz #comedy|format=video|work=w:TikTok|publisher=ByteDance|archiveurl=https://web.archive.org/web/20260611100828/https://www.tiktok.com/@lucischnucki/video/7630783777110887702 |passage={{lang|en|Luca}}: {{q-g|{{lang|en|shushes}}}} Er wird das nie checke.{{pb}}<!-- -->{{lang|en|Person talking through earbud}}: Das sind d'Lösige zum Tescht. Antwort äis isch vierzg. Bi de siebenesächzigscht '''Ufgab''', füüf plus füüf isch zwänzg.}}
2026 April 20, luusbueb [lucischnucki], “geheimagent odr so#kopfhörer #schule #prüfung #schweiz #comedy”, in TikTok[4] (video), ByteDance, archived from the original on 11 June 2026:
Luca: [shushes] Er wird das nie checke.Person talking through earbud: Das sind d'Lösige zum Tescht. Antwort äis isch vierzg. Bi de siebenesächzigscht Ufgab, füüf plus füüf isch zwänzg.
- (please add an English translation of this quotation)
quite ugly to need to wrap lang inside another lang block. the forced breaks (either with <br /> or {{pb}}) due to the lack of support for actual linebreaks, also.
what format should these quotations then take? as this has come up before in the template's talk page (for Chinese). Juwan 🕊️🌈 12:28, 11 June 2026 (UTC)
- I think the English-language one is a little needlessly complicated and redundant: if you list who the characters and performers are earlier, I don't think it helps understanding to list them again. For that matter, we should definitely discourage editorializing like the tone of the characters unless it's strictly necessary for understanding. ―Justin (koavf)❤T☮C☺M☯ 13:53, 11 June 2026 (UTC)
- on the redundancy, agreed. thinking about it, this may have been written in the docs before actor roles were present. regarding the tone, in this case in specific, the transcript I used included these markers; I would have to check whether this was in the subtitles or not (though I don't have a Disney+ subscription nor want to). Juwan 🕊️🌈 17:45, 11 June 2026 (UTC)
Workgroups and you
[edit]Some editors have made aware that workgroups are not commonly known, which is a disadvantage for those who are interested in those relevant discussions. To raise awareness, I optionally propose that others advertise it in their language's entry guideline page, stating something like: "If you want to be notified of any discussions pertaining to [LANGUAGE], add yourself to Module:workgroup ping/data"; as seen in the Chinese guideline page. TranqyPoo [💬 | ✏️] 17:31, 12 June 2026 (UTC)
Support. I've been editing here for a decade and I didn't realize until today that this is what people were talking about when they referred to workgroups. Andrew Sheedy (talk) 17:40, 12 June 2026 (UTC)
- I've been editing here since 2010 and this is the first time I've heard of them. Makes me wonder whether this reflects on their prominence and usefulness. — Sgconlaw (talk) 17:40, 13 June 2026 (UTC)
- I think what's not helping is that the template to ping a workgroup decomposes into pings of the individual editors, obscuring any difference to the poster arbitrarily deciding to ping some relevant people. Though the only pings I've gotten from my workgroups have been from language treatment stuff. PhoenicianLetters (talk) 09:19, 17 June 2026 (UTC)
- Do you think it would be helpful if the output was something like: (Notifying workgroup: @TranqyPoo, ...)? TranqyPoo [💬 | ✏️] 14:20, 17 June 2026 (UTC)
- At least a bit, yes. PhoenicianLetters (talk) 22:30, 17 June 2026 (UTC)
- Yes, this seems like a good idea. Ideally it would even output the name of the pinged workgroup. - -sche (discuss) 23:31, 17 June 2026 (UTC)
Done: I am not experienced enough to feature displaying the pinged workgroup. However, I emphatically support the idea. TranqyPoo [💬 | ✏️] 04:51, 18 June 2026 (UTC)- @TranqyPoo for the output, it would be even better if the output included which workgroup was being pinged, e.g. "Notifying Portuguese workgroup". Juwan 🕊️🌈 18:29, 1 July 2026 (UTC)
- Do you think it would be helpful if the output was something like: (Notifying workgroup: @TranqyPoo, ...)? TranqyPoo [💬 | ✏️] 14:20, 17 June 2026 (UTC)
- I think what's not helping is that the template to ping a workgroup decomposes into pings of the individual editors, obscuring any difference to the poster arbitrarily deciding to ping some relevant people. Though the only pings I've gotten from my workgroups have been from language treatment stuff. PhoenicianLetters (talk) 09:19, 17 June 2026 (UTC)
- I've been editing here since 2010 and this is the first time I've heard of them. Makes me wonder whether this reflects on their prominence and usefulness. — Sgconlaw (talk) 17:40, 13 June 2026 (UTC)
- I knew about them, though I am sad because I'm not in any —User:Vealhurl (talk 22:55, 17 June 2026 (UTC)
- You are in the exemplative one, rejoice. Catonif (talk) 18:56, 1 July 2026 (UTC)
What should we call старостинські округи?
[edit]Most of Ukraine is subdivided into oblasts (first established in 1932). Those oblasts are in turn subdivided into raions (first established in 1923). Those raions are in turn subdivided into hromady (first established in 2015). Those hromady are in turn (incompletely) subdivided into старостинські округи (first established in 2017). Старостинські округи in turn comprise individual settlements.
The {{place}} infrastructure currently has the holonyms oblast, raion, rural hromada, settlement hromada, and urban hromada for all but the last and lowest-level of these administrative–territorial units. It is sometimes necessary to specify so granular a subdivision as the старостинський округ because that is all that distinguishes two or more identically-named settlements from each other (examples include two villages called Shevchenkove in Ivankiv settlement hromada of Vyshhorod Raion of Kyiv Oblast, two villages called Ivanivka in Lozuvatka rural hromada of Kryvyi Rih Raion of Dnipropetrovsk Oblast, two villages and one rural settlement called Ivanivka in Barvinkove urban hromada of Izium Raion of Kharkiv Oblast, two villages called Ivanivka in Shevchenkove settlement hromada of Kupiansk Raion of Kharkiv Oblast, two villages called Mykolaivka in Hrodivka settlement hromada of Pokrovsk Raion of Donetsk Oblast, two villages called Mykolaivka in Chkalovske settlement hromada of Chuhuiv Raion of Kharkiv Oblast, and three villages called Mykolaivka in Lozova urban hromada of Lozova Raion of Kharkiv Oblast).
Unfortunately, старостинські округи are too obscure for there to be any established English name for them yet, which is why I have used the untransliterated Ukrainian term in this section so far. I'm sure that pretty much everyone will agree that using the untransliterated Cyrillic term in our definitions would be an unwarranted barrier to accessibility, so we must, in some way, establish our own usage here.
The most conservative thing we could do is simply transliterate the term; if we follow Ukrainian National transliteration, this gives us the singular starostynskyi okruh (старостинський округ) and the plural starostynski okruhy (старостинські округи). That is what I have been using in our entries so far. Doing this has the benefit that the adjective starostynskyi and the noun okruh are both attested in English, so even if the phrase starostynskyi okruh is not yet sufficiently well-attested to satisfy Wiktionary:Criteria for inclusion#Attestation, we would be innovating only minimally.
This topic came up on the English Wikipedia, where w:User:Shwabb1 and I have been having a fruitful discussion about this and other topics at w:Talk:Poliske settlement hromada since January. We have since reached the point that Shwabb1 opened up a discussion at w:Wikipedia talk:WikiProject Ukraine#Discussion regarding administrative divisions terminology to seek input from the wider editing community of Ukrainiana on that project. So far, however, the posting has not attracted any visible attention.
I recall the point Benwing2 made last year that “[g]enerally we try to avoid using special-purpose borrowings like ‘hromada’ in favor of more generic equivalents, but I don't know if hromadas are close enough to municipalities to make this equivalence”. What was a concern with regard to hromada must be even more pressing a concern with regard to starostynskyi okruh. Fortunately, Shwabb1 and I might well have found an English equivalent of the Ukrainian старостинський округ. As I wrote in w:Wikipedia talk:WikiProject Ukraine#Discussion regarding administrative divisions terminology, “alderman and староста are etymologically closely analogous and have been used to translate each other for centuries. An alderman's ward is an aldermanry, so if a староста is an alderman and an округ is his/her ward, then a старостинський округ, it follows, is an aldermanry”.
In conclusion, I think we should call старостинські округи either starostynski okruhy or aldermanries. Whatever we decide upon, be it either of them or something else, other terms would need to be added as aliases of the preferred term. Even if the displayed term is to be aldermanry I also suggest we have an abbreviating alias like starokruh to function like wcomm, so the term can link to the right Wikipedia article or Wiktionary (Glossary) entry. Finally, there are over seven thousand of these administrative–territorial units, so I think it would be worth categorising them by oblast, as we already do for hromady (since Wiktionary:Beer parlour/2025/May#Categorising Ukrainian toponyms by oblast). 0DF (talk) 16:09, 13 June 2026 (UTC)
Outsider support as non-Ukraianian editor. the rationale is very well explained. for the specific name, personally would lightly disagree with Ben in that the subdiv term we use should be the most common name or otherwise that editors agree upon (as in the Wikipedia discussion). likewise, nudging Ben for cleaning up the place module. Juwan 🕊️🌈 21:55, 18 June 2026 (UTC)
Do "alt forms" include just orthographic differences, or also word variants?
[edit]The current guidelines suggest "alt forms" is just for orthographic variation that has no correlate in speech. However, in practice I've seen it used for word variants such as bequadro and beqquadro, which reflect a difference in pronunciation. I changed those entries from alt to synonym, but the guideline isn't entirely clear on this, so I thought I should ask here. kwami (talk) 20:32, 14 June 2026 (UTC)
- @Kwami: the edges seem a little blurred here. If a word is pronounced the same way or very similarly to a lemma, I would use
{{alternative spelling of}}(adding an IPA transcription if the pronunciation differs). If there are more significant differences (see, for example, dance with the one that brought you), I would either use{{alternative form of}}or{{synonym of}}. The more different a word is from the lemma, the more likely it is that I would use the latter template. — Sgconlaw (talk) 20:41, 14 June 2026 (UTC)- I think if the pronunciation differs, then it's not an alt spelling, but a different word, though I suppose that's arguable. When the words are variants, the {alt of} template works for the secondary entries, but then it would be wrong to list them under an =Alternative forms= header without including the pronunciation, which I don't think I've ever seen anyone do. But doing that would mean that we would regularly have pronunciations outside the pronunciation section. kwami (talk) 21:16, 14 June 2026 (UTC)
- @Kwamikagami: oh no, I didn't mean that the pronunciations should be put in "Alternative forms" sections; I meant they should be indicated on the entry pages of the alternative forms. — Sgconlaw (talk) 21:55, 14 June 2026 (UTC)
- Okay, but then if we don't have an article for the alt form, the info we do provide would imply that it has the same pronunciation as the lemma. Even if we do have one, readers would have to follow the link to see that the pronunciation differs. Most people aren't going to bother to follow links for spelling variants just to verify that they're only spelling variants; I wouldn't even think of doing that. So I don't think we can afford to put words under the Alt Forms header if they differ in pronunciation -- they would need to be listed under Synonyms. kwami (talk) 02:29, 15 June 2026 (UTC)
- @Kwamikagami: oh no, I didn't mean that the pronunciations should be put in "Alternative forms" sections; I meant they should be indicated on the entry pages of the alternative forms. — Sgconlaw (talk) 21:55, 14 June 2026 (UTC)
- I think if the pronunciation differs, then it's not an alt spelling, but a different word, though I suppose that's arguable. When the words are variants, the {alt of} template works for the secondary entries, but then it would be wrong to list them under an =Alternative forms= header without including the pronunciation, which I don't think I've ever seen anyone do. But doing that would mean that we would regularly have pronunciations outside the pronunciation section. kwami (talk) 21:16, 14 June 2026 (UTC)
- Based on prior discussions (linked to in this short 2023 discussion), my understanding is that among editors who make a distinction between
{{altsp}}and{{altform}}and{{synonym of}}(instead of just using them haphazardly / interchangeably), "alternative spelling of" is for when only the spelling differs but the pronunciation is the same, "alternative form of" is for when the form of the words differs enough that the pronunciation also differs (but the word is still basically the same word) like e.g. kilikinick vs kinnikinnick, and{{synonym of}}is for when something is an entirely different word like myriad 1.1 as a synonym of decamillennium. (There was, as you see in those links, discussion of making the pronunciation-based distinction explicit in the text of the templates, but volunteer time to implement that has been lacking.) AFAICT bequadro and beqquadro are in the same boat as kilikinick and kinnikinnick: alt forms. (The distinction between "alt form" and "synonym" is less well maintained when it comes to phrases instead of single words: all over one like a cheap suit is currently listed as a synonym of all over someone like a rash whereas a few spanners short of a tool box is currently listed as an alt form of few sandwiches short of a picnic.) - -sche (discuss) 03:31, 15 June 2026 (UTC)- If we place biqquadro under an 'Alt form' heading at bequadro, and we don't want to use IPA outside the pronunciation section, wouldn't that mean we need to include the alt pronunciations in the pronunciation section? kwami (talk) 04:05, 15 June 2026 (UTC)
- ...no? The pronunciation of biqquadro goes in its own entry, biqquadro; AFAICT, the pronunciation of biqquadro does not need to be in the entry bequadro — or at least, Wiktionary's current and longtime practice is to present the pronunciation of each entry in its own entry, and not in other entries. (If we add the pronunciation of the plural bequadri it too would go in its own entry and not in bequadro.) If Italians commonly pronounce the spelling bequadro as if it were biqquadro, that would be something to mention in bequadro's pronunciation section (compare habanero, which some people—while still spelling it habanero—pronounce like habañero), but if not, then AFAIK current practice would be to only have biqquadro's pronunciation in its own entry, not also in bequadro. (This practice is not without drawbacks, and I do sometimes think a more compact approach like some other online dictionaries use — where you can see the pronunciation of not only the word you looked up but also its inflected forms, alt forms, derived adjectives and adverbs, etc, all in one place — would have some benefits (and different drawbacks), but changing that practice would be a separate and bigger discussion...) - -sche (discuss) 04:40, 15 June 2026 (UTC)
- Okay, but then it's not really an alt form, because the single given pronunciation is implied for the entire entry. Since biqquadro has a different pronunciation, IMO it would need to be treated as a synonym, unless we distinguish orthographic variants from spoken variants, which we seldom do.
- If this were a spoken dictionary, the problem wouldn't arise, but we're an orthographic dictionary, with different words lumped onto the same page just because they happen to be spelled the same. Though of course going too far down that rabbit hole runs into disputes of whether a jump and to jump are the same word or homophones. And of course this is a much greater problem with CJK characters, where we lump etymologically unrelated words together on the same page.
- But regardless, I think we do our readers a disservice if we don't distinguish orthographic variants from spoken variants. Kinnikinnick makes the difference pretty obvious, but that's rare. kwami (talk) 05:43, 15 June 2026 (UTC)
- I don't see a compelling reason to make such a switch to how things are handled. To me, at least, it makes sense that we consider spoken /ˌkɪnɪkəˈnɪk/ and /ˈkɪnɪkəˌnɪk/ as (alternative) forms of one word rather than synonyms (and already list both in kinnikinnick), and that we consider spoken /ˈkɪnɪkəˌnɪk/ and /ˈkɪlɪkəˌlɪk/ alternative forms (listing them in their separate entries), and that we consider written kinnikinnick and kinikinick and kilikinick alternative forms of that word rather than two separate, merely synonymous words (and that we list them as alt forms in their separate entries). I am not currently persuaded that these kinds of things would be better considered synonyms instead, but hopefully other people can weigh in, if they do think that would be better. For now, under current and longstanding practice, the thing to do is just change the "Synonyms" header in bequadro back to "Alternative forms", since "alternative forms" has always included forms that are pronounced differently here (scion, sient, cyons, etc; iluec / eleckes, eloc / hiluk / ileches / ileks / illus, illusques, etc; seien, san, sugcen / etc; motherfucker, motherfugga, etc) and the entries beqquadro and biqquadro are already correctly labelled as alternative forms. - -sche (discuss) 18:10, 15 June 2026 (UTC)
- Okay, reverted.
- I suppose, given spelling pronunciations, that it may not always be practical to distinguish alt spellings from alt forms in English. Maybe easier in French and Irish. kwami (talk) 19:14, 15 June 2026 (UTC)
- I don't see a compelling reason to make such a switch to how things are handled. To me, at least, it makes sense that we consider spoken /ˌkɪnɪkəˈnɪk/ and /ˈkɪnɪkəˌnɪk/ as (alternative) forms of one word rather than synonyms (and already list both in kinnikinnick), and that we consider spoken /ˈkɪnɪkəˌnɪk/ and /ˈkɪlɪkəˌlɪk/ alternative forms (listing them in their separate entries), and that we consider written kinnikinnick and kinikinick and kilikinick alternative forms of that word rather than two separate, merely synonymous words (and that we list them as alt forms in their separate entries). I am not currently persuaded that these kinds of things would be better considered synonyms instead, but hopefully other people can weigh in, if they do think that would be better. For now, under current and longstanding practice, the thing to do is just change the "Synonyms" header in bequadro back to "Alternative forms", since "alternative forms" has always included forms that are pronounced differently here (scion, sient, cyons, etc; iluec / eleckes, eloc / hiluk / ileches / ileks / illus, illusques, etc; seien, san, sugcen / etc; motherfucker, motherfugga, etc) and the entries beqquadro and biqquadro are already correctly labelled as alternative forms. - -sche (discuss) 18:10, 15 June 2026 (UTC)
- ...no? The pronunciation of biqquadro goes in its own entry, biqquadro; AFAICT, the pronunciation of biqquadro does not need to be in the entry bequadro — or at least, Wiktionary's current and longtime practice is to present the pronunciation of each entry in its own entry, and not in other entries. (If we add the pronunciation of the plural bequadri it too would go in its own entry and not in bequadro.) If Italians commonly pronounce the spelling bequadro as if it were biqquadro, that would be something to mention in bequadro's pronunciation section (compare habanero, which some people—while still spelling it habanero—pronounce like habañero), but if not, then AFAIK current practice would be to only have biqquadro's pronunciation in its own entry, not also in bequadro. (This practice is not without drawbacks, and I do sometimes think a more compact approach like some other online dictionaries use — where you can see the pronunciation of not only the word you looked up but also its inflected forms, alt forms, derived adjectives and adverbs, etc, all in one place — would have some benefits (and different drawbacks), but changing that practice would be a separate and bigger discussion...) - -sche (discuss) 04:40, 15 June 2026 (UTC)
- I know this discussion is somewhat stale by now, but I chanced across it and was surprised that this is not a unanimous topic. It’s always been second nature to me formatting things like -sche described – someone must have instructed me to do as much on Discord when I first started editing. My impression of things matches -sche’s exactly. — Polomo ⟨ oi! ⟩ · 01:55, 13 July 2026 (UTC)
- I've taken a stab at spelling out the difference (which is already spelled out in the glossary) in the documentation of the templates; please improve if needed. If anyone wants to revive the idea, discussed in the old BP section linked above, of actually spelling out the difference in the text the templates themselves produce in entries (i.e. have T:altsp spell out "alternative spelling of foobar with the same pronunciation" so as to remove the need for a pronunciation section), let's discuss that. Also: I know that at least a few people don't like to consider hyphenating or spacing differences (e.g. ray gun, ray-gun, raygun) to be differences of spelling; if this proves to be a major sticking point, might I suggest that we could create an "alternative hyphenation or spacing of" template for those cases (like we also have T:altcase). - -sche (discuss) 03:39, 13 July 2026 (UTC)
- If we place biqquadro under an 'Alt form' heading at bequadro, and we don't want to use IPA outside the pronunciation section, wouldn't that mean we need to include the alt pronunciations in the pronunciation section? kwami (talk) 04:05, 15 June 2026 (UTC)
- Would it be possible to merge
{{altform}}and{{altsp}}into a single template that decides whether to show "alternative form of" or "alternative spelling of" purely by comparing the parameter to the page title? (open question, but perhaps @Benwing2 or @Theknightwho could comment on this idea...) Ioaxxere (talk) 04:32, 6 July 2026 (UTC)- For languages where spelling is required to match pronunciation, that's a good idea. (In general / for all languages, I would think it would not be possible; suppose the template finds itself on an English page *febar, defining that as an alternative of *fibar: how would the template know whether *febar is pronounced the same as *fibar or differently? English spelling being what it is, they could both be /fi/ or both /fə/ or they could indicate some difference like /fi/ vs /fɪ/ or /fə/ vs /fi/, etc. But for languages where pronunciation is predictable from spelling, this could work.) - -sche (discuss) 03:39, 13 July 2026 (UTC)
Our terminology should, as far as practicable, reflect established linguistic scholarship, because that gives us a shared conceptual framework and avoids inventing definitions that differ from those used in the field. I do not think I have a strong enough theoretical background in linguistics to take a position on the nomenclature itself. In the meantime, regardless of what these relationships are ultimately called, I think the important thing is that these variant forms are identified through entry creation and properly documented with citations. (AI-assisted drafting) --Geographyinitiative (talk) 🎵 08:48, 6 July 2026 (UTC)
Where do we put English pronunciations of translingual entries?
[edit]An example is ln, which in English may be pronounced /ˈlɒn/ and /ˈlɪn/.
If we place this under the translingual entry, we'll have a mismatch between the language code of the IPA template and of the other templates.
If we place this under the English entry, we'll need an additional etymology, with a duplicated definition, just for pronunciation. kwami (talk) 20:38, 14 June 2026 (UTC)
- @Kwamikagami: on the assumption that a translingual entry is pronounced differently in different languages, I think it's reasonable to put the various pronunciations under that entry, marking them appropriately ("RP", "German", etc.).
- Could you explain your comment about the mismatch between the language code of the IPA template and other templates? Do you mean, for example, that using
{{IPA|en}}under a translingual entry would cause some problem? — Sgconlaw (talk) 22:03, 14 June 2026 (UTC)- Yes, I thought it might. My last block was for an ISO-code mismatch (forgetting to change the ISO code when I copied a template to a different language), so I thought I'd better check.
- I thought ISO mismatches auto-generated errors, but I don't see an error category at ln. So perhaps Translingual is an exception? kwami (talk) 02:39, 15 June 2026 (UTC)
- I think Translingual terms having English pronunciation info is allowed, at least de facto: Acacia has listed English pronunciations since 2013, Saccharomyces since 2019, and these are just two I was able to find quickly. Whether, in a case like ln, it would be better to put the English pronunciation in the Translingual section or create an English section, I am not sure. In practice, entries/editors don't seem to be terribly consistent about when to handle something as Translingual-only vs Translingual and English, e.g. we have an English section at @ that repeats many senses that are already present in the Translingual section. - -sche (discuss) 03:48, 15 June 2026 (UTC)
- I think those precedents answer it then.
- It would be rather silly to create separate English entries for every taxonomic name that has an established English pronunciation, just so we could list those pronunciations. kwami (talk) 03:56, 15 June 2026 (UTC)
- Sounds good. Asteraceae shows an example of that handling. I support having EN.Wikt show EN prons for taxonomic ranks and binomial names entered only as Translingual terms. Someone might point out a counterargument that if all languages are allowed under a Translingual Pron heading, then the list could be too huge. My response to that would be that each Wiktionary (i.e., each language's Wiktionary) could show the pron for that language, and there could be a collapsible box of links under en.Wikt > Translingual > Pron that jumped the user to whichever other language's pron they wanted. It could look something like "de•es•fr•pt" and be collapsed/hidden by default (and expand/show by opting in by clicking on it). Quercus solaris (talk) 21:04, 17 June 2026 (UTC)
- It would be a good idea to explain that in our guideline so that is what happens moving forward.
- Not just taxonomy, but other fields that have fixed international forms. e.g. mathematical abbreviations like sinh.
- Would it be a good idea to restrict this to established English pronunciations? Once we get into generic English conventions for pronouncing Latin, schools differ, and all the reader really needs to know is the vowel length in Latin, which should be in the etymology anyway. (Though we might need to add it for Greek, by adding Latin or otherwise remedying the defective Greek script.)
- I do think we should probably pick just the most common/anglicized pronunciation for suffixes such as -aceae. The suffix article shows the variation, and repeating it in every derivation would be a maintenance nightmare. Or maybe add a note to check the -aceae article for that variation -- I gave it a shot at Asteraceae. kwami (talk) 22:47, 17 June 2026 (UTC)
- Truly, I think we underuse "see foobar"-type wordings as a way of reducing how many places we have to duplicate and sync pronunciation information, and for my part, I would be fine with various -aceae terms just listing the most common pronunciation(s) and then saying something in the vein of "see -aceae for other pronunciations of that ending". Personally, I am also inclined to have things like special military operation just say "see special, military, operation", rather than the long list that presently repeats, in a way that may or may not be kept in sync, the information that each of those separate pages contains. I don't know how many other people support vs oppose that, though. - -sche (discuss) 23:38, 17 June 2026 (UTC)
- I support those same thoughts. That is the optimal approach regarding data normalization, SSOT, cross-reference (quick easy clickthrough) instead of fork/duplicate. Better for the maintainers and better for the users too; win-win. Quercus solaris (talk) 01:33, 18 June 2026 (UTC)
- Truly, I think we underuse "see foobar"-type wordings as a way of reducing how many places we have to duplicate and sync pronunciation information, and for my part, I would be fine with various -aceae terms just listing the most common pronunciation(s) and then saying something in the vein of "see -aceae for other pronunciations of that ending". Personally, I am also inclined to have things like special military operation just say "see special, military, operation", rather than the long list that presently repeats, in a way that may or may not be kept in sync, the information that each of those separate pages contains. I don't know how many other people support vs oppose that, though. - -sche (discuss) 23:38, 17 June 2026 (UTC)
- Sounds good. Asteraceae shows an example of that handling. I support having EN.Wikt show EN prons for taxonomic ranks and binomial names entered only as Translingual terms. Someone might point out a counterargument that if all languages are allowed under a Translingual Pron heading, then the list could be too huge. My response to that would be that each Wiktionary (i.e., each language's Wiktionary) could show the pron for that language, and there could be a collapsible box of links under en.Wikt > Translingual > Pron that jumped the user to whichever other language's pron they wanted. It could look something like "de•es•fr•pt" and be collapsed/hidden by default (and expand/show by opting in by clicking on it). Quercus solaris (talk) 21:04, 17 June 2026 (UTC)
- cc @Qwertygiy. Juwan 🕊️🌈 20:22, 18 June 2026 (UTC)
- I think Translingual terms having English pronunciation info is allowed, at least de facto: Acacia has listed English pronunciations since 2013, Saccharomyces since 2019, and these are just two I was able to find quickly. Whether, in a case like ln, it would be better to put the English pronunciation in the Translingual section or create an English section, I am not sure. In practice, entries/editors don't seem to be terribly consistent about when to handle something as Translingual-only vs Translingual and English, e.g. we have an English section at @ that repeats many senses that are already present in the Translingual section. - -sche (discuss) 03:48, 15 June 2026 (UTC)
June 2026 Wikimedia Café meetups regarding the English Wikipedia Editor Reflections project
[edit]Hello! There will be two Wikimedia Café discussion opportunities during the last weekend of June. Both sessions will focus on the English Wikipedia Editor Reflections project. The featured guest in the Café will be User:Clovermoss. Participants may attend either or both sessions.
- 27 June 2026 15:00 UTC (timestamp converter), at a time friendly to the Americas, Africa, and Europe
- 28 June 2026 03:00 UTC (timestamp converter), at a time friendly to Asia and the Pacific
Please see the Café page for more information, including how to register!
↠Pine (✉) 03:30, 15 June 2026 (UTC)
The code AU should be Australia
[edit]Until recently, the IPA transcriptions for Australia were only phonemic. These entries (phonemes) do not correspond to the phonetic realization in any of the common accents of Australia (General, Broad, and Cultivated), as you can see in w:Australian English phonology. However, when the code AU is used, it shows General Australian. Consequently, Orrigarmi has included some transcriptions corresponding to General Australian (for example in https://en.wiktionary.org/w/index.php?title=bow&diff=89872850&oldid=89634574https://en.wiktionary.org/w/index.php?title=bow&diff=89872850&oldid=89634574). I propose to change the aforementioned code to Australia to avoid confusing the editors. Greetings! Adelpine (talk) 14:46, 16 June 2026 (UTC)
- I recall this being brought up before, and getting little input and no action, at Wiktionary:Tea room/2025/August#AU should be Australia. On the face of it, I agree with you; if we want a label specifically for General Australian, we should have separate specific labels for General Australian, Broad Australian, and Cultivated Australian. I am minded to make this change soon unless people pipe up with reasoned opposition. (Ping me in a week if I forget to make the change and no-one else beats me to it.) - -sche (discuss) 18:11, 17 June 2026 (UTC)
Support. I would also argue the traditional tripartite model (broad/general/cultivated) is falling away. This article suggests its replacement by "ethnocultural, mainstream and Aboriginal". I can personally attest to observing the difference between the first two. This, that and the other (talk) 10:31, 18 June 2026 (UTC)
Support the initative, as not particularly knowledgeable in Australian English. but this matches up the standard of other dialect labels. Juwan 🕊️🌈 21:21, 18 June 2026 (UTC)
Monosyllabic lions and years
[edit]The Lightning Seeds' Three Lions rhymes "three lions on a shirt" with "no more years of hurt", lending each five syllables. I think this is an example of triphthong smoothing, present in many nonstandard English lects. We already document this phenomenon for days like year and clear. Should we document it for lion or Brian? If so, maybe a native speaker could add IPA. brittletheories (talk) 11:43, 18 June 2026 (UTC)
- @Brittletheories: I don't think songs are necessarily a good indication of how words are usually pronounced. Syllables are often elided to fit the meter. — Sgconlaw (talk) 13:18, 18 June 2026 (UTC)
- I can find evidence of this as a nonstandard Pakistani pronunciation: a paper by Hafiz Syed M. Yasir et al, "Phonological Shifts in Pakistani English (PakE): A Comparative Study under Standard British English (in xIlkogretim Elementary Education Online, 2021, vol. 20, issue 5, says "
'lion' it is pronounced in Pakistani English as /laɪn/ instead of /laɪən/
", and a paper by Rehab Ahmad Zakori et al., The Impact of Phonological Processes on Speech Intelligibility of Students at the University of Lakki Marwat (Advance Social Science Archive Journal, vol. 4, no. 1, July-September 2025) says "only one participant was able to pronounce “lion” accurately. Most participants replaced the RP diphthong /aɪə/ with /ɔː/ or simplified it to /laɪn/, dropping the schwa sound. These mistakes were due to vowel substitution, syllable elision, and the influence of L1 pronunciation rules. The word was often pronounced as it is spelt, disregarding the English diphthong structure
". I don't think the pronunciation is limited to Pakistani English, but (until we have more evidence of its use in other varieties) perhaps we could include it with a label like "nonstandard; especially Pakistan". - -sche (discuss) 17:42, 18 June 2026 (UTC)
Unnecessary Turkish Nonlemma forms
[edit](Notifying workgroup: Trimpulot, Bartanaqa, Wreaderick, Vox Sciurorum, Lambiam, Lagrium, Samubert96, ToprakM, Aydin Guluzada, Ardahan Karabağ, Whitekiko, Moonpulsar):
Currently I'm in the process of cleaning up entries that the template {{U:tr:first-person singular}} is used in. Once I'm done with that I will mark the temp for {{delete}}. The purpose of this template is to display a text like the one under Usage notes of "başkanım", and the reason why it was deemed necessary is the pronunciation difference of predicative inflections suffixed with -im, which, as the usage note states, gives it the "I am a/the ..." meaning. The problem, with the template, hence the Usage notes section, is that this is not only unseemly for Wiktionary's formality, but also simply unnecessary. The pronunciation difference (and incidentally the etymology difference) can easily be displayed as in kulağım. But some entries have lemma senses, like dizim, or kurulum, and some have Verb forms, like eklerim, so it's a little tricky. Up to now I have explained why the template needs to go, which I will take care of page by page.
The main concern of this discussion is the numerous nonlemma pages, created by same contributor, that have no beneficial reason to exist. They are possible and valid inflections but the meanings they convey are nonsensical and wouldn't ever be used in a natural sentence. Some examples for this are;
(These are all inflections of the lemma in brackets, as in "my (lemma)" and "I am a/the (lemma)"
güvenliliğim (“secureness”), enstrümantalim (“instrumental”), katamaranım (“catamaran”), mali yardımım (“financial aid”) (the lemma itself may be an SoP), dış ağım (“extranet”), iç ağım (“intranet”), model-view-controllerim (“model-view-controller”), konjenital defektim (“congentital defect”), köygöçürenim (“amanita phalloides; a deadly poisonous mushroom”), güneş dışı gezegenim (“exoplanet”), Katar'ım (“Qatar”), mimari desenim (“architectural pattern”) (may be SoP), lezyonum (“lesion”), soyut arayüzüm (“abstract interface”) etc.
There are some 430 pages left that I have yet to wipe this template off of. All traces of it must go, but what about all these inflections? Should I clean them all up and we keep them, no matter how meaningless they are, or should we pick the really outlandish ones and delete them? Note that there are some seriously weird ones, especially regarding computing and software terminology, where even the lemma is sketchy. Orexan (talk) 20:16, 18 June 2026 (UTC)
- Not at all surprised to see it was the same contributor as the one I mentioned in this thread earlier who created all these entries way back when. Thanks for reminding me of this, I will create a separate mass nomination for those within the next few days. A lot of these are horribly outlandish ("my exoplanet/I am an exoplanet"?). I don't know how exactly we should judge their outlandishness, but I'd say all multiword terms of this form likely qualify for deletion.
- I second getting rid of that template, I've always found it to be an eyesore. Can we also just get rid of any predicative definitions for inanimate objects and concepts? The only place I can imagine someone describing themselves as a book or a pencil is in like a children's TV show.
- You're also right that the lemmas of some of these terms are dubious. I think most Turkish speakers just use unadapted borrowings for technical computing and software terms. Wreaderick (talk) 21:08, 18 June 2026 (UTC)
- @Wreaderick His legacy will live on even after we're dead and gone from this world, and people will still be fixing his entries. I'm not entirely sure about getting rid of the ones we might somehow come to an agreement as unnatural or outlandish. I mean, yes it is very strange and foreign to my eyes and ears and I can't imagine any meaningful scenario where they can be used, but I don't know if that's grounds for removal, like Lambiam argues below. Someone could very well say "my congenital defect" but no congenital defect is gonna start speaking unless you're on drugs or something. Computing terms, similar thing, could be used in the possessive, almost impossible for them to ever be used in the predicative. But even if it was extremely low chances for either form to be encountered naturally, so what, right? I wouldn't go out of my way to create an entry for an unheard of (and will never be heard of) inflection of a niche terminology, but someone has created them and we may not have enough reason to delete them. Some of the lemmas of the inflections may need to be looked at, but thinking about it, they'd probably stay RFV indefinitely. I think the contributor in question may have been in the business to create these entries for computing terminology, so I wouldn't know which ones warrant keeping and which removing, and getting a council of Turkish speaking programmer Wiktionarians together may prove somewhat difficult. Thanks for your feedback. Orexan (talk) 08:57, 19 June 2026 (UTC)
- Can we agree to get rid of links to "synonyms" and "antonyms" suffixed with the negator -me such as the one listed under öldürtmek? They're just clutter. I even very clearly remember being taught in school that such forms do not constitute synonyms and antonyms. Very thankfully I don't see any negative infinitive forms that have entries. Is there any way to automate removing them? Cleaning up these "contributions", as you said, will last several lifetimes. Wreaderick (talk) 09:37, 19 June 2026 (UTC)
- Agree. Negations shouldn't be listed as synonyms or antonyms. And I don't think we need a consensus to take action on this, anyone who has free time on their hands can fix this whenever. I don't know much about automation, but I believe there are ways to do it. Orexan (talk) 19:31, 19 June 2026 (UTC)
- Can we agree to get rid of links to "synonyms" and "antonyms" suffixed with the negator -me such as the one listed under öldürtmek? They're just clutter. I even very clearly remember being taught in school that such forms do not constitute synonyms and antonyms. Very thankfully I don't see any negative infinitive forms that have entries. Is there any way to automate removing them? Cleaning up these "contributions", as you said, will last several lifetimes. Wreaderick (talk) 09:37, 19 June 2026 (UTC)
- @Wreaderick His legacy will live on even after we're dead and gone from this world, and people will still be fixing his entries. I'm not entirely sure about getting rid of the ones we might somehow come to an agreement as unnatural or outlandish. I mean, yes it is very strange and foreign to my eyes and ears and I can't imagine any meaningful scenario where they can be used, but I don't know if that's grounds for removal, like Lambiam argues below. Someone could very well say "my congenital defect" but no congenital defect is gonna start speaking unless you're on drugs or something. Computing terms, similar thing, could be used in the possessive, almost impossible for them to ever be used in the predicative. But even if it was extremely low chances for either form to be encountered naturally, so what, right? I wouldn't go out of my way to create an entry for an unheard of (and will never be heard of) inflection of a niche terminology, but someone has created them and we may not have enough reason to delete them. Some of the lemmas of the inflections may need to be looked at, but thinking about it, they'd probably stay RFV indefinitely. I think the contributor in question may have been in the business to create these entries for computing terminology, so I wouldn't know which ones warrant keeping and which removing, and getting a council of Turkish speaking programmer Wiktionarians together may prove somewhat difficult. Thanks for your feedback. Orexan (talk) 08:57, 19 June 2026 (UTC)
- Thanks for all the good work.
- I agree that "my instrumental" and "I am catamaran" are unnatural. But "I am instrumental" and "my catamaran" are meaningful and not unnatural. (Disclosure: I know someone who owns a catamaran.) And if someone wants to say, "my being secure is important", might they not use güvenliliğim instead of, say, güvende olmam? But, personally, I do not feel that naturalness of inflected forms warrants their inclusion. I'd only create an entry for such inflected forms, also when quite natural, if there is something special about them, like içim will have a page anyway since it is not only a possessive of iç but also a plain noun. And aklım should be included as being somewhat irregular, since there is no *akl of which it is a possessive. Otherwise, there is no end to it. Like, why is there no entry for patronumum, which is natural enough ("Artık kendi patronumum")? So we don't need entries for kralım in either of its natural senses. Someone entering this form in the search box will be directed to our entry for kral.
- If we can agree on actionable criteria, they can be described in Wiktionary:Turkish entry guidelines. ‑‑Lambiam 05:42, 19 June 2026 (UTC)
- @Lambiam Quick notes; enstrümantal doesn't have the senses "essential, critical, cental" etc. like in English, only the music sense, as in a song with no human voice, and the grammar sense, as in the instrumental case. I guess a melancholic poet could write "I am instrumental" to mean "I am a quiet person, I make art but I don't speak much" in a sophisticated way. :) And I agree, "my catamaran" is completely fine and encounterable. güvenliliğim brings up only 4 pages of results on Google, and only a few real instances of usage, where it is used incorrectly, like "Benim artık can güvenliliğim kalmadı" or "çok istiyorum evlenmek fakat ekonomik güvenliliğim maalesef yok." But güvenliliği has some good results, like "aşının güvenliliği ve etkililiğini değerlendirmek için çok önemli" or "yeni sistemlerin güvenliliği üzerine çalışmalar yapılacağı belirtildi." etc. It might be better to translate it as "reliability" but it seems to have a narrow scope. If someone asked me the direct translation for "reliability" I would probably answer güvenilirlik.
- Apart from that, you've pretty much summed up my hesitation to bring down the axe, I'm not entirely comfortable with the idea of deleting them either, as you said, naturalness of inflected forms isn't necessarily a criteria their inclusion. Even though most people will go their entire lives without saying or hearing someone say some of these either in possessive or predicative, they're still valid. I just wanted to hear what others think about them. Thanks for the feedback.
- By the way, cleaning up dizinim was pretty funny since it has multiple inflections on Etym 3 and 4, which mean "I am your knee" or "I am your sequence/TV series." Still valid, still valid inflections. Orexan (talk) 08:36, 19 June 2026 (UTC)
- I have long thought those "I am (fundamentally inanimate object)" entries to be odd. The Wikimedia search function lets you search within categories, such as
- incategory:"Turkish noun forms" intitle:/..im$/
- for all Turkish noun forms ending in "im" or
- incategory:"Turkish verb forms" negative imperative
- for negative imperatives. See Wiktionary:About_Turkish#Entry_layout_and_inclusion for guidance. Unlikely forms of nouns and verbs should not be included even if they are theoretically possible. Vox Sciurorum (talk) 16:42, 19 June 2026 (UTC)
- They most certainly are odd, we're all on the same page there. But the question is whether to remove them or not. It seems to me that according to the wording of the entry guidelines, these entries should not have been created in the first place, but it doesn't clearly urge their deletion once they are created. Besides, with the majority of these entries, the possessive form at least isn't exclusively a distant possibility in the world theoretical. It's not unimaginable for someone to say "my lesion, my abstract interface, my (beloved) Qatar" and things like that. Orexan (talk) 04:57, 21 June 2026 (UTC)
Cleaning up the clean up tasks
[edit]looking at the main pages for the todo tasks, noting how these are quite disorganised. I as a reader find it difficult to find specific tasks, spread over two pages with large walls of text and yet still missing tasks to be done. previous instead of dealing with this, personally just moved the problem away to another project altogether.
would there be any disagreement in merging these together and separating the subpages by language, renaming where needed to clear name.
these includes:
- Wiktionary:Task lists
- Wiktionary:Todo (and subpages)
- Wiktionary:Todo/Lists (and subpages)
pinging @This, that and the other who maintains a lot of the lists. Juwan 🕊️🌈 21:11, 18 June 2026 (UTC)
- I think it would be better to keep WT:Todo/Lists namespaced separately. These lists are automatically updated by a specific service, unlike everything else under WT:Todo which is manually updated and more or less ad hoc. Moreover the lists are not language-specific, except for a couple which are English only for the time being.
- I think there's merit in keeping two separate pages of tasks to do - one aimed at newer editors and focused on requests (a niche that WT:Task lists currently seems to be trying to fill), and one for more experienced users with lower-level cleanup tasks like syntax and categorisation errors (as WT:Todo currently aims to be). I could imagine renaming WT:Task lists to something less anodyne, removing the repetition, and fleshing it out with some basic instructions, a la WT: Webster's Dictionary, 1913. Perhaps it could be thought of as a user-friendly gateway to the Category:Requests by language tree.
- The subpages of WT:Todo could use a cleanup. Some may be obsolete. To the extent that they are language-specific I agree they should be renamed. I started to clean up the WT:Todo page itself once, but it is difficult, as most of the tasks people have noted down there over the years are probably still worthy of being noted down. But if you have any great ideas on how to tidy it up, I'd say go for it. This, that and the other (talk) 05:35, 20 June 2026 (UTC)
- thank you for the positive response. the plans that I have in mind are to sort through the todo lists and separate them out by language (for now,
enfor English,undfor no specific language asmulis taken) and rename them appropriately. warning you beforehand as I don't want to break your code by accident! the main page then can be structured by separating into 'creation' and 'maintance', with regular tasks as a subcat of this. the huge wall of text will be spread out into the subpages, with more space for instructions and also for including lists or query code to generate them. - the nature of some tasks being fully automated and some not is not too important to me. these may be simply cleared up with a banner in each subpage ("This task list is automatically maintained by FOOBOT. See the code at REPO on Git.") and a dedicated list
- to separate the "Task list" page as you suggested, then I could move it to a something with "Request" in the title. with the name freed up, I would also like to move "Todo" there. with all that, I would do, for example:
- Task lists
- Task lists/pt/Creation/Missing entries from Portuguese Wiktionary
- Task lists/pt/Maintenance/Entries incorrectly marked obsolete
- Task lists/und/Maintenance/Entries with obsolete IPA characters
- Task lists/und/Maintenance/Entries with incorrect language codes
- Task lists
- looking at cattree, a lot of these are marked as .../LIST and .../LIST/description. technical question, for fully automated lists, would it be okay to reverse this paradigm: the bots look into a new LIST/list and that is then transcluded to the main LIST page, where humans can edit. this isolates what the bots edit from everything else.
- another thing. to help with my effort, I would also create a few notice templates to tag some of these:
{{stale list}}— lists that haven't been updated in a long time or are all done{{automated list}}— lists that are regularly automatically updated by a bot{{semiautomated list}}— lists that are generated by generated with a script once in a while- these two should also have the code listed, indicated with a parameter
- what do you think of this? Juwan 🕊️🌈 10:20, 20 June 2026 (UTC)
- Thoughts:
- I am not a fan of nesting using the words "creation" and "maintenance". It makes the titles very long. Is there a reason why we can't put creation lists as subpages of the relevant WT:RE pages?
- I prefer the name "Todo" to "Task lists". It's shorter, and much better established, being the long-standing location for these pages. I'd want to see a strong rationale for moving.
- The main rationale for using a description subpage rather than a list subpage was so people can easily watchlist the lists in the usual way without having to click a special link. The bots only edit the main list page - so it seems already isolated? Is there another reason for swapping it around?
- The suite of templates seems like the Wikipedian's solution. I feel like we can achieve the same outcome with less bureaucracy. For example, getting into the habit of placing the last updated date at the top of the list subpages (which most of them already have) obviates the need for a "stale list" template. And as far as I know, only the "todo lists" project is the only source of automatically updating lists outside userspace.
- This, that and the other (talk) 03:31, 23 June 2026 (UTC)
- Thoughts:
- thank you for the positive response. the plans that I have in mind are to sort through the todo lists and separate them out by language (for now,
- cc also @Erutuon and @Chuck Entz. Juwan 🕊️🌈 10:21, 20 June 2026 (UTC)
This feature was implemented a couple of years ago, and, on the balance, has been quite useful. Basically, any template that links to or displays a term in any specific language as specified by a language code has for many years included script checking in order to display things properly. This goes a step further and compares the actual script with the ones defined for the language in our modules, and adds a maintenance category when they don't match.
I have my pet peeves:
- Diacritics that are perfectly okay in running text get flagged as nonstandard when they're stand-alone entries (probably having something to do with the placeholder used as a base).
- Syntactic variables such as ∅, which aren't really text at all, get flagged: there are whole languages such as Navajo where thousands of entries are flagged as nonstandard solely due to this.
- Pictures/emojis, which arguably aren't any script at all, get flagged
- Likewise, symbols used in astrology, alchemy and other arcane disciplines are flagged.
- If a language has no scripts defined due to being unattested in writing or because we haven't gotten around to setting any, everything gets flagged. The categories for those languages are pretty much useless: as the sayng goes, if everything is special, nothing is.
It does look like there's some kind of exemption for most proto-languages, but that's probably based on their being confined to the Reconstruction namespace. There are categories for Proto-Norse and Proto-Turkic, which both have some attestation in writing.Update: it looks like someone has added the Latin script to a lot of proto-languages since I looked into this last.
That brings me to another issue that's a bit less clear-cut: there are languages that are attested in various historical scripts, but the materials available online seem to be overwhelmingly in transliteration. I stumbled onto one of these today: Khotanese (kho). We list this as having two scripts: Brahmi (Brah) and Kharoshthi (Khar). Of the 300 or so pages thst use the "kho" code, only the following use anything but the Latin script:
- Old Armenian -աւէտ (-awēt). "compare
{{cog|kho|-𑀯𑀻𑀬}}" - Old Armenian ազբն (azbn). "Possibly connected with
{{cog|kho|𑀬𑁆𑀲𑁆𑀩|t=reed}}" - Hebrew פילגש. "(compare "...",
{{cog|kho|𑀧𑀮𑀻𑀓𑀸}}" - Sanskrit गृह्णाति (gṛhṇāti). "
{{desc|kho|گنِک|t=to seize|tr=ganik}}" - Sanskrit सत्तम (sattama). "Cognate with "...",
{{cog|kho|𑀳𑀲𑁆𑀢𑀫||best}}" - Sanskrit सप्तम (saptama). "Cognate with "...",
{{cog|kho|𑀳𑁄𑀤𑀫}}" - Sanskrit नवम (navama). "Cognate with "...",
{{cog|kho|𑀦𑁅𑀫}}" - Sanskrit आगत (āgata). "Cognate with
{{cog|kho|𑀆𑀢}}" - Ancient Greek κάνναβις (kánnabis). "Compare (within the Indo-European language family) "...",
{{cog|kho|𐨐𐨎𐨱|tr=kaṃha}},{{m|kho|𐨐𐨂𐨎𐨦𐨌|tr=kuṃbā}}" - Ancient Greek παλλακή (pallakḗ). "Other connections that have been proposed include "...",
{{cog|kho|𑀧𑀮𑀻𑀓𑀸}}" - Persian پری. "Connections that have been proposed include "...",
{{cog|kho|𑀧𑀮𑀻𑀓𑀸}} - English egg/translations. "Khotanese:
{{tt|kho|𐨀𐨱𐨀}}" - English water/translations. "Khotanese:
{{tt|kho|𐨀𐨂𐨟𐨿𐨕|tr=ūtca}}"
I should mention the variety Old Khotanese (kho-old), which is used on one page:
There are also 28 pages in Category:Khotanese lemmas and Category:Khotanese non-lemma forms, combined, but they're all in the Latin script (they have translations at English fox, English question, English snake, and as "Saka" at English Hotan and English Kashgar).
With the sparsity of Brahmi links (all in etymology and descendants sections except for 2 translations, and all redlinked) and the complete lack of entries in anything but the Latin script, I'm tempted to add "Latn" to the list of scripts for the language- but that opens other cans of worms. There have been times in the past that I've added scripts to the modules, but only in cases where I could find evidence that the languages in question actually used those scripts. Chuck Entz (talk) 22:58, 19 June 2026 (UTC)
Phonetic transcription
[edit]If someone speaks English, they already know, if only subconsciously, that the first t in trout is aspirated, retracted, and affricated, and the last t is unreleased. They know that the r is a devoiced coronal approximant and the ou spells the MOUTH diphthong. Rewriting trout phonemically as /tɹaʊt/ adds nothing to that, while people who don’t know English phonology need to have all that explained: they need a phonetic transcription.
Likewise, if someone speaks Japanese, then they already know that Mount Fuji is pronouced with a bilabial f, an unrounded or compressed u, and a palatal affricate j. If you don’t know Japanese, there’s no way for you to know any of that from the phonemic transcription: /huzi/.
What about the audio clips? Don’t they provide a definitive guide to pronunciation? Well, no - they’re also phonemic. Listeners automatically classify the sounds in the audios into their native phonemes – that’s how our brains work. English speakers will listen to München and hear Moon-chin – they can’t help it (and the German spelling doesn’t help). A Parisian will tell them Pas de problème, and they’ll hear bad problem. My Spanish friends can’t hear the difference between live and leave, or full and fool, without training. We need to be told what to listen for, and that’s the job of the transcription.
That’s why Wiktionary needs phonetic transcriptions. You could use the IPA in that role, but it’s not very accessible. In IPA, trout is [ˈt͡ʃʰɹ̠̥ä̆ʊ̯t̚ ], with 7 symbols and 9 diacritics, and Fuji is [ɸɯ̟ʑi]. In the words of Geoff Lindsey, those look “frighteningly unfamiliar and technical”. The IPA is not made for public use.
There is a better solution, a new alphabet called Musa. Like the IPA, Musa can write all the sounds of all the world’s languages. But unlike the IPA, it’s practical: English written in the IPA is an alphabet soup of diacritics, superscripts, and distorted letters, but English in Musa is clearer than the current spelling. Most important, Musa is much easier to learn, since the letters are featural and iconic: their shapes indicate the phonetic features of the sounds they represent. You can read more about it at www.musa.bet. When used to transcribe pronunciation, we use a version of Musa called the Universal Phonetic Alphabet. The UPA has its own website: www.upa.bet.
Musa is completely unfamiliar; it looks like Martian hieroglyphics! Users would have to learn a new alphabet ... just as with the IPA. The difference is that Musa has no diacritics, no false friends like c j q x y (not to mention ɔ ɟ ɥ ɯ ʌ ʎ), and because it’s featural, it’s much easier to learn. Don’t let the need to learn something new scare you off! We’ve never lived in a world in which so much was new, and everybody born after the internet is used to it.
Let me tell you the story of our numerals. For thousands of years, we wrote numbers using letters like MCXIV; nobody had to learn any weird symbols for numbers. But then in the 12th century, the Hindu numerals we now use – 0123456789 – first came to Europe, and over the next couple of centuries, they completely replaced Roman numerals except for some ornamental uses. How did that happen, when the Roman numerals were familiar letters that everyone knew? Well, it turns out that being better is much more important than being familiar or in widespread use. Without Hindu numerals, it's doubtful we'd have invented the calculus, double-entry bookkeeping, computers, or even license plate, telephone, and credit card numbers.
We’re now in a similar situation with our alphabets. They’re all pretty bad, and we pay the price in years spent learning them and in quirky spelling. Instead of coming up with workarounds like digraphs and diacritics to compensate for their flaws, we should just replace them with a better alphabet, just as we replaced typewriters with computer keyboards once they became available.
If this sounds like something you’d like to investigate further, we’re here to help. The Musa Academy is a Public Benefit Corporation dedicated to the development and promotion of the Musa Alphabet. We can provide almost everything you need: fonts, keyboards, transcribers, legends, and apps. Musa Spell explains Musa transcriptions using examples and audio clips, and it’s also available as a plug-in for your sites and apps. And we don’t charge anything for our services.
Why not take a serious look? Pcyrus (talk) 16:35, 20 June 2026 (UTC)
- Good writing, mate. Looks like a commercial.... let's see if this gains enough traction in 10 years' time. Til then, we'll keep the familiar IPA —User:Vealhurl (talk 22:36, 21 June 2026 (UTC)
- Thank you for the compliment on the writing of this entry. Did you also visit the site?
- Til then, we'll keep the familiar IPA : I checked the first three words of your comment. Only mate offers a phonetic transcription, and in fact it's just the phonemic transcription in [brackets]. You don't seem to share my conviction that the current approach is not an adequate solution. It's ironic, because Wiktionary is in general very thorough when it comes to etymology and definitions, but not pronunciation. Pcyrus (talk) 08:18, 25 June 2026 (UTC)
- Where's this Musa Academy located? Are there real people behind it? Are there peer reviewed publications? Has anyone filed any patents? Has Unicode Consortium already allocated codepoints for the Musa alphabet?
- "And we don’t charge anything for our services" - is there a formal non-revocable permissive license, so that everyone is reassured that there are no strings attached (such as future lawsuits, royalty payments and rent seeking)? --Ssvb (talk) 16:37, 23 June 2026 (UTC)
- https://www.musa.bet/legal.htm Pcyrus (talk) 07:28, 25 June 2026 (UTC)
- The Musa Academy, like Wiktionary, has collaborators all over the world. We think we're real people! :)
- There's a section on Unicode on the Questions page, also referenced in the Subjects index. Did you visit the site?
- If one wanted to make one's fortune by scamming people, there must be easier ways than organizing a collaboration to develop a new alphabet over 15+ years. Pcyrus (talk) 08:25, 25 June 2026 (UTC)
- So that's it? No serious consideration or discussion?
- Am I the only one who feels that Wiktionary's treatment of pronunciation is not up to the same level as the other parts of the entries? Nobody else feels that a direct phonetic transcription would be a clear benefit? Pcyrus (talk) 21:22, 8 July 2026 (UTC)
- Your dedication is admirable, but nobody here is interested in being the guinea pigs for a system without any kind of broad acceptance. — SURJECTION / T / C / L / 22:02, 8 July 2026 (UTC)
- This alphabet looks promising and some thought clearly went into it. I would love to see it gain more traction, but if our dictionary is to serve users, introducing a transcription system that is far more obscure than IPA is not the right move. At least with IPA, there are enough familiar letters to an English speaker that it isn't so hard to learn the rest. Andrew Sheedy (talk) 00:02, 9 July 2026 (UTC)
- Then really use the IPA! As I described above with the example of trout, the current Wiktionary entry offers only the phonemic transcription /tɹaʊt/, which doesn't spell out the pronunciation. In order to discover that the initial t is aspirated and the final t is unreleased - both of which distinguish trout from near-homophones like drowed (invented) - a reader would have to click the (key) link, then read down to the Fortis and Lenis section to learn the correct voicing. Is that really your best solution?
- The complete IPA phonetic transcription [ˈt͡ʃʰɹ̠̥ä̆ʊ̯t̚ ] would be sufficient, but it seems to me impenetrable despite including a few familiar letters. At least it can be displayed here in the browser without adding a Musa font.
- Musa IS obscure, but so is "real" IPA, which is why you're not using it. If you think it's the right solution, try using it! :) Pcyrus (talk) 10:03, 10 July 2026 (UTC)
- I think there are good reasons why the vast majority of dictionaries use phonemic rather than phonetic IPA transcriptions, and the phonetic transcription of trout you gave above is one of them—it’s too complicated and unnecessary for the average user.
- Also, I tried to look at the Musa website, and after no more than three or four clicks was warned that I should not visit certain webpages as they have the wrong certificates and so might be trying to scam me. Not particularly reassuring. And the home page of the website is written in the first person, suggesting this is an individual project rather than a project with wide public participation. All in all, as mentioned above, there isn’t a convincing reason at this time to abandon a well-established international system for an untested (if well-intentioned) one. — Sgconlaw (talk) 12:07, 10 July 2026 (UTC)
- Sorry you had a cert problem: AFAIK, our certs are up to date, and nobody else has mentioned it.
- The Musa Alphabet project doesn't enjoy wide public participation like Wiktionary, but it is a broad collaboration, as shown on the Acknowledgements page.
- Yes, IPA phonetic transcriptions are too complicated, and phonemic transcription - IPA or enPR - isn't sufficient. That's why we need a better phonetic transcription notation. The "good reasons" that reference works choose or the other of the lesser evils is simply that they haven't had a better alternative. But you don't have to abandon the current approach, just supplement it with the latest technology. You can have both the familiarity of phonemic transcription and the precision of phonetic transcription - the best of both worlds. Pcyrus (talk) 14:53, 10 July 2026 (UTC)
- @Pcyrus: the problematic webpage is https://upa.bet/index.htm. Still getting a certificate error despite using a different browser. — Sgconlaw (talk) 15:16, 10 July 2026 (UTC)
- Thank you! Pcyrus (talk) 17:16, 10 July 2026 (UTC)
- @Pcyrus: the problematic webpage is https://upa.bet/index.htm. Still getting a certificate error despite using a different browser. — Sgconlaw (talk) 15:16, 10 July 2026 (UTC)
Proto-Italic voiced *z
[edit]Currently, the PIt entries guidelines state that *z was an allophone of *s word-medially. This is, as far as I can tell, completely inaccurate. The Wikipedia article made the same claim, citing Silvestri (1998), p. 326, who actually only wrote that the development occurred intervocalically. All of the other major sources on the topic make no such claim about the development occurring word-medially in general: Meiser describes the "voicing of *s to z in intervocalic position or adjacent to liquids" in PIt,[1] Bakkum mentions a possible intervocalic voicing of /s/ in PIt,[2] Fortson also only describes the change intervocalically,[3] and Weiss describes rhotacism in Latin only occurring intervocalically.[4] Weiss, Fortson, and Bakkum all also doubt the operation of the sound law in any capacity during the PIt period, but still—even if it did occur—it would only have happened intervocalically and maybe adjacent to liquids. Graearms (talk) 18:20, 20 June 2026 (UTC)
- The sound [z] certainly didn't occur in all contexts word-medially: it remained [s] next to voiceless sounds, e.g. [st sk]. Between voiced sounds, the only place I remember [s] occurring is in [ns] (> L. ns), but there's evidence for the voiced sound in [rz], [lz] (> L. rr, ll), [zr] (> L. br), and [mzr] (> L. mbr). The change of final -rs to r might imply voicing also occurred in final position at some point, e.g. *fars was maybe [farz], although the change of *-ros to Latin -er suggests that sound change was active even at a relatively late point (after syncope of -ro- and e-epenthesis before syllabic r). As far as I can tell, there isn't clear evidence pointing to [sn sm msm] vs [zn zm mzm]. Overall, given the uncertainties, I see the benefit of using a phoneme-based transcription instead, although that would also have implications for how we transcribe other Proto-Italic fricative sounds.--Urszag (talk) 18:08, 22 June 2026 (UTC)
- Stuart-Smith explicitly argues that the voicing adjacent to liquids was an independent development of Latin and could not have occurred in Proto-Italic.[5] However, the evidence they cite is exclusively Umbrian tursitu (see tusetu), which reflects earlier *torsētōd. Interestingly enough, I did find an article by Weiss which claims that intervocalic voicing of *-s- is almost always accompanied by voicing of *s when is postvocalic and presonorant position, for which he does cite evidence in the form of αιζνιω (aizniō).[6] Based on what I've read, Weiss does not appear to treat this development as necessarily Proto-Italic; it need only have appeared at some point in the linguistic prehistory of Sabellic. In the OHCGL, he does ascribe the phonetic development of words such as ahēnus to "compensatory lengthening by loss of [z]."[7] So, based on all of this, we might be able to reconstruct for Proto-Italic [z] appearing when in postvocalic and presonorant position, but not in general when next to liquids. Graearms (talk) 19:07, 22 June 2026 (UTC)
- Also, just to complicate matters, Weiss actually argues that intervocalic *-z- sitll had a slightly different phonetic value than preconsonantal *z.[6] Graearms (talk) 19:09, 22 June 2026 (UTC)
- ^ Meiser, Gerhard (2017), “Chapter VIII: Italic”, in Klein, Jared S., Joseph, Brian D., Fritz, Matthias, editors, Handbook of Comparative and Historical Indo-European Linguistics: An International Handbook (Handbücher zur Sprach- und Kommunikationswissenschaft [Handbooks of Linguistics and Communication Science]; 41.2), volume 2, Berlin; Boston: De Gruyter Mouton, →ISBN, § The phonology of Italic, page 747
- ^ Bakkum, Gabriël C.L.M. (2009), The Latin dialect of the Ager Faliscus: 150 years of scholarship[2], Amsterdam University Press, →ISBN, page 68
- ^ Fortson IV, Benjamin W. (2017), “Chapter VIII: Italic”, in Klein, Jared S., Joseph, Brian D., Fritz, Matthias, editors, Handbook of Comparative and Historical Indo-European Linguistics: An International Handbook (Handbücher zur Sprach- und Kommunikationswissenschaft [Handbooks of Linguistics and Communication Science]; 41.2), volume 2, Berlin; Boston: De Gruyter Mouton, →ISBN, § The dialectology of Italic, page 839
- ^ Weiss, Michael L. (2009), Outline of the Historical and Comparative Grammar of Latin, Ann Arbor: Beech Stave Press, →ISBN, page 81
- ^ Stuart-Smith, Jane (17 June 2004), Phonetics and Philology: Sound Change in Italic[3], Oxford University Press, , →ISBN, pages 114-115
- ↑ 6.0 6.1 Weiss, Michael (2017), “An Italo-Celtic Divinity and a Common Sabellic Sound Change”, in Classical Antiquity, volume 36, number 2, University of California, page 381 of 370-389
- ^ Weiss, Michael L. (2009), Outline of the Historical and Comparative Grammar of Latin, Ann Arbor: Beech Stave Press, →ISBN, page 129
Rhyming reduplications vs. "chosen for the rhyme" compounds
[edit]As I was going through Category:English reduplications, I noticed a few entries that didn't seem like reduplication to me but rather compounds whose components where deliberately chosen to rhyme. Here is a list of those entries:
- arty-farty
- brain drain
- brain gain
- chunky monkey
- claptrap
- cockblock
- cuddle puddle
- double-trouble
- driver reviver
- gang bang
- gender bender
- ginger minger
- happy-clappy
- jelly belly
- jump hump
- motion of someone's ocean
- party hardy
- party hearty
- pill mill
- pocket rocket
- pooper scooper
- ragbag
- rocket docket
- Silly Billy
- slice and dice
- smock frock
- tear and wear
- trolley dolly
- walkie-talkie
- wear and tear
- wheeler-dealer
(Note that I already removed the category from some of them before I got unsure and decided to make this post.)
In my opinion, I would draw the following distinction:
- A rhyming reduplication is typically created by duplicating a term and altering the initial syllable onset of one of the copies. The altered copy is usually a completely new term without its own meaning, and the meaning of the entire compound is very similar to the meaning of the term that was duplicated.
- Examples: fuzzy-wuzzy, piggy wiggy, fancy-schmancy, cutesy-wootsy, easy peasy
- A "chosen for the rhyme" compound is made up out of components deliberately chosen to rhyme with each other, with each component contributing to the meaning of the compound. I would argue that these compounds shouldn't be classified as reduplications.
- Examples: stranger danger, crop top, culture vulture, tramp stamp, cheat sheet
- There are also compounds consisting of semantically meaningful components that seem to be chosen for (or at least reinforced by) their ablaut. I wouldn't classify those as reduplications either.
- Examples: kitty-cat, mixy-matchy, keepie-uppie
- (Edit: Found another one) Slip Slop Slap
- Examples: kitty-cat, mixy-matchy, keepie-uppie
So what do you think? Do you agree? Tc14Hd (aka Marc) (talk) 09:48, 21 June 2026 (UTC)
- I agree that compounds where both elements are words (or are derived from words) that add meaning to the compound don't feel like "reduplications", and there appear to be enough of them to justify a separate category. My initial reaction is that "mixy-matchy" and "keepie-uppie" seem like they could go in the same category as "party hardy", chunky monkey and Silly Billy, and this would also help us avoid swinging too far in the other direction and having too many subtly distinct categories (which people are unlikely to successfully keep separate over time). - -sche (discuss) 18:33, 21 June 2026 (UTC)
- We already have such a category: Category:English rhyming compounds (as well as Category:English rhyming phrases). PUC – 19:01, 21 June 2026 (UTC)
- With recently-improved wording; great; I support (re)moving these to that category (and out of the "reduplication" category), as Tc14Hd was doing. - -sche (discuss) 19:15, 21 June 2026 (UTC)
- I'm glad you agree. By the way, I was not trying to create a new category, I just wanted to remove the reduplication one from the listed entries. But if I were to add a new category, it would probably contain the "chosen for the ablaut" terms, essentially being the non-reduplication equivalent of Category:English apophonic reduplications, since I don't think that these terms fit into Category:English rhyming compounds together with the "chosen for the rhyme" terms. But having only
threefour examples for such a category, I will refrain from proposing it for now. Tc14Hd (aka Marc) (talk) 17:13, 22 June 2026 (UTC)
- I'm glad you agree. By the way, I was not trying to create a new category, I just wanted to remove the reduplication one from the listed entries. But if I were to add a new category, it would probably contain the "chosen for the ablaut" terms, essentially being the non-reduplication equivalent of Category:English apophonic reduplications, since I don't think that these terms fit into Category:English rhyming compounds together with the "chosen for the rhyme" terms. But having only
- With recently-improved wording; great; I support (re)moving these to that category (and out of the "reduplication" category), as Tc14Hd was doing. - -sche (discuss) 19:15, 21 June 2026 (UTC)
- We already have such a category: Category:English rhyming compounds (as well as Category:English rhyming phrases). PUC – 19:01, 21 June 2026 (UTC)
RFC about AI-generated content in Wikimedia Commons
[edit]You are invited to participate in a request for comment on Wikimedia Commons about a policy update for AI content. This may affect files that are uploaded to Wikimedia Commons for use on this project. Thank you. Codename Noreste (talk) 17:12, 23 June 2026 (UTC)
Multiple Derived terms tables for same PoS
[edit]At weasel#Derived terms, there are two tables, one for "Mustelidae", one for "other". When error-free, as this pair is not (cf. honey weasel), this arrangement well suits use cases in which the user is interested in one of the two groupings and needs to avoid distraction. It is not especially well-suited for cases where the user does not know in which table to look for a term or insert a new derived term. Nor is it well-suited for users unsure about "Mustelidae". In addition, it is not well-suited for cases in which a term might appear in both, esp. where one definition should have an {{&lit}} definition, which might make it a member of both (cf. wild weasel). These problems may not seem so bad for terms with only two such tables, but three or more tables for the same PoS are quite possible. We already often combine tables for different PoSes in the same Etymology section.
IMHO, users would be better off with single Derived-term tables, especially where the Derived terms tables are the only places where the term occurs. Also, a single table could not have mistaken placements by contributors, except of the most obvious kind. There would be less temptation for users to add additional tables for finer, more specialized categories.
Does anyone see overriding considerations favoring multiple Derived terms table for single PoSes? Should we only have single tables for multiple PoSes in the same etymology? Should we have one table for all etymologies, possibly augmented by PoS labels for each derived term? DCDuring (talk) 17:26, 23 June 2026 (UTC)
- @DCDuring: if there are many derived terms, it doesn't seem unreasonable to categorize the terms in some manner. Personally, for biology related terms, I put the names of some organisms (usually those of the same genus) under a "Hyponyms" heading, and do not repeat them under the "Derived terms" heading. I also use the "Derived terms" heading for organisms which are not of the same genus but happen to contain the entry term due to some superficial similarity.
- I'm not fond of the practice of combining derived terms from different parts of speech into one big table. I think it is clearer to have separate tables under each part of speech. — Sgconlaw (talk) 19:16, 23 June 2026 (UTC)
- I don't think Translingual taxonomy are much of a problem. I do them the way you do them.
- Vernacular names have many derived terms that don't belong as hyponyms. (Maybe that's the first thing to clean up!) I don't disagree with anything you have said. DCDuring (talk) 23:13, 23 June 2026 (UTC)
- Wonderfool works a loooot with DTs. Sometimes they're under 1 big section, like at across, coz who really cares if they come from prep or adv, right? More of this lazy DTing found with this search, currently 53 cases. Other sloppy cases are found all over. And everywhere there's unordered DT listing, easily fixed with
{{sort}}. Another sloppiness WF causes and solves is the blurring of Hyponyms and DTs. —User:Vealhurl (talk 13:04, 24 June 2026 (UTC)
- I support the convention that a term that is both a hyponym and a derived term is 100% allowed to be list at both spots, because it is each one in its own right — neither spot should be deprived of its proper population, because otherwise there's worthwhile information missing from that spot. Some years ago I was showing them only at hyponyms when hyponymy applied, but I no longer support that notion. As for related terms (not derived terms), it seems superfluous to list those if they're already listed in any of the syn/ant/hyper/hypo/hol/mer/cot/nearsyn spots. The difference is that derivation is worth showing in its own right, regardless. As for the OP question of this thread, I am sympathetic to all of the thinking that has been shared herein so far, and I don't yet have any strong leanings as to a preferred answer for grouping of derived terms versus not grouping them. Quercus solaris (talk) 01:48, 25 June 2026 (UTC)
- PS: God, the devil, a bishop, and a weasel walk into a bar, and hijinx ensue […] Quercus solaris (talk) 05:52, 25 June 2026 (UTC)
- I support the convention that a term that is both a hyponym and a derived term is 100% allowed to be list at both spots, because it is each one in its own right — neither spot should be deprived of its proper population, because otherwise there's worthwhile information missing from that spot. Some years ago I was showing them only at hyponyms when hyponymy applied, but I no longer support that notion. As for related terms (not derived terms), it seems superfluous to list those if they're already listed in any of the syn/ant/hyper/hypo/hol/mer/cot/nearsyn spots. The difference is that derivation is worth showing in its own right, regardless. As for the OP question of this thread, I am sympathetic to all of the thinking that has been shared herein so far, and I don't yet have any strong leanings as to a preferred answer for grouping of derived terms versus not grouping them. Quercus solaris (talk) 01:48, 25 June 2026 (UTC)
ſ (and possibly ꜩ) substitution in German
[edit]Hey guys! Could we get ⟨ſ⟩ and ⟨ꜩ⟩ to substitute ⟨s⟩ and ⟨tz⟩ the way that Latin substitutes macronless vowels in linking templates? As it stands, if you want to write something like {{l|de|ſpiꜩ}}, you have to instead do {{l|de|spitz|ſpiꜩ}}, which may not be per se bad, but could easily be cut down. Thank you!! Ow! That Hurts! (talk) 02:23, 24 June 2026 (UTC)
- Can you give any specific pages where this would be useful? In general, it is possible for spellings with "ſ" and "s" to be kept at different pages, so redirecting doesn't seem like a good idea to me, since ſpitz and spitz might both exist, and in that case someone might truly want to link to the former. (Wiktionary:German_entry_guidelines#Obsolete_spellings and Wiktionary:Quotations#Ligatures_and_archaic_letters are somewhat relevant.) If you instead want to link to the latter, I can envision only limited circumstances where it would be necessary/helpful to display the link as "ſpiꜩ"/"ſpitz". Is the Unicode character "ꜩ" even meant to be used for ordinary ligatures of "t" + "z"?--Urszag (talk) 15:12, 24 June 2026 (UTC)
- Wait, in the other direction, can you give examples of pages spelled with long s that we would want? As far as I know, the prevailing attitude here (contrary to my individual feelings, which fluctuate between ambivalent and weakly inclusionary) has been that a page like [[ſpitz]] should not exist (the only 'wanted' long-s pages I can call to mind are ſ and ſs) — the most recent discussion I recall was Wiktionary:Beer parlour/2025/July#long_s, and before that this, but the exclusion has been a thing for a long time (2011 mention of it, 2013 mention of it, more at WT: "long s"). (We don't even have Wachſtube, where long s actually carries semantic information and forms a minimal pair with the unrelated word Wachstube.)
That said, I agree with you that I'm struggling to think of many situations where we would want a link to display as "ſpiꜩ" or "ſpitz". - -sche (discuss) 17:25, 24 June 2026 (UTC)- I believe this was requested in order to add clickable links to all words in a quotation and having them link to the existing entries with greater ease. While I'm no fan of that practice, I do not oppose such a change because I believe that long s and ligatures such as the one proposed constitute more of a typographical convention rather than distinct obsolete spellings. Catonif (talk) 18:20, 24 June 2026 (UTC)
- @-sche Thanks, I see I was probably mistaken about the permissibility of pages like ſpitz. I had forgotten those discussions and didn't think to look at Wiktionary:Criteria for inclusion/Language-specific for typography practices. (That page seems pretty obscure and incomplete.) I assumed it would be allowed by comparison to the case with Latin v/u and i/j variants or Old English þ/ð variants. If there is actually a firm policy that German page names should never use ſ, then redirects would probably make sense.--Urszag (talk) 18:34, 24 June 2026 (UTC)
- I’m indifferent as to whether we keep or remove them. — Sgconlaw (talk) 04:40, 25 June 2026 (UTC)
- We currently consider at least ⟨ſ⟩ to be a typographic variant of ⟨s⟩. As for ⟨ꜩ⟩, it's up for debate. Ow! That Hurts! (talk) 02:24, 26 June 2026 (UTC)
- I’m indifferent as to whether we keep or remove them. — Sgconlaw (talk) 04:40, 25 June 2026 (UTC)
- Wait, in the other direction, can you give examples of pages spelled with long s that we would want? As far as I know, the prevailing attitude here (contrary to my individual feelings, which fluctuate between ambivalent and weakly inclusionary) has been that a page like [[ſpitz]] should not exist (the only 'wanted' long-s pages I can call to mind are ſ and ſs) — the most recent discussion I recall was Wiktionary:Beer parlour/2025/July#long_s, and before that this, but the exclusion has been a thing for a long time (2011 mention of it, 2013 mention of it, more at WT: "long s"). (We don't even have Wachſtube, where long s actually carries semantic information and forms a minimal pair with the unrelated word Wachstube.)
Manual Rhyme Indexes need to be phased out and eventually removed entirely
[edit]Hey all! I am quite new to posting on Beer parlour and Wiktionary meta discourse as a whole, so please forgive me if any jargon I use is incorrect or nonstandard. I think it is quite evident that rhyme categories make manual rhyme indexes largely obsolete and extremely tedious to edit. All languages with support for rhymes should have all entries converted from the manual index to categories and IPA templates should have rhyme categorization built in, similar to what @Fenakhay did for fa-IPA recently (thank you). Having two entirely separate locations for rhyming is quite confusing. Something like Rhymes: -aɪ for example should link to Category:Rhymes:English/aɪ/1_syllable instead of Rhymes:English/aɪ. Thoughts? SinaSabet28 (talk) 19:02, 24 June 2026 (UTC)
- For some background, Wiktionary:Beer_parlour/2021/August#Automatic_rhymes, Wiktionary:Beer_parlour/2021/August#Retiring_Rhymes:, Wiktionary:Grease_pit/2022/August#Rhymes_and_Category:Rhymes, Wiktionary:Beer_parlour/2023/May#Retiring_the_Rhymes_namespace and probably more. Most opposition has been that these entries sometimes contain additional nuance, as well as redlinks. I support retiring the namespace, we can surely put the redlinks in the category's description (maybe in a standardised format). The process cannot be entirely automated however, so it's a big endeavour. Maybe we can make this change gradually language-by-language, e.g. there are no existing pages in the Rhymes: namespace for Albanian, so changing
{{rhymes}}to link to CAT:Rhymes instead for Albanian specifically would surely raise no objections and require no additional work. Catonif (talk) 19:30, 24 June 2026 (UTC)- I support this fully. Ow! That Hurts! (talk) 00:05, 28 June 2026 (UTC)
Issues with upcoming FWOTD (June 26)
[edit]In a couple days, Ἀθηναΐς (Athēnaḯs) is set to be featured as FWOTD. I have a couple concerns about it and want to know what the rest of the community thinks.
- The substance of the entry (and what will be displayed as FWOTD) is a list of particular individuals. The community consensus has been quite clear that names of individuals are not dictionary material. Do we have a separate standard for LDLs or classical languages? I just removed a number of names at Latin Marinus. Should I not have done that? This kind of information looks nice visually, but feels very encyclopedic.
- The definition states that the name is equivalent to French Athénaïs. Is it really desirable to list an obscure French name as an equivalent in an English-language dictionary? Especially since the French name is already given as a borrowing under derived terms? I was about to boldly remove it, but then discovered that the text is templatized, which suggests that this is a common practice. Should it be?
Let me know your thoughts. I'm not super happy with this entry being featured, especially as it stands, but perhaps I am in the minority. Notifying both Ancient Greek and Latin workgroups, because at least the first issue arises in Latin entries as well. (Notifying workgroup: Mahagaja, Sartma, Theknightwho, Graearms, Exarchus, Fay Freak, Brutal Russian, Benwing2, Lambiam, Mnemosientje, Nicodene, Al-Muqanna, SinaSabet28, Imbricitor, Urszag, Kaloan-koko): Andrew Sheedy (talk) 23:15, 24 June 2026 (UTC)
- I've edited the definition to instead specify that it is equivalent to the English version of the name: Athenais. Graearms (talk) 23:35, 24 June 2026 (UTC)
- To not place undue urgency on its resolution, I will replace the FWOTD with μαστόδετον. TranqyPoo [💬 | ✏️] 00:01, 25 June 2026 (UTC)
- No, they should be removed. It clearly blurs and disintegrates the line between dictionary and encyclopedia or whatever Wikimedia project also may cover them.
- The reference to French was correctly fixed. We only do these kinds of things when they resolve ambiguity, as in the Armenian sociological concept of խմբակ (xmbak) or when we report technical descriptions as in سُكُرَّجَة (sukurraja) and doubt not losing the actual information otherwise or (but this has been circumvented) for words having ambiguous English glosses such as “spring” or “time”, which is naturally not the case with given names. Fay Freak (talk) 00:02, 25 June 2026 (UTC)
- I'm inclined to agree. Maybe one or two names could be mentioned as examples, but the current list is too long. Graearms (talk) 01:33, 26 June 2026 (UTC)
translingual euphemisms
[edit]i created pudenda muliebria as English in 2015, probably after something triggered a memory of the day i was looking up Εὐρώτας in a hardcover Ancient Greek dictionary and saw pudenda muliebria (probably italicized, but i dont have access to the book now) as one of the definitions, instead of an English definition like all the words around it. So I guess I figured it was an English expression even though clearly from Latin. membrum virile is also listed as English.
But just now I came across membrum puerile and penis puerilis, which are the same type of words, but we list them as Translingual. It seemed wrong to me at first, but it makes sense since they behave more or less like the scientific taxon names we're so used to which are based on Latin but can appear in running text in any language.
Should the earlier entries be moved to Translingual as well? —Soap— 03:26, 25 June 2026 (UTC)
- I think it depends somewhat on usage. Were these euphemisms used in any other language? If it's just a matter of English speakers of a certain era substituting Latin words because that's the language English-speaking adults would have been taught as part of their higher education, while speakers of other languages would have used something else, an argument could be made for it being a pseudo-loan within English. If the practice occurred in multiple languages indepently, it was probably translingual. Of course, culture played its part: the influence of the Puritans in England and New England made it necessary for generations of English speakers to code-switch into French and/or Italian to talk abut sex. Chuck Entz (talk) 05:27, 25 June 2026 (UTC)
- They were intended translingual – it depends not that much on usage – and used translingually, like names of muscles (e.g. in the form musculus detrusor vesicae), diseases (see in the translations of bejel, where things only found in German medical literature was still included as translingual, though I don’t fully trust Google Books to comprehensively cover literatures of other European countries), bacteria, apothecary concoctions, etc. Fay Freak (talk) 11:02, 25 June 2026 (UTC)
Sign language quotes (and also example sentences)
[edit]Neither sign language nor quotation guidelines specify how to dealt with the primary text--whether by glosses or each sign language word. To quote in sign language, we could have use these, arranged in descending accuracy: 1) entry names (such as 5@Chin-FingerAcross Twist) 2) sign gloss (such as MACAU) 3) literal translation (such as Macau, in some cases). Various questions have been answered in the talk page of TranqyPoo. My viewpoint is that since English translation is required for any non-English, we could use 1) entry names and 2) sign gloss to ease the learning for the readers. Beefwiki (talk) 04:09, 25 June 2026 (UTC)
- (providing context) In a nutshell, they are asking: What would the optimal sign language quotation look like? From my inexperienced viewpoint, it would be most ideal to have an image for each sign (as a person who signs the language would interpret it) as the original text, then its natural translation into English, followed by its literal translation using sign glosses. However, sign images as primary text is likely impossible at this point in time. So, I recommended instead to use Wiktionary's entry name as the primary text, where each term is wikilink'd. Feel free to disregard my viewpoint as its purpose is to get people thinking. Also, see this dictionary for further reading.
- Beefwiki, I hope that you find what you are seeking :) TranqyPoo [💬 | ✏️] 04:32, 25 June 2026 (UTC)
Implausible Turkish causatives for deletion
[edit](Notifying workgroup: İtidal, Fytcha, Vox Sciurorum, Lambiam, Whitekiko, Ardahan Karabağ, Orexan, Moonpulsar, Lagrium, Trimpulot, Sedataltundal): (Thanks to Orexan for this workgroup)
Here's my list for deletion of very strange Turkish causatives added years ago. The first list has double causatives (including causatives of passives and vice versa), and the second one has causatives for verbs I wouldn't ever see anyone causativize. I had to search for these using keywords because they were never properly categorized as words with causative suffixes, so I may have missed some. Please point out any that I may have missed.
yontturtmak, süslettirmek, ayarlattırmak, parçalattırmak, büktürtmek, alındırtmak, geciktirtmek, öldürtülmek, tüttürtmek, aşık attırtmak, sahnelettirmek, ışıttırmak, verdirtmek, ittirtmek, canlandırtılmak, süründürtmek, titrettirmek, yoldurtmak, saydırtmak, programlattırmak, ölçtürtmek, attırtmak, kaynaştırtmak, kaynaştırtılmak, birleştirtmek, ağlattırmak, yendirtmek
yontturmak, aşık attırmak, yoldurmak, programlatmak, sorumsuzlaştırmak
Some of these even have (badly written) usage examples somehow, including verdirtmek, which is apparently slang (!).
I didn't include yaptırtmak or yedirtmek because I think the base verbs are common enough that double causatives might be encountered. Wreaderick (talk) 10:15, 25 June 2026 (UTC)
- Alındırtmak is now defined as “to find somebody who lets someone resent by somebody else” 🤣🤣🤣.
- But alındırmak can be attested:
- These are natural, rather transparent, causatives of different senses of alınmak, including, in the last example, of the idiomatic sense of “to feel offended” – “have I made Him feel offended?”. We are better off without this entry than with the current definition of “to let someone made resent by somebody else” 😂, but the idiomatic sense may warrant inclusion.
- Also, uses of öldürtülmek as the passive of öldürtmek are easily attested:
- I’d translate these as “he was ordered killed” except in the last, pleaonastic example, where the order is already explicit by emriyle. Since this is an entirely transparent use of the passive (Navalni'yi Putin mi öldürttü? → Navalni Putin tarafından mı öldürtüldü?), I don’t think it will be missed, though.
- I see no problem in keeping yoldurmak, which is just a plain simple causative. ‑‑Lambiam 12:05, 25 June 2026 (UTC)
- Yeah soon after I sent my original message I realized that alındırmak can mean to offend, I'm taking that one back.
- Since we can attest öldürtülmek, despite how unwieldy it sounds, we should keep it. I can also see that yoldurmak is attested.
- I presume you agree with deleting the rest of the plain causatives? Wreaderick (talk) 12:14, 25 June 2026 (UTC)
- Sure. I spotted two uses of programlatmak,[11][12] but even if an unlikely few more can be attested, it is clearly not in widespread use. Removing this crud will be a considerable improvement, and particularly so in light of definitions in the style ”to have been found by someone to have let someone else to have it made become an entry in Wiktionary”. ‑‑Lambiam 15:20, 25 June 2026 (UTC)
- As a native speaker I find it hard to follow the plot for these structures, I find myself trying to figure it out like a math problem. Like,
- yontmak (“to chisel sth.”) → yontturmak (“to have sb. chisel sth.”) → yontturtmak (“to have sb. have sb. chisel sth.”)
- gecikmek (“to be late”) → geciktirmek (“to cause sb./sth. to be late”) → geciktirtmek (“to have sb. cause sb./sth. to be late”)
- acıkmak (“to be hungry”) → acıktırmak (“to make sb. hungry”) → acıktırtmak (“to have sb. make sb. hungry”) etc.
- The situations where a double causative is grammatically proper and sounds natural change from very unlikely to borderline impossible. In my opinion, most times a double causative is used, it is used incorrectly where the speaker simply meant "they had someone do it" rather than "they had someone have someone do it." Who hires a person to hire a person to carry out a task?
- There is something to be said about the frequency of some of the verbs themselves. "chiselling" for instance isn't exactly something any ordinary person does on a daily basis on their own, let alone have someone else do it. But there aren't many like that, most are common enough in their base forms. Also, there are some structurally causative verbs that have either completely lost their causative meaning over time and been established as transitive, or gained transitive senses over time like öldürmek (“to cause to die → to kill”), doğurmak (“to cause to birth → to birth, to give birth to”), doyurmak (“to cause to be full → to feed”) etc. So when they get the causative suffix, it's not necessarily double.
- Since the list is kinda long I feel it would be easier to determine the ones we should keep. Here's my white list I guess;
- öldürtülmek; As said before; it's semantically a single causative and the passive form of it is fine.
- saydırtmak; Same as above, saydırmak can be used transitively as "to throw a series of insults at," so the pseudo-double-causative "saydırtmak" can be encountered, but markedly rare.
- yoldurmak (“to have/make some pluck sth.”), programlatmak (“to have/make someone program”), sorumsuzlaştırmak (“to cause/make someone to become irresponsible”); These are single causatives, so if we go by principle we should keep them. They're all sort of rare but sorumsuzlaştırmak particularly has very little practical use to it. I don't think anyone would miss it. I think we can keep the first two, though.
- çıldırtmak; çıldırmak is structurally causative but semantically intransitive, definitely attestable.
- indirtmek; indirmek has transitive senses, can be encountered.
- işittirmek; This is a single causative for all intents and purposes.
- karıştırtmak; karıştırmak has transitive senses, can be encountered.
- koşturtmak; koşturmak has transitive senses, can be encountered.
- söndürtmek; söndürmek has transitive senses, can be encountered.
- Apart from the above, I think the rest may slip peacefully beyond the veil into the hereafter. Some thoughts about a few of the entries;
- şaşırttırmak; This one is a litle bizarre, seemingly şaşmak (“to marvel, to be baffled”) is the base form; şaşırmak (“to be surprised, to get mixed up”) is the causative, but used synonimously as the base form; şaşırtmak (“to surprise, to confuse sb.”) is the double causative, but has transitive senses evolved from the causative "to cause someone to be surprised/get mixed up." This is giving me a headache. So the triple causative sense "to make someone cause someone to be surprised/get mixed up" feels overly complicated. Apparently there is a folk song titled "Kahpe Felek Şaşırttırdı Yolumu" but like most double (or triple) causative usages, it is used incorrectly. Fickle Fate causes the speaker to get his path mixed up, rather than make someone else cause the speaker to get his path mixed up, at least it's not specified.
- aldırtmak; Clear case of double causative, even though the frequency of the base verb makes it confusing, I suspect most double causative usages that we may find are single in essence. It has to be a sentence where person A is making person B make person C take something.
- Orexan (talk) 20:55, 27 June 2026 (UTC)
- A while ago I tried to add some senses to şaşırmak but it was simply too confusing, with it being essentially interchangeable with şaşmak in some of them. I find the causative to be an especially versatile and mercurial suffix, and yesterday I added a definition of its use in more generic verbal derivation.
- And I agree with you when you say that most speakers who use the double causative use it "incorrectly", as in, they use it to mean the same thing a single causative would. I think we can even use this to argue for their inclusion, as long as we can attest them and not give their definitions as a causative of another inherent causative (except for perhaps the most common verbs such as yapmak and almak). It's just a bit of grammatical redundancy, something that happens all the time. We might even lemmatize -irt as a single suffix.
- Off topic: You certainly have a way with words. I find your comments colorful and interesting to read. Do you read a lot? Wreaderick (talk) 22:18, 28 June 2026 (UTC)
- Interesting idea to keep them as incorrect usages, I don't know if there's a precedence for it on the site. There are form of templates for misspellings, informal, nonstandard or alternative forms, but not specifically for incorrect usage of a word. As far as I know, misspellings and informal/nonstandard forms are included because they're shared by regional, colloquial, generational etc. dialects. I don't know the double causative structures have the user base that share a similar common trait. I think we should just get rid of them.
- I studied English Literature but my comments are the way they are because I'm a clinical overthinker. Thanks. Orexan (talk) 17:19, 30 June 2026 (UTC)
I found one, geciktirtmek, I'll add more if I find later. Lagrium (talk) 11:59, 25 June 2026 (UTC)
Here are some more double causatives:
- acıktırtmak, aldırtmak, anlattırmak, bitirttirmek, büktürtülmek, canlandırtmak, çıldırtmak, daldırtmak, ettirtmek, heyecanlandırtmak, indirtmek, ittirtilmek, işittirmek, kaplattırmak, karıştırtmak, kovdurtmak, koşturtmak, kırdırtmak, nefes aldırtmak, sattırtmak, söndürtmek, şaşırttırmak, yolundurtmak.
A few of these are certainly or probably worth keeping, but even among these some have unusable definitions. ‑‑Lambiam 16:32, 25 June 2026 (UTC)
- çıldırtmak is fine because *çıl- isn't a proper verb on its own anymore, çıldırmak is therefore not a causative. I also find aldırtmak to be likely okay because almak is a very common verb.
- Additionally, several of these have base verbs that are intransitive, so double causatives aren't implausible for all of them. The ones I find natural are söndürtmek, indirtmek, and koşturtmak. That last one is especially because koşturmak can also mean "to run around; to rush, to do something while rushing", so you could say something like
- Çocukların okul işleri beni bütün gün koşturttu. ― The kids' school stuff kept me rushing around all day.
- I find that the causative suffix often doesn't even create a causative but is in fact a more meaningless verbal derivation suffix and I plan on adding such a definition. Wreaderick (talk) 17:09, 25 June 2026 (UTC)
This discussion has died down, but something does have to be done about these terms. I propose defining morphologically double causatives (that are also ostensibly semantically a proper double causative, unlike öldürtmek) that have been attested in a semantically singular sense as alternative forms of morphologically singular causatives. So if we can attest aldırtmak to simply mean “to make take” we will define it as an alternative form of aldırmak. Otherwise they will be deleted. Single causatives that have already uncommon verbal roots, such as yontturmak, will also be deleted.
We will also define -irt as an alternative form of -ir. We may be able to attest yaptırtmak or other double causatives with similarly very common verbal roots as semantically double causatives.
I'll be glad to hear any objections anyone may have.
Deployment of Legal and Safety Contacts Link in the Footer of Your Wiki
[edit]Legal & Safety Contacts
Hello community, the Wikimedia Foundation has provided a single legal and safety contact page, to be linked in the footer of your wiki, to ensure access to accurate legal information. This is a regulatory requirement. We have already rolled out links to English, German, Italian, Spanish and other wikis and we will deploy to your wiki soon. Please read more on the project page and leave any comments in this thread or on the talk page.
-- User:Sannita (WMF) (talk) 13:31, 25 June 2026 (UTC)
Why did the amount of English gloss entries decrease?
[edit]It’s now listed as 707,993 at WT:Statistics. I swore it was like 900,000+. Inpacod2 (talk) 22:35, 26 June 2026 (UTC)
- As noted here, User:Jberkel changed the way that the entries are counted. ―Justin (koavf)❤T☮C☺M☯ 23:13, 26 June 2026 (UTC)
- So is it safe to say that the 707,993 count is more accurate? Inpacod2 (talk) 00:13, 27 June 2026 (UTC)
- ¯\_(ツ)_/¯ probably. JB is smarter than I am. ―Justin (koavf)❤T☮C☺M☯ 00:21, 27 June 2026 (UTC)
- So is it safe to say that the 707,993 count is more accurate? Inpacod2 (talk) 00:13, 27 June 2026 (UTC)
- (For some background—though admittedly not much—see Wiktionary:Beer parlour/2026/May#10 million, baby!.) - -sche (discuss) 17:58, 27 June 2026 (UTC)
- It has dropped because the definition of what constitutes an "entry" has changed: the old statistics counted each POS header (L3/L4) as a separate entry, and now each L2 language header == one entry (as it should be, and as defined in the glossary). If you want, I can re-add POS-level stats, but then the table quickly gets unwieldy. Jberkel 19:36, 27 June 2026 (UTC)
- I agree with the change overall. I do think it's a shame that different etymology sections (homographs) aren't counted as separate entries, which they would be in a print dictionary, but I don't know how involved it would be to count them. The update is definitely an improvement. Andrew Sheedy (talk) 06:01, 28 June 2026 (UTC)
- It can be done, maybe it could be integrated as a hidden hover/separate table. But the main count needs to be consistent with Module:EntryCount. Jberkel 07:40, 28 June 2026 (UTC)
- I agree with the change overall. I do think it's a shame that different etymology sections (homographs) aren't counted as separate entries, which they would be in a print dictionary, but I don't know how involved it would be to count them. The update is definitely an improvement. Andrew Sheedy (talk) 06:01, 28 June 2026 (UTC)
- It has dropped because the definition of what constitutes an "entry" has changed: the old statistics counted each POS header (L3/L4) as a separate entry, and now each L2 language header == one entry (as it should be, and as defined in the glossary). If you want, I can re-add POS-level stats, but then the table quickly gets unwieldy. Jberkel 19:36, 27 June 2026 (UTC)
Category: Year or Decade or Century of Earliest Attestation
[edit]I know the year of origin for some words, while for others I can determine only the decade or century. Many dictionaries already include this information in their entries, but Wiktionary currently has no systematic way to organize it.
Editors and readers currently have no way to browse English vocabulary chronologically despite this information already existing in many entries.
I propose creating categories for years, decades, and centuries of first attestation (or earliest known use), with entries categorized at the most precise level supported by reliable sources. For example, a word with a known first use in 1847 could be placed in "English terms first attested in 1847," while a word known only to date from the 1840s or the 19th century could be categorized accordingly.
Such a system would have several benefits. It would provide a bird's-eye view of lexical change over time, making it easier to see periods of rapid vocabulary growth, the emergence of scientific and technical terminology, or shifts associated with historical events and cultural movements. It would also make Wiktionary more useful for linguistic research, allowing editors and readers to browse words by period of origin rather than only by etymology or semantic field. Finally, it would encourage editors to document earliest attestations more consistently, improving the quality and completeness of etymological information across the dictionary.
Also, the category for the current year and immediately previous years would probably be heavily viewed and more closely investigated and better understood, including hot words. We could see the trends each year right in front of our eyes.
The exact categorization scheme could be discussed, but I think organizing entries by year, decade, and century, maybe dynasty for China, depending on the precision supported by the evidence, would be a useful addition. (AI-assisted drafting) --Geographyinitiative (talk) 🎵 23:20, 28 June 2026 (UTC)
- Looks like something that would be more efficient in a later generation, when there are more quotes filled, probably with the help of AI agents. Fay Freak (talk) 11:50, 29 June 2026 (UTC)
- No objection in general, but I think the data is lacking for most of our entries. Maybe categories could be added based on
{{defdate}}templates in entries, but as these are added according to sense rather than entries as a whole (which I think is correct), there would need to be a way to determine which is the earliest attestation. Another way is to base the categorization on{{etydate}}, though personally I think the use of this template is less useful than indicating the earliest attestation of each sense of a term. Perhaps{{etydate}}can be tweaked so that there is an option for it to produce no visible output but just categorize the term. — Sgconlaw (talk) 13:06, 29 June 2026 (UTC)
- No objection in general, but I think the data is lacking for most of our entries. Maybe categories could be added based on
@Fay Freak, Sgconlaw, Vealhurl Thank you for your replies. Could we maybe start with categories for words with earliest attestation in 2026, 2025, and 2024, and see how it goes? I guarantee if we can "start" it with the recent years, people will really get into it and it will become a resource unique in this world, and a powerful engine for studying the earliest attestations of words. --Geographyinitiative (talk) 🎵 23:47, 30 June 2026 (UTC)
User:Box16
[edit]I propose that Box16 (talk • contribs) be banned from this project for their edits, which not adhering to our policies based on established consensus, are tantamount to vandalism. There have been many issues raised in their userpage over the years, which they conveniently delete to hide their nefarious record— making it easy to fail to notice their behavior that seriously mars the quality of this project. I am particularly frustrated and appalled by their disruptive changing of {{alternative spelling}} to {{alternative form}} when it’s not applicable, which they do wholesale as they bear a personal aversion toward the alternative spelling template, inspite of my multiple headsups in their talkpage (where I once explained to them in which cases the appropriate templates belong) and edit summaries urging them not to engage in these edits. It is more than clear that this user, who also constantly alter their username very likely to elude vilgilance, disregard us and our policies, so it’s high time they had been spurned from the threshold of our community for good. Inqilābī 23:58, 28 June 2026 (UTC)
- Are you serious?! You're the editor who doesn't even understand the basic difference between an alt. form and and alt. spelling. box16 (talk) 00:12, 29 June 2026 (UTC)
- Btw, you just wrote in an edit summary telling me go to prison? Are you sure you're not the one who needs to be banned from the project? box16 (talk) 00:14, 29 June 2026 (UTC)
- Telling an editor "go to prison soon dude" is a direct threat and a violation of the community standards of Wiktionary. box16 (talk) 00:16, 29 June 2026 (UTC)
- I can sue you in court for draining my energy, time and health over containing your malicious actions to a project. So you literally belong in a prison for your deed anyway, no kidding or threats here sweety. Inqilābī 00:20, 29 June 2026 (UTC)
- I think you need to seek professional help. box16 (talk) 00:22, 29 June 2026 (UTC)
- I can sue you in court for draining my energy, time and health over containing your malicious actions to a project. So you literally belong in a prison for your deed anyway, no kidding or threats here sweety. Inqilābī 00:20, 29 June 2026 (UTC)
- Telling an editor "go to prison soon dude" is a direct threat and a violation of the community standards of Wiktionary. box16 (talk) 00:16, 29 June 2026 (UTC)
- @Benwing2 This longterm disruptive editor isn’t kind to us, they need a longer block at least. Inqilābī 16:44, 29 June 2026 (UTC)
- How am I not kind when I never engage in name-calling and threats (unlike some editors such as Equinox), while you literally told me yesterday "go to prison soon dude"? box16 (talk) 17:17, 29 June 2026 (UTC)
- It was but a harmless edit summary, and you’re exaggerating its gravity in a bid to divert attention from the issue raised here about you and create a separate drama targetting me- like you did when I clarified it and you elsewhere used the clarification to draw more attention. And you’re not kind for being a malicious vandal with the mental age of a preteen, no offense or personal attack intended but it’s the bitter truth I have concluded from your immaturity, hostility and insincerity. I wish you were not as bad and I was lying. Inqilābī 17:27, 29 June 2026 (UTC)
- No, telling someone to go to prison soon dude and that I literally belong in a prison for your deed is a pretty serious statement (and also quite juvenile) that is both unbecoming of a polite and professional Wiktionary editor and also a borderline violation of our community guidelines.
- I have never been nasty or unkind to any of the editors here, and in my humble opinion, still vehemently disagree with your position regarding the correct uses of the alt. form and alt. spelling. Adding a hyphen or a space to a word rarely renders it to be an orthographic variant (alt. spelling), but simply an alternate variant, (alt. form) conforming to Hftf's position. box16 (talk) 17:37, 29 June 2026 (UTC)
- Nope: diff. Inqilābī 17:43, 29 June 2026 (UTC)
- What he wrote is what I just stated in my response above. box16 (talk) 17:46, 29 June 2026 (UTC)
- He wrote: "Spelling is the way words are formed with letters. For other alternatives, use
{{alternative case form of}}or{{alternative form of}}". Hyphens and spaces aren't letters so this seems very clear cut." box16 (talk) 17:53, 29 June 2026 (UTC) - Spaced and hyphenated words constitute alternative spellings and never forms, and no one save you fails to acknowledge it. It’s misinterpretation on your part which you’ve chosen to stubbornly cleave to over the years— but it’s never late to admit your mistake, and we’ll forgive and forget. I was the one who added the welcome message when you joined Wikt. I'm always kind to everyone who deserves it. Nevertheless, failure to grasp a simple thing as this suggests you’re simply not fit to be a Wiktionary editor. Inqilābī 17:55, 29 June 2026 (UTC)
- Once again, that's the polar opposite of what Hftf wrote and you seem to have an issue with reading comprehension. I think it's best you just kindly leave me alone. box16 (talk) 18:10, 29 June 2026 (UTC)
- P.S. In your first post above, you stated that in the past I have conveniently deleted my talk page (or portions of it) to hide my nefarious record. That's quite ironic and bizarre for you to make such an accusation, when I have just perused your talk page, and it is literally pockmarked with you deleting past discussions in which you have engaged in very questionable and impertinent discussions with other editors, many of which clearly have a very provocative and somewhat aggressive tone and content. box16 (talk) 19:05, 29 June 2026 (UTC)
- I suppose the only way to end this crisis is through voting. And I’ve replied in my talk because you raised the same thing twice. Inqilābī 19:42, 29 June 2026 (UTC)
- Please just cease this discussion and move on. box16 (talk) 19:44, 29 June 2026 (UTC)
- This is not any discussion anymore because you cunningly converted it into a drama. I may start an abrupt vote in future even though there wasn’t enough fruitful discussions— best idea considering how the discussion got thwarted. Inqilābī 23:16, 29 June 2026 (UTC)
- I do not support box16's actions but I think you are mostly to blame for the derailing of the discussion. I've noticed Wiktionary has a very major toxicity problem and I would like for you to please refrain from threatening editors. BirchTainer (talk) 07:38, 1 July 2026 (UTC)
- This is not any discussion anymore because you cunningly converted it into a drama. I may start an abrupt vote in future even though there wasn’t enough fruitful discussions— best idea considering how the discussion got thwarted. Inqilābī 23:16, 29 June 2026 (UTC)
- Please just cease this discussion and move on. box16 (talk) 19:44, 29 June 2026 (UTC)
- I suppose the only way to end this crisis is through voting. And I’ve replied in my talk because you raised the same thing twice. Inqilābī 19:42, 29 June 2026 (UTC)
- He wrote: "Spelling is the way words are formed with letters. For other alternatives, use
- What he wrote is what I just stated in my response above. box16 (talk) 17:46, 29 June 2026 (UTC)
- Nope: diff. Inqilābī 17:43, 29 June 2026 (UTC)
- It was but a harmless edit summary, and you’re exaggerating its gravity in a bid to divert attention from the issue raised here about you and create a separate drama targetting me- like you did when I clarified it and you elsewhere used the clarification to draw more attention. And you’re not kind for being a malicious vandal with the mental age of a preteen, no offense or personal attack intended but it’s the bitter truth I have concluded from your immaturity, hostility and insincerity. I wish you were not as bad and I was lying. Inqilābī 17:27, 29 June 2026 (UTC)
- How am I not kind when I never engage in name-calling and threats (unlike some editors such as Equinox), while you literally told me yesterday "go to prison soon dude"? box16 (talk) 17:17, 29 June 2026 (UTC)
- What a coincidence. I point out that the same day I emprised to disentangle the distinctions between
{{alt form}}and{{alt sp}}practiced by different editors: User talk:Alexander Patmos#Hyphenation and spacing variants. Fay Freak (talk) 11:48, 29 June 2026 (UTC)- For those of us who are seeking to disentangle the distinction, it may be noted that the terms alternative forms/spellings themselves are possibly confusing, which unfortunately causes some to mistakenly assume that the difference itself is vague and trivial. To my mind, alterating the template namings to more well-defined vocabulary, namely “orthographic variation [of]” and “morphological variation [of]”, would resolve this crisis by ensuring the baffled editors finally fathom what’s the point we have already enforced through our guideline rules. Inqilābī 16:34, 29 June 2026 (UTC)
- @Fay Freak, Alexander Patmos, hftf: Inqilābī 16:36, 29 June 2026 (UTC)
- I did suggest in the past that if we want a distinction between these templates to be maintained, we need to spell it out in their actual wording: Wiktionary:Beer parlour/2024/February#alternative forms and alternative spellings; a plan to do that made some progress but then stalled because volunteer time is limited. - -sche (discuss) 16:49, 29 June 2026 (UTC)
- @-sche: Thank you, that has to our ultimate goal, by way of community consensus and reviving the discussion. Well for now, I’m very concerned about a newer and uninformed user like Box16 disregarding the difference by not bothering to understand older editors here, as they regularly revert my corrections of the pages they purposely modifies. I’m saying as someone who has spent time over the years fixing the wrong usage of the altform template in entries *created or edited by benevolent users*, and this problematic user is simply on a spree since 2024 affecting every entry with altsp they can spot, whether the ones corrected by me or already accurate ones. This I consider a grave issue not otherwise noticable to most and hence gone unchecked. Inqilābī 17:19, 29 June 2026 (UTC)
- I have no interest in this weird and exaggerated drama over alt templates (don't ping me again) but OP should be clearly BOOMERANGed; altering usernames to elude vigilance sounds like some long-term vandal thing anyway Hftf (talk) 19:12, 29 June 2026 (UTC)
- Thanks for the trollish comparison and slander. It was a different account of mine and I did minor disruptive edits with it for my lack of understanding of standard English. It may look like a drama because of how box16 changed the topic successfully, much to the detriment of an important issue. People of your mentality have daunted me here, so I have been inactive for some years. I came back only with grave misunderstanding of my words. I rightfully pinged you because you mentioned my discussion elsewhere and are actually involved in the template discussion. Inqilābī 19:27, 29 June 2026 (UTC)
- I have no interest in this weird and exaggerated drama over alt templates (don't ping me again) but OP should be clearly BOOMERANGed; altering usernames to elude vigilance sounds like some long-term vandal thing anyway Hftf (talk) 19:12, 29 June 2026 (UTC)
- @-sche: Thank you, that has to our ultimate goal, by way of community consensus and reviving the discussion. Well for now, I’m very concerned about a newer and uninformed user like Box16 disregarding the difference by not bothering to understand older editors here, as they regularly revert my corrections of the pages they purposely modifies. I’m saying as someone who has spent time over the years fixing the wrong usage of the altform template in entries *created or edited by benevolent users*, and this problematic user is simply on a spree since 2024 affecting every entry with altsp they can spot, whether the ones corrected by me or already accurate ones. This I consider a grave issue not otherwise noticable to most and hence gone unchecked. Inqilābī 17:19, 29 June 2026 (UTC)
- I don't think the discussion here ultimately supports a site ban. The underlying dispute concerns the interpretation of the Template:alternative form and Template:alternative spelling templates, which is an interesting content question but not one that, I think, justifies treating an editor as an actual rogue vandal. I'm not entirely certain of the distinction myself and am 100% open to revising my understanding. I'm kinda just "winging it" and just put what seems right. This appears to be a genuine content disagreement—or at least an area of uncertainty—especially given that multiple editors have observed that the distinction has been confusing enough to prompt repeated discussions about clarifying the documentation, including discussions I have participated in. Where the documentation has not been fully clarified and has been interpreted differently by reasonable editors, it is difficult to characterize edits made under a different interpretation as vandalism or bad faith.
- As for the tone of the discussion, I don't personally mind a certain amount of lighthearted name calling, but Wiktionary's norms generally call for discussions to remain civil and focused on content rather than contributors. I think everyone involved would be better served by avoiding personal remarks, it's all fun on a volunteer website.
- If there is concern that Box16 is making systematic errors, an appropriate solution is for the community to try to establish a clear consensus on the proper use of these templates (if that can be done). A site ban is generally reserved for persistent disruptive conduct that continues despite consensus and administrative intervention, not for a good-faith disagreement over how two totally unclear templates whose scope has not been fully clarified should be applied. (AI-assisted drafting)--Geographyinitiative (talk) 🎵 00:25, 30 June 2026 (UTC)
- Side note: despite multiple warnings from multiple people he is still changing British to American spellings, e.g. today [13]. ~2026-37577-52 (talk) 20:25, 30 June 2026 (UTC)
- Because the entry was not linked to the main article. That is a legitimate reason to change it. Btw, you make plenty of errors on here as well for which no action has ever been taken. box16 (talk) 20:31, 30 June 2026 (UTC)
- The entry flavouring is an alt. form entry, so that is unhelpful to the peruser. box16 (talk) 20:34, 30 June 2026 (UTC)
- I completely refrain from changing UK spellings to US spellings unless I see the given spelling links to an alt. entry article. box16 (talk) 20:53, 30 June 2026 (UTC)
- Because the entry was not linked to the main article. That is a legitimate reason to change it. Btw, you make plenty of errors on here as well for which no action has ever been taken. box16 (talk) 20:31, 30 June 2026 (UTC)
- He's also been asked many times to stop making small unnecessary changes that break the grammar or result in bad English. His English is not great. Today another example. It's so tiring after so many months of seeing this. [14] ~2026-37910-27 (talk) 08:41, 3 July 2026 (UTC)
- What are you talking about?
- “that which is mentioned” is redundant because “that” and “which” are both relative pronouns doing the same job
- “that which” together is usually reserved for more formal or archaic styles box16 (talk) 08:43, 3 July 2026 (UTC)
- @Box16: I think that definition was a bit difficult to parse (which is why you didn’t parse it correctly), so I’ve added commas for clarity. The definition is trying to say “One who [in other words, a person], or that which [in other words, a thing], is mentioned. — Sgconlaw (talk) 11:54, 3 July 2026 (UTC)
- @Box16: Your edit made the entry ungrammatical. It is not true that "that" was a relative pronoun in this definition, it was a demonstrative pronoun, which was functioning parallel to "one" (while "who" and "which" were functioning parallel to each other). The definition as you left it was comprehensible, but did not sound like native English. Andrew Sheedy (talk) 19:04, 3 July 2026 (UTC)
- Thank you for the clarification. I'm just wondering, why does Equinox continue to behave like an admin? He was desysopped over two years ago, no longer edits from his "Equinox" username, but regularly and routinely peruses our project in an admin-like fashion, (under an anonymous IP), and also has a very bad habit of following me around everywhere and scrutinizing every little thing I do. He still has not apologized to me for calling me a "c--t", apart from his other horrible insults (including insulting many other editors). box16 (talk) 19:13, 3 July 2026 (UTC)
- Seems like the nasty-minded specifically edit as IPs so they can have less regard on how they behave and hit below the belt. On the other end, real names are preferred by longstanding editors such as Chuck Entz here, opining that people lose posture in the “anonymity”, but evidently there is a sense behind constant pseudonyms whereby people build reputation and credibility, to detach oneself from the subjectivity of one’s personal position, which is proven to work in academia and journalism, where the name of a random researcher may not say much or be blinded in peer-review, and increasingly relevant in contrast to bots and keyboard armies sweeping the internet nowadays, so there is even more inhibition to be irresponsibly inaccurate or injurious: unlike when one is precariously engaged by a marketing department, an intelligence agency, a shortselling institution or one’s own vanity to spread FUD and fill one's pockets or self-conceit about how great one is at making a fool out of everyone. I had more respect for the insufferable personalities having skin in the game, no need to go all in with full nudes. Fay Freak (talk) 11:12, 4 July 2026 (UTC)
- FF,
- I admire and appreciate your level 99 prolix purple prose, but the crux of the issue is Equinox and his adminlike behavior (including constantly monitoring and scrutinizing all of my edits) which he should have peeled away and forsaken back in 2024 when he was desysopped and abandoned his handle entirely. box16 (talk) 18:28, 4 July 2026 (UTC)
- Or he should continue age-appropriately and reregister under a new branding, because he has a benign addiction. For, the oddity I pointed out is that, by speciously breaking out of the system of this collaborative work, the behaviours suggesting the resignation remained, to the most noteworthy, while the supportive appearance, to others and himself, to make progress in a constructive lifework, vowing to take other editors by the hand and guide them through it, was cast away, in exchange for snivelling hues and cries. Fay Freak (talk) 19:28, 4 July 2026 (UTC)
- Here is a sample of his behavior:
- https://www.reddit.com/r/WikipediaAdmins/comments/14kv582/wiktionary_admin_equinox_tells_user_to_fuck_off/
- screenshot of the insult:
- https://postimg.cc/MnN58H34
- He was NEVER banned or even sanctioned for this at all during his time here. In fact, he continues to occasionally hurl insults at me and other editors. box16 (talk) 19:53, 4 July 2026 (UTC)
- Or he should continue age-appropriately and reregister under a new branding, because he has a benign addiction. For, the oddity I pointed out is that, by speciously breaking out of the system of this collaborative work, the behaviours suggesting the resignation remained, to the most noteworthy, while the supportive appearance, to others and himself, to make progress in a constructive lifework, vowing to take other editors by the hand and guide them through it, was cast away, in exchange for snivelling hues and cries. Fay Freak (talk) 19:28, 4 July 2026 (UTC)
- Seems like the nasty-minded specifically edit as IPs so they can have less regard on how they behave and hit below the belt. On the other end, real names are preferred by longstanding editors such as Chuck Entz here, opining that people lose posture in the “anonymity”, but evidently there is a sense behind constant pseudonyms whereby people build reputation and credibility, to detach oneself from the subjectivity of one’s personal position, which is proven to work in academia and journalism, where the name of a random researcher may not say much or be blinded in peer-review, and increasingly relevant in contrast to bots and keyboard armies sweeping the internet nowadays, so there is even more inhibition to be irresponsibly inaccurate or injurious: unlike when one is precariously engaged by a marketing department, an intelligence agency, a shortselling institution or one’s own vanity to spread FUD and fill one's pockets or self-conceit about how great one is at making a fool out of everyone. I had more respect for the insufferable personalities having skin in the game, no need to go all in with full nudes. Fay Freak (talk) 11:12, 4 July 2026 (UTC)
- I don't think the discussion, as it currently stands, supports a site ban.
- The original rationale was that Box16's edits amounted to vandalism because they reflected a persistent disregard for established consensus. However, the subsequent discussion has instead highlighted that there has been longstanding uncertainty over the distinction between Template:alternative form and Template:alternative spelling, to the point that multiple editors have discussed clarifying the documentation itself- see discussion above diff. That makes it difficult for me to conclude that edits based on a different interpretation are equivalent to vandalism or bad-faith editing.
- Several additional examples have since been raised. Some appear to be ordinary editing mistakes; others concern style or judgment calls. Editors, including even myself, inevitably make mistakes, and the appropriate response is usually discussion, explanation, or correction. A pattern of errors, by itself, is not sufficient to justify a site ban unless accompanied by a refusal to follow clear community consensus after repeated attempts at resolution.
- My own interactions with Box16 have not been uniformly negative. In a recent discussion over an RFV, for example, Box16 responded by locating several durable citations that materially improved the entry. We also had a productive discussion about why I had marked the entry for verification, and although we approached the issue differently, the exchange remained focused on improving the dictionary. That experience is inconsistent with the characterization of Box16 as someone acting solely to disrupt the project: the editor helped. Same on other pages as well.
- The tone of this discussion has also become increasingly personal on multiple sides, which has made it harder to evaluate the underlying issues objectively. I would prefer that we separate concerns about civility from concerns about content and assess each on its own merits.
- If there is concern that Box16 is making systematic errors, the appropriate remedy is to continue discussing specific edits- edit by edit-, establish clearer consensus where the documentation is ambiguous, and, if necessary, use the ordinary editorial and administrative processes available on the project. Based on what has been presented here, I do not think the threshold for a community ban has been met: the editor is helping to build the dictionary. (AI-assisted drafting) --Geographyinitiative (talk) 🎵 20:12, 4 July 2026 (UTC)
Euphemisms as replacements for extended senses of slang terms
[edit]There are two hits on Google Books for "you're such a penis!" and I think I've heard someone say "you're a literal vagina" (pussy) before. Are these lexicalized? How would we list them, since penis and vagina correspond to more than one slang term each, and those terms have more than one meaning each? —Soap— 00:37, 29 June 2026 (UTC)
- for these, if you are unsure, then adding say
{{lb|en|colloquial|euphemisic}}{{syn of|en|dick}}etc. serves as a starting point. Juwan 🕊️🌈 18:00, 14 July 2026 (UTC)
over-Wikidataing
[edit]Firstly, WD is really dull. Secondly, it's probably generally useful. Jlwoodwa gave nice insights to it w.r.t. WT, and is my nomination for WD-WT ambassador. all along gives WD overprominence, due to Template:wikidata lexeme, the culprit here. Let's look...lots of links are given to WD, and in (probably, I checked a dozen) every case the WD link is the most prominent item at the top of the page. For example absconder links to this WD item, which contains arcane information (syntactic dependency/head relationship/motivating word[WTF???]). I'd simply "hide" WD stuff in a margin like interwiki links. Whadda we reckon how WT can smoothly integrate WD-WT? —User:Vealhurl (talk 18:29, 29 June 2026 (UTC)
- By interwiki links, do you mean the links to other Wiktionaries? Those are per-page, while each language section on a page will have its own lexeme(s), so I don't think that approach will work at all here. jlwoodwa (talk) 18:38, 29 June 2026 (UTC)
- We could easily put WD links with interwikis, or hide them like
{{senseid}}. —User:Vealhurl (talk 18:44, 29 June 2026 (UTC)- Again, please clarify what you mean by "interwikis". Do you mean the links to other Wiktionaries that appear outside the page body, or do you mean templates like
{{pedia}}? jlwoodwa (talk) 18:48, 29 June 2026 (UTC)
- Again, please clarify what you mean by "interwikis". Do you mean the links to other Wiktionaries that appear outside the page body, or do you mean templates like
- We could easily put WD links with interwikis, or hide them like
- I don't think we should prominently link to Wikidata. I would be better to have a more subtle link along the lines of
{{pedia}}. Andrew Sheedy (talk) 19:35, 29 June 2026 (UTC)- I'm not strongly opposed to this, but how would homographs be handled? In the current approach, each lexeme can be straightforwardly associated with a specific part-of-speech or etymology section. This wouldn't work if the lexeme links were all moved to the same "further reading" section. jlwoodwa (talk) 19:47, 29 June 2026 (UTC)
- What if we placed a WD template in between the headword and any term labels? I'm unsure of the styling, but something that is either in line with the headword style or more subtle. TranqyPoo [💬 | ✏️] 19:58, 29 June 2026 (UTC)
- For example: beautiful (comparative more beautiful, superlative most beautiful)

- Where the icon is a clickable link to the Wikidata Lexeme. What do you guys think? TranqyPoo [💬 | ✏️] 22:40, 29 June 2026 (UTC)
- Please, please, please do not clutter the headword with links to other projects (or anything other than basic lexical information). Wikidata often works at cross purposes to us and is not helpful to most users, so links to it should be placed wherever less useful links usually go. Andrew Sheedy (talk) 02:03, 30 June 2026 (UTC)
- For example: beautiful (comparative more beautiful, superlative most beautiful)
- What if we placed a WD template in between the headword and any term labels? I'm unsure of the styling, but something that is either in line with the headword style or more subtle. TranqyPoo [💬 | ✏️] 19:58, 29 June 2026 (UTC)
- I'm not strongly opposed to this, but how would homographs be handled? In the current approach, each lexeme can be straightforwardly associated with a specific part-of-speech or etymology section. This wouldn't work if the lexeme links were all moved to the same "further reading" section. jlwoodwa (talk) 19:47, 29 June 2026 (UTC)
- How does WT easily link to WD? wd:something doesn#t work —User:Vealhurl (talk 21:15, 29 June 2026 (UTC)
- The interwiki prefixes are d: and wikidata:. jlwoodwa (talk) 21:19, 29 June 2026 (UTC)
- Let's delete
{{wikidata lexeme}}. On mobile it is literally above everything else which is kind of ridiculous. There's no reason to send our readers to Wikidata given that the whole point of it is to be read by machines. Ioaxxere (talk) 06:46, 30 June 2026 (UTC)- Wikidata is not solely meant to be read by machines. Where did you get that impression? jlwoodwa (talk) 06:54, 30 June 2026 (UTC)
- I agree with @Andrew Sheedy. The general reader is not going to find visiting Wikidata particularly useful (unlike, say, interwiki links to Wikipedia or Wikispecies), so if possible we should find a way to link to Wikidata silently without a big box being displayed. For example, our templates
{{root}}and{{word}}create categories in the background without any visible output. — Sgconlaw (talk) 21:39, 29 June 2026 (UTC) - I agree. - -sche (discuss) 06:16, 30 June 2026 (UTC)
- Wikidata lexemes almost always have links to other dictionaries, in addition to structured data. So I think they are at least slightly useful to general readers. jlwoodwa (talk) 06:37, 30 June 2026 (UTC)
- At [[all along]] WikiData provides links to OED (paywalled), MWOnline, and a Slovenian dictionary, whereas “all along”, in OneLook Dictionary Search. provides links to half a dozen English dictionaries. Ie,
{{R:OneLook}}is more useful than the WD link and it is less demanding of user attention as an inline link template. DCDuring (talk) 17:44, 30 June 2026 (UTC)- “bulbousness”, in OneLook Dictionary Search. doesn't include the OED entry that d:L:L1567631 does. And I often find that OneLook has links to "entries" that just say "we don't have an entry for this, did you misspell your query?". It is not strictly better than Wikidata. jlwoodwa (talk) 04:05, 1 July 2026 (UTC)
- The Dictionary.com link on “abderian”, in OneLook Dictionary Search. is another example of how it can go wrong. jlwoodwa (talk) 07:28, 1 July 2026 (UTC)
- “bulbousness”, in OneLook Dictionary Search. doesn't include the OED entry that d:L:L1567631 does. And I often find that OneLook has links to "entries" that just say "we don't have an entry for this, did you misspell your query?". It is not strictly better than Wikidata. jlwoodwa (talk) 04:05, 1 July 2026 (UTC)
- At [[all along]] WikiData provides links to OED (paywalled), MWOnline, and a Slovenian dictionary, whereas “all along”, in OneLook Dictionary Search. provides links to half a dozen English dictionaries. Ie,
- Most users probably don't know what Wikidata is and how it works, and that they can click on the links to go to the dictionaries. Maybe it would be better to pull this data to render dictionary links directly on Wiktionary. absconder mentioned above has no external dictionary links here, but the WD item has 7 references. Jberkel 08:31, 30 June 2026 (UTC)
- I feel that Wikidata, with its use of "statements", is not easy for the general user to understand, so it is not a particularly useful way to find links to other Wiktionaries. In any case, don't all our entries already have a "languages" link at the top right corner providing links to corresponding entries in other Wiktionaries which exist? — Sgconlaw (talk) 14:24, 30 June 2026 (UTC)
- Wikidata doesn't link to other Wiktionaries, it links to other dictionaries. jlwoodwa (talk) 17:06, 30 June 2026 (UTC)
- I have to agree with everyone else in this discussion. The average reader will likely not consider the link useful (although objectively, it is) due to its not-so-friendly UI and how it conveys information. It is more useful for the editor. Fenakhay did add support to sense IDs, where they can link to their respective Wikidata lexeme senses. This is probably the best option that maintains POS distinction, while not showing such prominence. TranqyPoo [💬 | ✏️] 17:28, 30 June 2026 (UTC)
- I feel that Wikidata, with its use of "statements", is not easy for the general user to understand, so it is not a particularly useful way to find links to other Wiktionaries. In any case, don't all our entries already have a "languages" link at the top right corner providing links to corresponding entries in other Wiktionaries which exist? — Sgconlaw (talk) 14:24, 30 June 2026 (UTC)
- I do find it odd that we have a template that takes this much space and is so prominent for something that in its current state is unlikely to be useful to the vast majority of readers. — SURJECTION / T / C / L / 15:01, 30 June 2026 (UTC)
- As a human, I have never found Wikidata useful at all. I assume it's useful to machines. ~2026-37577-52 (talk) 17:30, 30 June 2026 (UTC)
- I mean, it’s how interwikis nowadays work. Still, I would not seek out the lexeme part. I had more luck connecting specific entities with it: places and people, plants and statutes. For anything that does not entirely overlap Wikidata will stay unfinished. Fay Freak (talk) 17:37, 30 June 2026 (UTC)
- For the kind of content it contains it is generally unhelpful to human readers, it appears. It’s always the same Oxford dictionaries, Merriam Webster and American Heritage, and Slovene Termania because they have particular support for digitized language records in Slovenia. Funny enough, the very first one I clicked was a dead link, Sõnaveeb entry ID for diagnose. If kept for power users (the existence of which even I doubt, since this is more like a failed attempt at data hoarding) there should only be very dainty boxes that do not invite anyone who does not know to click on them. Not sure why Wikidata tries for all this redundancy to Wiktionary though or at which state it would more likely be useful. Fay Freak (talk) 17:33, 30 June 2026 (UTC)
- A link to WikiData might be acceptable as an inline link in "Further reading" or "References", though the way it is used in Translingual entries (not by me), seems okay. DCDuring (talk) 17:44, 30 June 2026 (UTC)
- @DCDuring: I think that would be fine. — Sgconlaw (talk) 12:00, 3 July 2026 (UTC)
- Wikidata Lexemes clearly serves a very restricted niche. However, within this niche, there is me, as I'm planning to rely on lexemes codes for consistent lemma-linking in my TEI transcriptions of texts, and having a stronger correlation between lexemes and Wiktionary entries would be very helpful. I assume there are other uses, considering the attention it's getting. I very much like @TranqyPoo's suggestion, and I would suggest those icons to be invisible by default and toggleable by gadget, so that all the normal people that don't care about this have no clutter at all. On the technical side, it's just a matter of turning
{{wikidata lexeme}}into that small icon and moving all instances after the headword template. On the aesthetic side, I would not make it superscript. Catonif (talk) 15:13, 1 July 2026 (UTC)- On thinking about it more,
{{wikidata lexeme}}is more akin to{{commons inline}}({{R:commons}}),{{wikipedia inline}}({{pedia}}), and{{wikispecies inline}}({{specieslite}}) in that its usefulness lies in providing a link to other projects, than to templates like{{root}}and{{word}}which generate categorization. Thus, would someone like to edit the template so that it can generate an inline version for use in "Further reading" sections? (@NaomiAmethyst?) — Sgconlaw (talk) 22:19, 8 July 2026 (UTC)- Are you thinking a parameter to
{{wikidata lexeme}}? The separate{{wikidata inline}}exists already, or am I missing something? — Naomi Amethyst 23:14, 8 July 2026 (UTC)- @NaomiAmethyst: ah, sorry, I didn’t know the inline version of the template existed. Silly me. — Sgconlaw (talk) 14:40, 10 July 2026 (UTC)
- Are you thinking a parameter to
- On thinking about it more,
"Pseudo-forms" for protolanguages
[edit]- Discussion moved from Wiktionary:Grease pit/2026/June#"Pseudo-forms" for protolanguages.
There seem to be different opinions on what should count as what language when it comes to reconstructed languages that are themselves ancestors of reconstructed languages. @Victar/Sokkjo (talk • contribs) has pointed out his nonstandard use of pre-{{m+|gem-pro||…}}, sometimes for text that is normally considered a late stage of PIE for Wiktionary's purposes (pre–Grimm's law, pre–Kluge's law, but apparently post-laryngeal, which seems like a questionable and arbitrary dividing line). This user is proving impossible to talk to, so I bring it up here to get hopefully more opinions. Surely if the notation is for a different language, such as using PIE notation ⟨gʷʰ⟩ in text templatized as Proto-Germanic, the choice of language code is incorrect. — Ganjabarah (talk) 21:34, 24 June 2026 (UTC)
- This is a topic for BP, not GP. Saph (talk) 01:41, 25 June 2026 (UTC)
- Moved. Juwan 🕊️🌈 18:06, 1 July 2026 (UTC)
July 2026
Poor contributions from User:Златарка
[edit]I want to bring to the attention of admins the poor contributions of User:Златарка to the Bulgarian project within Wiktionary.
This editor has been working on the kinship vocabulary in Bulgarian, however, her edits are full with inaccuracies and, in some cases, blatant mistakes.
Shortlist of inaccuracies:
- Small pragmatic inaccuracies: Labelling diminutive/augmentative forms as synonyms (e.g. щерка (šterka) as synonym of дъщеря (dăšterja)). Technically, such words have different connotation.
- Imprecise lexical classification: Classifying cohyponyms as antonyms where that is not justified (e.g. gender pairs such as син (sin) : дъщеря (dăšterja), мъж (măž) : жена (žena), etc.). Humans may have no trouble recognising the context in which such pairs behave as antonyms, however, it is inaccurate to perceive them as such in general. For instance, if AI is trained to treat мъж (măž, “male adult”) as an all-around antonym of жена (žena, “female adult”), then it would misclassify момиче (momiče, “female child”) as мъж (măž), because момиче (momiče) and жена (žena) also form a complementary pair albeit with respect to another criterion.
- Unencyclopedic style of expression: User:Златарка prefers to use prescriptive, rather than descriptive diction. Occasionally, she employs ostensive and sophistic means to insinuate particular sense that does not exist (e.g. she insisted to provide an ostensive list of "positive characteristics" that Bulgarians approve of under българин (bălgarin)). While technically not a mistake, her style of writing causes confusion.
- Quantity over quality: Her edits are often repetitive and redundant, repeating the same sort of information (e.g. Related terms) under numerous lemmas for no reason. Technically not a mistake, but still wasteful.
Shortlist of mistakes:
- Typological errors: Classifying different types of nouns as synonyms (e.g. agent nouns such as родител (roditel, “parent”) and condition/event nouns such as причина (pričina, “cause”)). This is wrong. In Bulgarian, substituting an agent noun for an event noun (most of the time) changes the meaning of a message.
- Wrong lexical classification: Classifying nouns of different semantic strata as synonyms (e.g. дете (dete, “child”) as a synonym of потомък (potomăk, “offspring”)). This is wrong. In Bulgarian, substituting a hyponym for a hyperonym or vice versa (always) alters the meaning of a message.
- Wrong morphological parsing: On a few occasions, she has provided wrong etymologies of lemmas (e.g. умнокрасивитет (umnokrasivitet): the second and third etymology are hers and are entirely butchered).
I should mention that I and other members have tried to fix some of the mistakes that User:Златарка commits, but she doesn't let anyone refine her edits. Once she contributes to an entry, it becomes her property and she reverts any future corrections.
I do not know what actions are supposed to be taken in cases like those, but some reaction is required. User:Златарка lacks in-depth knowledge in the field that she contributes to and she refuses to acknowledge any critique that is addressed to her. If she is left to continue in the same manner, eventually, she will just cause a bigger mess. Any suggestions how to proceed? Безименен (talk) 08:03, 1 July 2026 (UTC)
- Has any of this actually been brought up with this user directly, or is this the first time? Because on their user discussion page I only see a few posts, and (translating them with deepl) it doesn't seem like any of them relate to what you mention here, aside from general accusations of propagandizing from an anonymous temp account. But I don't know if there are edits where users are tagged or something.
- The point about antonyms doesn't seem severe to me, Wiktionary is written for people to read, not as the basis for AI to learn ontology. Thesaurus:boy also has both "girl" and "man" as antonyms. For clarity they could be specified wrt what axis the antonymy is on (age, gender, etc). It's an inherent issue with antonyms in many cases.
- I don't really get what you're arguing about the wastefulness of having the same Related terms section under numerous lemmas? While it could maybe be centralized under one term and referred to with a "see <here>", that is how the related terms header is often used, and since Wiktionary Is Not On Paper, we don't need to be minimalist, nor does every edit have to be unique. PhoenicianLetters (talk) 18:53, 2 July 2026 (UTC)
- @Bezimenen I do sympathize with and respect you as an editor, but on this one, I am not sure I am convinced. Maybe you could array your points with some diffs or any proof that the user is being deliberately harmful despite being told how to improve?
- On the contrary, I think @SixtyShips has generally been happy with Zlatarka's edits, which have been benevolent, as far as I have seen. If you are saying they are not as competent as you would like, that's different, but we all start somewhere - perhaps it would be best to try to teach this user how things stand, if you think there is a problem? Yet, you have been quite combative with your handling, as far as I can see. I know there was that case with the political matters earlier, as well, which may have scared you, too, but insofar as editing here, I think we should not let those kinds of things (which may even just be misunderstandings) distract from the editing work.
- Like PhoenicianLetters said, I think more discussion is really needed first. I noticed Zlatarka did revert you some times, with no justification, which I see fit into the pattern of your argument, but if you haven't had a real organized discussion then it's hard to judge who's right on such cases.
- I haven't noticed a problem with Zlatarka myself, but I wanted to reply here out of respect for your concern for Bulgarian entries - RSVP if you have any further points to add or any thoughts. Kiril kovachev (talk・contribs) 23:30, 15 July 2026 (UTC)
Context labels & standard spellings
[edit]I challenge the following statement from Template:standard spelling of: Normally this [identification of spelling system] is done using the
|from= parameter (see below), or, optionally, with no |from= field set but with context tags used.
Compare the following senses:
- (British spelling) Standard spelling of behaviorism
- British standard spelling of behaviorism
I believe the first sense states “a British spelling, which is coincidentally the standard spelling of TERM” and the second states what is (literally) intended. I say this because I made the mistake when I saw it. I propose to remove the underlined portion. TranqyPoo [💬 | ✏️] 04:55, 2 July 2026 (UTC)
- I agree and I support the direction of your edit as the clearer way to express the meaning. A spelling is either standard in one geolectal variety albeit not in another (such as BrE-versus-AmE differences) or standard across all varieties; there's no such thing as a geolectal difference that is a panvarietally standard-versus-nonstandard difference (which is to say, there's no such thing as one geolect being panvarietally "correct" and another being panvarietally "incorrect"). Quercus solaris (talk) 22:27, 2 July 2026 (UTC)
Support. Davi6596 (talk) 17:27, 5 July 2026 (UTC)- I support changing this, in the template documentation and in entries. - -sche (discuss) 21:27, 5 July 2026 (UTC)
Proto-Turkic Romanization, Again
[edit]I do remember having another shot at this earlier, yet I could not locate the previous discussion, so I'll start anew:
After a (brief) talk in the Discord server, @Yorınçga573 and I decided that current Proto-Turkic romanization scheme (directly copied from EDAL (StarLingDB), mind you; with some manual changes on an entry-by-entry basis) falls short on actually conveying the academical and Turcological standards. This is partly because there hasn't been much published on Proto-Turkic, specifically, and partly because the works that are published all follow varying standards. We are not bound by any standards, and can choose what'd reflect the Proto-Tukic the best.
I'll divide this into two separate proposals:
Proposal 1:
- *ạ > [REMOVED] (This was not a phoneme in PT, purported by EDAL to support Altaic compranada, and was removed from all entries around a year or so ago by help of the user Surjection.)
- *ẹ > *e
- *e > *ä
(The first proposal aims to unify the umlaut-for-front/back distinction system in PT. We already have *ï/i, *o/*ö and *u/*ü; and this system is immediately clear to any learner/editor. We currently have *ẹ (*e) (historically and phonologically closer to *i than *a/*ä) and *e (*ä). If this proposal goes through, we'd have *a/ä; *e/i/ï; *o/*ö and *u/*ü, and would finally divorce the PT vowel system from the standart used in EDAL.)
Proposal 2:
- *ń > *ɲ
- *l > *l₁
- *ĺ > *l₂
- *r > *r₁
- *ŕ > *r₂
- *č > *ç
- *š > *ş
(The second proposal aims to, likewise, use Turcological signs used more commonly in contemporary/recent Turcological studies, with the exception of *ń > *ɲ, which is my addition, and *č > *ç and *š > *ş, which are bundled with these changes to further align with a diacritic-less, CTA-influenced system. I reckon this will be a more controversial change. *l₁ and *l₂, and *r₁ and *r₂ are distinguished like so, and not with an acute diacritic, because the latter implies a palatal variant of PT lateral and rhotic consonants, which is not the case, [and not-surprisingly] also adopted from EDAL, which assumes such to purport [yet again] Altaic comparanda. You can cast your vote separately on these, if you want.)
Pinging Proto-Turkic and historical Turkic languages editors. @Ardahan Karabağ, @Bartanaqa, @BurakD53, @Rizozoda34, @Rttle1, @Yerkishisi, @Yorınçga573 (and more that I definitely missed.)
AmaçsızBirKişi (talk) 13:43, 3 July 2026 (UTC)
- *ẹ > *e, *e > *ä
Weak oppose If this is based on Proto-Turkic *e/*ä becoming Chuvash a, I don't think that's the case. It looks more like a later, language-specific change. The alternative shift (*ẹ > *é) is cleaner and makes more sense structurally. - *ń > *ɲ
Oppose The symbol *ń already clearly marks palatalization, so switching to *ɲ is redundant. - *l > *l₁; *ĺ > *l₂; *r > *r₁; *ŕ > *r₂
Weak support The distinction itself is fine, but using numbers is clunky compared to standard diacritics. - *č > *ç, *š > *ş
Oppose *č and *š are already widely accepted standard symbols for these sounds, so changing the characters doesn't add any extra clarity. – BurakD53 (talk) 14:39, 3 July 2026 (UTC)
- Instead of using the abstract indices *l₁, *l₂ and *r₁, *r₂, it would be more suitable to adopt a notation like *l, *lᶴ/lˢ̌ and *r, *rˢ.
- About *l, *lᶴ/lˢ̌
- According to Wikipedia, the most commonly used symbol for the voiceless alveolar lateral fricative [ɬ] seems to be ł, as seen in Caucasian languages like Adyghe (Circassian) and Avar. One might object to using *ɫ because it represents a different sound in the IPA and could cause confusion. However, if we strictly follow the IPA, we could directly use the explicit ɬ sound—a symbol actually utilized in Proto-Yeniseian reconstructions. But it seems we have no intention of using IPA. Phonetic processes like crasis can shed light on these phonemic developments. We know clearly that the /l₂/ sound developed from the combination of the /l/ and /t͡ʃ/ sounds. Therefore, the sound we are looking for must be one that falls between these two. As you are aware, sources identify this sound as /ɬ(t͡ʃ)/. Instead of using numerals, we can explicitly express this sound as *lᶴ or *lˢ̌.
- About *r, *rˢ
- For instance, the *z phoneme itself could be a product of crasis as well, as suggested by alternations like *koňuz vs. *komursga, *boguz vs. *bogursdak, and *ti:z vs. *tirsgek, etc. This structural behavior even brings to mind the word for "shoulder"—omuz—and its Latin equivalent umerus and its descendant in Romanian umăr. Of course, this does not prove that it arose as a result of crasis; however, the r sound closest to the z sound is the one pronounced like s. For this reason, I propose representing this phoneme as *rˢ.
- Conclusion
- Therefore, replacing abstract indices with phonetically descriptive symbols provides a clearer, more grounded framework for these reconstructions. And you will se similar reconstructions in other reconstructed languages such as Proto-Indo-European in Wiktionary.
- I am of the opinion that these revisions should also be considered during the voting:
- *l > *l
- *ĺ > *lᶴ or *lˢ̌
- *r > *r
- *ŕ > *rˢ
- – BurakD53 (talk) 22:30, 3 July 2026 (UTC)
- The change proposal of ẹ > e, e > ä doesn't really depend on Chuvash, I don't get that argument. These two are phonemic in PT anyway, this is in actuality just a cosmetic change.
- AmaçsızBirKişi (talk) 10:00, 4 July 2026 (UTC)
- I’ve changed my mind; I’m neutral on this. Both 'e/ä' and 'é/e/ are fine by me, though the need to move so many pages—even with a bot—is bothersome. – BurakD53 (talk) 10:44, 4 July 2026 (UTC)
- As a result, after the discussions on Discord, my view has shifted as follows:
Support for *ẹ > *e, *e > *ä (ẹ really needs to go)
Support for nʸ
Support for *lˢ,*rˢ or *lˢ̌,*rᶻ;
Weak support for *l,*r,*l₂,*r₂;
Weak oppose for *l₁,*r₁,*l₂,*r₂
Oppose to š>ş, č>ç. That's really unnecessary. BurakD53 (talk) 10:59, 4 July 2026 (UTC)
- I’ve changed my mind; I’m neutral on this. Both 'e/ä' and 'é/e/ are fine by me, though the need to move so many pages—even with a bot—is bothersome. – BurakD53 (talk) 10:44, 4 July 2026 (UTC)
- *ẹ > *e, *e > *ä
- I agree with the idea of changing vowels to the new forms as shown above, however I prefer that the consonants remain the same, most of which are already claimed by the community. Ardahan Karabağ (talk) 18:16, 3 July 2026 (UTC)
- I support the vowel changes. If they are changed for Proto-Turkic, they should probably also be changed for Qarakhanid, and perhaps Old Anatolian Turkish as well. Chaghatay transliteration already uses ä/e rather than e/ẹ. Rizozoda34 (talk) 19:01, 3 July 2026 (UTC)
- Yes, they should. I'm already doing my part changing é > e and e > ä in Old Turkic stuff.
- AmaçsızBirKişi (talk) 05:27, 10 July 2026 (UTC)
- Yes, they should. I'm already doing my part changing é > e and e > ä in Old Turkic stuff.
- i definitely agree with vowels because it makes the e/ä distinction clearer to read, also nicely fitting among other umlauts
Strong support - i think im neutral about any 'ny' (maybe slightly less for ñ),
Abstain - also about ç and ş, carons also very neat. i don't see any of them representing respective sounds better or worse.
Abstain - numbered r and l's will definitely take time to get used to but i love its ambiguity. one could say r1 and l1 is a bit eyesore, and could look overcrowding, but since reconstructions are just one/two word-length, not full-on texts, i don't see it as a problem. tho at the end l, r acute are not in any way terrible imo. *l > *l₁; *ĺ > *l₂; *r > *r₁; *ŕ > *r₂
Support Yerkishisi (talk) 19:29, 3 July 2026 (UTC) - 1. *ẹ > *e, *e > *ä
Support - 2. *ń > *ɲ
Weak oppose should be nʸ as suggested on Discord - 3.
Support phonetically we simply don't know,
Strong oppose on lˢ and rˢ for that same reason. You're reinventing ĺ and ŕ. This is like ascribing PIE h1, h2 and h3 values... - 4.
Abstain I know I suggested this but eh, it doesn't really matter. Yorınçga573 (talk) 00:16, 4 July 2026 (UTC)
- I also support nʸ. – BurakD53 (talk) 00:42, 4 July 2026 (UTC)
- I will not cast a vote, but I note that changing *ĺ to *l₂ and *ŕ > *r₂ does not imply that *l and *r must necessarily be changed to a numbered letter as well. For a precedent compare the well-established Linear B transliteration, where numbered syllables exist but the base ones do not acquire a 1. I think a scheme with *l *r *l₂ *r₂ would strike a good balance between maintaining good readability for the plain series while keeping the realisation of the fricative series intentionally vague. Furthermore, I wholeheartedly support ä, have no opinion of č š, but I'd rather we kept ń over any of its proposed alternatives. Catonif (talk) 11:30, 4 July 2026 (UTC)
- (PS:) I should note that for the laterals and rhotics, we (BurakD53, Yorınçga573 and I) decided to adopt this:
- *l > *l
- *ĺ > *lᶴ
- *r > *r
- *ŕ > *rᶻ
- AmaçsızBirKişi (talk) 12:42, 4 July 2026 (UTC)
- Thank you all for your inputs. Based on the popular demand, the new romanisation scheme will look like this:
- *ạ > [X]
- *ẹ > *e
- *e > *ä
- *ń > *nʸ
- *l > *l
- *ĺ > *lᶴ
- *r > *r
- *ŕ > *rᶻ
- *č > *č (stays the same)
- *š > *š (stays the same)
- AmaçsızBirKişi (talk) 17:48, 7 July 2026 (UTC)
WT:FICTION verbiage loophole
[edit]The verbiage in WT:FICTION seemingly has a logical loophole that can be easily misunderstood. According to the policy, if the term originates within the fictional universe, it must survive a special criteria. However, the special criteria can be bypassed if the term does not originate within the fictional universe and refers to some object/concept only existing in the fictional universe. Compare: Darth Vaderian & “peanut”, in Appendix:SCP Foundation. Darth Vaderian (coined outside the Star Wars works) refers to some concept within the the Star Wars universe and not something existing in the real-world (or across the science fiction genre). peanut is in the same criteria. Both are considered compliant per CFI.
Currently, this is halting a RFD process. As any change to the CFI policy requires a vote, this discussion serves to check community consensus. Is what I stated consistent with the community's consensus or shall its verbiage be updated? If it shall be changed, I recommend: Terms
. Feel free to propose alternative phrasing.
originating referring to objects or concepts in fictional universes […]
In the event editors state that the current policy's verbiage is consistent with the community consensus, I ask: what is the purpose of WT:FICTION and why is the distinction only if the term was coined within an original work of the fictional series? Why is it acceptable to feature a term in the main namespace if it only refers to something within (and does not in any way extend past) its fictional universe? TranqyPoo [💬 | ✏️] 19:45, 3 July 2026 (UTC)
- @TranqyPoo: I think that was intentional. If a term is coined by the author of a work of fiction and only used in the fictional universe, the Wiktionary community has decided that the term should not be included in the dictionary. If it were otherwise, there would be a proliferation of in-universe terms from books, films, TV shows, video games, etc., in the dictionary. On the other hand, a term which relates to a work of fiction in a meta-sense but is not in-universe does seem to merit inclusion in the dictionary as a term which is not purely fictional but has some real-world application, for example, adjectives relating to fictional characters (Darth Vaderian), nicknames for such characters, and terms originating from fictional works which are used out of context to refer to real-world things. — Sgconlaw (talk) 20:36, 3 July 2026 (UTC)
- Compare: lightsaber. This term was featured during the voting process for WT:FICTION. Despite the term originating in Star Wars, it has a more generic meaning that doesn't require the Star Wars context (in this case, within science fiction). I admit, it is odd that the term's prominent quotations mostly reference Star Wars, but the supplemental policy primarily displays quotes without a reference to the franchise. I'm assuming it is to clarify what makes a term (originating from a fictional universe) worthy of main namespace inclusion. I would not see Darth Vaderian as problematic if its sense has extended to something that does not require the context of a specific fictional universe (like Jedi Master).
- Is it sensible to record non-original, fictional terms in the main namespace and then refer to an appendix for more context of the original terms? Sticking with the examples, peanut can be added to the main namespace, but its main entry (SCP-173) must be in the appendix. I'm not seeing the consistency.
- Note: I am solely concerned with terms that are not meaningful outside the fictional universe. For example, I do not see an issue with Buffyhead as it is referring to a real-world person. TranqyPoo [💬 | ✏️] 21:10, 3 July 2026 (UTC)
- Think of it this way: If someone uses a term like "Darth-Vaderian" in a non-Star Wars context, they are expecting the reader to know what it means apart from the context of Star Wars. If it occurred on the lips of a character in a Star Wars novel, then the author is only assuming that those who are familiar with that universe will know the term. Thus, if a term is citeable outside of the context of that universe, that indicates a certain level of cultural (and hence lexical) significance that not all fictional-entity-related terms will achieve. Andrew Sheedy (talk) 21:53, 3 July 2026 (UTC)
- Also note that there is no loophole. The same criterion applies to both a term like Jedi (originating from within the universe) and Darth Vaderian (originating from outside it): that is, both terms must be citeable apart from that universe. No loophole, no double-standard. Andrew Sheedy (talk) 21:55, 3 July 2026 (UTC)
- I appreciate that you are engaging with me, but I must humbly continue dissecting. I agree that a term originating from a Star Wars context can be used in a non-Star Wars context, but it must be used without expecting the reader to have knowledge of the Star Wars context. Again, following your understanding, peanut (not originating within the fictional universe) could be featured on the main namespace, but its main entry (SCP-173, originating within the fictional universe) could not. That is a loophole, as peanut does not refer to anything else except "SCP-173". To show the loophole more clearly:
- ❌ "I can't sleep with my eyes shut anymore; I'm afraid SCP-173 will find and kill me."
- ✅ "I can't sleep with my eyes shut anymore; I'm afraid peanut will find and kill me."
- TranqyPoo [💬 | ✏️] 22:19, 3 July 2026 (UTC)
- Upon re-understanding your last point, the verbiage states in a manner that would only apply to terms originating within a universe. This is precisely why I started this discussion. TranqyPoo [💬 | ✏️] 23:25, 3 July 2026 (UTC)
- @TranqyPoo: I also understand the policy in the same way as Sgconlaw and everyone else has stated, nor did it seem ambiguous or unclear to me. J3133 (talk) 06:37, 4 July 2026 (UTC)
- Very well, then. A slightly different approach: if a set of citations makes no reference to the Star Wars franchise and each implicitly refers to the Star Wars franchise, how is a lexicographer able to deduce its meaning if they were unaware of the Star Wars context? Is it not strange that the reader is supposed to already know where it came from? It would make more sense that the lexicographer would try to deduce its meaning based on what concept/quality the author is trying to illustrate without the Star Wars context (i.e. Darth Vaderian → "powerful, mind-altering with villainous connotations"). My particular issue is when a definition refers to a specific fictional universe object/concept when none of the quotations don't (and can't) say such a thing.
- Please consider carefully what I am saying as I don't think I am being understood. TranqyPoo [💬 | ✏️] 12:39, 4 July 2026 (UTC)
- In doing that — i.e., considering carefully, regarding the pair of questions "how is a lexicographer able to deduce its meaning if they were unaware of the [canon-referent] context?" and "Is it not strange that the reader is supposed to already know where it came from?" — the line of thought prompted me to think about another class of examples (of the same underlying phenomenon): classical mythology (i.e., Greco-Roman mythology) and Abrahamic holy books, and their influence on English, as seen in English words like Titan, titan, titanic, judas, and many others. The lexicography can function OK without detailed knowledge of the lore, but the lexicography will inevitably be at least aware that the canon-referent aspect/history exists (i.e., not wholly ignorant of its existence). The reader's reading comprehension can function OK without detailed knowledge of the lore, and it can even function OK in cases where the reader's ignorance of the canon-referent aspect/history is either total or near-total. (In fact it is often even amazing and miraculous how fluently people can speak natural language while being astoundingly ignorant of history, geography, canonical literary influences, etc.) I'm thinking that to wrestle with how this pair of questions is answered for the modern-fiction class (e.g., Star Wars, Star Trek, the Marvel universe), it is helpful to ponder the classes that came before (classical mythology and Abrahamic holy books) and how the lexicography and reading go (e.g., relative degrees of awareness) with words related thereto. One man's canon is another man's cool story bro lol. Quercus solaris (talk) 15:50, 4 July 2026 (UTC)
- PS: Kind of serendipitous that an hour or two after I was thinking about the above ("amazing and miraculous"), I read a news article about how it came to be that humans agreed to put down white lines to mark the lane edge on major roads. A quote in the article is that "To this day, you depend on it without knowing anything about it."(Ben Cohen in the Wall Street Journal on 2026-07-03) But to say it more precisely, "To this day, we depend on it while knowing the important things about it, even without knowing anything about its history." Where "the important things about it" are the ways in which it is practical and useful, and the fact that most of us would complain if it were taken away. In the same way, users of natural language, both lexicographers and readers, can depend on words such as titan and judas and lightsaber and kryptonite while knowing the important things about them, even without knowing anything about their history (although plenty do know at least a bit about the history). Quercus solaris (talk) 17:40, 4 July 2026 (UTC)
- The mythological & religious terms are likely considered acceptable here (without special criteria) not solely on its cultural significance, but because they are/were believed to be true. That is, its influence has expanded beyond its internal, hypothetical scope. For example, see this RFD discussion of Slender Man. Another interesting and relevant RFD discussion is Flying Spaghetti Monster, where it combines religious and fictional elements (it was kept for unknown reasons, implying the scenario requires further exploration to determine suitable criteria). I admit that I do not know where to draw the line and would love to hear others' opinions. TranqyPoo [💬 | ✏️] 18:42, 4 July 2026 (UTC)
- (spitballing) Do you guys think a safe dividing line would be whether the term could be classified as a myth (false belief)? TranqyPoo [💬 | ✏️] 22:40, 4 July 2026 (UTC)
- It seems to me that dictionary inclusion is more about wide-enough use of the term, rather than true/false or myth/fact. For example, abominable snowman is myth but is widely enough used that it needs a general-dictionary entry. Quercus solaris (talk) 23:50, 4 July 2026 (UTC)
- Did a bit of digging in the archives, see:
- Vote comment and its remaining thread.
- Intro of the original BP discussion.
- Further intent of the vote's proposer.
- Granted, these comments are all from the same person (@BD2412, the vote's proposer), but I gather the spirit of this vote was to not include terms referring to purely fictional entities or concepts of a specific universe. Particularly, these statements resonate with me:
[…] Harry Potter, James Bond, and Captain Kirk would presumably all be aware of Osiris as a mythic figure and unicorns as mythic beasts […] There is a line between the mythic and the fictional.
I believe it is this line that we need to put into words that would make this a simple, objective test. TranqyPoo [💬 | ✏️] 05:04, 5 July 2026 (UTC)
- Did a bit of digging in the archives, see:
- It seems to me that dictionary inclusion is more about wide-enough use of the term, rather than true/false or myth/fact. For example, abominable snowman is myth but is widely enough used that it needs a general-dictionary entry. Quercus solaris (talk) 23:50, 4 July 2026 (UTC)
- (spitballing) Do you guys think a safe dividing line would be whether the term could be classified as a myth (false belief)? TranqyPoo [💬 | ✏️] 22:40, 4 July 2026 (UTC)
- In doing that — i.e., considering carefully, regarding the pair of questions "how is a lexicographer able to deduce its meaning if they were unaware of the [canon-referent] context?" and "Is it not strange that the reader is supposed to already know where it came from?" — the line of thought prompted me to think about another class of examples (of the same underlying phenomenon): classical mythology (i.e., Greco-Roman mythology) and Abrahamic holy books, and their influence on English, as seen in English words like Titan, titan, titanic, judas, and many others. The lexicography can function OK without detailed knowledge of the lore, but the lexicography will inevitably be at least aware that the canon-referent aspect/history exists (i.e., not wholly ignorant of its existence). The reader's reading comprehension can function OK without detailed knowledge of the lore, and it can even function OK in cases where the reader's ignorance of the canon-referent aspect/history is either total or near-total. (In fact it is often even amazing and miraculous how fluently people can speak natural language while being astoundingly ignorant of history, geography, canonical literary influences, etc.) I'm thinking that to wrestle with how this pair of questions is answered for the modern-fiction class (e.g., Star Wars, Star Trek, the Marvel universe), it is helpful to ponder the classes that came before (classical mythology and Abrahamic holy books) and how the lexicography and reading go (e.g., relative degrees of awareness) with words related thereto. One man's canon is another man's cool story bro lol. Quercus solaris (talk) 15:50, 4 July 2026 (UTC)
- FWIW, while I share everyone(?)'s understanding of what the policy currently is, I agree with OP that it leads to weird results, and have opined that before.
I think Darth Vaderian (with the cites it has, where it's used to describe real-world people with no [other] reference to Star Wars) vs. peanut (or perhaps more clearly, Lumity) (with the cites it currently has, where the referent is a thing from a work of fiction) are different, and though grouped by OP, aren't in the same boat. Indeed, I think Darth Vaderian is in much the same boat as Wookiee, in that even if Darth Vaderian had been coined in the Star Wars films, the cites show figurative use, uses that are in reference to non-Star Wars people/things and don't [otherwise] mention Star Wars, that rely on someone knowing the characteristics of Darth Vader (or a hairy Wookiee, etc) without necessarily having watched the films all the way through. I do think we should expand the definition to spell out what qualities of Darth Vader it invokes (dark villainousness? his peculiar vocal qualities?), like Wookiee mentions "hairy".
To me, it seems like cases where the referent is a thing in the fictional world, like Lumity and arguably peanut, are different, and I do think it's at least a little bit weird that (per current policy) we include Lumity and might include peanut.
But it seems like almost anywhere we could draw lines, to try to exclude some things and include others, will be messy: as various internet philosophers have opined, proprietary fictional works today do a lot of the things that folklore and mythology did in the past, e.g. creating fictional species of creature — how different is a chupacabras from a nauga, for example? They're both recently-made-up, non-real creatures that some people now refer to as if they existed.
Also, given the nature of SCP fiction being written by myriad people, could we really say that peanut was coined outside of that fiction, just because it was coined (I take it) outside of one specific website and one specific way of writing about SCP concepts? (And/or how are we distinguishing SCP terms as being "coined within fiction" and being different from e.g. terms that people on a folklore or ghost-hunting (etc) website might coin for types of ghost they imagine?) Also, especially looking at how many of the cites of peanut are actually of Peanut and are naming a specific entity, wouldn't Peanut and SCP-173 run into NSE rules, anyway, as would e.g. Luz or Amity (the individuals who make up Lumity)? We don't include e.g. Anohni or Aurora, why would we include Luz or Peanut? - -sche (discuss) 17:33, 4 July 2026 (UTC)- In retrospect, you are correct that my grouping was flawed (and by extension, its following arguments). I completely agree that terms like Darth Vaderian (with its given citations) should be recorded in the main namespace, but its definition must change to match the citations. I do not have an issue with an etymology adding the Star Wars context as it is relevant, but the term's illustrative usage does not require knowledge of Star Wars. I suggest that terms (and its citations; regardless of its origins) that refer only to specific in-universe concepts should be recorded in an appendix for that given culture group. When a term emerges beyond its fictional universe, the figurative, non-contextual sense can be recorded in the main namespace while retaining its literal (and original) meaning in its appendix. One can even point to the Star Wars appendix term in the main entry. What do you guys think? TranqyPoo [💬 | ✏️] 18:59, 4 July 2026 (UTC)
- @TranqyPoo: I'm sceptical about having appendices filled with fiction terms which otherwise do not satisfy WT:FICTION. I don't think it's Wiktionary's job to become a giant Pokédex. — Sgconlaw (talk) 19:23, 4 July 2026 (UTC)
- I may have failed to elaborate; appendix terms should still meet the separate works criteria, although separate works should be more broadly stated as "official works of the fictional universe is considered a single work of multiple parts". This would include fictions such as SCP & Creepypasta. Feel free to express your thoughts, especially any opinions on how things should be from your perspective. TranqyPoo [💬 | ✏️] 19:37, 4 July 2026 (UTC)
- @TranqyPoo: I'm sceptical about having appendices filled with fiction terms which otherwise do not satisfy WT:FICTION. I don't think it's Wiktionary's job to become a giant Pokédex. — Sgconlaw (talk) 19:23, 4 July 2026 (UTC)
- In retrospect, you are correct that my grouping was flawed (and by extension, its following arguments). I completely agree that terms like Darth Vaderian (with its given citations) should be recorded in the main namespace, but its definition must change to match the citations. I do not have an issue with an etymology adding the Star Wars context as it is relevant, but the term's illustrative usage does not require knowledge of Star Wars. I suggest that terms (and its citations; regardless of its origins) that refer only to specific in-universe concepts should be recorded in an appendix for that given culture group. When a term emerges beyond its fictional universe, the figurative, non-contextual sense can be recorded in the main namespace while retaining its literal (and original) meaning in its appendix. One can even point to the Star Wars appendix term in the main entry. What do you guys think? TranqyPoo [💬 | ✏️] 18:59, 4 July 2026 (UTC)
- Nah, Wikipedia catalogues pop culture adequately. ~2026-38329-29 (talk) 14:15, 5 July 2026 (UTC)
- I agree with everyone else about the policy as is. I write separately to mention that dictionaries in general are supposed to include words that have entered the everyday lexicon, and other dictionaries show that. For example, the RFD lists Jedi, but that is a concept and term that has made a lasting cultural impact to the highest levels of speakers. See also: Jedi census phenomenon. As such, other English dictionaries also list Jedi as a term: MW, Cambridge and Collins.
- The OED's entry is especially of note: they put the term in band 4, meaning that it's as common in written English as words like bipartisan, rodeo, productively, and skyrocket. They define the term as follows:
In the fictional universe of the Star Wars films: a member of an order of heroic, skilled warrior monks who are able to harness the mystical power of the Force (see force n.1 Additions). Also in extended and allusive use; esp. someone (humorously) credited with great skill or preternatural powers. Also more fully Jedi knight, Jedi master.
- They follow it with quotes from outside the universe (except for the initial quote) as any typical entry. Other entries that they have include: lightsabre | lightsaber, Padawan, dark side, (the) force, Jedi mind trick, and carbonite. Star Wars is one of those fictional universes whose terms have genuinely become part of the English language, and I don't think that we should be having a mass-RFD for them. AG202 (talk) 13:31, 6 July 2026 (UTC)
- Thank you for your elaboration, especially with references. I absolutely agree that we should include words originating (coined) from a specific fictional universe, but only if its meaning has emerged beyond its fictional universe. The example you listed, Jedi, currently only points to the Star Wars context (that is: it only applies to something within its internal universe, not extending to another universe or reality). The references you mentioned predominantly state its definition as attributive qualities or figurative uses of the term.
- MW:
a person who shows extraordinary skill or expertise in a specified field or endeavor
. - Cambridge:
a philosophy based on the beliefs of the Jedi characters in the Star Wars movies
(although, this definition needs some work) &a person who follows a philosophy based on the beliefs of the Jedi characters in the Star Wars movies
. - Collins:
a person who claims to live according to a philosophy based on that of the fictional Jedi
- In OED, all non-specific-universe quotations are figurative uses and therefore, state more (or differently) than calling someone a member of some fictional universe order that has harnessed a mythical power.
- MW:
- These definitions are sensible to include in a dictionary because they are describing objects or concepts extending beyond its specific fictional universe (MW's definition is ideal; the others pass but need improvement IMO). I admit that the current RFD list is flawed as I only reviewed their definitions and not their supporting quotations. In this case, I re-direct to this comment and its chain of responses. Now, since you said you agreed with everyone, I must ask some clarying questions as the others have not recognized the problems I am trying to identify:
- WT:FICTION's verbiage mentions if a term originates from a specific fictional universe. This means that if a term does not originate within a specific fictional universe, but only refer to a specific universe object/concept, then WT:FICTION does not apply and can be treated like any other term. This is particularly troubling for nicknames of specific fictional characters coined by fans. The term is discriminated based on who coined it, versus how it is used. Why is this a legitimate criteria and should it be?
- A dictionary is intended to help the reader understand (or translate) the term that the reader came across. Technically,
One of a fictional order of beings from the Star Wars universe who are gifted with heightened awareness of the Force
is a definition, but is it meaningful? Would a person unaware of Star Wars, after being called a Jedi, understand what was meant? I don't think so and IMO that is a failure of definition. Therefore, there ought to be clearer (or more sensible) criteria for when to include a fictional term in a dictionary. We cannot assume that everyone has watched the Star Wars franchise and it would be odd for the speakers to assume that as well. In your opinion, should definitions like these be included on Wiktionary's main namespace and if so, why?
- TranqyPoo [💬 | ✏️] 17:16, 6 July 2026 (UTC)
- Thank you for your elaboration, especially with references. I absolutely agree that we should include words originating (coined) from a specific fictional universe, but only if its meaning has emerged beyond its fictional universe. The example you listed, Jedi, currently only points to the Star Wars context (that is: it only applies to something within its internal universe, not extending to another universe or reality). The references you mentioned predominantly state its definition as attributive qualities or figurative uses of the term.
- Thinking out loud about what kinds of things there are to include or exclude ...
- Proper names (proper nouns) of specific fictional characters seem like something we should (and currently do, AFAICT) only include if figurative use exists: we have Darth Vader (in its figurative sense), but not SCP-173. This is also how we treat non-fictional names, right? We have Albert Einstein (a full first and last name combo) only because figurative use exists; we don't have Rishi Sunak, and we don't have modern mononyms (Anohni, Aurora, Suharto) except as generic "a given name". (We do have Adolf at Hitler, but he has attained a singular degree of infamy and association with his name, and widespread use in allusions.)
- Following some inconclusive RFDs, we currently include nicknames of specific real people, e.g. Talk:RPattz. I don't recall if there has been discussion about whether or not people like including names of groups of specific individuals; we do have e.g. Gang of Four (real), and we have Dynamic Duo (fictional, but was that coined inside or outside the fiction?), but we don't have the Avengers; we do have Lumity. Does this all make sense; should anything be changed? When the reference/denotation is just "the fictional thing/person/people", does it make sense to treat fan-coined terms differently than fiction-coined ones?
- IMO fictional place names should be (and AFAICT are) handled like fictional personal names, we have Mordor (used figuratively), we don't have Hoenn or Endor (though it might be possible to find figurative uses of the later and add it, as a place where hairy short people are found, for example).
- Fictional places' nicknames / unofficial (fan-coined) names... I tend to think these should only be included on the same basis as fictional people's names.
- Terms for items that exist in the fictional universe are in an interesting position. Real-world (non-functional or even functional) replicas may exist, and at the moment, we have a second sense at lightsaber for "A real-world toy, prop, or device fashioned after the fictional lightsaber", but I'm not sure how much sense that makes, as a corresponding sense could just as well (or just as ill) be added for sonic screwdriver (for example, Hacksmith Industries has built a real one that has many of the abilities of the fictional one, and of course cosplayers build nonfunctional ones), and for any number of other fictional things we don't currently have entries for (Hacksmith and others have built functional versions of many other fictional things), so it's not clear to me that we should necessarily be considering the existence of "a real-world replica of X" to a second sense of X, and I'm unsure whether or not it makes sense to consider it to be proof that X is used figuratively / meets FICTION / should be included.
- Terms for specific species of creature seem to me most likely to escape fiction (though there is no guarantee of this). In some cases, like Wookiee, people allude to the characteristics of the creature in ways that establish what we currently call figurative use. In some cases, like hobbit, people use the term in many other, unrelated / unconnected works of fiction, which I think makes them includable (like also e.g. verbs like transmat): some users, though aware hobbits aren't real, might not even realize the creature is from a specific fictional work as opposed to folklore (in contrast to e.g. Pikachu, which I think everyone realizes is from a particular work of fiction). And in some cases, a term itself will predate fiction, but its current meaning will come from specific fictional works, like much of the modern understanding of elves is due to Tolkien. Some more obscure folkloric creatures that were, as best I can tell, coined outside fiction (although it's often unclear to me where they were coined), e.g. gumberoos, are almost as infrequently mentioned as things like naugas, coined for an ad campaign (...I'm not entirely sure whether that counts as being coined in a work of fiction, and it seems reasonable to me that we include it), or e.g. ents, coined by Tolkien, but used by others (and therefore currently included, which seems reasonable to me), and maybe xenomorphs.
- - -sche (discuss) 14:41, 6 July 2026 (UTC)
- Disclaimer: I am answering as to what I think is right and therefore, all statements are subjective opinions (not the status quo).
- To further your thoughts:
- According to this comment in the Hitler RFD dicussion, it survived because of its extended sense (Adolf Hitler → Nazi imperialism and dictatorship). The buck was passed to RFV (which passed with flying colors) and the proposed sense was never added. IMO, alternative names of individuals should be kept because it may not be obvious to someone what the term is referring to. However, this is a separate issue as it is not limited to fictional characters. That is, even if we determine all purely fictional terms should go to their appendix page and keep all alternative names of individuals, the alternative names of the fictional term would still go under its appendix.
- I believe the alternative names of an individual issue would also apply to groups. That is, the official group name (or what they call themselves) should not be recorded, but alternative names should be.
- I completely agree.
- I also completely agree.
- lightsaber is the ideal entry IMO. The first sense is not limited to a specific fictional universe; instead it is more broadly stated belonging to science fiction with quotations showing usage without the implication of Star Wars (although hidden away in its
Citations:page). An example for why the toy sense is acceptable: “I bought my kid a lightsaber for his birthday and he has been playing with it non-stop.” It is impossible to think the speaker is referring to the object that exists in fiction (because it does not exist in reality and therefore, the kid cannot be playing with it). However, I do not think that the toy sense provides justification to add the specific universe sense to establish the link; it can very well be placed in the etymology section. Additionally, we can point the in-universe term to its respective appendix page. Frankly, I wouldn't even consider that the toy/replica sense should fall under WT:FICTION; if it refers to something in reality (even it is heavily based on fiction), it should be included granted that it meets normal attestation requirements. - If this creature (or species) only exists within a fictional universe, then it should not be included in Wiktionary's main namespace unless there is figurative/attributive sense (and therefore, extending its meaning beyond its fictional universe). So Nauga would either be put in its respective appendix or deleted.
- I hope this clears things up and I really want to know your thoughts on this. Is what I am saying too exclusive and simplified? Is this not where to draw the line?
- OPTIONAL TO READ: I have given up that this BP discussion will make it to a vote, but I believe it is due to my poor introduction, framing and initial understanding of the issue. I will take what has been said, draft a new BP post to re-engage in a few months. What we are talking about is not for nothing. TranqyPoo [💬 | ✏️] 04:23, 7 July 2026 (UTC)
- To correct, nauga would fall under the alternative names of an individual issue. Whether we decide to keep or remove alternative names in general would affect how we handle alternative names of fictional beings as well. TranqyPoo [💬 | ✏️] 04:47, 7 July 2026 (UTC)
- I agree with others here that terms such as Darth Vaderian and lightsaber are correctly included here – the former, because it did not originate (as a term) in fiction and is thus not subject to WT:FICTION; the latter, because it is used
independent of reference to that universe
(later worded in policy asout of context
). In the latter case, the way the WT:FICTION policy works has principles consistent also with WT:BRAND (which I have come to see with better eyes). As such, I think it is a neat piece of policy. - Of course, this means WT:FICTION does not exclude some things that may nevertheless be undesirable. For example, I would think peanut does not deserve inclusion (although it seems hard to argue it "did not originate in fiction", seeing as SCP has no centralized media afaik, so anything by the community could well be considered fiction – putting that aside...) because it refers to a specific entity... and this is already covered by WT:NSE. For the same reason (i.e., being strongly against the inclusion of "specific entities" on Wiktionary) I have argued for the deletion of senses for people (e.g., at Sunak), which should be formatted as collocations instead; @-sche made this change to that entry, but I regret we did not have a broader discussion that would allow larger changes.
- And should we have the names of ships and such? Well, I think those are far from the most useful entries in our project; but it’s most definitely part of our goal to include slang (sense 2), and fandom is not a domain whose slang is any less worthy of inclusion.
- Regarding the definition of Darth Vaderian, I believe that the connotations associated with Darth Vader should be additionally described after a semicolon
.— Polomo ⟨ oi! ⟩ · 19:50, 6 July 2026 (UTC);- WT:NSE states
Names of fictional people and places are subject to the “Fictional universes” section of this page
. Taken at face value, one could argue that nicknames or unofficial abbreviations coined by fans are still a name of a fictional person or place and thus, the loophole is exploited. Therefore, the verbiage requires clarification. If one argues the spirit of that verbiage, then one should also think about the spirit of WT:FICTION, which I direct you to my findings. I realize that peanut is a poor illustration of what I am trying to convey; Palpy is a better example (i.e. not originating in-universe, but only refers to an in-universe person). According to our current policy, there are no grounds for removal but yet, we agree that it should be removed. I am advocating to further refine our policy to cover this class of terms. - Regarding fandom: The issue is when the slang term purely refers to an object or concept within a specific fictional universe. This would primarily classify the word as within the Star Wars domain with a subdomain of fandom slang, not vice versa. The moment a definition extends beyond that (i.e. to another fictional universe or into reality), then it is acceptable for inclusion IMO. For example, a term for a fan of a specific fictional universe is valid because it is referring to a real human being (like Buffyhead). Comparing with relationships of fictional characters, this is still entirely within a specific fictional universe as nothing outside of it would affect the relationship.
- I agree that further elaboration is needed for Darth Vaderian and I refer you to my 2nd numbered list item in this comment. Anything else I did not mention from your comment are things on which I completely agree with you. TranqyPoo [💬 | ✏️] 21:00, 6 July 2026 (UTC)
- The issue is broader... we also (and still) have lots of nicknames on Wiktionary, with their own category
[[Category:en:Nicknames for individuals]]. For some relating to Donald Trump, see Orange Head and Forrest Trump. It is my opinion that all 310 of these nicknames should all be deleted (unless someone convinces me otherwise), seeing as Wiktionary is not a catalog of insults to real individuals — if the individuals themselves cannot have entries because Wiktionary is not an encyclopedia, then why do offenses towards them deserve a page?? Anyway the issue extends beyond fiction. — Polomo ⟨ oi! ⟩ · 21:16, 6 July 2026 (UTC)- I share your frustration, but I do not want to tangent away from the main purpose of this discussion. A nickname can apply to reality, mythical creatures, fiction along with other contexts and therefore, deserves its own discussion. This discussion's purpose is to identify what criteria should specific in-universe fictional terms be applied, regardless whether they are nicknames, slang or coined within (or outside of) its universe. TranqyPoo [💬 | ✏️] 21:42, 6 July 2026 (UTC)
- I’m afraid I don’t see any flaws specifically in the WT:FICTION policy, so I and others have addressed a few non-issues with policy. Any remaining gaps seem to me like a much broader-scope problem. — Polomo ⟨ oi! ⟩ · 21:51, 6 July 2026 (UTC)
- I regret that I am unable to show you what I (and -sche, perhaps Quercus solaris?) see. As a last ditch effort, I have selected terms that would survive CFI whose classes have not been discussed. Please consider why the following entries should be included on Wiktionary: Mulderangst, boatsex, Bang That Was Promised, Targling, seeker(sense 3; the term existed before the fiction. Does WT:FICTION mean sense instead of term?) TranqyPoo [💬 | ✏️] 22:40, 6 July 2026 (UTC)
- I didn't understand what "loophole" you were talking about, because fandom slang is a distinct subcategory from terms like Darth Vaderian. My initial inclination is to agree with you that fandom slang is undesirable. It should not be considered citeable unless it also spreads beyond discussions of the universe in question. Andrew Sheedy (talk) 04:47, 7 July 2026 (UTC)
- Er, are you saying this about all fandoms or only fandoms situated in fictional works and universes? Hftf (talk) 04:59, 7 July 2026 (UTC)
- I mean fandom slang that refers to specific in-universe events, people, things, etc. Terms whose referents are strictly within a fictional universe. In these cases, it seems to me to parallel the terms the authors/creators themselves use within their work. However, my opinion isn't settled. I'm just more sympathetic to TranqyPoo on this point than on the others they raised. Andrew Sheedy (talk) 02:06, 8 July 2026 (UTC)
- I may have misrepresented my position, but
terms whose referents are strictly within a fictional universe
is precisely what I'm talking about. It doesn't matter if it is fandom slang, normal slang, nicknames, abbreviations, etc. Any term whose meaning does not extend outside of its specific fictional universe should be subject to WT:FICTION. If you do not agree, I request your position so that I may understand clearly what the blockage is. TranqyPoo [💬 | ✏️] 22:31, 9 July 2026 (UTC)- I don't understand why this is a "loophole". The line was drawn somewhere, and it goes between different words that both ""refer'" to concepts in some universe, and the dictionary and RFD has operated along those lines for a while already. The fact that there are some "fictional" or "fiction-originating" or "fiction-referring" words that have entered the lexicon makes our job subjective and harder and because of this it's better for the line to be inclusive. My two slips of latinum (refers to something only in a fictional universe btw) Hftf (talk) 22:45, 9 July 2026 (UTC)
- I believe the line (the letter of the policy) does not match the spirit, as I have identified here. The "loophole" is we are discriminating words based on who created them, not what they mean. Fan-created words referring to the same thing and using non-independent references are allowed in the main namespace. So, an author-created word is under much stricter criteria than a fan created one where both are semantically identical. Does our current criteria make sense and if so, why does it matter who created it? To be clear, I do not have an issue if a term's meaning has extended beyond its specific fictional scope (like: Jedi Master; note how we do not have its original fictional meaning). Please also see line #2 in this comment. If anything, I'm trying to make this line clearer and more consistent.
- Yes, the dictionary has been operating under these guidelines for a while (since 2008!), but that doesn't mean it is set in stone nor should we be appealing to tradition. TranqyPoo [💬 | ✏️] 23:03, 9 July 2026 (UTC)
- I forgot to mention the important part! That term you just referenced should stay in the mainspace. Maybe its sum of parts meaning refers to something in-universe, but the term's function does not refer strictly to the specific fictional universe because its meaning applies to reality. Hence why it is a synonym. TranqyPoo [💬 | ✏️] 23:13, 9 July 2026 (UTC)
- I don't understand why this is a "loophole". The line was drawn somewhere, and it goes between different words that both ""refer'" to concepts in some universe, and the dictionary and RFD has operated along those lines for a while already. The fact that there are some "fictional" or "fiction-originating" or "fiction-referring" words that have entered the lexicon makes our job subjective and harder and because of this it's better for the line to be inclusive. My two slips of latinum (refers to something only in a fictional universe btw) Hftf (talk) 22:45, 9 July 2026 (UTC)
- To elaborate,
subject to WT:FICTION
is to place these terms within an its specific fictional universe appendix. The moment a term has been used figuratively or attributively to refer to another fictional universe or reality (and can survive attestation) should it then be recorded in main namespace with the new defintion. The specific fictional universe sense should remain in the appendix. TranqyPoo [💬 | ✏️] 22:45, 9 July 2026 (UTC)
- I may have misrepresented my position, but
- I mean fandom slang that refers to specific in-universe events, people, things, etc. Terms whose referents are strictly within a fictional universe. In these cases, it seems to me to parallel the terms the authors/creators themselves use within their work. However, my opinion isn't settled. I'm just more sympathetic to TranqyPoo on this point than on the others they raised. Andrew Sheedy (talk) 02:06, 8 July 2026 (UTC)
- Er, are you saying this about all fandoms or only fandoms situated in fictional works and universes? Hftf (talk) 04:59, 7 July 2026 (UTC)
- I didn't understand what "loophole" you were talking about, because fandom slang is a distinct subcategory from terms like Darth Vaderian. My initial inclination is to agree with you that fandom slang is undesirable. It should not be considered citeable unless it also spreads beyond discussions of the universe in question. Andrew Sheedy (talk) 04:47, 7 July 2026 (UTC)
- I regret that I am unable to show you what I (and -sche, perhaps Quercus solaris?) see. As a last ditch effort, I have selected terms that would survive CFI whose classes have not been discussed. Please consider why the following entries should be included on Wiktionary: Mulderangst, boatsex, Bang That Was Promised, Targling, seeker(sense 3; the term existed before the fiction. Does WT:FICTION mean sense instead of term?) TranqyPoo [💬 | ✏️] 22:40, 6 July 2026 (UTC)
- I’m afraid I don’t see any flaws specifically in the WT:FICTION policy, so I and others have addressed a few non-issues with policy. Any remaining gaps seem to me like a much broader-scope problem. — Polomo ⟨ oi! ⟩ · 21:51, 6 July 2026 (UTC)
- I share your frustration, but I do not want to tangent away from the main purpose of this discussion. A nickname can apply to reality, mythical creatures, fiction along with other contexts and therefore, deserves its own discussion. This discussion's purpose is to identify what criteria should specific in-universe fictional terms be applied, regardless whether they are nicknames, slang or coined within (or outside of) its universe. TranqyPoo [💬 | ✏️] 21:42, 6 July 2026 (UTC)
- The issue is broader... we also (and still) have lots of nicknames on Wiktionary, with their own category
- WT:NSE states
Handling ᚨᚾᚾ
[edit](Feel free to move this if this is the wrong forum.) ᚨᚾᚾ was originally listed as Frankish, but was moved to Old Dutch after Frankish was made an etymology-only variant of Proto-West Germanic, including retroactively at FWOOD. However, seeing as ᚲᚨᛒᚨ is attested and listed as Proto-West Germanic, is it OK to move this back to Proto-West Germanic, with a Frankish context label (as per entries in Category:Frankish)? Pinging @Mahagaja, who made the move originally. -Brainulator9 (TALK) 22:11, 3 July 2026 (UTC)
- It would require unmarking Proto-West Germanic as a reconstructed language so that it could have an entry in main space instead of RC: space. That's not unprecedented for a proto-language; both Proto-Brythonic and Proto-Norse have entries in main space and are not marked as reconstructed. But it feels a bit excessive for a single word, or even a handful of words from a single inscription. I just don't see any harm in continuing to call it Old Dutch. —Mahāgaja · talk 08:48, 4 July 2026 (UTC)
- I think we already have the infrastructure to support this, and indeed already have a mainspace PWG entry in the form of ᚲᚨᛒᚨ (discussed here), don't we? IIRC a change was made to allow individual entries in a mostly-reconstructed language to link to / be in mainspace by adding !! in front of the term instead of *. Whether it makes sense to consider the particular entry in question above to be PWG or not I leave to other people. - -sche (discuss) 17:41, 4 July 2026 (UTC)
User:Blahh (bot) for bot status
[edit]I would like to see if we can use User:Blahh (bot) for bot status.
There exists a template called Template:cmn-ear-l. It is used to show the Early Mandarin reconstructed readings of various Chinese characters. You can look at Template:cmn-ear-l/documentation to get a feel of how it works.
But the way that it works is, if you click this link:
There needs to be a separate Module page for each separate character, in the format of
Module:zh/data/pulleyblank/"Chinese_character"
and then the Module:zhx-pulleyblank-pron, which is the core of how the template works, searches for the Module for each separate Chinese character, and obtains their reconstruction data. Unfortunately, doing this manually will be very slow, inefficient, and painful. This is a prime example of a tool that a robot/bot can accomplish better than a human.
I have privately already tested the capabilities. You can look up the recent edits of User:Blahh (bot) to see how well it does at this job. Blahhmosh (talk) 09:28, 4 July 2026 (UTC)
- @Blahhmosh I think this looks like a cool data source to add. The bot task seems quite structured, too, so probably not likely to go wrong. Do you also have the source code available?
- P.S. You might like to copy this to Wiktionary:Votes, where you can get the bot approved. Kiril kovachev (talk・contribs) 23:37, 15 July 2026 (UTC)
- My bad, didn't see you already had one created! Kiril kovachev (talk・contribs) 23:40, 15 July 2026 (UTC)
- I have changed my mind. A bot is not necessary here. Blahhmosh (talk) 02:00, 16 July 2026 (UTC)
- My bad, didn't see you already had one created! Kiril kovachev (talk・contribs) 23:40, 15 July 2026 (UTC)
Bantu noun classes
[edit]Many Bantu languages have large numbers of noun classes and their noun entries are categorized accordingly (see Category:Nouns by class by language). However, few of these languages seem to have appendices (like Appendix:Swahili noun classes) detailing how their noun classes work.
As far as I can tell, classes and their numberings are largely inherited from Proto Bantu, and there is decent overlap between all of them. How could this information best be shown? RajanD100 (talk) 20:29, 4 July 2026 (UTC)
Mass-addition of Oromo entries
[edit]First I noticed some module errors from {{om-IPA}} feeding bad characters to the IPA module. Then I noticed that some of those were because of entries for Oromo interjections (under an erroneous "Noun" POS) ending in exclamation marks. There were also some other odd characters, so I thought I would check older entries for comparison. After spot-checking in Category:Oromo lemmas, I discovered that pretty much everything was added very recently by the same user, @EyobAbebe7. Of their 14,689 edits, all but a dozen or so have been since noon on July 3- even if they were editing nonstop for those 4 1/2 days, that would still average out to more than 2 per minute.
As @Thadh mentioned on EyobAbebe7's talk page, our Oromo coverage before this was pretty thin, but- aside from the blatant violation of our bot policy- you have to wonder if the content was illegally scraped from some website or app.
Any thoughts on what we should do about this? Chuck Entz (talk) 05:21, 8 July 2026 (UTC)
- I noticed that this morning. It's worth pointing out many definitions contain obvious typos (e.g. xoofoo, "calourful cow") as well, and some are unclear (e.g. handaaqqoo wayaa, "hen of cloth", presumably a literal translation of an idiom). Also, they cite a dictionary whose author appears to be that same user. HootenannyHoller (talk) 05:29, 8 July 2026 (UTC)
- Hi everyone,
- Thanks for raising these concerns.
- I'd like to clarify a few things.
- The recent large number of edits is because I've been spending a lot of time expanding Oromo coverage, which was previously very limited. Most of the content was generated from my own lexical data and then reviewed before being added. I agree that the edit rate was too high, and I now understand that this conflicts with Wiktionary's bot policy. I apologize for not requesting bot approval first.
- Regarding the source of the data, it was not scraped from another website. I am the author of the Buurana Oromo–English Dictionary application, and the lexical database used for these entries is my own work. I therefore hold the rights to the underlying data.
- As for the module errors, thank you for pointing them out. They appear to be caused by some entries passing unexpected characters (such as punctuation in interjections) to
- I am already working on fixing the module and cleaning up the affected entries.
- If necessary, I'm happy to slow down my editing, seek bot approval before making large-scale automated edits, and work with the community to ensure everything complies with Wiktionary's policies.
- Thank you for your patience and for helping improve Oromo on Wiktionary. EyobAbebe7 (talk) 05:33, 8 July 2026 (UTC)
- Hi again @Chuck Entz and @HootenannyHoller,
- To follow up on my previous message and show my commitment to resolving these issues, here is my immediate plan of action:
- 1. Copyright Verification: To verify my identity as the data collector and author of the source material, my contact email is eyobdessalegn15@gmail.com. If I need to send a formal declaration of consent through the Wikimedia VRT (Volunteer Response Team) process to clear up any copyright concerns, please let me know and I will do so right away.
- 2. Bot Approval: I will review the bot policies and submit a formal Bot Request for approval before I do any further large-scale or automated editing.
- 3. Manual Cleanup: I am immediately starting the manual cleanup process today. I will be going through my recent contributions to correct typos, fix literal translations of idioms, and remove the punctuation from interjections that caused the IPA module errors.
- Thank you again for your guidance. I am getting right to work on the manual corrections! EyobAbebe7 (talk) 07:20, 8 July 2026 (UTC)
- Hi EyobAbebe7,
- it's always nice to see new people working on languages with small coverages here on Wiktionary. Thank you for that!
- That being said, I can't shake the feeling that many of your contributions are AI-generated. Large edits on modules with hundreds of lines of code added, the use of emojis in declension tables and the general tone and styling of your comments seem quite AI-like to me. I'm sorry if this isn't the case, but if it is, we would ask you to disclose that.
- Anyway, here are two ways you can improve your edits:
- Happy editing! Tc14Hd (aka Marc) (talk) 14:58, 8 July 2026 (UTC)
- Hi Marc,
- Thank you for your kind words and for the helpful suggestions.
- Yes, I do use AI tools as an assistant for some tasks, especially when writing Lua modules, documentation, or improving my English. However, I review, test, and modify the code and content before publishing it. The Oromo linguistic data, templates, and language-specific rules are my own work.
- Thank you also for pointing out { {lb} } and { {lit} }. I wasn't aware of those templates, and I'll start using them where appropriate. I appreciate the guidance and will continue improving my edits to better follow Wiktionary's conventions.
- Thanks again, and happy editing! EyobAbebe7 (talk) 17:55, 8 July 2026 (UTC)
Problematic edits by a user
[edit]Edits by CockroachHunter. It feels bad but I'd revert most of his new definitions. The added senses are just needless repetition or specific usage examples. What do you think? Aloysius Jr (talk) 10:43, 8 July 2026 (UTC)
- The edits have mostly been reverted. That's that then. Aloysius Jr (talk) 21:09, 8 July 2026 (UTC)
- I've reverted most of their edits as they were engaging in biased blogging. box16 (talk) 21:20, 8 July 2026 (UTC)
Korean vowel length
[edit]I find the current treatment of Korean vowel length unclear, and hence unsatisfactory. Template Template:ko-IPA generates the phrase
- Though still prescribed in Standard Korean, most speakers in both Koreas no longer distinguish vowel length.
and I have no major problem with that statement.
What bothers me is the selective use of the vowel length mark (which triggers the appearance of the above statement). Two simple examples are 벌 and 별.
In the case of 벌, I believe that the sense bee traditionally has a long vowel, and the sense punishment traditionally has a short vowel. (I have no knowledge of traditional vowel length in the other three senses/etymologies.) For bee the length mark is parenthesised, and even shown in red in the phonetic hangul, namely [pɘ(ː)ɭ] and [벌(ː)]. No length mark is shown for the other senses/etymologies.
- By reading the pronunciations, it would appear that in the bee sense the vowel can be pronounced either long or short (or anywhere in between) in today's Korean, but the other senses/etymologies must be pronounced with short vowels.
- Yet the claim that "most speakers in both Koreas no longer distinguish vowel length" — if true — implies that vowels in EVERY Korean word can be pronounced with any length according to the speaker's whim. (If so, then parenthesised length marks should follow EVERY Korean vowel, and the explanation should be included for EVERY sense/etymology.)
So the first problem is that the pronunciation guides and the explanation lead to two different conclusions.
The second problem is that the parenthesised length mark does not directly tell the reader what the traditional vowel length was (as still used by some minority of speakers today) — which I presume was the intended purpose.
Of course, by examining the other senses/etymologies and with some prior background knowledge one can play the role of Sherlock Holmes and make a deduction. But I don't think that is an ideal way for Wiktionary to be set up.
Put it another way. The current convention in Wiktionary results in
- 벌 [pɘ(ː)ɭ] / [벌(ː)] bee (Vowel length is not distinguished by many modern speakers.)
- 벌 [pʌ̹ɭ] / [벌] punishment
But there seems to be just as much (bad) logic in an alternative convention of
- 벌 [pɘːɭ] / [벌ː] bee
- 벌 [pʌ̹(ː)ɭ] / [벌(ː)] punishment (Vowel length is not distinguished by many modern speakers.)
I think most logical would be to drop the parentheses entirely and rely on the explanation.
- 벌 [pɘːɭ] / [벌ː] bee (Traditional vowel length is shown, but it's not distinguished by many modern speakers.)
- 벌 [pʌ̹ɭ] / [벌] punishment (Traditional vowel length is shown, but it's not distinguished by many modern speakers.)
별 and almost every other Korean entry in English-language Wikipedia suffer the same issue(s).
BTW, I've also seen the tilde used in a small fraction of Wiktionary entries to indicate vowel length. (Unfortunately I cannot find an example right now. At the time that I saw them I was blocked from doing anything as a victim of collateral damage from a long, broad IP-range block, with a dynamically assigned IP address. Which has become an increasingly common occurrence. [No, I do not want to create an account, thank you.] I am currently, temporarily, on a different IP address.) This should not be happening on English-language Wiktionary; it is a convention from Korean, so it may be tolerable on Korean-language Wiktionary.
—DIV (~2026-38747-31 (talk) 14:14, 9 July 2026 (UTC))
- Upon further inspection, the technical source of the current Wiktionary convention may be Module:ko-pron.
- —DIV (~2026-38747-31 (talk) 14:17, 9 July 2026 (UTC))
- See w:Korean_phonology#Loss_of_vowel_length_contrast. Although "for most of the speakers who still utilize vowel length contrastively, long /ʌː/ is actually [ɘː]," I have ignored that in the above discussion; please don't let that become the focus of discussion here. —DIV (~2026-38747-31 (talk) 14:24, 9 July 2026 (UTC))
- While the transcriptions are shown in phonetic brackets, I believe they need to be interpreted according to the phonemic principle. (That's often the case for the symbol "ː" anyways, since phonetic length is a gradient variable, whereas "ː" is a discrete symbol. That is, I would say the point of using "ː" is usually to imply that we're talking about some perceptually significant category of duration, whether phonemic or allophonically conditioned, and not just meaningless free variation in length, which is always possible to some extent.)
- Therefore, I think it's overly literal to interpret "[pɘ(ː)ɭ]" and "[pʌ̹ɭ]" as "long or short" vs. "mandatorily short", respectively. What it means is instead this: "long" vs. "short" for an accent that makes length distinctions, no length distinction for an accent that makes no length distinctions. I can see the theoretical potential for confusion, but I think the proposed convention of dropping the brackets doesn't really make the transcription clearer in practice. It's true that leaving short vowels unmarked and with no note is ambiguous; however, including the note beneath every Korean transcription might use up a lot of space, so I think it's a tradeoff.--Urszag (talk) 16:05, 9 July 2026 (UTC)
- Well put Hftf (talk) 22:47, 9 July 2026 (UTC)
Uzbek orthography
[edit]The Uzbek parliament passed an act a few days ago, to officially change ⟨ch⟩, ⟨sh⟩, ⟨gʻ⟩ and ⟨oʻ⟩ to ⟨ç⟩, ⟨ş⟩, ⟨ğ⟩ and ⟨ö⟩ respectively. Should we start moving pages now to align with that standard? So shtat would be ştat for example, and özbek instead of oʻzbek. Dijacz (talk) 13:02, 10 July 2026 (UTC)
- Are we here to implement what governments decide or to document what human people do? Vahag (talk) 13:20, 10 July 2026 (UTC)
- I guess what you're saying then is to wait and see whether Uzbeks are enthusiastic about using this new orthography, and then only moving it when/if Uzbek people have widely adopted it? Dijacz (talk) 13:22, 10 July 2026 (UTC)
- I am saying we shouldn't call the entity in Anatolia Türkiye just because the Angora government has asked to. Vahag (talk) 13:24, 10 July 2026 (UTC)
- I mean, far be it from me to comment on Uzbek politics, and I don't discount the possibility that this is the ruling party trying to consolidate voter support in imitation of some sort of "pan-Turkic" identity, but this has literally zero impact on the wider world outside of Uzbekistan and Uzbek-related content. English transliteration will probably spell Xusanov as Khusanov until the end of time. I'm fine with staying put for a bit with this new orthography, but surely some sort of criteria should be set, the fulfillment of which should necessitate updating the Wiktionary entries accordingly. Dijacz (talk) 13:37, 10 July 2026 (UTC)
- That's ridiculous and not actually related to what we're talking about. ―Justin (koavf)❤T☮C☺M☯ 23:18, 10 July 2026 (UTC)
- I am saying we shouldn't call the entity in Anatolia Türkiye just because the Angora government has asked to. Vahag (talk) 13:24, 10 July 2026 (UTC)
- I guess what you're saying then is to wait and see whether Uzbeks are enthusiastic about using this new orthography, and then only moving it when/if Uzbek people have widely adopted it? Dijacz (talk) 13:22, 10 July 2026 (UTC)
- I'm not Özbek (obviously) but this is a part of a modern 'Common Turkic' trend, similar to a number of recent(ish) reforms [Azerbaijani in 1994, Turkmen in late 1990s and early 2000s, Tatar in 2000s (later shot down by Putin), Kazakh in 2010s, also recently Kyrgyz, and so on] that propose a unified spelling for Turkic languages (CTA).
- My vote is irrelevant, but I'd support changing it if it was.
- AmaçsızBirKişi (talk) 14:11, 10 July 2026 (UTC)
- Unfortunately, these spelling reform dictates sometimes take hold and sometimes don't, so it's hard to know if this will be a lasting change. My suspicion is that it will be largely implemented, due to how Uzbek is not a very widely distributed language and a government authority could realistically impose these spellings on a lot of legal documents and public broadcasting pretty quickly, which will have a definite impact on everyday use. As you can see from a local search, spelling reform and orthography issues have come up for several languages and they usually seem to be resolved on a case-by-case basis, so there also isn't a site-wide policy that we can apply here.
- All that said, as someone who doesn't know this language, I imagine that the best practice will likely be to keep older forms, mark them as "pre-2026 reform spellings", categorize them as variants, and make the post-reform spellings the standard lemmas. ―Justin (koavf)❤T☮C☺M☯ 23:24, 10 July 2026 (UTC)
- As an example of the above list of search results: Wiktionary:Tajik_entry_guidelines#1998_spelling_reform. ―Justin (koavf)❤T☮C☺M☯ 23:26, 10 July 2026 (UTC)
- Thanks for the response. Yes, the Tajik example I think serves as a good guideline for how to go about handling the Uzbek spelling reform. And I reckon a good benchmark for when to start doing this stuff is if/when online Uzbek dictionaries start adopting the new spelling en masse. Dijacz (talk) 00:03, 11 July 2026 (UTC)
- What are the details of this change? Is there a transition period? Are the former spellings completely outlawed?
- Depending on these details, it may or may not be prudent to move the entries. If the former spellings will be considered misspellings at some point, then surely we shouldn’t have the lemma entries placed there, just like User:Koavf suggests. I suspect the best way to gauge this is indeed probably by following what Uzbek dictionaries do, though, as User:Dijacz said. — Polomo ⟨ oi! ⟩ · 17:07, 11 July 2026 (UTC)
- Reading the article I linked earlier, it doesn't seem like too many implementation details are set in stone yet, and the bill still has to be approved by the Senate of Uzbekistan, but what I gleaned from it is that the transition will be a gradual one. According to another Russian-language article that I was sent, they will only determine a date for its enforcement after being signed by the president.
- As for what actually changes, well, those four letters - ⟨ch⟩, ⟨sh⟩, ⟨gʻ⟩ and ⟨oʻ⟩ being changed to ⟨ç⟩, ⟨ş⟩, ⟨ğ⟩ and ⟨ö⟩. That's it. The English-langauge article I linked also listed some words which contained sequences that would be spelt with сҳ (sh) in Uzbek Cyrillic, which had to be written as s’h in Latin because of the existence of ⟨sh⟩, e.g. Is’hoq and As’hobiddin, whose misspellings Ishoq and Ashobiddin would now presumably be considered the correct spellings. Although the glottal stop, spelled ъ in Uzbek Cyrillic and written as an apostrophe(?) in Uzbek Latin, would still remain unchanged after the letter ⟨o⟩, and the article lists mo'tabar (моътабар), mo'tadil (моътадил) and mo'jiza (моъжиза) as examples of this.
- And a quote from the Russian-language article: "the new letters will be introduced gradually and without a campaign: existing documents, national currency and securities, signs, and seals of government agencies will remain valid until the government-set deadline, so as not to replace everything at once and burden the budget". (Courtesy of Google Translate.) And we don't even know when the bill is going to go to the Senate and President yet, so I guess we needn't be worrying about this for the moment. Just a heads up for the future, I suppose. Dijacz (talk) 19:07, 11 July 2026 (UTC)
Wiktionary:Votes/2026-07/Unblocking Apisite
[edit]
Support @Apisite I would like to respectfully propose unblocking Apisite. I do not know how to do a vote, so if anyone agrees and has the capability, I would like you to take over. But I think it makes sense. cf. Lazar Kaganovich: "Releasing people we arrested three days ago will make us look cretinous." [16] at 2:08 --Geographyinitiative (talk) 🎵 23:08, 10 July 2026 (UTC) (modified)
- Courtesy links: Apisite (talk • contribs • global account info • deleted contribs • nuke • abuse filter log • page moves • block • block log • active blocks) ―Justin (koavf)❤T☮C☺M☯ 23:18, 10 July 2026 (UTC)
Oppose: this is handled by an unblock request, not by a vote. This is their 6th block, the 4th for this exact reason, with increasing durations each time, so it’s a natural progression. Unless you think their edits are all fine, and that the issues that were seldom addressed on their talk page don’t exist? — Polomo ⟨ oi! ⟩ · 16:58, 11 July 2026 (UTC)
Oppose the vote per WT:Blocking policy. No opinion on the unblock request. DCDuring (talk) 17:18, 11 July 2026 (UTC)
Oppose the editor was only blocked six days ago. What are the convincing reasons justifying unblocking? What has changed about their behaviour since then? — Sgconlaw (talk) 19:13, 11 July 2026 (UTC)
- I hope the editor will not give up on Wiktionary because of this block. I tried something to help, if I could- see above- but I am ultimately powerless. I will support the editor in any unblock because I believe they are ultimately contributing in a positive way to the dictionary project. In history, many people have been temporarily stopped from pursuing their passions for a short time. I look forward to the editor's return and give them and above editors my best wishes. --Geographyinitiative (talk) 🎵 21:45, 11 July 2026 (UTC)
- Did you look at WT:Blocking policy? DCDuring (talk) 00:17, 12 July 2026 (UTC)
Pannonian Rusyn etymon text
[edit]Can Pannonian Rusyn (rsk) be added to Module:etymon/data/text allowed? Since it's a descendant of Old Slovak (zlw-osk), and Slovak (sk) is already in that list. Thanks. Dijacz (talk) 11:05, 12 July 2026 (UTC)
Is the Columbia Gazetteer actually documenting English pronunciations of non-nativized place names, or approximating their foreign pronunciations with an English soundset?, conlanging, and other tales
[edit]I read the short frontmatter of this geographical reference work (viewable on archive.org) but it's not shedding much unambiguous light on what its pronunciation guides really intend to accomplish (prescriptivism? descriptivism?) and how strongly it gives evidence for the actual English pronunciation of words, so I thought this community could take a look. I accept that a competent English speaker has unintuitive grapheme–phoneme mappings in marginal/particular words or in a subsystem of words (say, pronouncing /v/ for <w> and /ts/ for <z> in German-derived words, nasal vowels in French-derived words, etc.) and that this phenomenon varies across world Englishes with contact; I am also willing to accept that some English speakers and some particular English words happen to use the English /p/ in a word that happens to be spelled with <b> ("Beijing" etc.). It just seems difficult to accept that at this stage (1) actual descriptive data is available on the actual English-language pronunciations of thousands of Chinese place names that an extremely tiny number of English-language speakers have ever heard of, (2) that English speakers are seeing a word written with <r> and actually pronouncing it /l/ according to a frankensteined assemblage of parts of other words whose native Modern Standard Mandarin pronunciation contains a related syllable or a word with <si> and pronouncing it /ʃ/ or a word with a "silent y" in <yi> (Yibin) and so on. See one of many prior discussions at Wiktionary:Beer_parlour/2025/August#Mandarin-esque_pronunciations_on_English_entries. Similarly I can't imagine we actually want AIs which don't even understand enPR evaluating those pronunciations for us. Is it better for the dictionary to abstain from authoritatively citing/claiming unintuitive pronunciations as English pronunciations where actual descriptive evidence is weak? Hftf (talk) 23:21, 12 July 2026 (UTC)
- @Hftf, If you want to throw out Columbia Gazetter and its pronunciations, we'll need to get rid of Merriam Webster and International Geographic Encyclopedia and Atlas, too: see for example: Fuzhou diff, Guangxi diff, Beijing, Gansu, Xinjiang, Jiangsu, Wuhan, Guangzhou, Qinghai, Hunan, Hangzhou, Qingdao, Hsinchu, etc., where two or all three align on one pronunciation when converted into Wiktionary's enPR system. It's a thorny area, and there are multiple variant pronunciations for almost every word, and Wiktionary is perfectly designed for exploring and documenting this area. I will try to explain my edits.
Do not let the perfect be the enemy of the good, or the preliminary. In fact, go ahead and add pronunciations. I think the key question you're raising whether these pronunciations should be evaluated on a case-by-case basis, or whether the position is that Columbia Gazetteer pronunciations should generally be excluded. I think there may be a distinction between something being "authoritative" and something being a useful starting point or preliminary source. For example, Zhengzhou examines a broader range of possible pronunciations rather than relying solely on the Columbia Gazetteer. The entry also draws from sources such as The International Geographic Encyclopedia and Atlas (Boston: Houghton Mifflin Company, 1979), Merriam-Webster, and other dictionaries.
- (1) The fact that a place name has very limited usage does not necessarily mean that no systematic pronunciation exists. Many scientific terms and technical names have relatively rare usage but still develop predictable pronunciation patterns among English speakers. Similarly, some geographic names may exhibit systematic hyperforeign features or other adaptations rather than purely spontaneous or irregular pronunciations.
- (2) Transliteration systems and recurring syllable patterns can create predictable ranges of possible English pronunciations. The question is therefore not whether every listed pronunciation represents a widely attested spontaneous usage, but whether the sources are documenting established patterns, conventional adaptations, or plausible English renderings. Entries such as Zhengzhou attempt to explore that range rather than assume a single outcome.
- The Columbia Gazetteer and The International Geographic Encyclopedia and Atlas are not generally being questioned as sources for non-Chinese geographic names, and their treatment of Chinese names should likewise be evaluated on a case-by-case basis. AI should also be treated as a tool, comparable to a calculator: useful when it assists human editors and reduces unnecessary labor, but not a substitute for human judgment, verification, or source evaluation. (AI-assisted drafting) --Geographyinitiative (talk) 🎵 00:12, 13 July 2026 (UTC)
- I'd take it with a grain of salt and common sense even if it is mostly right (I have not tried to evaluate whether it is or not), if it ever says that e.g. the written form Beijing is pronounced with /p/, or that a written r is pronounced /l/. As I said in the August 2025 discussion, as someone who finds it amusing that everyone uses spelling pronunciations of Chinese when speaking English, it is nonetheless my experience that everyone—including Chinese government officials—pronounces Beijing with /b/ when speaking English (and not only English, by the way: Hindi could reproduce the aspiration distinction Mandarin makes, but it doesn't, it too goes for a pinyin spelling pronunciation), so a claim that people pronounce b with /p/ would need good (ideally multiple and modern) sources. It's not inconceivable that speakers might pronounce r as /l/, cf. Wiktionary:Tea room/2026/February#Taoism (pronunciation) where multiple people said that yes, they pronounce d as /t/ and/or t and /d/, but I'd want to see good evidence. - -sche (discuss) 00:21, 13 July 2026 (UTC)
- Please note that the Columbia Gazetteer, as cited on Beijing, does not give a /p/ pronunciation for Beijing. Rather, it gives bāʹjĭngʹ, which corresponds to an English pronunciation beginning with /b/. That pronunciation has its own issues, but it is at least close to what many English speakers actually say.
- As for forms such as Gwoyeu Romatzyh, I don't think the existence of an /l/ pronunciation is as implausible as suggested. In the underlying romanization, Ro represents Mandarin luo, so there is a straightforward historical explanation for why some pronunciation guides approximate it with English /l/. That does not mean every English speaker will do so. Many readers encountering the spelling for the first time will naturally pronounce the written <r> as /r/, and that pronunciation is already represented. The question is whether a more Mandarin-approximating pronunciation should also be documented where reputable sources support it.
- I agree that Gwoyeu Romatzyh is an unusual case and should not be taken as representative of Chinese loanwords generally. If evidence ultimately shows that one pronunciation predominates, then that should be reflected. Until then, I think documenting a supported range of pronunciations is reasonable. However, out of respect fo your opinion, I'm just taking down the enPR temporarily. It's unquestionably a freak edge case and should not be the focus of this kind of discussion.
- On Yibin, I share some of the skepticism. The <yi> → "ye" pronunciation is strange and seems less convincing to em as a possible variant pronunciation than I previously gave weight. Unless additional evidence turns up, I have hidden and am inclined to remove that pronunciation until it can be better supported.
- Regarding Gong Xi Fa Cai, there is at least one cited recording of an English speaker pronouncing the expression- see the Citations page. I took down the enPR out of respect for your opinion.
- More broadly, my goal is simply to document how these words are pronounced in English as accurately as possible. Where the evidence is strong, we should present it confidently. Where it is weak or conflicting, we should acknowledge the uncertainty. I don't think that uncertainty justifies discarding pronunciation evidence from sources such as the Columbia Gazetteer, The International Geographic Encyclopedia and Atlas, or other dictionaries altogether.
- If you could somehow invalidate all the pronunciations from Columbia Gazetteer, including even the ones with support from other dictionaries, that's still okay with me in theory; the whole set of pronunciations could therefore be subject to deletion if invalidated. However, I think the Columbia Gazetteer is an important resource on the pages it is currently linked to, so even if the whole collection of pronunciations is removed from Wiktionary, I would want to keep those Columbia Gazetteer links in Further reading, in my opinion. (AI-assisted drafting) --Geographyinitiative (talk) 🎵 00:41, 13 July 2026 (UTC)
July 2026 Wikimedia Café meetups regarding Wikimedia governance and options for reform
[edit]Hello! There will be two Wikimedia Café discussion opportunities in July. Both sessions will focus on Wikimedia governance, including possible follow-ups to the Movement Charter and options for reform. Participants may attend either or both Café sessions.
This month, to deconflict the Café meetups from Wikimania, the meetups will be held one day later than usual.
- 26 July 2026 15:00 UTC (timestamp converter), at a time friendly to the Americas, Africa, and Europe
- 27 July 2026 03:00 UTC (timestamp converter), at a time friendly to Asia and the Pacific
Please see the Café page for more information, including how to register!
↠Pine (✉) 03:43, 13 July 2026 (UTC)
Use of letters ଵ ('va') and ୱ ('wa') for Odia
[edit]According to the Unicode specifications (version 17.0) the letter ଵ ('va') was invented to distinguish original 'v' and 'b' in Sanskrit (both becoming /b/ in Odia). "The letter va is not in common use today", so we probably shouldn't use it in page names (there's currently ଦ୍ଵୀପ (dvipa) and ଵିଷ୍ଣୁ (viṣṇu), the latter already has the duplicate ବିଷ୍ଣୁ (biṣṇu)).
As to the letter ୱ, Unicode indicates it is "sometimes used in Perso-Arabic or English loan words for [w]". But we are also using it for words of Sanskrit origin in positions (after consonants) where it is pronounced [w]. This is how they are spelled by the Praharaj dictionary (from 1931-40), but from what Unicode says the letter ବ 'ba' is used there.
So should we switch all uses of ଵ and ୱ (outside of Perso-Arabic and English loans) to ବ ?
Pinging Odia editors is useless, as they are either inactive or blocked. Exarchus (talk) 16:58, 13 July 2026 (UTC)
- Apparently Odia Wikipedia also uses spellings like ସ୍ୱାଧୀନତା (with ୱ), so if that is actually common, then only uses of ଵ would have to be changed. Exarchus (talk) 17:36, 13 July 2026 (UTC)
Encourage the use of practical quotations, instead of meta discussion
[edit]As the title states, I would like to propose some policy to use practical quotations over meta discussing ones. Let me explain. The point would be to have the quotation be a real life practical use of the word, instead of a meta discussion of it, through some garbage tier news article or sham academic paper. Basically, prioritise posts made by people and book passages for quotation, which use the world earnestly, as it would be in real life, over academic papers discussing the meta of a word. This is very noticeable with most of the slang. An example, for "pajeet", one of the quotations is some garbage tier academic paper called "Myth, Mysticism, India, and the Alt-Right, in The International Alt-Right: Fascism for the 21st Century?", that discusses the word in a meta manner. It fails to show the real word usage of the word. Instead, a quote should show someone practically and earnestly using the word, in a Reddit post, for example. A quote should show how a word is being used by people in real life, so the reader can understand how it is being used. Using meta discussing academic papers and article misses the point. ~2026-36117-21 (talk) 03:35, 14 July 2026 (UTC)
- It seems like you are discussing the use-mention distinction. If someone is organically using a word in a way that is illustrative of what that word means and does so in durably-cited media, that is definitely a better quotation to include than one where someone is discussing the word as a word. In those cases, please include
brackets=yesas you can see at (e.g.){{quote-journal}}. If I'm missing your point, pardon me. If you have any proposed changes to our existing Wiktionary:Quotations, please let us know. ―Justin (koavf)❤T☮C☺M☯ 04:08, 14 July 2026 (UTC) - @~2026-36117-21: your suggestion to "encourage the use of practical quotations" is already Wiktionary policy. The quotation you mentioned in pajeet is bracketed, indicating that it doesn't even count as a use (see Wiktionary:Criteria for inclusion § Conveying meaning). However, the quotation is still useful to the entry as it proves that the term was already widely used by 2020. Ioaxxere (talk) 18:18, 14 July 2026 (UTC)
- If it is already a policy, why is there so much of meta discussion quotes like this, for most of modern slang? Seems like a weak policy, to me. Also, the brackets are useless. Only a seasoned editor knows what it means. An average reader who goes up to look up a word, won't know what it means. Either delete those types of quotes (preferred), or mark it with a more explicit note, that informs the reader about it. ~2026-36117-21 (talk) 21:44, 14 July 2026 (UTC)
why is there so much of meta discussion quotes like this, for most of modern slang?
Your premise is incorrect, only a tiny fraction of quotes are mentions like the one on pajeet. I know this because I have created many entries of "modern slang" and I rarely include such quotes. There is no reason to delete the existing ones as they are typically helpful for understanding the term. Ioaxxere (talk) 21:58, 14 July 2026 (UTC)
- If it is already a policy, why is there so much of meta discussion quotes like this, for most of modern slang? Seems like a weak policy, to me. Also, the brackets are useless. Only a seasoned editor knows what it means. An average reader who goes up to look up a word, won't know what it means. Either delete those types of quotes (preferred), or mark it with a more explicit note, that informs the reader about it. ~2026-36117-21 (talk) 21:44, 14 July 2026 (UTC)
How to quote terrorist content
[edit]On some entries, it may be useful to quote extremist sources in an entry. For example, they may be the first recorded usage of a particular term. This shouldn't be a problem for the U.S.-based Wikimedia Foundation, whose content is protected by the First Amendment. However, other countries, including mine, give the government sweeping powers to censor and criminalize hate speech.
Example 1: The Turner Diaries was published in 1978, making it an important historical source for neo-Nazi/white supremacist ideology. It should arguably be quoted in our entry for day of the rope. However, this could be risky in certain countries, particularly the U.K.
Example 2: The Christchurch shooter manifesto is fully banned in New Zealand. The government can imprison editors adding quotations from it and then block Wiktionary.
Proposal
[edit]Since Wiktionary is WT:NOTCENSORED, we shouldn't cater to repressive governments beyond the minimum required by U.S. law. However, since we are interested in minimizing cases of countries blocking Wiktionary, I suggest implementing a gadget that automatically hides "objectionable" quotes within certain countries. Editors would be able to disable this gadget at their own risk, just like Iranian and Chinese editors use VPNs at their own risk.
Of course, it goes without saying that extremist content should only be quoted in entries that deal with extremist language. Quotations shouldn't be offensive or controversial unless it actually shows typical usage of the term. Ioaxxere (talk) 19:20, 14 July 2026 (UTC)
- Wouldn't a legal disclaimer be possible to add instead of a hiding gadget in certain cases? Davi6596 (talk) 23:17, 14 July 2026 (UTC)
- If the goal were to stop Wiktionary from getting banned, unfortunately the disclaimer would probably not be enough, I suppose... Kiril kovachev (talk・contribs) 22:27, 15 July 2026 (UTC)
- @Ioaxxere What do you find to be the likelihood that Wiktionary actually be banned because of hosting excerpts of terrorist content? I admire your idea, it's good - just I wonder whether it's really something we need to worry about.
- Also, from another angle, I don't know if it really counts as hiding the terrorist content in that country unless it becomes impossible to access it via the website at all, whereas, since everyone is able to read the source code of entries, whilst they may find it hard to e.g. discover the manifesto from within New Zealand, they will definitely be able to access the excerpt if they go to edit the page. So, can that method really be said to be in compliance with the law?
- I think, personally, Wiktionary is not on any governments' radars, and it is one of the last places they will go to look when implementing crackdowns on things like these. I don't know if us doing anything about this is feasible. Kiril kovachev (talk・contribs) 22:26, 15 July 2026 (UTC)
- I believe the only possible solution (but most likely only within WMF's control) is the ability to block specific entry pages. However, I agree with your statement. Hiding a quote would not satisfy a government's ban on the material; it must be inaccessible. TranqyPoo [💬 | ✏️] 00:35, 16 July 2026 (UTC)
Consistent transliteration of Bulgarian ъ/ь
[edit]P.S. Sorry if you get pinged twice, I just accidentally put my post in GP first. (Notifying workgroup: Atitarev, Benwing2, Bogorm, Bezimenen, Chernorizets, Loccaall, SimonWikt, SixtyShips, Илья А. Латушкин, Chihunglu83, Psi-Lord, Kaloan-koko): Following the previous post about fixing the symbol used for transliterating Ъ in Bulgarian, I would also like to bring a related matter to the fore, which is how Ъ and Ь (usually transliterated Ă and J/ʹ) should be handled in the word-final position.
In Module:bg-translit, Ъ is special-cased when it appears at the end of words, in order to account for the pre-1945 orthography. In this orthography, Ъ and Ь often appear at the end of words ending in a consonant, including feminine nouns (Ь) and masculine nouns (Ъ or Ь depending on the declension), indicating some lexical properties, but they are not pronounced. (In the modern orthography, these have just been removed, but in the Alternative forms section we often link to them from the modern entry.)
Because they are not pronounced, a special case has been added to make word-final Ъ transliterate to nothing, but the same special case has not been added to Ь, which leaves things inconsistent.
I think we should make whatever treatment we make complete, so, my proposals are these:
- Uniformly transliterate both Ъ and Ь, i.e. remove the special case. Ъ will always be Ă, and Ь will be J or ' if not paired with an О (just how it generally works up until now).
- Also special-case word-final Ь, so that it also transliterates to nothing, since it's not pronounced.
- Use different symbols for the old orthography's word-final Ъ and Ь: the Old Church Slavonic standard is ŭ and ĭ respectively, which would also reflect the fact that the symbols are not used for their usual vowel values, but as markers indicating some grammatical features of the word.
Of these, my favorite is probably 1 or 3, since they more accurately represent the fact that a letter is present in the original Bulgarian spelling, whereas the method of erasing it obscures that, although it does more clearly indicate the lack of pronunciation.
3 is perhaps a good compromise, because the alternative symbol does not suggest the same usual vowel pronunciation (of Ъ), but it does also show the present vowel.
Finally, we should also consider that in some dialectal spellings, e.g. those found in Bulgarian Etymological Dictionary, ъ may be found in spellings in a word-final position, sometimes even with stress, so it is actually pronounced as its usual vowel self. This could make the method of spelling it using ă more useful, since it would correctly handle these cases by default.
Just for now, I think I would support option 1. But, perhaps there are more sides that I haven't considered - so please let me know your thoughts. Kiril kovachev (talk・contribs) 18:56, 15 July 2026 (UTC)
- @Kiril kovachev Pre-1945 Bulgarian had lots of extra symbols, not just ъ and ь at the end of a word, so we should consider how to transliterate them as well. Under хиляда (hiljada) there's a quote of the text
падшїх бѡ множьство обрѣте сѧ исчитаемо до .ле. '''хилїадъ''' въсѣхъthat transliterates it aspadšïh bō množьstvo obrěte sę isčitaemo do [35] '''hilïadъ''' vъsěhъ. Maybe we could take a cue from this. I could imagine, for example, transliterating word-final ъ and ь as-is. Benwing2 (talk) 19:11, 15 July 2026 (UTC)- @Benwing2 That quotation is, I suppose, indeed pre-1945, but it's from Middle Bulgarian, so I figure that's more like an entirely different language :) I would say this example would be more apt under a Middle Bulgarian heading than Bulgarian, so the transliterations used there could be quite different - e.g. using the OCS values for yers would probably be very reasonable.
- But I do agree we should also figure out those Latin equivalents for the other letters, for Middle Bulgarian and for later archaic modern orthographies (pre-1899, e.g.).
- I hadn't thought of just using Ъ as-is. I do think it could work well, though. The disadvantage I see is that, in those dialect examples where Ъ is used as for its usual value word-finally, using that convention would make it look different from Ъ in its usual position and pronunciation, whereas it would actually be pronounced the same. Contrarily, using A with breve would suggest the final ъ is also pronounced the same, which is good for modern uses but bad for the old orthography. I would still find an explicit showing preferable to emptiness, though, since it shows the marked placement of the yer.
- And, like Chernorizets says, I find using yers may best be kept a last resort, since we should try picking from the Latin alphabet as much as possible in transliterations.
- Re: the other letters, I think those values in the module are satisfactory, and on top of that we could also add reasonable transliterations for other archaic Cyrillic letters, e.g. omega or iota. I can try have a look at that and talk about it in a separate topic, perhaps, but thanks for raising that as well. Kiril kovachev (talk・contribs) 21:55, 15 July 2026 (UTC)
- @Kiril kovachev there are a few things at play:
- there are actually several pre-1945 orthographies that were used in Bulgaria post-Independence, and they don't all have the same set of letters, but they do agree on word-final ъ/ь being silent.
- pre-Independence texts (15th c. - first half of 19th c.) use a variety of conventions, since the orthography was not standardized.
- I'm not sure whether transliteration is supposed to maintain a 1-to-1 mapping between the letters of the original alphabet and the letters of the target alphabet. If yes, then either option 1 or option 3 would work. I wouldn't transliterate those word-final letters using the actual Cyrillic ъ/ь, because 1) that's not transliteration in the real sense, and 2) it incorrectly overlaps with OCS transliteration where these letters could have fully-pronounced vowel qualities (strong yers).
- At any rate, the bg-translit module should allow overrides, in cases where the default isn't the correct thing to do. This should be rare, but it addresses e.g. dialectal spellings. Chernorizets (talk) 20:14, 15 July 2026 (UTC)
- @Chernorizets Okay, good to know the others are all the same in their treatment of yers at the end - I had read about that too a little bit, but good to be sure. For now, let us not worry about the pre-standardized texts, since we can always do our best job of those in the module and indeed do overrides ad hoc if we need to.
- In my view, it is preferable to have a 1-1 conversion between Bulgarian and Latinized spellings: it most-accurately represents the original text whilst being readable to those who don't know Cyrillic. A knowledge of the language is still needed to know precisely how to pronounce it (although the transliteration by itself will be good enough even if you don't know it well), so if there is a small discrepancy with yers at the end of words, I don't find that discrepancy to be a harm compared to the good of showing in the transliteration that the old orthography is different from the modern spelling in its ending.
- I'm also not sure what exactly the weighting between indicating pronunciation and reflecting spelling should be, but on this occasion I think I lean towards the second one.
- Thanks for your point of not using ъ/ь directly - I also think not using Cyrillic letters if we don't have to is ideal. Kiril kovachev (talk・contribs) 22:11, 15 July 2026 (UTC)
- @Kiril, thank you for bringing this up! Fixing this inconsistency is a great idea. It will make our work with old and dialectal Bulgarian words much cleaner.
Here are my thoughts on your three options: ★ On Option 2 (Make word-final Ь transliterate to nothing) I think hiding silent letters in transliteration is a bad idea. A transliteration should show exactly what is written in the original text. If we make these letters disappear, readers cannot tell what the original spelling was. It hides the reality of the pre-1945 writing system. ★ Option 1 (Transliterate both normally: Ъ \to ă, Ь \to j/ʹ) I strongly prefer this option. Here is why: Dialects: As you pointed out, some dialects actually pronounce a stressed ъ at the end of words. If we always hide the letter, we break the transliteration for these real words. Clarity: Keeping a simple 1-to-1 match helps users who do not read Cyrillic. They can clearly see the difference between the old spelling and the new one (like old градъ vs. modern град). ★ Option 3 (Use Old Church Slavonic ŭ and ĭ) This is a clever idea, but I worry it might make the module code too complicated. If we use ŭ and ĭ, we add a second set of rules. We would have to write code that tries to guess if an ending is "silent" (historical) or "pronounced" (dialect). That will probably cause bugs. SixtyShips (talk) 19:39, 15 July 2026 (UTC)
- @SixtyShips Good point about #3 needing to discriminate uses of dialect vs silent. I guess only #1 and #2 are viable, then. Glad we also agree on #1 :) Kiril kovachev (talk・contribs) 22:14, 15 July 2026 (UTC)
High-speed vandal Special:Contributions/~2026-36379-90
[edit]Please mass-revert. ~2026-39870-67 (talk) 20:15, 15 July 2026 (UTC)
- Thank you for the report. In the future, please post in WT:VIP for a more rapid response. TranqyPoo [💬 | ✏️] 21:32, 15 July 2026 (UTC)
Ceasing inclusion of Homophones in Appendix:Variations of X
[edit]Why bother with homophones in these appendices? The whole point of them is to redirect to other orthographies, homophones are a phonological/phonetic, language-specific domain pertaining which feel off in such appendices. I am proposing deleting all homophones from Appendix:Variations of X. SinaSabet28 (talk) 22:16, 15 July 2026 (UTC)
- Please excuse typos, idk how "pertaining" snuck in there. SinaSabet28 (talk) 22:17, 15 July 2026 (UTC)
- I think that seems reasonable, unless there is a lot of precedent for using those for homophones in the past, but I'm not aware of that much... Kiril kovachev (talk・contribs) 23:58, 15 July 2026 (UTC)
- It's not that the "whole point" of them is only to redirect to "other orthographies". As I mentioned in Wiktionary:Grease_pit/2026/June#Automation suggestion for 'also' forms, "Variations" appendices fill the same sort of idea as a "nearby/similar entries" navigation box, where "nearby/similar" can be conceived along a whole variety of axes (shape, shape in cursive form, sound, etc.), for navigating the giant intertwingled graph that is a wiki dictionary, since there is no natural or compelling "language-independent" way to define what is "similar". That said, the Variations appendices are quite an annoying and bloated to navigate mess and "homophones" sort of have less place there, being that the place already for them is in a specific language entry. Hftf (talk) 01:29, 16 July 2026 (UTC)
Reference template for Bopath IPD, a Pāli dictionary — feedback wanted
[edit]Hello. I've created {{R:pi:Bopath}}, a reference template linking to Bopath IPD, an online Pāli dictionary that I maintain — disclosing my conflict of interest up front. Unlike Pāli–English references, it gives definitions in nine target languages (English, French, German, Spanish, Italian, Dutch, Czech, Norwegian, Vietnamese). On e.g. dhamma it renders:
- dhamma, in Bopath IPD (International Pāli Dictionary)
with the headword linking to the matching entry (which serves the reader's language).
How it is produced — and its limits, stated plainly. Entries are built by an AI-assisted lexicographic pipeline, not a word-for-word machine translation: it works from the Pāli canon and existing scholarship to derive sense distinctions, etymology, morphological and root (dhātu) analysis, and canonical attestations. I want to be equally transparent about the trade-off: human review is currently limited. I'm raising this here precisely so the community — not me — can judge whether it is acceptable as a reference, and under what conditions.
A few more points:
- It is freely accessible and CC BY-NC-SA licensed. Many Pāli entries currently have no dictionary reference at all.
- Definitions in nine languages may be useful to editors and readers beyond English.
- If it's welcome, should I add the template broadly, or would you prefer I limit it / hold off?
- Any concerns with the template itself?
So far I've added it to a single entry (dhamma) as a test, and I won't add more until I hear back. Reverts and criticism very welcome. Jfromang (talk) 01:26, 16 July 2026 (UTC)
Hittite words with unattested lemma forms
[edit]Currently, the Hittite entry guidelines mandate that any term that is not attested in the ideal lemma form ought to be treated as a reconstruction (e.g. *𒀭𒋫𒊏𒀸 (/*antaraš/)). Ignoring the fact that the {{cuneiform|hit}} template doesn't work properly on reconstructed pages, this policy also seems problematic given we already have a separate template for this exact issue: {{unattested}}. I propose that we put Hittite words such as *𒀭𒋫𒊏𒀸 (/*antaraš/) in the mainspace and just use the {{unattested}} template. Graearms (talk) 18:03, 16 July 2026 (UTC)