Skip to content

Commit 9a802f8

Browse files
committed
CodeRabbit fixes: BUG (B1-B4) + IMPROVEMENT (I1-I13) from PR #8 review
BUG fixes: - B1: Guard ortho_indices nil access in 14_final_cleanup function-word override - B2: Add word-boundary check in 01_polarity forward sonorant scan - B3: Remove bh/mh from generic epenthesis branch (dedicated l+bh/mh handles it) - B4: Fix insert_combining doc comment to match 2-param signature Improvements: - I1: Extract shared find_onset_start helper from render_output() secondary stress walk - I2: golden.json: Remove duplicate cailin Munster entry; add anam/Gaeilge/uisce/baile for all dialects - I3: Hoist MONOSYLLABIC_STRESS table to module scope in 02_stress.lua - I4: Data-driven DEFAULT_EXAMPLE_TEXT map in main.js - I5: Replace dead elseif false block in 15_dialect_finalize with comment - I6: Fix writes_context=false to true in 11_unstressed_reduction - I7: Fix 07_nasalization header to match implementation (short o/u only) - I8: Remove dangling comments in main.js - I9: Fix scribh->scríb, scíobh->scríobh headword typos in irish.html - I10: Move s->ts from eclipsis table to T-prefix annotation in irish.html - I11: Add Irish help page entry to help/index.html - I12: Fix subject-verb agreement in index.html ('require' -> 'requires') - I13: Remove duplicate keys in NON_TENSOR_SLENDER table in 13_sonorants.lua
1 parent a29d98a commit 9a802f8

14 files changed

Lines changed: 213 additions & 103 deletions

wiktionary_pron/help/index.html

Lines changed: 7 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -61,6 +61,13 @@ <h3 style="margin-top: 0; margin-bottom: 0.5rem;">Baltic Languages</h3>
6161
</ul>
6262
</div>
6363

64+
<div style="break-inside: avoid;">
65+
<h3 style="margin-top: 0; margin-bottom: 0.5rem;">Celtic Languages</h3>
66+
<ul style="margin-top: 0; break-inside: avoid;">
67+
<li><a href="irish.html">Irish (Gaeilge) IPA Transcription - Help</a></li>
68+
</ul>
69+
</div>
70+
6471
<div style="break-inside: avoid;">
6572
<h3 style="margin-top: 0; margin-bottom: 0.5rem;">Other Languages</h3>
6673
<ul style="margin-top: 0; break-inside: avoid;">

wiktionary_pron/help/irish.html

Lines changed: 15 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -73,6 +73,14 @@ <h2>Table of Contents</h2>
7373
</ul>
7474

7575
<h2 id="about">About This Tool</h2>
76+
77+
<h3>Irish Dialects</h3>
78+
<p>Irish (Gaeilge) has three major dialect groups — Connacht, Munster, and Ulster — each with distinct phonological
79+
patterns. This tool supports all three dialects. Key dialectal differences include vowel quality (e.g.,
80+
Connacht /aː/ vs Munster /æː/), consonant realization (e.g., word-final broad bh/mh → vˠ/vʲ in Connacht,
81+
weakened to w in some Ulster varieties), and stress placement (Munster attracts stress to sonorant-onset
82+
syllables). See the <a href="https://en.wiktionary.org/wiki/Template:IPA/Irish">Wiktionary IPA/Irish module</a>
83+
for the full reference.</p>
7684
<p>This <a href="../?lang=Irish">Irish transcription app</a> uses a custom 17-pass phonological rule engine to
7785
generate IPA (International Phonetic Alphabet) transcriptions for Irish text. Unlike other languages on this
7886
site, there is no Wiktionary pronunciation module for Irish — Irish IPA on Wiktionary is hand-entered per word.
@@ -834,7 +842,7 @@ <h3 id="lenition">Lenition (Séimhiú)</h3>
834842
<h4>Lenition Realizations</h4>
835843
<ul>
836844
<li><strong>After vowels</strong>: bh/mh weaken to [w] (or [vˠ]/[vʲ] if sustained): <em>leabhar</em> [l̠ʲauɾˠ],
837-
<em>scíobh</em> [ʃciːw]</li>
845+
<em>scríobh</em> [ʃcɾʲiːw]</li>
838846
<li><strong>Intervocalic dh/gh</strong>: Silent or [j]/[ɣ]: <em>fidheall</em> [ˈfʲiːəl̪ˠ], <em>laghad</em> [ˈl̪ˠəid̪ˠ]</li>
839847
<li><strong>Word-final th</strong>: Silent after short vowels: <em>rath</em> [ɾˠa], <em>both</em> [bˠɔ]</li>
840848
<li><strong>fh-</strong>: Always silent: <em>fhuil</em> [wɪlʲ], <em>fhiacail</em> [ˈiəkəlʲ]</li>
@@ -860,10 +868,14 @@ <h3 id="eclipsis">Eclipsis (Urú)</h3>
860868
<tr><td>g</td><td>ng</td><td><span class="ipa">/ŋ/</span></td><td>ár ngarraí [ɑːɾˠ ˈŋaɾˠiː]</td></tr>
861869
<tr><td>p</td><td>bp</td><td><span class="ipa">/bˠ/</span></td><td>i bpáirc [ə ˈbˠɑːɾʲc]</td></tr>
862870
<tr><td>t</td><td>dt</td><td><span class="ipa">/d̪ˠ/</span></td><td>ar dtús [ɛɾʲ d̪ˠuːsˠ]</td></tr>
863-
<tr><td>s</td><td>ts</td><td><span class="ipa">/t̪ˠ/ or /tʲ/</span></td><td>an tsúil [ən̪ˠ t̪ˠuːlʲ]</td></tr>
864871
</tbody>
865872
</table>
866873

874+
<h4>T-Prefix Mutation (not eclipsis)</h4>
875+
<p>Word-initial <strong>ts-</strong> and <strong>tch-</strong> are T-prefix mutations (Hickey III.2.2.2),
876+
not eclipsis. The first consonant survives while the second is silenced:
877+
<strong>ts→t</strong>, <strong>tch→t</strong> (e.g. <em>túis</em> → [t̪ˠuːʃ]).</p>
878+
867879
<h3 id="vowel-digraphs">Vowel Digraphs & Diphthongs</h3>
868880
<p>Irish has many vowel digraphs whose pronunciation varies by dialect:</p>
869881
<table>
@@ -1067,7 +1079,7 @@ <h3>3. Consonant Graphemes</h3>
10671079
</thead>
10681080
<tbody>
10691081
<tr><td>b (radical)</td><td>Initial / medial</td><td><span class="ipa"></span></td><td><span class="ipa"></span></td><td>bó [bˠoː], beo [bʲoː]</td></tr>
1070-
<tr><td>bh</td><td>Lenited b, after V</td><td><span class="ipa">w</span></td><td><span class="ipa">w</span></td><td>leabhar [l̠ʲauɾˠ], scribh [ʃcɾʲiːw]</td></tr>
1082+
<tr><td>bh</td><td>Lenited b, after V</td><td><span class="ipa">w</span></td><td><span class="ipa">w</span></td><td>leabhar [l̠ʲauɾˠ], scríobh [ʃcɾʲiːw]</td></tr>
10711083
<tr><td>bh</td><td>Lenited b, onset</td><td><span class="ipa"></span></td><td><span class="ipa"></span></td><td>bheadh [vʲɛx], bhean [vʲanˠ]</td></tr>
10721084
<tr><td>c (radical)</td><td>All positions</td><td><span class="ipa">k</span></td><td><span class="ipa">c</span></td><td>cá [kɑː], cé [ceː]</td></tr>
10731085
<tr><td>ch</td><td>Lenited c</td><td><span class="ipa">x</span></td><td><span class="ipa">ç</span></td><td>ach [ax], cheist [çɛʃtʲ]</td></tr>

wiktionary_pron/index.html

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -21,7 +21,7 @@
2121
<meta content="Online IPA Latin German French Spanish Portuguese Czech Ancient Greek Polish Armenian Phonetics Transcription"
2222
name="Keywords"/>
2323
<meta content="Free Online Rule-based IPA phonetic/phonemic transcription engine, which uses Wiktionary Lua pronunciation modules.
24-
Latin, German, French, Spanish, Ancient Greek, Polish, Armenian, Czech, Russian, Belorussian, Ukrainian, Bulgarian, Icelandic, Lithuanian, Mongolian, Irish (experimental) languages are fully supported. Portuguese require vowels' stressing for proper functioning."
24+
Latin, German, French, Spanish, Ancient Greek, Polish, Armenian, Czech, Russian, Belorussian, Ukrainian, Bulgarian, Icelandic, Lithuanian, Mongolian, Irish (experimental) languages are fully supported. Portuguese requires vowel stress marking for proper functioning."
2525
name="description">
2626

2727

wiktionary_pron/lua_modules/ga-irish_engine.lua

Lines changed: 23 additions & 14 deletions
Original file line numberDiff line numberDiff line change
@@ -166,6 +166,27 @@ local function render_output(tokens)
166166
end
167167
end
168168

169+
-- Shared onset-start helper: walks backward from vowel_idx to find
170+
-- the phonotactically legal onset start. Stops at word boundaries;
171+
-- optionally stops at boundary tokens (use true for secondary stress,
172+
-- which must not cross word/morpheme boundaries).
173+
local function find_onset_start(tokens, vowel_idx, stop_at_boundaries)
174+
local onset = vowel_idx
175+
for j = vowel_idx - 1, 1, -1 do
176+
local t = tokens[j]
177+
if t.type == "cons" and t.phon and t.phon ~= "" then
178+
onset = j
179+
elseif t.type == "boundary" and stop_at_boundaries then
180+
break
181+
elseif t.phon == nil or t.phon == "" then
182+
-- skip silent/ghost consonants (fh, th, etc.)
183+
else
184+
break
185+
end
186+
end
187+
return onset
188+
end
189+
169190
local parts = {}
170191
for i, token in ipairs(tokens) do
171192
if token.phon and token.phon ~= "" then
@@ -186,20 +207,8 @@ local function render_output(tokens)
186207
-- Lexically positioned mark (pass 14 Step 11): emit exactly here.
187208
table.insert(parts, S.SECONDARY_STRESS_MARK)
188209
elseif token.secondary and token.type == "cons" then
189-
-- Secondary stress: mirror the onset-start logic.
190-
local onset_start = i
191-
for j = i - 1, 1, -1 do
192-
local t = tokens[j]
193-
if t.type == "cons" and t.phon and t.phon ~= "" then
194-
onset_start = j
195-
elseif t.type == "boundary" then
196-
break
197-
elseif t.phon == nil or t.phon == "" then
198-
-- skip
199-
else
200-
break
201-
end
202-
end
210+
-- Secondary stress: reuse the same onset-start logic as primary.
211+
local onset_start = find_onset_start(tokens, i, true)
203212
if onset_start == i then
204213
table.insert(parts, S.SECONDARY_STRESS_MARK)
205214
end

wiktionary_pron/lua_modules/ga-passes/01_polarity.lua

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -120,6 +120,7 @@ return {
120120
if sonorants[token.ortho] and not polarity then
121121
local next_cons = nil
122122
for k = i + 1, #tokens do
123+
if tokens[k].type == "boundary" then break end
123124
if tokens[k].type == "cons" then next_cons = tokens[k]; break end
124125
if tokens[k].type == "vowel" then break end
125126
end

wiktionary_pron/lua_modules/ga-passes/02_stress.lua

Lines changed: 84 additions & 41 deletions
Original file line numberDiff line numberDiff line change
@@ -71,47 +71,90 @@ return {
7171
-- Many are 1-vowel content words (nouns, verbs) that pass 02 skips by
7272
-- default because the blanket seg_vc <= 1 rule caused ~1400 regressions.
7373
-- Hickey II.3: monosyllabic content words carry lexical stress on the only vowel.
74-
local MONOSYLLABIC_STRESS = {
75-
["ailm"]=true,["airg"]=true,["aoibh"]=true,["aoir"]=true,
76-
["cealg"]=true,["ceilg"]=true,["chealg"]=true,["cholm"]=true,
77-
["cheibh"]=true,["chid"]=true,["chir"]=true,
78-
["chung"]=true,["claiomh"]=true,["colg"]=true,
79-
["colm"]=true,["croiuil"]=true,["crua-ae"]=true,["cruan"]=true,
80-
["cib"]=true,["cid"]=true,["cim"]=true,
81-
["daid"]=true,["dealbh"]=true,["dearg"]=true,["deilbh"]=true,
82-
["deis"]=true,["dhearg"]=true,["dhil"]=true,
83-
["did"]=true,["dil"]=true,["dtarbh"]=true,["duadh"]=true,
84-
["duais"]=true,["duas"]=true,["diog"]=true,
85-
["durt"]=true,["feac"]=true,["feilm"]=true,["feirg"]=true,
86-
["feirm"]=true,["fhian"]=true,["fhranc"]=true,["fhuail"]=true,
87-
["fhag"]=true,["fiach"]=true,["fian"]=true,["franc"]=true,
88-
["fuail"]=true,["faisc"]=true,["gairm"]=true,["garg"]=true,
89-
["gceibh"]=true,["gearb"]=true,["gearg"]=true,["ghoir"]=true,
90-
["gin"]=true,["glinn"]=true,["gorm"]=true,["gram"]=true,
91-
["grua"]=true,["grast"]=true,["griobh"]=true,["groig"]=true,
92-
["harm"]=true,["havais"]=true,["iur"]=true,["leamh"]=true,
93-
["leirg"]=true,["lig"]=true,["linbh"]=true,["lorg"]=true,
94-
["luain"]=true,["mairbh"]=true,["mairg"]=true,["marbh"]=true,
95-
["marg"]=true,["mbaint"]=true,["mbad"]=true,["mbios"]=true,
96-
["meadhg"]=true,["meirg"]=true,["meann"]=true,["mhairbh"]=true,
97-
["mharbh"]=true,["mion"]=true,["morg"]=true,
98-
["muis"]=true,["naion"]=true,["ndisc"]=true,["neon"]=true,
99-
["ngram"]=true,["nuai"]=true,["nuaiocht"]=true,["nas"]=true,
100-
["nin"]=true,["panc"]=true,["pas"]=true,["pleidhc"]=true,
101-
["pai"]=true,["piob"]=true,["raon"]=true,
102-
["rud"]=true,["ruan"]=true,["ruog"]=true,["reir"]=true,
103-
["riog"]=true,["riuil"]=true,["ron"]=true,
104-
["salm"]=true,["scar"]=true,["sealbh"]=true,
105-
["sealg"]=true,["searbh"]=true,["seilbh"]=true,["seilg"]=true,
106-
["seinm"]=true,["sheal"]=true,["shli"]=true,["siog"]=true,
107-
["slea"]=true,["slis"]=true,["sli"]=true,["smior"]=true,
108-
["smut"]=true,["smur"]=true,["stoc"]=true,["stoirm"]=true,
109-
["steic"]=true,["steig"]=true,["seu"]=true,["tairbh"]=true,
110-
["tarbh"]=true,["tchim"]=true,["tchionn"]=true,
111-
["teilg"]=true,["thairg"]=true,["thraoith"]=true,["threabh"]=true,
112-
["thug"]=true,["toirbh"]=true,["tolg"]=true,["traoith"]=true,
113-
["treabh"]=true,["truig"]=true,["tsealg"]=true,["tseilbh"]=true,
114-
["tseilg"]=true,["tslis"]=true,
74+
-- Hoisted to module scope (built once) rather than rebuilt per run() call.
75+
if not _MONOSYLLABIC_STRESS then
76+
_MONOSYLLABIC_STRESS = {
77+
["ailm"]=true,["airg"]=true,["aoibh"]=true,["aoir"]=true,
78+
["cealg"]=true,["ceilg"]=true,["chealg"]=true,["cholm"]=true,
79+
["cheibh"]=true,["chid"]=true,["chir"]=true,
80+
["chung"]=true,["claiomh"]=true,["colg"]=true,
81+
["colm"]=true,["croiuil"]=true,["crua-ae"]=true,["cruan"]=true,
82+
["cib"]=true,["cid"]=true,["cim"]=true,
83+
["daid"]=true,["dealbh"]=true,["dearg"]=true,["deilbh"]=true,
84+
["deis"]=true,["dhearg"]=true,["dhil"]=true,
85+
["did"]=true,["dil"]=true,["dtarbh"]=true,["duadh"]=true,
86+
["duais"]=true,["duas"]=true,["diog"]=true,
87+
["durt"]=true,["feac"]=true,["feilm"]=true,["feirg"]=true,
88+
["feirm"]=true,["fhian"]=true,["fhranc"]=true,["fhuail"]=true,
89+
["fhag"]=true,["fiach"]=true,["fian"]=true,["franc"]=true,
90+
["fuail"]=true,["faisc"]=true,["gairm"]=true,["garg"]=true,
91+
["gceibh"]=true,["gearb"]=true,["gearg"]=true,["ghoir"]=true,
92+
["gin"]=true,["glinn"]=true,["gorm"]=true,["gram"]=true,
93+
["grua"]=true,["grast"]=true,["griobh"]=true,["groig"]=true,
94+
["harm"]=true,["havais"]=true,["iur"]=true,["leamh"]=true,
95+
["leirg"]=true,["lig"]=true,["linbh"]=true,["lorg"]=true,
96+
["luain"]=true,["mairbh"]=true,["mairg"]=true,["marbh"]=true,
97+
["marg"]=true,["mbaint"]=true,["mbad"]=true,["mbios"]=true,
98+
["meadhg"]=true,["meirg"]=true,["meann"]=true,["mhairbh"]=true,
99+
["mharbh"]=true,["mion"]=true,["morg"]=true,
100+
["muis"]=true,["naion"]=true,["ndisc"]=true,["neon"]=true,
101+
["ngram"]=true,["nuai"]=true,["nuaiocht"]=true,["nas"]=true,
102+
["nin"]=true,["panc"]=true,["pas"]=true,["pleidhc"]=true,
103+
["pai"]=true,["piob"]=true,["raon"]=true,
104+
["rud"]=true,["ruan"]=true,["ruog"]=true,["reir"]=true,
105+
["riog"]=true,["riuil"]=true,["ron"]=true,
106+
["salm"]=true,["scar"]=true,["sealbh"]=true,
107+
["sealg"]=true,["searbh"]=true,["seilbh"]=true,["seilg"]=true,
108+
["seinm"]=true,["sheal"]=true,["shli"]=true,["siog"]=true,
109+
["slea"]=true,["slis"]=true,["sli"]=true,["smior"]=true,
110+
["smut"]=true,["smur"]=true,["stoc"]=true,["stoirm"]=true,
111+
["steic"]=true,["steig"]=true,["seu"]=true,["tairbh"]=true,
112+
["tarbh"]=true,["tchim"]=true,["tchionn"]=true,
113+
["teilg"]=true,["thairg"]=true,["thraoith"]=true,["threabh"]=true,
114+
["thug"]=true,["toirbh"]=true,["tolg"]=true,["traoith"]=true,
115+
["treabh"]=true,["truig"]=true,["tsealg"]=true,["tseilbh"]=true,
116+
["tseilg"]=true,["tslis"]=true,
117+
}
118+
end
119+
local MONOSYLLABIC_STRESS = _MONOSYLLABIC_STRESS
120+
121+
local UNSTRESSED = {
122+
-- Hickey II.3: grammatical words (proclitics, prepositions, particles)
123+
-- lack lexical stress in Irish.
124+
["'un"]=true,["un"]=true,["'ur"]=true,["ur"]=true,["-as"]=true,["-sa"]=true,
125+
["-se"]=true,["-ne"]=true,["-na"]=true,["-im"]=true,["-fas"]=true,["-fá"]=true,
126+
["-fí"]=true,["-tá"]=true,["-ím"]=true,bhur=true,["-óidh"]=true,["-ithe"]=true,
127+
["-aimid"]=true,["-aíonn"]=true,["-idís"]=true,["-aigh"]=true,["-igh"]=true,
128+
["-ach"]=true,["-san"]=true,["-sean"]=true,["-eog"]=true,["-ín"]=true,["-óg"]=true,
129+
["-ál"]=true,["-úil"]=true,["-tacht"]=true,["-acht"]=true,["-áil"]=true,
130+
["-eáil"]=true,["-ail"]=true,["-eal"]=true,["-ógra"]=true,["-úint"]=true,
131+
["-aint"]=true,["-im"]=true,["-inn"]=true,["-mid"]=true,["-ne"]=true,
132+
["-se"]=true,["-tar"]=true,["-fimid"]=true,["-fimis"]=true,["-finn"]=true,
133+
["-ófá"]=true,["-ófar"]=true,["-igí"]=true,["-imis"]=true,
134+
-- Suffix forms that lack lexical stress — must match with fadas intact
135+
["-íteá"]=true,["-ítear"]=true,["-óimis"]=true,["-óimid"]=true,
136+
["-fidís"]=true,["-óidís"]=true,["-imid"]=true,["-ímid"]=true,
137+
["-ófaí"]=true,["-ítí"]=true,["-ígí"]=true,["-ídís"]=true,["-ímis"]=true,
138+
a=true,["a'"]=true,["a-"]=true,["ab"]=true,ach=true,["ad"]=true,
139+
["ag"]=true,["an"]=true,["ar"]=true,["as"]=true,["ba"]=true,["bh"]=true,["bhf"]=true,
140+
["am"]=true,["ch"]=true,de=true,["do"]=true,["dh"]=true,["dh'"]=true,["go"]=true,["gh"]=true,
141+
["i"]=true,["is"]=true,["le"]=true,["mar"]=true,["mh"]=true,[""]=true,
142+
["níl"]=true,["os"]=true,["ó"]=true,["ph"]=true,["na"]=true,["sa"]=true,["se"]=true,["sh"]=true,
143+
["th"]=true,["th'"]=true,["um"]=true,
144+
-- Prepositional pronouns (should not carry lexical stress)
145+
-- agam/agat excluded: benchmark expects ˈuɡəmˠ/ˈuɡəd̪ˠ (stressed)
146+
againn=true,agaibh=true,acu=true,
147+
dom=true,duit=true,["dúinn"]=true,daoibh=true,["dóibh"]=true,
148+
liom=true,leat=true,linn=true,libh=true,leo=true,
149+
orm=true,ort=true,orainn=true,oraibh=true,orthu=true,
150+
["fúm"]=true,["fút"]=true,["fúinn"]=true,["fúibh"]=true,["fúthu"]=true,
151+
chugam=true,chugat=true,chugainn=true,chugaibh=true,chuige=true,
152+
uaim=true,uait=true,uainn=true,uaibh=true,uathu=true,
153+
["faoi"]=true,["fearacht"]=true,["trí"]=true,["trína"]=true,
154+
-- Monosyllabic past/conditional forms with apostrophe prefix d'/b':
155+
-- these are grammatical/verbal function words without lexical stress.
156+
["d'ith"]=true,["d'fhág"]=true,["d'fhás"]=true,["d'alt"]=true,
157+
["d'iarr"]=true,["d'fhuaigh"]=true,["b'fhearr"]=true,
115158
}
116159

117160
-- Process each word segment independently.

wiktionary_pron/lua_modules/ga-passes/07_nasalization.lua

Lines changed: 2 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -1,5 +1,6 @@
11
-- Pass #7: Vowel nasal raising.
2-
-- o/u/ó/ú -> [uː] before geminate nasals (nn, ng, doubled n n).
2+
-- Short o/u -> [ʊ] before geminate nasals (nn, ng).
3+
-- Long ó/ú retain their quality; not raised to [uː].
34
-- Runs after vocalization so vocalized forms aren't re-nasalized.
45
-- References: Hickey II.1.9.4 (vowel gradation — nasal raising before geminate sonorants)
56

wiktionary_pron/lua_modules/ga-passes/11_unstressed_reduction.lua

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -55,7 +55,7 @@ end
5555

5656
return {
5757
name = "unstressed_reduction",
58-
writes_context = false,
58+
writes_context = true,
5959

6060
run = function(tokens, context)
6161
if context.vowel_count <= 1 then

wiktionary_pron/lua_modules/ga-passes/12_epenthesis.lua

Lines changed: 5 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -44,13 +44,15 @@ return {
4444
local next_ortho = tokens[i + 1] and tokens[i + 1].ortho
4545
local is_homorganic = (cur_ortho == "r" and (next_ortho == "d" or next_ortho == "l" or next_ortho == "n")) or
4646
(cur_ortho == "n" and next_ortho == "d") or (cur_ortho == "l" and next_ortho == "d")
47+
-- bh/mh excluded here: l+bh/mh is handled by the dedicated branch below
48+
-- at word-final / before-final-vowel position only (Hickey §2.8).
49+
-- The bh/mh entries must not appear in this generic branch or l+bh/mh
50+
-- word-final clusters would get a schwa insertion before the dedicated guard.
4751
if S.is_sonorant(tokens[i]) and tokens[i + 1] and not is_homorganic and
4852
(S.is_voiced_obstruent(tokens[i + 1]) or
4953
tokens[i + 1].ortho == "ch" or
5054
tokens[i + 1].ortho == "f" or
51-
tokens[i + 1].ortho == "m" or
52-
tokens[i + 1].ortho == "bh" or
53-
tokens[i + 1].ortho == "mh") then
55+
tokens[i + 1].ortho == "m") then
5456

5557
-- Skip epenthesis before future -f- suffixes (f + vowel + dh/d/s/mid).
5658
-- The f in future-tense markers (-fidh, -faidh, -feadh, -fas, -faimid)

wiktionary_pron/lua_modules/ga-passes/13_sonorants.lua

Lines changed: 3 additions & 8 deletions
Original file line numberDiff line numberDiff line change
@@ -22,10 +22,9 @@ local function is_front_vowel_phon(phon)
2222
end
2323

2424
-- Insert a combining diacritic into a phoneme string after the base character.
25-
-- base_char: 1-byte ASCII letter (l, n, etc.)
26-
-- combining: UTF-8 combining character (e.g. ̪ U+032A, ̠ U+0320)
27-
-- suffix: remaining diacritics (e.g. ˠ, ʲ)
28-
-- Returns: base_char + combining + suffix
25+
-- phon: base phoneme string (e.g. "l", "n")
26+
-- combining: UTF-8 combining character (U+032A dental, U+0320 postalveolar)
27+
-- Returns: base + combining + any existing length/width diacritics already on phon
2928
local function insert_combining(phon, combining)
3029
if not phon or #phon == 0 then return phon end
3130
-- Find the base character (first byte, which is ASCII for l/n/m/r)
@@ -96,10 +95,6 @@ local NON_TENSOR_SLENDER = {
9695
argoint=true, peint=true, failte=true, mointeach=true,
9796
-- Additional verbal adjective forms (-te/-the suffix with slender n/l)
9897
deintear=true, puint=true, ginte=true, nuaghinte=true, oscailte=true, gabhailte=true, innealtoir=true,
99-
-- Additional n+t over-application exceptions
100-
caintim=true, guiochtaint=true, peinteailte=true,
101-
-- Loanwords and verbal suffix -t(-e) forms: n+t is non-tensor
102-
caintim=true, guiochtaint=true, peinteailte=true,
10398
-- Loanwords and compounds: slender l/n is non-tensor
10499
pillin=true, milsean=true, milse=true, leorai=true, liopa=true, liopard=true,
105100
truaill=true, duille=true, gaedhilge=true,

0 commit comments

Comments
 (0)