Archive/Cross-Family Stem-Suffix Boundary Segmentation in Turkic and Uralic Languages: A ByT5 Study of Azerbaijani, Turkish, Hungarian, and Finnish
Cross-Family Stem-Suffix Boundary Segmentation in Turkic and Uralic Languages: A ByT5 Study of Azerbaijani, Turkish, Hungarian, and Finnish
Kamran Ibiyev
31 de julio de 2026
en

Abstract

Turkic and Uralic languages share almost everything except ancestry: both are agglutinative, both stack long suffix chains onto a stem, both use vowel harmony, and both put the verb last. That coincidence is what makes the pairing worth testing. If a neural model learns where to place a stem–suffix boundary in one family, does the skill carry over to the other, given that nothing but structure is shared? We fine-tuned ByT5-small, a byte-level encoder–decoder, on segmentation pairs built from UniMorph 4.0 and Universal Dependencies for Azerbaijani, Turkish, Hungarian, and Finnish. To our knowledge, no earlier segmentation benchmark covers these four languages together, and the SIGMORPHON 2022 shared task included only Hungarian among them. These targets contain exactly one boundary at most, a stem plus one undivided suffix block, so the supervised task is single-boundary segmentation, not full morpheme decomposition. On this task, in-domain exact match is high across a 4×4 train-on-one, test-on-another matrix run over three seeds, 0.91 to 0.98, against a copy-input baseline of 0.02 to 0.20. Cross-family transfer is not: Finnish-to-Azerbaijani reaches 0.640 against 0.208 the other way (p<10−66), and an exploratory reading of the twelve pairs suggests surface typology tracks this gap more closely than genealogy does, though one high-overlap pair dominates and the sample is too small to settle the question. The central finding comes from a multi-boundary stress test: on 800 words we hand-annotated from Wikipedia (model-correctness agreement was perfect, Cohen’s κ=1.00), the models stall at two morphemes, and the 37% of words with three or more come out almost entirely wrong (one correct prediction in 897). Boundary-level scoring shows the shape of the failure: precision stays high, up to 0.97, while recall drops to 0.39–0.52, so the boundaries the models place are mostly right, they just place far too few. This ceiling is consistent with the two-part training targets rather than with the byte-level architecture, although only one architecture under one supervision regime was tested, so confirming the attribution will take retraining on fully decomposed targets. We release all code, data, and annotations to support that step.

IPC Classification

G06

Keywords

cross-familystem-suffixboundarysegmentationturkicuraliclanguagesbyt5azerbaijaniturkishhungarianfinnishinformationsharealmosteverythingexceptancestrybothagglutinativestacklongsuffixchains
Citar esta publicación

€ 4.00