日本語
Wakun Usage Frequency

Wakun Usage Frequency #

Ikeda Shoju

January 12–13, 2025, revised April 10, 2025 (English translation added)

Summary #

This page is an English-language summary of 和訓の使用頻度, a Japanese case study organizing the usage frequency of wakun (Japanese readings, 和訓) attested in the Kanchi-in manuscript of the Ruiju Myōgishō (KRM), based on krm_wakun.tsv as of December 11, 2024. The full frequency tables, tabulating every attested wakun form down to a single occurrence, are available only in Japanese; this page presents only the abstract.

Sixteen wakun forms occur 100 times or more, led by “miru” (みる, 191), “tsuku” (つく, 166), and “toru” (とる, 159). Related inflectional and adverbial forms cluster together: “akiraka” (あきらか, 116), “akiraka ni” (あきらかに, 16), and “akiraka nari” (あきらかなり, 14) total 146 occurrences between them, suggesting that variant forms of a single stem should be consolidated before frequency counts are compared across words. At the opposite extreme, 4,623 distinct wakun forms are attested only once, showing that the great majority of wakun in KRM are low-frequency and that high-frequency forms are a small, disproportionately useful subset for annotation work.

The frequency counts also expose an unresolved homograph problem: among the 166 occurrences counted under “tsuku” (つく), some are presumed to actually belong to the near-homographic “tsugu” (つぐ, 56 occurrences), but the two have not yet been systematically disambiguated in the underlying data.

The Japanese version concludes with three recommendations for future wakun annotation work: prioritize investigation of high-frequency forms first; consolidate variant forms ending in “nari” (なり) or “ni” (に) with their base stem before counting; and remain alert to homograph conflation of the kind seen with “tsuku”/"tsugu" throughout the investigation.