日本語
KRM Documentation

KRM Documentation #

Scholarly and Technical Documentation for the KRM Database

About This Documentation #

KRM Documentation is the official scholarly and technical reference documentation for the KRM Database — the full-text database of the Kanchi-in manuscript of the Ruiju Myōgishō. It documents the source manuscript, the published data files, the entry data model, data entry conventions, annotation policy, the typesetting environment, project progress, and related case studies.

This documentation is intended for readers working with the KRM data directly — researchers in Japanese historical linguistics, lexicography, and digital humanities — as well as readers who want to understand how the KRM Database was built and is maintained.

  • Author: Shōju Ikeda, Professor Emeritus, Hokkaido University
  • Documentation Version: 0.9 (draft; Version 1.0 planned on completion of this reorganization)
  • Publication Date: draft — to be finalized at Version 1.0
  • Last Updated: draft — to be finalized at Version 1.0
  • Project Website: https://shikeda.github.io/
  • Documentation Website: https://shikeda.github.io/docs/krm/
  • Documentation License: CC BY-SA 4.0
  • Suggested Citation: (draft — finalized at Version 1.0) Ikeda, Shōju. KRM Documentation. Version 0.9. https://shikeda.github.io/docs/krm/.
  • Relationship: KRM Documentation documents the KRM Database, which is distributed through GitHub and Zenodo (see Resource Documented below).

Note that while the explanation in this documentation overlaps in part with what is stated in the paper by Shōju Ikeda, Liu Guanwei, Jung Munho, Zhang Xinfang, and Li Yuan, “Full-text Database of Ruiju Myōgishō, Kanchi-in MS : A Look at Development Methods and Calculating the Number of Headwords." (Kuntengo to Kuten Shiryō 144, 2020), it has been completely overhauled and rewritten by the first author, Ikeda, who organized the terminology and substantially added subsequent research findings.

Resource Documented #

This documentation describes KRM: Database of the Kanchi-in Manuscript of the Ruiju Myōgishō (KRM) — a full-text digitization of the Kanchi-in manuscript, together with location data, textual collation, and source studies. KRM is not only a dataset: its repository also includes processing scripts and a local search web application.

Repository Contents:

PathContents
krm_*.tsv, krm_*.jsonPublished KRM data files (headwords, definitions, notes, readings, etc.)
scripts/Python utility scripts for data conversion and maintenance (MIT License)
webapp/A local full-text search application (Next.js; MIT License)
docs/Repository documentation (distinct from this Documentation site)
examples/, images/, diff/Supporting examples, images, and change-tracking material

Relationship to the HDIC Project #

KRM is one of the Hanzi dictionary databases that make up the Integrated Database of Hanzi Dictionaries in Early Japan (HDIC). For an introduction to the HDIC Project as a whole — its background, its other constituent databases, and the HDIC Viewer search tool — see the HDIC Project home page.

About the Ruiju Myōgishō #

The KRM Database is based on the Kanchi-in manuscript of the Ruiju Myōgishō (類聚名義抄), a twelfth-century Sino-Japanese character dictionary compiled by a Shingon Buddhist monk. The Kanchi-in manuscript is the only complete extant witness to the work’s revised-compilation lineage, and is a valuable resource for research in the history of the Japanese lexicon, the historical phonology of Sino-Japanese character readings, and the history of Chinese character forms as used in Japan.

For a full description see Chapter 1: Overview of the Ruiju Myōgishō.

How This Documentation Is Organized #

Chapters 1–5 make up the Core Documentation — the primary reference for the KRM data model, input conventions, and annotation policy. Chapters 6–9 provide supporting material: typesetting workflow, project records, and case studies.

This documentation is organized into the following chapters:

  1. Overview of the Ruiju Myōgishō — the source manuscript: its textual traditions, compiler, date, significance, and structure.
  2. Overview of Published Data — the KRM data files and their structure.
  3. Entry Data Model — the conceptual model behind KRM entries.
  4. Input of Entry Data — headwords, IDs, character encoding, and transcription conventions.
  5. Basic Policy for Annotation Creation — annotation policy and methodology, with worked examples.
  6. Typesetting Configuration — the typesetting environment for transcriptions and annotations.
  7. Project Progress — development records (Japanese only).
  8. Case Studies — applied research examples (Japanese only).
  9. Development History — how the KRM Database was constructed.

Use the navigation sidebar to move between chapters and pages.

Acknowledgements #

The construction and publication of the full-text database of the Ruiju Myōgishō of the Kanchi-in manuscript are being carried out with special permission from the authorities of Tenri Library, and we have also received exceptional consideration from Yagi Shoten, the publisher of the Tenri Library Rare Books Series. We hereby express our gratitude for this.

This work was supported by JSPS KAKENHI Grant Numbers 25370506, 16H03422, 19H00526, 23K17500, 25K00466 and 26K21717.