Wiktionary Schema - Chinese Edition
The Chinese edition is implemented in both Rust and Go.
Rust
- Entry point:
zh::WordData— a single word entry from the Chinese Wiktionary data dump. - API responses wrap the word data in
zh::WrappedWordData, which adds asourceobject with provenance information (urlandlicense).
Go
- Entry point:
zh.WordDatainzh/schema.go. - API responses wrap the word data in
WrappedWordData[zh.WordData](packagewrapper), which adds the samesourceobject. - Uses Go’s
json/v2package; compile withGOEXPERIMENT=jsonv2.
Data Model
The schema mirrors the Pydantic models used by wiktextract for the Chinese edition (the WordEntry model and its dependencies in wiktextract/extractor/zh/models.py), kept in references/zh.json in the repository. The top-level type WordData represents a single word entry: each JSONL line of the zh-edition wiktextract dump deserializes into one instance.
All fields are optional (Option<T> in Rust, pointers in Go) except the identification fields (word, lang, lang_code) and pos.
Note: Hard-redirect pages in the raw dump are emitted as plain
{"title": ..., "redirect": ..., "pos": "hard-redirect"}objects that do not conform to this schema; consumers streaming the raw dump should skip lines without awordfield.
Main Types
| Type | Description |
|---|---|
WordData | Top-level word entry: word, language, part of speech, and all lexical data. |
Sense | A single sense (numbered definition) with glosses, examples, classifiers, and lexical relations. |
Classifier | A classifier (measure word) used with a noun sense, e.g. 个 in 一个苹果. |
Example | Example sentence with Chinese translation, romanization, ruby, and bold-span offsets. |
Form | Inflected / alternative surface form (may include a Hiragana reading for Japanese entries). |
Sound | Pronunciation data: zh_pron (Pinyin, Bopomofo, Jyutping, etc.), IPA, audio URLs, Hangeul, homophones. |
Translation | Translation of a sense into another language. |
Linkage | A lexical relation (synonym, antonym, hypernym, hyponym, compound, etc.). |
Descendant | Descendant word in another language, forming a recursive etymological tree. |
AttestationData / ReferenceData | Historical attestations with dates and bibliographic references. |
Hyphenation | Syllable-level hyphenation. |
AltForm | Alternative / base form (e.g. simplified/traditional variants, inflected forms). |
WordData Fields
| Category | Fields |
|---|---|
| Identification | word, lang, lang_code, pos, pos_title, original_title, title |
| Senses | senses |
| Forms & variants | forms |
| Pronunciation | sounds |
| Etymology | etymology_texts, etymology_examples, descendants |
| Lexical relations | synonyms, antonyms, hypernyms, hyponyms, holonyms, meronyms, troponyms, coordinate_terms, derived, related, paronyms, abbreviations, anagrams, proverbs, compounds, various |
| Classifiers | classifiers |
| Translations | translations |
| Hyphenation | hyphenations |
| Links & metadata | categories, notes, redirect, redirects, literal_meaning, raw_tags, tags |
Example
A minimal wrapped entry as returned by the API:
{
"word": "猫",
"lang": "汉语",
"lang_code": "zh",
"pos": "noun",
"pos_title": "名詞",
"senses": [
{
"glosses": ["一種貓科動物,俗稱貓咪。"]
}
],
"sounds": [
{
"zh_pron": "māo"
}
],
"source": {
"url": "https://zh.wiktionary.org/wiki/猫",
"license": {
"name": "CC BY-SA 4.0",
"url": "https://creativecommons.org/licenses/by-sa/4.0/"
}
}
}
Reference
- Field-level documentation:
zh::WordDataand its dependencies on docs.rs. - Generated JSON Schema: https://tatuylonen.github.io/wiktextract/zh.json, also kept in
references/zh.jsonin the wiktionary-schema-rs repository.