Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

FreeDictionary

FreeDictionary is an open-source organization providing free dictionary lookup services. We are committed to making high-quality online dictionary tools freely accessible to everyone.

Our Goals

  • Provide completely free online dictionary services
  • Ads-free
  • No tracking
  • No access restrictions
  • Built on open data sources such as Wiktionary
  • All code and tools are released as open source

Projects

ProjectDescriptionLanguage
pagesOrganization websiteMarkdown
pick-up-desktopDesktop dictionary appRust (iced)
load2db-goLoad dictionary data into PostgreSQLGo
dictionary-server-rsDictionary server implemented in RustRust
wiktionary-schema-goWiktionary data structure definitionsGo
wiktionary-schema-rsWiktionary data structure definitionsRust
wiktionary-schema-tsWiktionary data structure definitionsTypeScript, ArkTS compatible
split-jsonlSplit JSONL files into smaller chunksPython

Services

Dictionaries

Server

Contributing

Contributions are welcome! Here’s how you can get involved:

  1. Star and fork projects under the FreeDictionary organization on Codeberg
  2. Submit an Issue to report bugs or suggest new features
  3. Submit a Pull Request to contribute code

License

The content of this project is licensed under the Creative Commons Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) license.

Wiktionary

The Wiktionary raw data are obtained from https://kaikki.org/dictionary/rawdata.html (extracted with wiktextract).

Data structures are defined in Rust and Go.

A download mirror (which doesn’t provide all language editions or files) is available at https://download.freedictionary.cn/wiktionary-raw-data/.

Editions

Two Wiktionary editions are currently supported, each with its own JSON structure (schemas):

EditionLanguageImplemented inServer route
enEnglishRust, Go/v1/wiktionary/en/...
zhChineseRust, Go/v1/wiktionary/zh/...

Both editions can be queried through the dictionary server and imported into the database with load2db-go.

Wiktionary Schema

Wiktionary includes multiple language editions, and each edition has its own JSON structure (the fields differ from one edition to the next). Schemas are defined in Rust and Go.

EditionRust moduleGo packageDetails
en (English)wiktionary_schema::enenEnglish Edition
zh (Chinese)wiktionary_schema::zhzhChinese Edition

In both languages, API responses wrap the word data in a WrappedWordData type that adds a source object carrying provenance information (url and license).

Wiktionary Schema - English Edition

The English edition is implemented in both Rust and Go.

Rust

  • Entry point: en::WordData — a single word entry from the English Wiktionary data dump.
  • API responses wrap the word data in en::WrappedWordData, which adds a source object with provenance information (url and license).

Go

  • Entry point: en.WordData in en/schema.go.
  • API responses wrap the word data in WrappedWordData[en.WordData] (package wrapper), which adds the same source object.
  • Uses Go’s json/v2 package; compile with GOEXPERIMENT=jsonv2.

Data Model

The schema mirrors the Python TypedDict definitions used by wiktextract for the English edition, kept in references/en.py in the repository. The top-level type WordData represents a single word entry: each JSONL line of the wiktextract dump deserializes into one instance.

All fields are optional (Option<T> in Rust, pointers in Go) except the identification fields (word, lang, lang_code) and a few core fields such as LinkageData::word. Fields deprecated upstream in wiktextract are annotated with #[deprecated] (e.g. english → translation, hyphenation → hyphenations).

Main Types

TypeDescription
WordDataTop-level word entry: word, language, part of speech, and all lexical data.
SenseDataA single sense (numbered definition) with glosses, examples, lexical relations, and Wikidata/Wikipedia links.
ExampleDataExample sentence with translations, romanizations, ruby annotations, and bold-span offsets.
FormData / FormOfInflected / alternative surface forms and the lemmas they belong to.
SoundDataPronunciation data: IPA, audio files, rhymes, dialect tags, etc.
TranslationDataTranslation of a sense into another language.
LinkageDataA lexical relation (synonym, antonym, hypernym, hyponym, derived term, etc.).
DescendantDataDescendant word in another language, forming a recursive etymological tree.
AttestationData / ReferenceDataHistorical attestations with dates and bibliographic references.
TemplateDataA Wikitext template invocation (head line, inflection tables, info boxes).
Hyphenation / DeprecatedHyphenationSyllable-level hyphenation, plus the legacy list[str] format.
EtymologyExampleExample sentence appearing in an etymology section.
AltOfThe target word a sense is an alternative form of.
RubyDataRuby (furigana) annotation for CJK characters.
LinkDataA wiki-internal link tuple [target, gloss].
PlusObjTemplateData / ExtraTemplateDataStructured arguments of certain templates (e.g. {{compound+}}).
StringOrInt / TemplateArgsTemplate argument keys: positional indices (i64) or named parameters (String).

WordData Fields

CategoryFields
Identificationword, lang, lang_code, pos, original_title
Sensessenses
Forms & variantsforms, form_of, alt_of, instances
Pronunciationsounds
Etymologyetymology_text, etymology_number, etymology_examples, descendants
Lexical relationssynonyms, antonyms, hypernyms, hyponyms, holonyms, meronyms, troponyms, coordinate_terms, derived, related, abbreviations, anagrams, proverbs
Translationstranslations
Hyphenationhyphenation (deprecated), hyphenations
Templateshead_templates, inflection_templates, info_templates
Links & metadatacategories, wikidata, wikipedia, redirects, literal_meaning

Example

A minimal wrapped entry as returned by the API:

{
  "word": "hello",
  "lang": "English",
  "lang_code": "en",
  "pos": "intj",
  "senses": [
    {
      "glosses": ["A greeting (hello!)"]
    }
  ],
  "sounds": [
    {
      "ipa": "/həˈloʊ/"
    }
  ],
  "source": {
    "url": "https://en.wiktionary.org/wiki/hello",
    "license": {
      "name": "CC BY-SA 4.0",
      "url": "https://creativecommons.org/licenses/by-sa/4.0/"
    }
  }
}

Reference

  • Field-level documentation: en::WordData and its dependencies on docs.rs.
  • Reference Python type definitions: references/en.py in the wiktionary-schema-rs repository.

Wiktionary Schema - Chinese Edition

The Chinese edition is implemented in both Rust and Go.

Rust

  • Entry point: zh::WordData — a single word entry from the Chinese Wiktionary data dump.
  • API responses wrap the word data in zh::WrappedWordData, which adds a source object with provenance information (url and license).

Go

  • Entry point: zh.WordData in zh/schema.go.
  • API responses wrap the word data in WrappedWordData[zh.WordData] (package wrapper), which adds the same source object.
  • Uses Go’s json/v2 package; compile with GOEXPERIMENT=jsonv2.

Data Model

The schema mirrors the Pydantic models used by wiktextract for the Chinese edition (the WordEntry model and its dependencies in wiktextract/extractor/zh/models.py), kept in references/zh.json in the repository. The top-level type WordData represents a single word entry: each JSONL line of the zh-edition wiktextract dump deserializes into one instance.

All fields are optional (Option<T> in Rust, pointers in Go) except the identification fields (word, lang, lang_code) and pos.

Note: Hard-redirect pages in the raw dump are emitted as plain {"title": ..., "redirect": ..., "pos": "hard-redirect"} objects that do not conform to this schema; consumers streaming the raw dump should skip lines without a word field.

Main Types

TypeDescription
WordDataTop-level word entry: word, language, part of speech, and all lexical data.
SenseA single sense (numbered definition) with glosses, examples, classifiers, and lexical relations.
ClassifierA classifier (measure word) used with a noun sense, e.g. 个 in 一个苹果.
ExampleExample sentence with Chinese translation, romanization, ruby, and bold-span offsets.
FormInflected / alternative surface form (may include a Hiragana reading for Japanese entries).
SoundPronunciation data: zh_pron (Pinyin, Bopomofo, Jyutping, etc.), IPA, audio URLs, Hangeul, homophones.
TranslationTranslation of a sense into another language.
LinkageA lexical relation (synonym, antonym, hypernym, hyponym, compound, etc.).
DescendantDescendant word in another language, forming a recursive etymological tree.
AttestationData / ReferenceDataHistorical attestations with dates and bibliographic references.
HyphenationSyllable-level hyphenation.
AltFormAlternative / base form (e.g. simplified/traditional variants, inflected forms).

WordData Fields

CategoryFields
Identificationword, lang, lang_code, pos, pos_title, original_title, title
Sensessenses
Forms & variantsforms
Pronunciationsounds
Etymologyetymology_texts, etymology_examples, descendants
Lexical relationssynonyms, antonyms, hypernyms, hyponyms, holonyms, meronyms, troponyms, coordinate_terms, derived, related, paronyms, abbreviations, anagrams, proverbs, compounds, various
Classifiersclassifiers
Translationstranslations
Hyphenationhyphenations
Links & metadatacategories, notes, redirect, redirects, literal_meaning, raw_tags, tags

Example

A minimal wrapped entry as returned by the API:

{
  "word": "猫",
  "lang": "汉语",
  "lang_code": "zh",
  "pos": "noun",
  "pos_title": "名詞",
  "senses": [
    {
      "glosses": ["一種貓科動物,俗稱貓咪。"]
    }
  ],
  "sounds": [
    {
      "zh_pron": "māo"
    }
  ],
  "source": {
    "url": "https://zh.wiktionary.org/wiki/猫",
    "license": {
      "name": "CC BY-SA 4.0",
      "url": "https://creativecommons.org/licenses/by-sa/4.0/"
    }
  }
}

Reference

Dictionary Server

The dictionary server (dictionary-server-rs) is a server programme implemented in Rust that offers word enquiry service. It serves Wiktionary data that has been extracted by wiktextract and loaded into a PostgreSQL database (e.g. by load2db-go).

A live instance is available at https://api.freedictionary.cn/.

Building

cargo build --release

The CI pipeline additionally checks formatting and hints:

cargo fmt --check
cargo clippy --all-targets -- -D warnings

Usage

Run the FreeDictionary server

Usage: serve-dict [OPTIONS] --pg-username <PG_USERNAME> --pg-passwd <PG_PASSWD> <--pg-addr <PG_ADDR>|--pg-socket <PG_SOCKET>>

Options:
      --pg-addr <PG_ADDR>               Address/url of PostgreSQL
      --pg-socket <PG_SOCKET>           Path to the **directory** containing PostgreSQL socket
      --pg-port <PG_PORT>               Port to PostgreSQL. sqlx defaults to 5432
      --pg-db <PG_DB>                   PostgreSQL Database in which FreeDictionary data is stored [default: FreeDictionary]
      --pg-username <PG_USERNAME>       PostgreSQL user name which has read access to the Database
      --pg-passwd <PG_PASSWD>           PostgreSQL password of the account
      --pg-pool-size <PG_POOL_SIZE>     Maximum number of connections in the PostgreSQL connection pool [default: 50]
      --listen <LISTEN>                 Listen address for the server [default: 0.0.0.0]
  -p, --port <PORT>                     Listen port for the server [default: 8080]
      --verbose                         Enable debug-level logging; overridden by the RUST_LOG environment variable
      --cache-max-mb <CACHE_MAX_MB>     Maximum total size of the word lookup cache in MiB (0 disables caching) [default: 256]
      --cache-ttl-secs <CACHE_TTL_SECS> Time-to-live for cache entries in seconds [default: 3600]
  -h, --help                            Print help (see more with '--help')
  -V, --version                         Print version

The connection to PostgreSQL is established either through a TCP address (--pg-addr) or through a Unix socket directory (--pg-socket); exactly one of the two must be given.

Endpoints

The server exposes the following routes (see the API Reference for details):

RouteDescription
GET /Landing page (static HTML).
GET /v1/wiktionary/{edition}/words/{word}Exact word lookup, with optional lang_code / lang filters.
GET /v1/wiktionary/editionsLists all supported Wiktionary editions.
GET /v1/wiktionary/{edition}/wordsFuzzy word search (reserved; returns 501 Not Implemented).

Database

The server expects a PostgreSQL database (default name FreeDictionary) that has been initialised with load2db-go. For each Wiktionary edition it must contain a table named WiktionaryData_{edition} (e.g. WiktionaryData_en, WiktionaryData_zh) with at least the following columns:

ColumnDescription
wordThe headword / lemma.
langLocalised language name of the word (e.g. 汉语).
lang_codeWiktionary language code (e.g. en, zh).
dataThe full entry as JSON.

The server refuses to start if the table WiktionaryData_en is missing.

The connection pool is capped at --pg-pool-size connections (default 50).

Word Lookup Cache

Successful word-lookup responses are cached in memory to avoid repeated database round-trips. The cache is implemented with moka; entries are stored zstd-compressed and keyed by (edition, word, lang, lang_code).

OptionDefaultDescription
--cache-max-mb256Maximum total size of the cache in MiB; 0 disables caching.
--cache-ttl-secs3600Time-to-live for cache entries in seconds.

Memory Allocator

The global allocator is selectable at compile time through Cargo features:

FeatureDefaultAllocator
mimallocyesmimalloc
jemallocnojemalloc

With no flags the default feature set builds with mimalloc. To use the system allocator instead, disable default features:

cargo build --release --no-default-features

To use jemalloc:

cargo build --release --features jemalloc

If both features are enabled, jemalloc takes precedence. On platforms where a selected allocator is unsupported (for example jemalloc on Windows), the build silently falls back to the system allocator.

Logging

Logging is initialised with env_logger. The default log level is info (startup and per-request logs); pass --verbose to raise it to debug (which additionally prints per-request response details).

The RUST_LOG environment variable, when set, overrides this default entirely:

RUST_LOG=warn serve-dict --pg-addr localhost --pg-username user --pg-passwd secret

Benchmark

A Go load-testing tool is provided to measure the server’s throughput, latency distribution, and error rate under high concurrency.

Tip: When benchmarking remotely, use the Go client — it uses goroutines across all cores with minimal client-side overhead, making it suitable for finding the server’s true throughput ceiling.

Go (bench-go)

Requires Go 1.21+:

cd bench-go
go build -o bench .

Constant-concurrency mode (default: 100 concurrent connections for 30 s):

./bench --url http://localhost:8080
./bench --url http://localhost:8080 -c 200 -d 60

Ramp mode — steps concurrency from 10 to 1000 (step 50), holding each level for 30 s:

./bench --url http://localhost:8080 --mode ramp --ramp-from 10 --ramp-to 1000 --ramp-step 50 --ramp-duration 30

Useful options:

FlagDescription
--timeout <s>Per-request timeout in seconds (default 10). Requests exceeding this are counted as client-side timeout errors.
--json <path>Export a structured JSON report.
--words <file>Custom word list (one word per line).
--editions <e,...>Editions to test, comma-separated (default en,zh).
--no-compressDisable HTTP compression (send Accept-Encoding: identity).
--no-warmupSkip the warm-up phase.
--no-query-paramsDo not generate requests with lang_code query parameters.

The tool exercises both editions (en, zh), with and without lang_code query parameters, non-existent words (empty-result path), and unsupported editions (error path). A warm-up phase runs before measurement; pass --no-warmup to skip it.

Note: “errors” reported by the tools include client-side timeouts — the server may still return 200 for those requests after the client has given up. Use a larger --timeout to measure raw capacity, or a smaller one (e.g. --timeout 2) to measure effective capacity within an SLA budget.

Project Layout

FileDescription
src/main.rsServer bootstrap and route registration.
src/allocator.rsCompile-time global allocator selection.
src/cli.rsCommand-line argument parsing and PostgreSQL URL construction.
src/errors.rsError types and their mapping to HTTP status codes.
src/wiktionary.rsDatabase queries.
static/index.htmlLanding page served at /.
bench-go/Load-testing / benchmark tool (Go, zero dependencies).

License

Licensed under the GNU Affero General Public License v3.0.

API Reference

The dictionary server exposes a REST API under the /v1 prefix. A live instance is available at https://api.freedictionary.cn/.

GET /v1/wiktionary/{edition}/words/{word}

Returns the Wiktionary entries for the given word from the designated edition.

Path Parameters

  • edition: the code of the Wiktionary edition, e.g. en or zh. A route is registered for every WiktionaryData_* table found in the database; requesting an edition that is not supported returns 404 Not Found.

Query Parameters

Both parameters are optional and narrow the search when provided:

ParameterDescription
lang_codeThe language code of the word, e.g. en.
langThe localised language name of the word, e.g. 汉语.

If both are provided, both are applied as filters.

Examples

  • GET /v1/wiktionary/en/words/hello?lang_code=en — English-edition Wiktionary entry for the English word “hello”.
  • GET /v1/wiktionary/en/words/hallo?lang_code=de — English-edition Wiktionary entry for the German word “hallo”.
  • GET /v1/wiktionary/zh/words/你好?lang_code=zh — Chinese-edition Wiktionary entry for the Chinese word “你好”.

Response

  • 200 OK: a JSON array of entries for the word; an empty array means the word was not found in that Wiktionary. Each entry is the wrapped word-data object produced by the wiktionary-schema crate: the word data itself (see en::WordData and zh::WordData) plus a source object carrying provenance information — url (the Wiktionary page the data came from) and license, itself an object with name and url (defaults to CC BY-SA 4.0).

Error Responses

Error responses are JSON objects with the following shape:

{
  "status": 404,
  "message": "Unimplemented wiktionary edition: fr"
}
FieldTypeDescription
statusu16The HTTP status code of the error.
messageStringA human-readable error message.
StatusMeaning
404 Not FoundThe requested edition is not supported.
500 Internal Server ErrorStored data could not be decoded, or an unexpected database/server error occurred.
503 Service UnavailableThe database is temporarily unreachable; retrying the request may succeed.

GET /v1/wiktionary/editions

Returns a JSON array of all supported Wiktionary edition codes discovered in the database (e.g. ["en", "zh"]).

Response

  • 200 OK: a JSON array of edition code strings.

GET /v1/wiktionary/{edition}/words

Fuzzy word search. Returns a list of headwords matching the given pattern (e.g. som* matches “some”, “somebody”, “something”).

Note: This endpoint is reserved and not yet implemented. It currently returns 501 Not Implemented.

Query Parameters

ParameterRequiredDescription
qyesThe search pattern. Use * as a wildcard suffix for prefix matching (e.g. q=som*).
lang_codenoLanguage code filter.
langnoLocalised language name filter.
limitnoMaximum number of results to return.
offsetnoNumber of results to skip (for pagination).

Response (Planned)

  • 200 OK: a JSON array of matching headword strings.
  • 501 Not Implemented: the endpoint is not yet available.

Downloads

Pre-built binaries for the dictionary server are available on the Codeberg releases page.

Latest Release

The latest stable release is v0.5.1 (August 2, 2026).

Supported Platforms

Binaries are built for the following targets. Starting from v0.5.0, each binary is packaged as a tarball (.tar.xz on Linux/macOS, .zip on Windows).

PlatformArchitecture VariantsFormat
Linux (glibc) — x86_64-unknown-linux-gnux86-64, x86-64-v2, x86-64-v3, x86-64-v4.tar.xz
Linux (musl) — x86_64-unknown-linux-muslx86-64, x86-64-v2, x86-64-v3, x86-64-v4.tar.xz
Linux (glibc) — aarch64-unknown-linux-gnuv8a, v8.2a, v9a.tar.xz
Linux (musl) — aarch64-unknown-linux-muslv8a, v8.2a, v9a.tar.xz
macOS (Apple Silicon) — aarch64-apple-darwinv8a, v8.2a, v9a.tar.xz
macOS (Intel) — x86_64-apple-darwinx86-64, x86-64-v3.tar.xz
Windows — x86_64-pc-windows-gnux86-64, x86-64-v2, x86-64-v3, x86-64-v4.zip
Linux (LoongArch) — loongarch64-unknown-linux-muslloongarch64, lasx.tar.xz
Linux (RISC-V) — riscv64gc-unknown-linux-muslrv64gc, rv64gcv.tar.xz

Choosing an Architecture Variant

For x86-64 targets, multiple microarchitecture levels are provided:

VariantCPUs
x86-64Baseline x86-64 (compatible with all x86-64 CPUs).
x86-64-v2Requires SSE4.2 and POPCNT (Nehalem+, ~2008).
x86-64-v3Requires AVX, AVX2, BMI1/2, FMA (Haswell+, ~2013).
x86-64-v4Requires AVX-512F/BW/DQ/CD/VL (Skylake-X+, ~2017).

For ARM64 targets:

VariantCPUs
v8aBaseline ARMv8-A (compatible with all ARM64 CPUs).
v8.2aARMv8.2-A with optional FP16 and RDM extensions.
v9aARMv9-A with SVE2 and Scalable Matrix Extension (SME).

Usage

After downloading and extracting the archive, run the binary with the required PostgreSQL connection options:

# Linux/macOS
tar xf serve-dict-x86_64-unknown-linux-musl-x86-64.tar.xz
./serve-dict --pg-addr localhost --pg-username user --pg-passwd secret

# Windows
unzip serve-dict-x86_64-pc-windows-gnu-x86-64.zip
serve-dict.exe --pg-addr localhost --pg-username user --pg-passwd secret

See Server for the full list of command-line options.

Building from Source

If your platform is not listed, you can build from source:

git clone https://codeberg.org/FreeDictionary/dictionary-server-rs.git
cd dictionary-server-rs
cargo build --release

The resulting binary will be at target/release/serve-dict.