FreeDictionary
FreeDictionary is an open-source organization providing free dictionary lookup services. We are committed to making high-quality online dictionary tools freely accessible to everyone.
Our Goals
- Provide completely free online dictionary services
- Ads-free
- No tracking
- No access restrictions
- Built on open data sources such as Wiktionary
- All code and tools are released as open source
Projects
| Project | Description | Language |
|---|---|---|
| pages | Organization website | Markdown |
| pick-up-desktop | Desktop dictionary app | Rust (iced) |
| load2db-go | Load dictionary data into PostgreSQL | Go |
| dictionary-server-rs | Dictionary server implemented in Rust | Rust |
| wiktionary-schema-go | Wiktionary data structure definitions | Go |
| wiktionary-schema-rs | Wiktionary data structure definitions | Rust |
| wiktionary-schema-ts | Wiktionary data structure definitions | TypeScript, ArkTS compatible |
| split-jsonl | Split JSONL files into smaller chunks | Python |
Services
Dictionaries
Server
Contributing
Contributions are welcome! Here’s how you can get involved:
- Star and fork projects under the FreeDictionary organization on Codeberg
- Submit an Issue to report bugs or suggest new features
- Submit a Pull Request to contribute code
License
The content of this project is licensed under the Creative Commons Attribution-ShareAlike 4.0 International (CC BY-SA 4.0) license.
Related Links
- Organization website: https://freedictionary.codeberg.page
- Organization home: https://codeberg.org/FreeDictionary
- Data source: https://kaikki.org/dictionary/rawdata.html
- Data mirror: https://download.freedictionary.cn
Wiktionary
The Wiktionary raw data are obtained from https://kaikki.org/dictionary/rawdata.html (extracted with wiktextract).
Data structures are defined in Rust and Go.
A download mirror (which doesn’t provide all language editions or files) is available at https://download.freedictionary.cn/wiktionary-raw-data/.
Editions
Two Wiktionary editions are currently supported, each with its own JSON structure (schemas):
| Edition | Language | Implemented in | Server route |
|---|---|---|---|
en | English | Rust, Go | /v1/wiktionary/en/... |
zh | Chinese | Rust, Go | /v1/wiktionary/zh/... |
Both editions can be queried through the dictionary server and imported into the database with load2db-go.
Wiktionary Schema
Wiktionary includes multiple language editions, and each edition has its own JSON structure (the fields differ from one edition to the next). Schemas are defined in Rust and Go.
| Edition | Rust module | Go package | Details |
|---|---|---|---|
en (English) | wiktionary_schema::en | en | English Edition |
zh (Chinese) | wiktionary_schema::zh | zh | Chinese Edition |
In both languages, API responses wrap the word data in a WrappedWordData type that adds a source object carrying provenance information (url and license).
Wiktionary Schema - English Edition
The English edition is implemented in both Rust and Go.
Rust
- Entry point:
en::WordData— a single word entry from the English Wiktionary data dump. - API responses wrap the word data in
en::WrappedWordData, which adds asourceobject with provenance information (urlandlicense).
Go
- Entry point:
en.WordDatainen/schema.go. - API responses wrap the word data in
WrappedWordData[en.WordData](packagewrapper), which adds the samesourceobject. - Uses Go’s
json/v2package; compile withGOEXPERIMENT=jsonv2.
Data Model
The schema mirrors the Python TypedDict definitions used by wiktextract for the English edition, kept in references/en.py in the repository. The top-level type WordData represents a single word entry: each JSONL line of the wiktextract dump deserializes into one instance.
All fields are optional (Option<T> in Rust, pointers in Go) except the identification fields (word, lang, lang_code) and a few core fields such as LinkageData::word. Fields deprecated upstream in wiktextract are annotated with #[deprecated] (e.g. english → translation, hyphenation → hyphenations).
Main Types
| Type | Description |
|---|---|
WordData | Top-level word entry: word, language, part of speech, and all lexical data. |
SenseData | A single sense (numbered definition) with glosses, examples, lexical relations, and Wikidata/Wikipedia links. |
ExampleData | Example sentence with translations, romanizations, ruby annotations, and bold-span offsets. |
FormData / FormOf | Inflected / alternative surface forms and the lemmas they belong to. |
SoundData | Pronunciation data: IPA, audio files, rhymes, dialect tags, etc. |
TranslationData | Translation of a sense into another language. |
LinkageData | A lexical relation (synonym, antonym, hypernym, hyponym, derived term, etc.). |
DescendantData | Descendant word in another language, forming a recursive etymological tree. |
AttestationData / ReferenceData | Historical attestations with dates and bibliographic references. |
TemplateData | A Wikitext template invocation (head line, inflection tables, info boxes). |
Hyphenation / DeprecatedHyphenation | Syllable-level hyphenation, plus the legacy list[str] format. |
EtymologyExample | Example sentence appearing in an etymology section. |
AltOf | The target word a sense is an alternative form of. |
RubyData | Ruby (furigana) annotation for CJK characters. |
LinkData | A wiki-internal link tuple [target, gloss]. |
PlusObjTemplateData / ExtraTemplateData | Structured arguments of certain templates (e.g. {{compound+}}). |
StringOrInt / TemplateArgs | Template argument keys: positional indices (i64) or named parameters (String). |
WordData Fields
| Category | Fields |
|---|---|
| Identification | word, lang, lang_code, pos, original_title |
| Senses | senses |
| Forms & variants | forms, form_of, alt_of, instances |
| Pronunciation | sounds |
| Etymology | etymology_text, etymology_number, etymology_examples, descendants |
| Lexical relations | synonyms, antonyms, hypernyms, hyponyms, holonyms, meronyms, troponyms, coordinate_terms, derived, related, abbreviations, anagrams, proverbs |
| Translations | translations |
| Hyphenation | hyphenation (deprecated), hyphenations |
| Templates | head_templates, inflection_templates, info_templates |
| Links & metadata | categories, wikidata, wikipedia, redirects, literal_meaning |
Example
A minimal wrapped entry as returned by the API:
{
"word": "hello",
"lang": "English",
"lang_code": "en",
"pos": "intj",
"senses": [
{
"glosses": ["A greeting (hello!)"]
}
],
"sounds": [
{
"ipa": "/həˈloʊ/"
}
],
"source": {
"url": "https://en.wiktionary.org/wiki/hello",
"license": {
"name": "CC BY-SA 4.0",
"url": "https://creativecommons.org/licenses/by-sa/4.0/"
}
}
}
Reference
- Field-level documentation:
en::WordDataand its dependencies on docs.rs. - Reference Python type definitions:
references/en.pyin the wiktionary-schema-rs repository.
Wiktionary Schema - Chinese Edition
The Chinese edition is implemented in both Rust and Go.
Rust
- Entry point:
zh::WordData— a single word entry from the Chinese Wiktionary data dump. - API responses wrap the word data in
zh::WrappedWordData, which adds asourceobject with provenance information (urlandlicense).
Go
- Entry point:
zh.WordDatainzh/schema.go. - API responses wrap the word data in
WrappedWordData[zh.WordData](packagewrapper), which adds the samesourceobject. - Uses Go’s
json/v2package; compile withGOEXPERIMENT=jsonv2.
Data Model
The schema mirrors the Pydantic models used by wiktextract for the Chinese edition (the WordEntry model and its dependencies in wiktextract/extractor/zh/models.py), kept in references/zh.json in the repository. The top-level type WordData represents a single word entry: each JSONL line of the zh-edition wiktextract dump deserializes into one instance.
All fields are optional (Option<T> in Rust, pointers in Go) except the identification fields (word, lang, lang_code) and pos.
Note: Hard-redirect pages in the raw dump are emitted as plain
{"title": ..., "redirect": ..., "pos": "hard-redirect"}objects that do not conform to this schema; consumers streaming the raw dump should skip lines without awordfield.
Main Types
| Type | Description |
|---|---|
WordData | Top-level word entry: word, language, part of speech, and all lexical data. |
Sense | A single sense (numbered definition) with glosses, examples, classifiers, and lexical relations. |
Classifier | A classifier (measure word) used with a noun sense, e.g. 个 in 一个苹果. |
Example | Example sentence with Chinese translation, romanization, ruby, and bold-span offsets. |
Form | Inflected / alternative surface form (may include a Hiragana reading for Japanese entries). |
Sound | Pronunciation data: zh_pron (Pinyin, Bopomofo, Jyutping, etc.), IPA, audio URLs, Hangeul, homophones. |
Translation | Translation of a sense into another language. |
Linkage | A lexical relation (synonym, antonym, hypernym, hyponym, compound, etc.). |
Descendant | Descendant word in another language, forming a recursive etymological tree. |
AttestationData / ReferenceData | Historical attestations with dates and bibliographic references. |
Hyphenation | Syllable-level hyphenation. |
AltForm | Alternative / base form (e.g. simplified/traditional variants, inflected forms). |
WordData Fields
| Category | Fields |
|---|---|
| Identification | word, lang, lang_code, pos, pos_title, original_title, title |
| Senses | senses |
| Forms & variants | forms |
| Pronunciation | sounds |
| Etymology | etymology_texts, etymology_examples, descendants |
| Lexical relations | synonyms, antonyms, hypernyms, hyponyms, holonyms, meronyms, troponyms, coordinate_terms, derived, related, paronyms, abbreviations, anagrams, proverbs, compounds, various |
| Classifiers | classifiers |
| Translations | translations |
| Hyphenation | hyphenations |
| Links & metadata | categories, notes, redirect, redirects, literal_meaning, raw_tags, tags |
Example
A minimal wrapped entry as returned by the API:
{
"word": "猫",
"lang": "汉语",
"lang_code": "zh",
"pos": "noun",
"pos_title": "名詞",
"senses": [
{
"glosses": ["一種貓科動物,俗稱貓咪。"]
}
],
"sounds": [
{
"zh_pron": "māo"
}
],
"source": {
"url": "https://zh.wiktionary.org/wiki/猫",
"license": {
"name": "CC BY-SA 4.0",
"url": "https://creativecommons.org/licenses/by-sa/4.0/"
}
}
}
Reference
- Field-level documentation:
zh::WordDataand its dependencies on docs.rs. - Generated JSON Schema: https://tatuylonen.github.io/wiktextract/zh.json, also kept in
references/zh.jsonin the wiktionary-schema-rs repository.
Dictionary Server
The dictionary server (dictionary-server-rs) is a server programme implemented in Rust that offers word enquiry service. It serves Wiktionary data that has been extracted by wiktextract and loaded into a PostgreSQL database (e.g. by load2db-go).
A live instance is available at https://api.freedictionary.cn/.
Building
cargo build --release
The CI pipeline additionally checks formatting and hints:
cargo fmt --check
cargo clippy --all-targets -- -D warnings
Usage
Run the FreeDictionary server
Usage: serve-dict [OPTIONS] --pg-username <PG_USERNAME> --pg-passwd <PG_PASSWD> <--pg-addr <PG_ADDR>|--pg-socket <PG_SOCKET>>
Options:
--pg-addr <PG_ADDR> Address/url of PostgreSQL
--pg-socket <PG_SOCKET> Path to the **directory** containing PostgreSQL socket
--pg-port <PG_PORT> Port to PostgreSQL. sqlx defaults to 5432
--pg-db <PG_DB> PostgreSQL Database in which FreeDictionary data is stored [default: FreeDictionary]
--pg-username <PG_USERNAME> PostgreSQL user name which has read access to the Database
--pg-passwd <PG_PASSWD> PostgreSQL password of the account
--pg-pool-size <PG_POOL_SIZE> Maximum number of connections in the PostgreSQL connection pool [default: 50]
--listen <LISTEN> Listen address for the server [default: 0.0.0.0]
-p, --port <PORT> Listen port for the server [default: 8080]
--verbose Enable debug-level logging; overridden by the RUST_LOG environment variable
--cache-max-mb <CACHE_MAX_MB> Maximum total size of the word lookup cache in MiB (0 disables caching) [default: 256]
--cache-ttl-secs <CACHE_TTL_SECS> Time-to-live for cache entries in seconds [default: 3600]
-h, --help Print help (see more with '--help')
-V, --version Print version
The connection to PostgreSQL is established either through a TCP address (--pg-addr) or through a Unix socket directory (--pg-socket); exactly one of the two must be given.
Endpoints
The server exposes the following routes (see the API Reference for details):
| Route | Description |
|---|---|
GET / | Landing page (static HTML). |
GET /v1/wiktionary/{edition}/words/{word} | Exact word lookup, with optional lang_code / lang filters. |
GET /v1/wiktionary/editions | Lists all supported Wiktionary editions. |
GET /v1/wiktionary/{edition}/words | Fuzzy word search (reserved; returns 501 Not Implemented). |
Database
The server expects a PostgreSQL database (default name FreeDictionary) that has been initialised with load2db-go. For each Wiktionary edition it must contain a table named WiktionaryData_{edition} (e.g. WiktionaryData_en, WiktionaryData_zh) with at least the following columns:
| Column | Description |
|---|---|
word | The headword / lemma. |
lang | Localised language name of the word (e.g. 汉语). |
lang_code | Wiktionary language code (e.g. en, zh). |
data | The full entry as JSON. |
The server refuses to start if the table WiktionaryData_en is missing.
The connection pool is capped at --pg-pool-size connections (default 50).
Word Lookup Cache
Successful word-lookup responses are cached in memory to avoid repeated database round-trips. The cache is implemented with moka; entries are stored zstd-compressed and keyed by (edition, word, lang, lang_code).
| Option | Default | Description |
|---|---|---|
--cache-max-mb | 256 | Maximum total size of the cache in MiB; 0 disables caching. |
--cache-ttl-secs | 3600 | Time-to-live for cache entries in seconds. |
Memory Allocator
The global allocator is selectable at compile time through Cargo features:
With no flags the default feature set builds with mimalloc. To use the system allocator instead, disable default features:
cargo build --release --no-default-features
To use jemalloc:
cargo build --release --features jemalloc
If both features are enabled, jemalloc takes precedence. On platforms where a selected allocator is unsupported (for example jemalloc on Windows), the build silently falls back to the system allocator.
Logging
Logging is initialised with env_logger. The default log level is info (startup and per-request logs); pass --verbose to raise it to debug (which additionally prints per-request response details).
The RUST_LOG environment variable, when set, overrides this default entirely:
RUST_LOG=warn serve-dict --pg-addr localhost --pg-username user --pg-passwd secret
Benchmark
A Go load-testing tool is provided to measure the server’s throughput, latency distribution, and error rate under high concurrency.
Tip: When benchmarking remotely, use the Go client — it uses goroutines across all cores with minimal client-side overhead, making it suitable for finding the server’s true throughput ceiling.
Go (bench-go)
Requires Go 1.21+:
cd bench-go
go build -o bench .
Constant-concurrency mode (default: 100 concurrent connections for 30 s):
./bench --url http://localhost:8080
./bench --url http://localhost:8080 -c 200 -d 60
Ramp mode — steps concurrency from 10 to 1000 (step 50), holding each level for 30 s:
./bench --url http://localhost:8080 --mode ramp --ramp-from 10 --ramp-to 1000 --ramp-step 50 --ramp-duration 30
Useful options:
| Flag | Description |
|---|---|
--timeout <s> | Per-request timeout in seconds (default 10). Requests exceeding this are counted as client-side timeout errors. |
--json <path> | Export a structured JSON report. |
--words <file> | Custom word list (one word per line). |
--editions <e,...> | Editions to test, comma-separated (default en,zh). |
--no-compress | Disable HTTP compression (send Accept-Encoding: identity). |
--no-warmup | Skip the warm-up phase. |
--no-query-params | Do not generate requests with lang_code query parameters. |
The tool exercises both editions (en, zh), with and without lang_code query parameters, non-existent words (empty-result path), and unsupported editions (error path). A warm-up phase runs before measurement; pass --no-warmup to skip it.
Note: “errors” reported by the tools include client-side timeouts — the server may still return 200 for those requests after the client has given up. Use a larger --timeout to measure raw capacity, or a smaller one (e.g. --timeout 2) to measure effective capacity within an SLA budget.
Project Layout
| File | Description |
|---|---|
src/main.rs | Server bootstrap and route registration. |
src/allocator.rs | Compile-time global allocator selection. |
src/cli.rs | Command-line argument parsing and PostgreSQL URL construction. |
src/errors.rs | Error types and their mapping to HTTP status codes. |
src/wiktionary.rs | Database queries. |
static/index.html | Landing page served at /. |
bench-go/ | Load-testing / benchmark tool (Go, zero dependencies). |
License
Licensed under the GNU Affero General Public License v3.0.
API Reference
The dictionary server exposes a REST API under the /v1 prefix. A live instance is available at https://api.freedictionary.cn/.
GET /v1/wiktionary/{edition}/words/{word}
Returns the Wiktionary entries for the given word from the designated edition.
Path Parameters
edition: the code of the Wiktionary edition, e.g.enorzh. A route is registered for everyWiktionaryData_*table found in the database; requesting an edition that is not supported returns404 Not Found.
Query Parameters
Both parameters are optional and narrow the search when provided:
| Parameter | Description |
|---|---|
lang_code | The language code of the word, e.g. en. |
lang | The localised language name of the word, e.g. 汉语. |
If both are provided, both are applied as filters.
Examples
GET /v1/wiktionary/en/words/hello?lang_code=en— English-edition Wiktionary entry for the English word “hello”.GET /v1/wiktionary/en/words/hallo?lang_code=de— English-edition Wiktionary entry for the German word “hallo”.GET /v1/wiktionary/zh/words/你好?lang_code=zh— Chinese-edition Wiktionary entry for the Chinese word “你好”.
Response
200 OK: a JSON array of entries for the word; an empty array means the word was not found in that Wiktionary. Each entry is the wrapped word-data object produced by thewiktionary-schemacrate: the word data itself (seeen::WordDataandzh::WordData) plus asourceobject carrying provenance information —url(the Wiktionary page the data came from) andlicense, itself an object withnameandurl(defaults toCC BY-SA 4.0).
Error Responses
Error responses are JSON objects with the following shape:
{
"status": 404,
"message": "Unimplemented wiktionary edition: fr"
}
| Field | Type | Description |
|---|---|---|
status | u16 | The HTTP status code of the error. |
message | String | A human-readable error message. |
| Status | Meaning |
|---|---|
404 Not Found | The requested edition is not supported. |
500 Internal Server Error | Stored data could not be decoded, or an unexpected database/server error occurred. |
503 Service Unavailable | The database is temporarily unreachable; retrying the request may succeed. |
GET /v1/wiktionary/editions
Returns a JSON array of all supported Wiktionary edition codes discovered in the database (e.g. ["en", "zh"]).
Response
200 OK: a JSON array of edition code strings.
GET /v1/wiktionary/{edition}/words
Fuzzy word search. Returns a list of headwords matching the given pattern (e.g. som* matches “some”, “somebody”, “something”).
Note: This endpoint is reserved and not yet implemented. It currently returns
501 Not Implemented.
Query Parameters
| Parameter | Required | Description |
|---|---|---|
q | yes | The search pattern. Use * as a wildcard suffix for prefix matching (e.g. q=som*). |
lang_code | no | Language code filter. |
lang | no | Localised language name filter. |
limit | no | Maximum number of results to return. |
offset | no | Number of results to skip (for pagination). |
Response (Planned)
200 OK: a JSON array of matching headword strings.501 Not Implemented: the endpoint is not yet available.
Downloads
Pre-built binaries for the dictionary server are available on the Codeberg releases page.
Latest Release
The latest stable release is v0.5.1 (August 2, 2026).
Supported Platforms
Binaries are built for the following targets. Starting from v0.5.0, each binary is packaged as a tarball (.tar.xz on Linux/macOS, .zip on Windows).
| Platform | Architecture Variants | Format |
|---|---|---|
Linux (glibc) — x86_64-unknown-linux-gnu | x86-64, x86-64-v2, x86-64-v3, x86-64-v4 | .tar.xz |
Linux (musl) — x86_64-unknown-linux-musl | x86-64, x86-64-v2, x86-64-v3, x86-64-v4 | .tar.xz |
Linux (glibc) — aarch64-unknown-linux-gnu | v8a, v8.2a, v9a | .tar.xz |
Linux (musl) — aarch64-unknown-linux-musl | v8a, v8.2a, v9a | .tar.xz |
macOS (Apple Silicon) — aarch64-apple-darwin | v8a, v8.2a, v9a | .tar.xz |
macOS (Intel) — x86_64-apple-darwin | x86-64, x86-64-v3 | .tar.xz |
Windows — x86_64-pc-windows-gnu | x86-64, x86-64-v2, x86-64-v3, x86-64-v4 | .zip |
Linux (LoongArch) — loongarch64-unknown-linux-musl | loongarch64, lasx | .tar.xz |
Linux (RISC-V) — riscv64gc-unknown-linux-musl | rv64gc, rv64gcv | .tar.xz |
Choosing an Architecture Variant
For x86-64 targets, multiple microarchitecture levels are provided:
| Variant | CPUs |
|---|---|
x86-64 | Baseline x86-64 (compatible with all x86-64 CPUs). |
x86-64-v2 | Requires SSE4.2 and POPCNT (Nehalem+, ~2008). |
x86-64-v3 | Requires AVX, AVX2, BMI1/2, FMA (Haswell+, ~2013). |
x86-64-v4 | Requires AVX-512F/BW/DQ/CD/VL (Skylake-X+, ~2017). |
For ARM64 targets:
| Variant | CPUs |
|---|---|
v8a | Baseline ARMv8-A (compatible with all ARM64 CPUs). |
v8.2a | ARMv8.2-A with optional FP16 and RDM extensions. |
v9a | ARMv9-A with SVE2 and Scalable Matrix Extension (SME). |
Usage
After downloading and extracting the archive, run the binary with the required PostgreSQL connection options:
# Linux/macOS
tar xf serve-dict-x86_64-unknown-linux-musl-x86-64.tar.xz
./serve-dict --pg-addr localhost --pg-username user --pg-passwd secret
# Windows
unzip serve-dict-x86_64-pc-windows-gnu-x86-64.zip
serve-dict.exe --pg-addr localhost --pg-username user --pg-passwd secret
See Server for the full list of command-line options.
Building from Source
If your platform is not listed, you can build from source:
git clone https://codeberg.org/FreeDictionary/dictionary-server-rs.git
cd dictionary-server-rs
cargo build --release
The resulting binary will be at target/release/serve-dict.