Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Dictionary Server

The dictionary server (dictionary-server-rs) is a server programme implemented in Rust that offers word enquiry service. It serves Wiktionary data that has been extracted by wiktextract and loaded into a PostgreSQL database (e.g. by load2db-go).

A live instance is available at https://api.freedictionary.cn/.

Building

cargo build --release

The CI pipeline additionally checks formatting and hints:

cargo fmt --check
cargo clippy --all-targets -- -D warnings

Usage

Run the FreeDictionary server

Usage: serve-dict [OPTIONS] --pg-username <PG_USERNAME> --pg-passwd <PG_PASSWD> <--pg-addr <PG_ADDR>|--pg-socket <PG_SOCKET>>

Options:
      --pg-addr <PG_ADDR>               Address/url of PostgreSQL
      --pg-socket <PG_SOCKET>           Path to the **directory** containing PostgreSQL socket
      --pg-port <PG_PORT>               Port to PostgreSQL. sqlx defaults to 5432
      --pg-db <PG_DB>                   PostgreSQL Database in which FreeDictionary data is stored [default: FreeDictionary]
      --pg-username <PG_USERNAME>       PostgreSQL user name which has read access to the Database
      --pg-passwd <PG_PASSWD>           PostgreSQL password of the account
      --pg-pool-size <PG_POOL_SIZE>     Maximum number of connections in the PostgreSQL connection pool [default: 50]
      --listen <LISTEN>                 Listen address for the server [default: 0.0.0.0]
  -p, --port <PORT>                     Listen port for the server [default: 8080]
      --verbose                         Enable debug-level logging; overridden by the RUST_LOG environment variable
      --cache-max-mb <CACHE_MAX_MB>     Maximum total size of the word lookup cache in MiB (0 disables caching) [default: 256]
      --cache-ttl-secs <CACHE_TTL_SECS> Time-to-live for cache entries in seconds [default: 3600]
  -h, --help                            Print help (see more with '--help')
  -V, --version                         Print version

The connection to PostgreSQL is established either through a TCP address (--pg-addr) or through a Unix socket directory (--pg-socket); exactly one of the two must be given.

Endpoints

The server exposes the following routes (see the API Reference for details):

RouteDescription
GET /Landing page (static HTML).
GET /v1/wiktionary/{edition}/words/{word}Exact word lookup, with optional lang_code / lang filters.
GET /v1/wiktionary/editionsLists all supported Wiktionary editions.
GET /v1/wiktionary/{edition}/wordsFuzzy word search (reserved; returns 501 Not Implemented).

Database

The server expects a PostgreSQL database (default name FreeDictionary) that has been initialised with load2db-go. For each Wiktionary edition it must contain a table named WiktionaryData_{edition} (e.g. WiktionaryData_en, WiktionaryData_zh) with at least the following columns:

ColumnDescription
wordThe headword / lemma.
langLocalised language name of the word (e.g. 汉语).
lang_codeWiktionary language code (e.g. en, zh).
dataThe full entry as JSON.

The server refuses to start if the table WiktionaryData_en is missing.

The connection pool is capped at --pg-pool-size connections (default 50).

Word Lookup Cache

Successful word-lookup responses are cached in memory to avoid repeated database round-trips. The cache is implemented with moka; entries are stored zstd-compressed and keyed by (edition, word, lang, lang_code).

OptionDefaultDescription
--cache-max-mb256Maximum total size of the cache in MiB; 0 disables caching.
--cache-ttl-secs3600Time-to-live for cache entries in seconds.

Memory Allocator

The global allocator is selectable at compile time through Cargo features:

FeatureDefaultAllocator
mimallocyesmimalloc
jemallocnojemalloc

With no flags the default feature set builds with mimalloc. To use the system allocator instead, disable default features:

cargo build --release --no-default-features

To use jemalloc:

cargo build --release --features jemalloc

If both features are enabled, jemalloc takes precedence. On platforms where a selected allocator is unsupported (for example jemalloc on Windows), the build silently falls back to the system allocator.

Logging

Logging is initialised with env_logger. The default log level is info (startup and per-request logs); pass --verbose to raise it to debug (which additionally prints per-request response details).

The RUST_LOG environment variable, when set, overrides this default entirely:

RUST_LOG=warn serve-dict --pg-addr localhost --pg-username user --pg-passwd secret

Benchmark

A Go load-testing tool is provided to measure the server’s throughput, latency distribution, and error rate under high concurrency.

Tip: When benchmarking remotely, use the Go client — it uses goroutines across all cores with minimal client-side overhead, making it suitable for finding the server’s true throughput ceiling.

Go (bench-go)

Requires Go 1.21+:

cd bench-go
go build -o bench .

Constant-concurrency mode (default: 100 concurrent connections for 30 s):

./bench --url http://localhost:8080
./bench --url http://localhost:8080 -c 200 -d 60

Ramp mode — steps concurrency from 10 to 1000 (step 50), holding each level for 30 s:

./bench --url http://localhost:8080 --mode ramp --ramp-from 10 --ramp-to 1000 --ramp-step 50 --ramp-duration 30

Useful options:

FlagDescription
--timeout <s>Per-request timeout in seconds (default 10). Requests exceeding this are counted as client-side timeout errors.
--json <path>Export a structured JSON report.
--words <file>Custom word list (one word per line).
--editions <e,...>Editions to test, comma-separated (default en,zh).
--no-compressDisable HTTP compression (send Accept-Encoding: identity).
--no-warmupSkip the warm-up phase.
--no-query-paramsDo not generate requests with lang_code query parameters.

The tool exercises both editions (en, zh), with and without lang_code query parameters, non-existent words (empty-result path), and unsupported editions (error path). A warm-up phase runs before measurement; pass --no-warmup to skip it.

Note: “errors” reported by the tools include client-side timeouts — the server may still return 200 for those requests after the client has given up. Use a larger --timeout to measure raw capacity, or a smaller one (e.g. --timeout 2) to measure effective capacity within an SLA budget.

Project Layout

FileDescription
src/main.rsServer bootstrap and route registration.
src/allocator.rsCompile-time global allocator selection.
src/cli.rsCommand-line argument parsing and PostgreSQL URL construction.
src/errors.rsError types and their mapping to HTTP status codes.
src/wiktionary.rsDatabase queries.
static/index.htmlLanding page served at /.
bench-go/Load-testing / benchmark tool (Go, zero dependencies).

License

Licensed under the GNU Affero General Public License v3.0.