About
How this dictionary is built
粵拼 Jyutping (jyutpi.ng) merges open data about the Chinese writing system into one Cantonese-first reference: every CJK character with Jyutping and Pinyin readings, definitions in three languages, structural decomposition, stroke order, and example sentences. Inspired by CUHK Lexis-Cantonese.
By the numbers
Sources & licenses
character backbone: codepoints, strokes, radicals, definitions, variants
Jyutping readings for characters and words, reading weights
Mandarin words: pinyin, English definitions, simplified forms, classifiers
Cantonese-specific words and readings
character decomposition (⿰⿱… structures) and components
stroke-order diagrams (derived from Arphic Technology's KaitiM GB and UKai fonts)
example sentences (Cantonese and Mandarin, with English translations)
corpus word frequencies for ranking
Cantonese definitions, English glosses, curated example sentences, attested Jyutping readings
Notices
- Every entry lists which sources contributed to it. Entry data derived from CC-CEDICT and CC-Canto is itself available under the same CC BY-SA licenses.
- Example sentences are contributed by individual Tatoeba users (CC BY 2.0 FR); each example links to its source sentence and credits its author.
- Stroke-order diagrams are derived from Arphic Technology's KaitiM GB and UKai fonts via Make Me a Hanzi, under the Arphic Public License.
- Unihan data © 1991–2026 Unicode, Inc., used under the Unicode License. Jyutping data by the 粵語計算語言學基礎建設組 (rime-cantonese / LSHK romanization).
- Word frequencies come from wordfreq, whose data (CC BY-SA 4.0) incorporates SUBTLEX, Google Books Ngrams, OpenSubtitles, Wikipedia, and other corpora.
- Dictionary content from words.hk (粵典) is © 2015–2025 Hong Kong Lexicography Limited, used under the NC Open Data License 1.0; the credits list and full terms are on that page. Only entries words.hk has published are included, and every entry credits words.hk as its source.
Methodology
- One entry per traditional form
- Sources disagree about word identity, so entries are merged by their traditional spelling: a word's Cantonese readings, Mandarin readings, and definitions from every source live on one page.
- Readings marked inferred
- When no source attests a word's pronunciation, one is synthesized from each character's most common reading and clearly flagged. Polyphonic characters can make an inferred reading wrong — treat it as a starting point, not gospel.
- Frequency ranking
- Search results are ordered by a blend of real corpus frequency (wordfreq) and structural signals (character commonness, word length, number of attesting sources), so common words surface first even when a query matches thousands of entries.
- Example sentences
- Tatoeba sentences are matched to words by containment, Cantonese first, capped per word. They are crowd-sourced — natural, but not editorially reviewed.