Data sources and licensing

Sources, methodology, and licenses

Vocabulary source

The vocabulary data is adapted from Hanping Chinese HSK (1–6) by EmberMitre Ltd, licensed under the Creative Commons Attribution-ShareAlike 4.0 International license. The source deck states that its definitions are taken from CC-CEDICT, a Creative Commons Attribution-ShareAlike dictionary.

What HanziSpace changed

We normalized source fields, converted numbered pinyin to tone marks, generated stable identifiers and URLs, separated definitions for display, and applied documented reviewed corrections. The original creators do not endorse HanziSpace.

Mandarin pronunciation audio

Pronunciations use self-hosted, native-speaker recordings from the audio-cmn collection, not a browser text-to-speech voice. Most pages use a complete word recording by Yue Tan, a speaker from Liaoning, China, originally recorded for the Shtooka Project and licensed under CC BY-SA 3.0 US. Where no complete word recording exists, HanziSpace combines the longest available recorded subwords with tone-matched recordings from the collection's Chen Wang syllable set according to the card's pinyin. Those fallback files were selected, combined, and re-encoded by HanziSpace and remain under the source collection's CC BY-SA terms.

Character writing data

Interactive stroke-order animations and tracing use Hanzi Writer, released under the MIT License. Its character data is derived from Make Me a Hanzi and Arphic Technology font data and is redistributed under the Arphic Public License.

Character components and radical families

The complete 214-category index, each character’s primary radical, alternate radical assignments, and residual stroke counts come from Unicode Standard Annex #38 and the versioned Unicode 17.0.0 Unihan data. Official radical mappings and names come from Unicode’s versioned CJKRadicals.txt and UnicodeData.txt.

Character definitions, pinyin, and structural decompositions are selected from Make Me a Hanzi at pinned commit bddc96d, whose dictionary data is available under LGPL-3.0-or-later. Formation explanations prefer the explicit semantic, phonetic, pictographic, and simplification roles from the open Dong Chinese Lexicon at pinned commit de64ca4; this cross-check covers 3,058 of the 3,684 course characters, with Make Me a Hanzi retained as the documented fallback.

HanziSpace treats the first kRSUnicode value as the primary category, preserves alternate assignments in the published data, and links every record to local HSK words. Radical classification is presented separately from semantic and phonetic roles. Those roles are shown only when identified by the cited formation source; historical pronunciation and modern meaning may differ from their earlier forms. See the component data notices and modification record.

Reuse

The adapted card dataset at /assets/cards.json is available under CC BY-SA 4.0. You may copy, adapt, and use it commercially if you provide attribution, link to the license, indicate your changes, and distribute adaptations under the same or a compatible license.

This data license does not cover third-party marks, pronunciation audio, character-writing data, or the HanziSpace site code and design; those retain their respective terms described above.