# HanziSpace character-component data

`components.json` is a selected and transformed character-reference dataset for
the 3,684 Han characters used by HanziSpace. It is intentionally licensed
separately from `cards.json`.

## Unicode data

Primary Kangxi radical classifications, residual stroke counts, radical
mappings, official radical names, selected character definitions, and Mandarin
readings are derived from the Unicode Character Database and Unihan 17.0.0:

- https://www.unicode.org/Public/17.0.0/ucd/Unihan.zip
- https://www.unicode.org/Public/17.0.0/ucd/CJKRadicals.txt
- https://www.unicode.org/Public/17.0.0/ucd/UnicodeData.txt
- https://www.unicode.org/reports/tr38/

Copyright © 1991–2025 Unicode, Inc. Unicode data is used under the Unicode
Terms of Use: https://www.unicode.org/terms_of_use.html

## Make Me a Hanzi dictionary data

Character definitions, pinyin, structural decompositions, and formation fields
are selected from Make Me a Hanzi `dictionary.txt` at commit
`bddc96d41bef78427ed0e034e9f7e31d71fd1b92`:

https://github.com/skishore/makemeahanzi/tree/bddc96d41bef78427ed0e034e9f7e31d71fd1b92

The upstream `dictionary.txt` is available under the GNU Lesser General Public
License, version 3 or (at your option) any later version. The license and
upstream notices are available at:

https://github.com/skishore/makemeahanzi/blob/bddc96d41bef78427ed0e034e9f7e31d71fd1b92/LGPL

## Dong Chinese lexicon

Formation explanations and explicit meaning, sound, iconic, simplification, and
unknown component roles prefer the Chinese Lexicon built for Dong Chinese at
commit `de64ca4c5d3fef6694a1270f943726c5f622bb03`:

https://github.com/peterolson/chinese-lexicon/tree/de64ca4c5d3fef6694a1270f943726c5f622bb03

The repository declares the ISC license. The upstream Dong Chinese character
wiki also publishes its community dictionary under CC BY-SA 4.0. HanziSpace
records the exact repository revision used for reproducibility.

## HanziSpace modifications

HanziSpace selected only characters present in its HSK corpus, mapped each
character to the first (primary) Unicode `kRSUnicode` value, normalized
formation whitespace, preferred Dong Chinese component roles for 3,058
characters, retained Make Me a Hanzi formation data as a documented fallback,
derived visible common radical forms from decomposition data, connected
characters to course levels and word counts, and generated navigable category
pages. Alternate Unicode radical assignments remain in the machine-readable
data. The import logic is in
`scripts/import_character_components.mjs`.
