forked from MrCorn0-0/hunspell_ja_JP
删掉了来自汉直输入法词库中的人名地名的混书。人名地名一般不太可能会出现混ぜ書き。
增加了来自《三◯堂国语辞典》的语词的混书。大概有万条?
《词悦》的作者实现了两次调用 Hunspell,于是把 ja_JP.dic 写成这样:
引き籠もる/5A5I5T5E5OBDUYUBubKN
ひきこもる/5A5I5T5E5OBDUYUBubKN
引きこもる/5A5I5T5E5OBDUYUBubKN
引きこもる st:引き籠もる
和原来手动写出动词六活用形
引き籠もる/5A5I5T5E5OBDUYUBubKN
ひきこもる/5A5I5T5E5OBDUYUBubKN
引きこもっ st:引き籠もる
引きこもら st:引き籠もる
引きこもり st:引き籠もる
引きこもる st:引き籠もる
引きこもれ st:引き籠もる
引きこもろ st:引き籠もる
相比,可维护性更好。
9月26日补充:
原来Hunspell支持
引きこもる/5A5I5T5E5OBDUYUBubKN st:引き籠もる
这样写……对.dic全面修订中
英、德文的Hunspell拼写法文件来自
但是就我实际体验,英文有些重读闭音节转换有误(比如hoped的原形应该是hope,但是却也会显示hop),德文缺少某些可分动词的变化,故用DeepSeek Harness驱动DF4.1修正之。
AI总结的修改en_US拼写法文件的内容:
2025 morphology cleanup (GoldenDict-ng)
The D/G/R/T affixes in this file do not double a final consonant, so a base
carrying them generated a wrong surface form that collided with another lemma
(hop/D → hoped, while “hoped st:hope” is the correct link). All such tags were
removed from the base; the doubled forms stay explicit (hopped st:hop).
- 212 base words lost 438 wrong D/G/R/T (and one Z) tags.
- bar(red/ring), cub(bed/bing), gap(ped/ping), snip(ped/ping) got the missing
st: links; halved/halving now point at halve instead of half.- sky lost R/Z, ski gained R.
- Derivational tags were added to 2001 words, only where the derived form is
already a headword (nation/a → national, hope/f → hopeful, …).- New suffixes in en_US.aff: f -ful, l -less, o -ous, i -ish.
- The new -ic flag is “Q”, NOT “c”: “ONLYINCOMPOUND c” already reserves lower
case c, which silently removed academy/acid/artist from the dictionary.- Added “googleable st:google” (google/B only generated “googlable”).
See REPORT.md and CHANGES.tsv for details.
2025 ODE merge (GoldenDict-ng)
The Oxford Dictionary of English headword list (ODE.csv, 160 984 entries) was
merged into this dictionary. Derived words are not listed when a chain of at
most two suffixes derives them from a base that is itself an ODE word of at
least four letters (Hunspell’s two-affix limit); the flag goes on the base.
- 16 941 ODE words were compressed away, e.g.
reconstruct/V covers reconstructive, reconstructively, reconstructiveness
reconstruct/j covers reconstruction, reconstructional, reconstructionary,
reconstructionist
reconstruct/B/k/v cover reconstructable / reconstructible / reconstructor- 104 229 new headwords were added (including 3 508 hyphen parts such as
“headedness” needed by compounds like “hard-headedness”).- All 159 744 addable ODE words are accepted by hunspell; 0 regressions against
the previous dictionary.- New affixes: j -ion, k -ible, v -or, e -ist, r -ism, z -ize, x -ise,
s -ship, b -hood, d -dom, w -ward, y -ify, g -ate, h -ary, O non-, W anti-.- WORDCHARS now contains ‘-’ and ‘.’, so hyphenated entries and abbreviations
such as “a.m.” are matched as single words.- Compression only uses suffixes; prefixed ODE words (reconstruct, nonabrasive,
…) stay headwords so that they can be the base of their own derivatives.See ODE_compression.tsv for the full compressed list and REPORT.md for details.
支持Hunspell的辞典程序:
另外:
Hunspell还能用来实现中文的简体繁体通查。
zh_CN.aff 不需要任何实质内容,指定编码为 UTF‑8 就行:
SET UTF-8
LANG zh
FLAG long
# 空壳
zh_CN.dic 的内容类似这样,第一行是词条数目,然后每行都是st:(本来是hunspell用来标记不规则动词原形的doge):
12
簡體 st:简体
简体 st:簡體
繁體 st:繁体
繁体 st:繁體
发 st:髮
发 st:發
髮 st:发
發 st:发
復制 st:复制
複製 st:复制
复制 st:復制
复制 st:複製
在GoldenDict-ng的某本词典上按右键,然后选择「词典词条」,再选择「导出」。
如果导出词条是传统汉字或者传统和简化写法并存的,那就可以直接用,如果是纯简化字,用OpenCC转换成传统字之后手动校对一对多的汉字。
然后运行以下脚本:
#!/usr/bin/env python3
# -*- coding: utf-8 -*-
"""
gen_zh_dic.py
讀取一行一個詞的詞彙表,生成 Hunspell .dic,用於 Goldendict 簡繁通查。
依賴:
pip install opencc-python-reimplemented
用法:
python3 gen_zh_dic.py words.txt zh_CN.dic
"""
import sys
import os
from collections import OrderedDict, defaultdict
from opencc import OpenCC
# OpenCC t2s 漏掉的一簡對多繁異體字,逐字手動補齊。
CHAR_FIX = {
"懽": "欢",
"歓": "欢",
"歡": "欢",
#"讙": "欢",
}
def make_normalizer(cc):
def normalize(word: str) -> str:
s = cc.convert(word).strip()
if not s:
return s
return "".join(CHAR_FIX.get(ch, ch) for ch in s)
return normalize
def main():
if len(sys.argv) != 3:
print("用法: python3 gen_zh_dic.py 詞彙表.txt 輸出.dic")
sys.exit(1)
in_path = sys.argv[1]
out_path = sys.argv[2]
# 繁體 -> 簡體,作為規範形式
cc = OpenCC("t2s")
normalize = make_normalizer(cc)
words = []
seen = set()
with open(in_path, "r", encoding="utf-8") as f:
for line in f:
w = line.strip()
if not w or w.startswith("#"):
continue
if w not in seen:
seen.add(w)
words.append(w)
std_of = OrderedDict() # 原詞 -> 規範形式
variants = defaultdict(list) # 規範形式 -> [原詞, ...]
for w in words:
#std = cc.convert(w).strip()
std = normalize(w)
if not std:
continue
std_of[w] = std
if w not in variants[std]:
variants[std].append(w)
lines = []
emitted_std = set()
for w in words:
std = std_of.get(w)
if std is None or std in emitted_std:
continue
emitted_std.add(std)
others = [v for v in variants[std] if v != std]
if not others:
# 繁簡相同、無變體,對 Goldendict 詞幹還原無用,跳過
continue
for v in others:
lines.append(f"{v} st:{std}")
for v in others:
lines.append(f"{std} st:{v}")
with open(out_path, "w", encoding="utf-8", newline="\n") as f:
f.write(str(len(lines)) + "\n")
if lines:
f.write("\n".join(lines) + "\n")
# 順便生成空殼 .aff;如果不想自動生成,可刪掉下面這段
aff_path = os.path.splitext(out_path)[0] + ".aff"
with open(aff_path, "w", encoding="utf-8", newline="\n") as f:
f.write("SET UTF-8\n")
f.write("LANG zh\n")
f.write("FLAG long\n")
print(f"已生成: {out_path}")
print(f"已生成: {aff_path}")
if __name__ == "__main__":
main()
这比GoldenDict直接调用opencc没准还好,因为OpenCC的一个输入只能对应一个输出,而hunspell是可以一个输入对应多个输出的。挂载Hunspell之后,輸入「复制」,GoldenDict能同时查到「復制」和「複製」。
de_DE_1.2.zip (1.1 MB)
en_US_1.1.zip (760.9 KB)



