改进了用于查词的Hunspell拼写法文件(日文、英文、德文、中文繁简通查)

forked from MrCorn0-0/hunspell_ja_JP
删掉了来自汉直输入法词库中的人名地名的混书。人名地名一般不太可能会出现混ぜ書き。
增加了来自《三◯堂国语辞典》的语词的混书。大概有万条?
《词悦》的作者实现了两次调用 Hunspell,于是把 ja_JP.dic 写成这样:

引き籠もる/5A5I5T5E5OBDUYUBubKN
ひきこもる/5A5I5T5E5OBDUYUBubKN
引きこもる/5A5I5T5E5OBDUYUBubKN
引きこもる st:引き籠もる

和原来手动写出动词六活用形

引き籠もる/5A5I5T5E5OBDUYUBubKN
ひきこもる/5A5I5T5E5OBDUYUBubKN
引きこもっ st:引き籠もる
引きこもら st:引き籠もる
引きこもり st:引き籠もる
引きこもる st:引き籠もる
引きこもれ st:引き籠もる
引きこもろ st:引き籠もる

相比,可维护性更好。
9月26日补充:
原来Hunspell支持
引きこもる/5A5I5T5E5OBDUYUBubKN st:引き籠もる
这样写……对.dic全面修订中
英、德文的Hunspell拼写法文件来自

但是就我实际体验,英文有些重读闭音节转换有误(比如hoped的原形应该是hope,但是却也会显示hop),德文缺少某些可分动词的变化,故用DeepSeek Harness驱动DF4.1修正之。


AI总结的修改en_US拼写法文件的内容:

2025 morphology cleanup (GoldenDict-ng)

The D/G/R/T affixes in this file do not double a final consonant, so a base
carrying them generated a wrong surface form that collided with another lemma
(hop/D → hoped, while “hoped st:hope” is the correct link). All such tags were
removed from the base; the doubled forms stay explicit (hopped st:hop).

  • 212 base words lost 438 wrong D/G/R/T (and one Z) tags.
  • bar(red/ring), cub(bed/bing), gap(ped/ping), snip(ped/ping) got the missing
    st: links; halved/halving now point at halve instead of half.
  • sky lost R/Z, ski gained R.
  • Derivational tags were added to 2001 words, only where the derived form is
    already a headword (nation/a → national, hope/f → hopeful, …).
  • New suffixes in en_US.aff: f -ful, l -less, o -ous, i -ish.
  • The new -ic flag is “Q”, NOT “c”: “ONLYINCOMPOUND c” already reserves lower
    case c, which silently removed academy/acid/artist from the dictionary.
  • Added “googleable st:google” (google/B only generated “googlable”).

See REPORT.md and CHANGES.tsv for details.

2025 ODE merge (GoldenDict-ng)

The Oxford Dictionary of English headword list (ODE.csv, 160 984 entries) was
merged into this dictionary. Derived words are not listed when a chain of at
most two suffixes derives them from a base that is itself an ODE word of at
least four letters (Hunspell’s two-affix limit); the flag goes on the base.

  • 16 941 ODE words were compressed away, e.g.
    reconstruct/V covers reconstructive, reconstructively, reconstructiveness
    reconstruct/j covers reconstruction, reconstructional, reconstructionary,
    reconstructionist
    reconstruct/B/k/v cover reconstructable / reconstructible / reconstructor
  • 104 229 new headwords were added (including 3 508 hyphen parts such as
    “headedness” needed by compounds like “hard-headedness”).
  • All 159 744 addable ODE words are accepted by hunspell; 0 regressions against
    the previous dictionary.
  • New affixes: j -ion, k -ible, v -or, e -ist, r -ism, z -ize, x -ise,
    s -ship, b -hood, d -dom, w -ward, y -ify, g -ate, h -ary, O non-, W anti-.
  • WORDCHARS now contains ‘-’ and ‘.’, so hyphenated entries and abbreviations
    such as “a.m.” are matched as single words.
  • Compression only uses suffixes; prefixed ODE words (reconstruct, nonabrasive,
    …) stay headwords so that they can be the base of their own derivatives.

See ODE_compression.tsv for the full compressed list and REPORT.md for details.

支持Hunspell的辞典程序:

另外:
Hunspell还能用来实现中文的简体繁体通查。

zh_CN.aff 不需要任何实质内容,指定编码为 UTF‑8 就行:

SET UTF-8
LANG zh
FLAG long
# 空壳

zh_CN.dic 的内容类似这样,第一行是词条数目,然后每行都是st:(本来是hunspell用来标记不规则动词原形的doge):

12
簡體 st:简体
简体 st:簡體
繁體 st:繁体
繁体 st:繁體
发 st:髮
发 st:發
髮 st:发
發 st:发
復制 st:复制
複製 st:复制
复制 st:復制
复制 st:複製

在GoldenDict-ng的某本词典上按右键,然后选择「词典词条」,再选择「导出」。



如果导出词条是传统汉字或者传统和简化写法并存的,那就可以直接用,如果是纯简化字,用OpenCC转换成传统字之后手动校对一对多的汉字。
然后运行以下脚本:

#!/usr/bin/env python3
# -*- coding: utf-8 -*-
"""
gen_zh_dic.py

讀取一行一個詞的詞彙表,生成 Hunspell .dic,用於 Goldendict 簡繁通查。

依賴:
    pip install opencc-python-reimplemented

用法:
    python3 gen_zh_dic.py words.txt zh_CN.dic
"""

import sys
import os
from collections import OrderedDict, defaultdict
from opencc import OpenCC

# OpenCC t2s 漏掉的一簡對多繁異體字,逐字手動補齊。
CHAR_FIX = {
    "懽": "欢",
    "歓": "欢",
    "歡": "欢",
    #"讙": "欢",
}


def make_normalizer(cc):
    def normalize(word: str) -> str:
        s = cc.convert(word).strip()
        if not s:
            return s
        return "".join(CHAR_FIX.get(ch, ch) for ch in s)
    return normalize


def main():
    if len(sys.argv) != 3:
        print("用法: python3 gen_zh_dic.py 詞彙表.txt 輸出.dic")
        sys.exit(1)

    in_path = sys.argv[1]
    out_path = sys.argv[2]

    # 繁體 -> 簡體,作為規範形式
    cc = OpenCC("t2s")
    normalize = make_normalizer(cc)

    words = []
    seen = set()

    with open(in_path, "r", encoding="utf-8") as f:
        for line in f:
            w = line.strip()
            if not w or w.startswith("#"):
                continue
            if w not in seen:
                seen.add(w)
                words.append(w)

    std_of = OrderedDict()        # 原詞 -> 規範形式
    variants = defaultdict(list)  # 規範形式 -> [原詞, ...]

    for w in words:
        #std = cc.convert(w).strip()
        std = normalize(w)
        if not std:
            continue
        std_of[w] = std
        if w not in variants[std]:
            variants[std].append(w)

    lines = []
    emitted_std = set()

    for w in words:
        std = std_of.get(w)
        if std is None or std in emitted_std:
            continue
        emitted_std.add(std)

        others = [v for v in variants[std] if v != std]
        if not others:
            # 繁簡相同、無變體,對 Goldendict 詞幹還原無用,跳過
            continue

        for v in others:
            lines.append(f"{v} st:{std}")
        for v in others:
            lines.append(f"{std} st:{v}")

    with open(out_path, "w", encoding="utf-8", newline="\n") as f:
        f.write(str(len(lines)) + "\n")
        if lines:
            f.write("\n".join(lines) + "\n")

    # 順便生成空殼 .aff;如果不想自動生成,可刪掉下面這段
    aff_path = os.path.splitext(out_path)[0] + ".aff"
    with open(aff_path, "w", encoding="utf-8", newline="\n") as f:
        f.write("SET UTF-8\n")
        f.write("LANG zh\n")
        f.write("FLAG long\n")

    print(f"已生成: {out_path}")
    print(f"已生成: {aff_path}")


if __name__ == "__main__":
    main()

这比GoldenDict直接调用opencc没准还好,因为OpenCC的一个输入只能对应一个输出,而hunspell是可以一个输入对应多个输出的。挂载Hunspell之后,輸入「复制」,GoldenDict能同时查到「復制」和「複製」。


de_DE_1.2.zip (1.1 MB)
en_US_1.1.zip (760.9 KB)

用汉语大词典构建的dic有40w条,8M大小。
抛开体积不谈,这种做法完全抛弃了模糊匹配,正常他应该匹配「木红」开头的词,但hunspell做不到,他只能完全匹配


hunspell的拼写建议是按照西文来的,所以他会觉得「红木」跟「木红」更接近啊。类似西文中正好把两个字母按反了,搜preisdent他会给你建议president。对中日文来说,他这个拼写建议确实比较鸡肋。

我还在想中文有没有Hunspell呢,网络上确实没有,好像只有你做了?哈哈

嗯,网上应该没有。我是想到似乎可行,试了一下还真的可行。不过由于中文不分词,不同辞典生成的zh_CN.dic差距巨大,比如《大辞海》就有大量像「奧伯斯佯謬 st:奥伯斯佯谬」这样的条目,导致zh_CN.dic无法像西文那样一个hunspell.dic适配绝大多数辞典,最好是根据自己实际使用的辞典条目生成。所以我就只放了AI生成的python代码没有放.dic本身。PC上最好跟GoldenDict-ng自带的OpenCC一起使用,hunspell.dic只负责一简对多繁。

查询supersized没查到supersize,遂打开Codebuddy蹬了一轮腾讯的免费智能体,把德文和英文的Hunspell文件修正了一番。已上传到主楼。