# 有没有什么简明的 JSON 词典模板？

**URL:** <https://forum.freemdict.com/t/topic/41990>\
**Category:** 技术交流与词典编修\
**Created:** [2025 年11 月 2 日 16:11 UTC](https://forum.freemdict.com/t/topic/41990 "2025-11-02T16:11:44Z")\
**Posts on this page:** 19\
**Page:** 1

<div class="post-metadata">

作者： ![last\_idol](https://forumcdn.freemdict.com/user_avatar/forum.freemdict.com/last_idol/32/2205_2.png) [@last\_idol](https://forum.freemdict.com/u/last_idol)\
发布日期： [2025 年11 月 2 日 16:11 UTC](https://forum.freemdict.com/t/topic/41990/1 "2025-11-02T16:11:45Z")

</div>

想找一套简明的 JSON 模板能覆盖英语、汉语、日语词典。

找了牛津、韦氏的 JSON 样例还有朗文的 HTML，看起来太复杂了，互相没法兼容，可能需要抽出公共新的结构来，只需要覆盖各大词典里常见的元素就可以了，比如子词条，词性，标签，义项，例句。

---

<div class="post-metadata">

作者： ![last\_idol](https://forumcdn.freemdict.com/user_avatar/forum.freemdict.com/last_idol/32/2205_2.png) [@last\_idol](https://forum.freemdict.com/u/last_idol)\
发布日期： [2025 年11 月 2 日 16:13 UTC](https://forum.freemdict.com/t/topic/41990/2 "2025-11-02T16:13:27Z")

</div>

还有索引也是，辨析词典多个词头怎么搞。日语词典的索引要怎么设计，也没想好。

---

<div class="post-metadata">

作者： ![last\_idol](https://forumcdn.freemdict.com/user_avatar/forum.freemdict.com/last_idol/32/2205_2.png) [@last\_idol](https://forum.freemdict.com/u/last_idol)\
发布日期： [2025 年11 月 3 日 07:10 UTC](https://forum.freemdict.com/t/topic/41990/4 "2025-11-03T07:10:55Z")

</div>

我需要通用的 JSON 模板，到时喂给 AI ，帮我挑选最恰当的义项、例句和跳转。

---

<div class="post-metadata">

作者： ![First\_Last](https://forumcdn.freemdict.com/user_avatar/forum.freemdict.com/first_last/32/5_2.png) [@First\_Last](https://forum.freemdict.com/u/First_Last)\
发布日期： [2025 年11 月 3 日 08:38 UTC](https://forum.freemdict.com/t/topic/41990/5 "2025-11-03T08:38:44Z")

</div>

有看过 Yomitan 吗？

---

<div class="post-metadata">

作者： ![wynick27](https://forumcdn.freemdict.com/letter_avatar_proxy/v4/letter/w/958977/32.png) [@wynick27](https://forum.freemdict.com/u/wynick27)\
发布日期： [2025 年11 月 3 日 08:59 UTC](https://forum.freemdict.com/t/topic/41990/6 "2025-11-03T08:59:32Z")

</div>

yomitan不是真正结构化的，旧版有少量几个字段，其他基本照搬html标签。

---

<div class="post-metadata">

作者： ![mixivivo](https://forumcdn.freemdict.com/letter_avatar_proxy/v4/letter/m/8491ac/32.png) [@mixivivo](https://forum.freemdict.com/u/mixivivo)\
发布日期： [2025 年11 月 3 日 09:46 UTC](https://forum.freemdict.com/t/topic/41990/7 "2025-11-03T09:46:03Z")

</div>

词典没有通用模板。一个词汇，它有三项本质构成，形、音、义。形和音比较固定简短，但释义部分，等于是一段或者一篇文章了，想怎么写就怎么写。不过通常释义部分还是有一定固定结构的，分词性（中文可以没有）、义项、例句这几个元素。此外的东西，像同义反义、衍生词、短语、词源、辨析等，都可以根据体例随意增减。

要最大的兼容性，就只分词头和释义，mdx那样。稍微扩展，就是词头、发音、〈词性〉、义项、例句、其他这几类。继续细分，就花样百出，没有标准了。

---

<div class="post-metadata">

作者： ![last\_idol](https://forumcdn.freemdict.com/user_avatar/forum.freemdict.com/last_idol/32/2205_2.png) [@last\_idol](https://forum.freemdict.com/u/last_idol)\
发布日期： [2025 年11 月 3 日 12:03 UTC](https://forum.freemdict.com/t/topic/41990/8 "2025-11-03T12:03:08Z")

</div>

Yomitan 原本设计的很好，是完全结构化的 JSON + 标签化的系统，以方便统一检索和样式化，但词典作者制作的时候完全没按这个思路来。

---

<div class="post-metadata">

作者： ![amob](https://forumcdn.freemdict.com/user_avatar/forum.freemdict.com/amob/32/14630_2.png) [@amob](https://forum.freemdict.com/u/amob)\
发布日期： [2025 年11 月 3 日 12:14 UTC](https://forum.freemdict.com/t/topic/41990/9 "2025-11-03T12:14:28Z")

</div>

参考xml规范自己创造一个，xml改造成json应该是可以的。使用xml的具体实例有kindle和苹果，改成json用的就是yomichan嘛。

> [@E-dictionaries XML 规范](https://forum.freemdict.com/t/topic/33556):
>
> 有志于开发自己格式或者对mdx的HTML处理精益求精者，不妨看看世界上的专家们制订的规范。 [Interchange format for e-dictionaries.pdf](https://forum.freemdict.com/uploads/short-url/4qgANTTY2lUUGkd9bnxE7jvu52C.pdf) (4.6 MB) [LeXML310.pdf](https://forum.freemdict.com/uploads/short-url/io0hKktLu8r6f5ILPzKlPNemGZi.pdf) (590.4 KB) [GBT238292009.pdf](https://forum.freemdict.com/uploads/short-url/uUpbwvV3EA2OeWxJ0q1QjevukPa.pdf) (2.7 MB)

---

<div class="post-metadata">

作者： ![last\_idol](https://forumcdn.freemdict.com/user_avatar/forum.freemdict.com/last_idol/32/2205_2.png) [@last\_idol](https://forum.freemdict.com/u/last_idol)\
发布日期： [2025 年11 月 3 日 15:45 UTC](https://forum.freemdict.com/t/topic/41990/10 "2025-11-03T15:45:46Z")

</div>

看过 LeXML 对日语词头的处理有点疑问，不知道为什么核心词头选择的是假名形式，而汉字表记形这里只是用作功能词头，我觉得反过来会更正常一些，P170：

```auto
<dic-item id="ABCD00000100">
<head>c
<headword>みだし・ご</headword>
<key>みだしご</key>
<headword type="表記">見出し語</headword>
<key type="表記">見出し語</key>
</head>
<meaning>語義語釈、解説など</meaning>
<example>用例</example>
<subheadword type="子見出し">子[小]見出し（派生語、複合語、成句な
ど）</subheadword>
<key type="子見出しかな">こみだし</key>
<key type="子見出し表記">子見出し</key>
<key type="子見出し表記">小見出し</key>
</dic-item>

```

---

<div class="post-metadata">

作者： ![seid](https://forumcdn.freemdict.com/letter_avatar_proxy/v4/letter/s/b9bd4f/32.png) [@seid](https://forum.freemdict.com/u/seid)\
发布日期： [2025 年11 月 3 日 19:56 UTC](https://forum.freemdict.com/t/topic/41990/11 "2025-11-03T19:56:33Z")

</div>

不知是否理解你的要求，我之前转换 “牛津现代英汉双解词典第9版” 时就发现它的XML标签非常规整，你可以导出mdx的内容看看，这是我整理的标签说明。

```
'a' #超链接，类似这样<a href="entry://pants">pants</a>
'also' #源文件中只有这么一种形式：<also><cn>亦称作</cn></also>
'apd' #Appendix，附录，额外的释义，一般被apx包裹
'apx' #Appendix，附录
'c' #词条单词
'cn' #一般被zh包裹，中文释义
'def' #一个释义块
'div' #div
'eg' #具体的某一个例子
'en' #英语释义
'eps' #Example sentence, 例句，标记具体的示例句
'epx' #Expanded, 扩展，用于标记扩展的解释
'ex' #Example，例子区域，可能包含多个例子 ex > exp > eps > eg > en|zh
'exp' #例子区域
'gg' #Gloss, 解释，通常是某个词或短语的解释或注释
'gra' #Grammatical notation
'gs' #Glossary, 词汇表，用于表示词汇注释或解释
'h' #Highlight，高亮，标记特别需要强调的词或短语
'hwd' #Headword, 词头，标记词条的主词，包含c/sup等
'link' #css链接
'nvl' #Novel, 新，可能用于表示新的释义或新词
'ori' #Origin, 词源，用于显示词的来源或历史，里面可能包含a
'orx' #Origin explanation, 词源解释，与词源相关
'ot' #Other，标记词汇的其他类别或语境
'pd' #短语，一般里面包含一个h
'phd' #Phrases & Derivatives，phd > phr > pd > 
'pho' #Phonetics, 语音，表示音标的具体部分
'phr' #Phrase, 短语，用于标记短语或固定搭配
'phx' #Phonetic transcription, 单词发音区块，可能包含pho,pos,ot,gg等发音标签
'pos' #Part of speech，词性
'prx' #Prefix, 前缀，标记前缀部分或在词义前的额外信息
'rw' #row
'sa' #Sense area, 义项区，标记一个义项部分
'sb' #Subsection, 小节，用于标记词义或说明的分段，包含sc
'sc' #Subclassification, 子分类，标记词义下的子分类
'sqa' #Sub-qualifier,子限定符，标记更详细的限定部分，如例子中的a、b、c等分支义项
'sqn' #Sequence number, 序列号，标记词义的编号
'sub' #Subscript，下标
'sup' #Superscript，上标
'tm' #Term, 术语，标记特定的术语
'uge' #Usage, 用法，标记词语的用法说明
'uit' #Usage in text, 文本用法，标记某个词的用法解释
'x' #Example or additional info, 例子或附加信息，标记示例或补充信息
'xr' #Cross reference, 交叉引用，用于链接到相关词条，一般里面包含一个a
'zh' #Chinese translation 中文翻译

```

---

<div class="post-metadata">

作者： ![Vim](https://forumcdn.freemdict.com/user_avatar/forum.freemdict.com/vim/32/8408_2.png) [@Vim](https://forum.freemdict.com/u/Vim)\
发布日期： [2025 年11 月 4 日 00:41 UTC](https://forum.freemdict.com/t/topic/41990/12 "2025-11-04T00:41:32Z")

</div>

找出典型词典案例，让AI提炼一个？

---

<div class="post-metadata">

作者： ![last\_idol](https://forumcdn.freemdict.com/user_avatar/forum.freemdict.com/last_idol/32/2205_2.png) [@last\_idol](https://forum.freemdict.com/u/last_idol)\
发布日期： [2025 年11 月 4 日 03:05 UTC](https://forum.freemdict.com/t/topic/41990/14 "2025-11-04T03:05:07Z")

</div>

喂给 AI 的数据，通常需要结构化、扁平化的处理，JSON 更简洁、无冗余标签，所以工程上更常用，而 XML 读写都很困难。

这里有一份韦氏大学的 JSON 说明文档，可以参考下：

[https://dictionaryapi.com/products/json](https://dictionaryapi.com/products/json)

样例：（实际会更复杂了，还需要精简

```auto
"hom":1,
"hwi":{
  "hw":"ba*lo*ney",
  "prs":[
    {
      "mw":"b\u0259-\u02c8l\u014d-n\u0113",
      "sound":{"audio":"bologn01","ref":"c","stat":"1"}
    }
  ]
},
"fl":"noun",
"cxs":[
  {
    "cxl":"less common spelling of",
    "cxtis":[
      {"cxt":"bologna"}
    ]
  }
],
"def":[
  {
    "sseq":[
      [
        [
          "sense",
          {
            "dt":[
              ["text","{bc}a large smoked sausage of beef, veal, and pork"]
            ],
            "sdsense":{
              "sd":"also",
              "dt":[
                ["text","{bc}a sausage made (as of turkey) to resemble bologna"]
              ]
            }
          }
        ]
      ]
    ]
  }
]

```

---

<div class="post-metadata">

作者： ![last\_idol](https://forumcdn.freemdict.com/user_avatar/forum.freemdict.com/last_idol/32/2205_2.png) [@last\_idol](https://forum.freemdict.com/u/last_idol)\
发布日期： [2025 年11 月 4 日 03:48 UTC](https://forum.freemdict.com/t/topic/41990/15 "2025-11-04T03:48:58Z")

</div>

请各位再看下这份词头的设计是否合适：（我觉得汉语和日语的很别扭，不知道怎么搞

英语1：

```auto
"lang": "en-GB",
"title": "take",
"headwords": [
    {"hw": "take", "values": ["takes", "taking"]},
]

```

英语2：

```auto
"lang": "en-GB",
"title": "take over",
"headwords": [
    {"hw": "take over"},
]

```

汉语1：

```auto
"lang": "zh-CN",
"title": " 词典",
"headwords": [
    {"hw": "词典", "values": ["詞典"]},
    {"hw": "cí diǎn", "values": ["ci2 dian3"], "ref": "词典", "type": "pinyin"}
]

```

日语1:

```auto
"lang": "ja-JP",
"title": "とり‐け・す【取（り）消す】",
"headwords": [
    {"hw": "取り消す", "values": ["取消す"]},
    {"hw": "とりけす", "values": [], "ref": "取り消す", "type": "reading"},
]

```

---

<div class="post-metadata">

作者： ![amob](https://forumcdn.freemdict.com/user_avatar/forum.freemdict.com/amob/32/14630_2.png) [@amob](https://forum.freemdict.com/u/amob)\
发布日期： [2025 年11 月 4 日 04:07 UTC](https://forum.freemdict.com/t/topic/41990/16 "2025-11-04T04:07:26Z")

</div>

给AI的话好像还是太复杂了，AI根本不听话。但离便于使用还差远了。

个人认为只需要title和variants列表并列，variants列表里想列几个都可以。别区分那么多类型，AI太累了。以及，个人认为ID设计是必须的，不过AI不适合加，可以通过程序批量加。

---

<div class="post-metadata">

作者： ![seid](https://forumcdn.freemdict.com/letter_avatar_proxy/v4/letter/s/b9bd4f/32.png) [@seid](https://forum.freemdict.com/u/seid)\
发布日期： [2025 年11 月 4 日 08:58 UTC](https://forum.freemdict.com/t/topic/41990/17 "2025-11-04T08:58:35Z")

</div>

牛津韦氏等词典官网API可以返回json格式，注册一个免费账号，调用几次API就可以通过其格式改进自己的模型

---

<div class="post-metadata">

作者： ![endnote](https://forumcdn.freemdict.com/letter_avatar_proxy/v4/letter/e/9de0a6/32.png) [@endnote](https://forum.freemdict.com/u/endnote)\
发布日期： [2025 年11 月 4 日 14:57 UTC](https://forum.freemdict.com/t/topic/41990/18 "2025-11-04T14:57:56Z")

</div>

> [@last\_idol](#):
>
> 喂给 AI 的数据，通常需要结构化、扁平化的处理，JSON 更简洁、无冗余标签，所以工程上更常用

确实。

最近看到不少人建议，prompt要使用JSON格式，这样AI的回答比较结构化且稳定、减少随机的风格变换

---

<div class="post-metadata">

作者： ![First\_Last](https://forumcdn.freemdict.com/user_avatar/forum.freemdict.com/first_last/32/5_2.png) [@First\_Last](https://forum.freemdict.com/u/First_Last)\
发布日期： [2025 年11 月 5 日 10:18 UTC](https://forum.freemdict.com/t/topic/41990/19 "2025-11-05T10:18:35Z")

</div>

我也记得是这样。  
楼主只要模版，  
还是可以参考的。

> <https://github.com/yomidevs/yomitan/blob/master/docs/making-yomitan-dictionaries.md#read-the-schemas>

---

<div class="post-metadata">

作者： ![jdiary](https://forumcdn.freemdict.com/letter_avatar_proxy/v4/letter/j/ad7895/32.png) [@jdiary](https://forum.freemdict.com/u/jdiary)\
发布日期： [2025 年11 月 8 日 11:26 UTC](https://forum.freemdict.com/t/topic/41990/20 "2025-11-08T11:26:17Z")

</div>

个人觉得这类“大一统”的方式不现实。大一统的结果是抹杀特征，但一本词典恰恰是有用的特征让它更有价值。

但是呢，小规模，不是为了统一的取较大公约数还是可以做一些有用的事情的。比如，有一个为 外汉双解词典设计的JSON结构，那么比如：  
基于一本【0】法汉双解词典，可以做出【1】法法词典 （用于培养阅读法文解释的习惯）【2】法汉词典（用于快速获得词义）【3】逆向的汉法词典【4】通过中文某个释义为连接点构成的法语同义词词典。  
然后重点来了，这些都并不需要让数据重复，而只是词典软件通过算法便能基于一份数据自动构建。

---

<div class="post-metadata">

作者： ![last\_idol](https://forumcdn.freemdict.com/user_avatar/forum.freemdict.com/last_idol/32/2205_2.png) [@last\_idol](https://forum.freemdict.com/u/last_idol)\
发布日期： [2025 年11 月 8 日 12:03 UTC](https://forum.freemdict.com/t/topic/41990/21 "2025-11-08T12:03:02Z")

</div>

想要的太多就搞复杂了。因为是喂给 AI 的，格式统一结构越简单越好。
