# DWDS ( Collaboration to create .mdx)

**URL:** <https://forum.freemdict.com/t/topic/7373>\
**Category:** 资源求助\
**Tags:** 求助\
**Created:** [2021 年7 月 19 日 23:16 UTC](https://forum.freemdict.com/t/topic/7373 "2021-07-19T23:16:29Z")\
**Posts on this page:** 5\
**Page:** 1

<div class="post-metadata">

作者： ![tovaremeterio](https://forumcdn.freemdict.com/user_avatar/forum.freemdict.com/tovaremeterio/32/15709_2.png) [@tovaremeterio](https://forum.freemdict.com/u/tovaremeterio)\
发布日期： [2021 年7 月 19 日 23:16 UTC](https://forum.freemdict.com/t/topic/7373/1 "2021-07-19T23:16:29Z")

</div>

Is someone also interested in scraping this German Dictionary ?

> **[DWDS – Digitales Wörterbuch der deutschen Sprache](https://www.dwds.de)**
>
> DWDS − Der deutsche Wortschatz von 1600 bis heute.

I am looking for volunteers to work together.

Please send a message if interested.  
Email: [tovaremeterio.56l9g@simplelogin.fr](mailto:tovaremeterio.56l9g@simplelogin.fr)

An .mdx of DWDS was made but is not perfect. Works very well on Android “MDict” but not so well on GoldenDict PC. If interested send a PM.

---

<div class="post-metadata">

作者： ![James1](https://forumcdn.freemdict.com/letter_avatar_proxy/v4/letter/j/c67d28/32.png) [@James1](https://forum.freemdict.com/u/James1)\
发布日期： [2021 年7 月 20 日 05:07 UTC](https://forum.freemdict.com/t/topic/7373/2 "2021-07-20T05:07:39Z")

</div>

How to make a web scraping on this website with Beautiful Soup in python to make a dictionary?[https://pypi.org/project/beautifulsoup4/](https://pypi.org/project/beautifulsoup4/)

---

<div class="post-metadata">

作者： ![tovaremeterio](https://forumcdn.freemdict.com/user_avatar/forum.freemdict.com/tovaremeterio/32/15709_2.png) [@tovaremeterio](https://forum.freemdict.com/u/tovaremeterio)\
发布日期： [2021 年8 月 23 日 16:42 UTC](https://forum.freemdict.com/t/topic/7373/3 "2021-08-23T16:42:36Z")

</div>

A user from Telegram Group provided this info useful for anyone who would like to scrape and create an .mdx 😃

> Download all entries as web pages then edit them through notpad++ then merge all html. You make a txt file after combining the edited html files then convert with mdx builder.  
> You have to know a little bit something about regex. Notepad ++ can edit thousands of files at the same time. To download all the entries you need wget. To get the webwite links for entries you view source the Web pages that contain the entries then delete another infos with regex (you keep only the links to the entries) You may need to do this for the sub pages. To merge the files at the end you need powershell or third party Programm. You can also download the audio and link it with wget and regex. You don’t need to be programmer. Sorry if my description isn’t clear but these are the basic steps.

---

<div class="post-metadata">

作者： ![tovaremeterio](https://forumcdn.freemdict.com/user_avatar/forum.freemdict.com/tovaremeterio/32/15709_2.png) [@tovaremeterio](https://forum.freemdict.com/u/tovaremeterio)\
发布日期： [2021 年8 月 23 日 16:48 UTC](https://forum.freemdict.com/t/topic/7373/4 "2021-08-23T16:48:42Z")

</div>

Here is the raw data from DWDS if someone would like to make his/her own version for .mdx 😃

> [@DWDS ( German - Deutsch )](https://forum.freemdict.com/t/topic/7504):
>
> [http://dwds.de/](http://dwds.de/) was converted into .mdx format for GoldenDict:

---

<div class="post-metadata">

作者： ![James1](https://forumcdn.freemdict.com/letter_avatar_proxy/v4/letter/j/c67d28/32.png) [@James1](https://forum.freemdict.com/u/James1)\
发布日期： [2021 年9 月 9 日 11:46 UTC](https://forum.freemdict.com/t/topic/7373/5 "2021-09-09T11:46:28Z")

</div>

Thanks for your beautiful guideline. I don’t know what wget and regex is, though.
