This is
"naver-clova-ix/donut-base"
but with all non-ascii tokens removed. This means the model is good for basic English use cases where the text is primarily a-zA-Z0-9 and basic punctuation.
The original model,
"naver-clova-ix/donut-base"
, did not have a token for
"1"
, so that has also been added. The notebook
remove-donut-tokens.ipynb
details the whole process.
This has not been trained any more than the original model.
I did a quick speed test for generation against the default model and using
bad_words_ids
. The
bad_words_ids
was only 12k tokens instead of the 30k that were removed and it was still noticeably slower.
donut-base-ascii huggingface.co is an AI model on huggingface.co that provides donut-base-ascii's model effect (), which can be used instantly with this nbroad donut-base-ascii model. huggingface.co supports a free trial of the donut-base-ascii model, and also provides paid use of the donut-base-ascii. Support call donut-base-ascii model through api, including Node.js, Python, http.
donut-base-ascii huggingface.co is an online trial and call api platform, which integrates donut-base-ascii's modeling effects, including api services, and provides a free online trial of donut-base-ascii, you can try donut-base-ascii online for free by clicking the link below.
nbroad donut-base-ascii online free url in huggingface.co:
donut-base-ascii is an open source model from GitHub that offers a free installation service, and any user can find donut-base-ascii on GitHub to install. At the same time, huggingface.co provides the effect of donut-base-ascii install, users can directly use donut-base-ascii installed effect in huggingface.co for debugging and trial. It also supports api for free installation.