Based on
Qwen2-0.5B
, the tokenizer has been replaced with
BilingualTokenizer-8K
to reduce the number of parameters. The total parameters have been reduced from 0.5B to 365M.
Details
To recover some performance and facilitate fine-tuning for downstream tasks, I chose to freeze the backbone parameters and only train the embedding part after replacing the tokenizer. Training was conducted for 40,000 steps on
wikipedia-zh
and
cosmopedia-100k
.
NanoLM-365M-Base huggingface.co is an AI model on huggingface.co that provides NanoLM-365M-Base's model effect (), which can be used instantly with this Mxode NanoLM-365M-Base model. huggingface.co supports a free trial of the NanoLM-365M-Base model, and also provides paid use of the NanoLM-365M-Base. Support call NanoLM-365M-Base model through api, including Node.js, Python, http.
NanoLM-365M-Base huggingface.co is an online trial and call api platform, which integrates NanoLM-365M-Base's modeling effects, including api services, and provides a free online trial of NanoLM-365M-Base, you can try NanoLM-365M-Base online for free by clicking the link below.
Mxode NanoLM-365M-Base online free url in huggingface.co:
NanoLM-365M-Base is an open source model from GitHub that offers a free installation service, and any user can find NanoLM-365M-Base on GitHub to install. At the same time, huggingface.co provides the effect of NanoLM-365M-Base install, users can directly use NanoLM-365M-Base installed effect in huggingface.co for debugging and trial. It also supports api for free installation.