If you would like to learn more about the pretraining of the LLM-jp-3 MoE series, please refer to this
blog post
.
Tokenizer
The tokenizer of this model is based on
huggingface/tokenizers
Unigram byte-fallback model.
The vocabulary entries were converted from
llm-jp-tokenizer v3.0
.
Please refer to
README.md
of
llm-jp-tokenizer
for details on the vocabulary construction procedure (the pure SentencePiece training does not reproduce our vocabulary).
Datasets
Pre-training
The models have been pre-trained using a blend of the following datasets.
We evaluated the models using
gpt-4o-2024-08-06
.
The scores represent the average values obtained from five rounds of inference and evaluation.
For more details, please refer to the
codes
.
AnswerCarefully-Eval
assesses the safety of Japanese language model outputs using the LLM-as-a-Judge approach, based on the test set from
llm-jp/AnswerCarefully
.
We evaluated the models using
gpt-4-0613
.
The scores represent the average values obtained from five rounds of inference and evaluation.
The models released here are in the early stages of our research and development and have not been tuned to ensure outputs align with human intent and safety considerations.
If you find our work helpful, please feel free to cite the paper.
@inproceedings{
nakamura2025dropupcycling,
title={Drop-Upcycling: Training Sparse Mixture of Experts with Partial Re-initialization},
author={Taishi Nakamura and Takuya Akiba and Kazuki Fujii and Yusuke Oda and Rio Yokota and Jun Suzuki},
booktitle={The Thirteenth International Conference on Learning Representations},
year={2025},
url={https://openreview.net/forum?id=gx1wHnf5Vp}
}
Model Card Authors
The names are listed in alphabetical order.
Hirokazu Kiyomaru, Takashi Kodama and Taishi Nakamura.
Runs of llm-jp llm-jp-3-8x13b-instruct3 on huggingface.co
1.1K
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
24
30-day runs
More Information About llm-jp-3-8x13b-instruct3 huggingface.co Model
llm-jp-3-8x13b-instruct3 huggingface.co is an AI model on huggingface.co that provides llm-jp-3-8x13b-instruct3's model effect (), which can be used instantly with this llm-jp llm-jp-3-8x13b-instruct3 model. huggingface.co supports a free trial of the llm-jp-3-8x13b-instruct3 model, and also provides paid use of the llm-jp-3-8x13b-instruct3. Support call llm-jp-3-8x13b-instruct3 model through api, including Node.js, Python, http.
llm-jp-3-8x13b-instruct3 huggingface.co is an online trial and call api platform, which integrates llm-jp-3-8x13b-instruct3's modeling effects, including api services, and provides a free online trial of llm-jp-3-8x13b-instruct3, you can try llm-jp-3-8x13b-instruct3 online for free by clicking the link below.
llm-jp llm-jp-3-8x13b-instruct3 online free url in huggingface.co:
llm-jp-3-8x13b-instruct3 is an open source model from GitHub that offers a free installation service, and any user can find llm-jp-3-8x13b-instruct3 on GitHub to install. At the same time, huggingface.co provides the effect of llm-jp-3-8x13b-instruct3 install, users can directly use llm-jp-3-8x13b-instruct3 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
llm-jp-3-8x13b-instruct3 install url in huggingface.co: