Caution!: While it has been confirmed that the performance of LLM-jp-3 172B alpha1 and alpha2 is significantly lower than previously released models, we believe they can still be useful for research purposes and are making them available to the public.
For more information, please visit
this link
.
Required Libraries and Their Versions
torch>=2.3.0
transformers>=4.40.1
tokenizers>=0.19.1
accelerate>=0.29.3
flash-attn>=2.5.8
Usage
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("llm-jp/llm-jp-3-172b-alpha2")
model = AutoModelForCausalLM.from_pretrained("llm-jp/llm-jp-3-172b-alpha2", device_map="auto", torch_dtype=torch.bfloat16)
text = "自然言語処理とは何か"
tokenized_input = tokenizer.encode(text, add_special_tokens=False, return_tensors="pt").to(model.device)
with torch.no_grad():
output = model.generate(
tokenized_input,
max_new_tokens=100,
do_sample=True,
top_p=0.95,
temperature=0.7,
repetition_penalty=1.05,
)[0]
print(tokenizer.decode(output))
Model Details
Model type:
Transformer-based Language Model
Total seen tokens:
:
alpha1: 0.7T
alpha2: 1.4T
beta1: 0.7T
Params
Layers
Hidden size
Heads
Context length
172b
96
12288
96
4096
Tokenizer
The tokenizer of this model is based on
huggingface/tokenizers
Unigram byte-fallback model.
The vocabulary entries were converted from
llm-jp-tokenizer v3.0
.
Please refer to
README.md
of
llm-jp-tokenizer
for details on the vocabulary construction procedure (the pure SentencePiece training does not reproduce our vocabulary).
Datasets
Pre-training
The models have been pre-trained using a blend of the following datasets.
The models released here are in the early stages of our research and development and have not been tuned to ensure outputs align with human intent and safety considerations.
llm-jp-3-172b-alpha2 huggingface.co is an AI model on huggingface.co that provides llm-jp-3-172b-alpha2's model effect (), which can be used instantly with this llm-jp llm-jp-3-172b-alpha2 model. huggingface.co supports a free trial of the llm-jp-3-172b-alpha2 model, and also provides paid use of the llm-jp-3-172b-alpha2. Support call llm-jp-3-172b-alpha2 model through api, including Node.js, Python, http.
llm-jp-3-172b-alpha2 huggingface.co is an online trial and call api platform, which integrates llm-jp-3-172b-alpha2's modeling effects, including api services, and provides a free online trial of llm-jp-3-172b-alpha2, you can try llm-jp-3-172b-alpha2 online for free by clicking the link below.
llm-jp llm-jp-3-172b-alpha2 online free url in huggingface.co:
llm-jp-3-172b-alpha2 is an open source model from GitHub that offers a free installation service, and any user can find llm-jp-3-172b-alpha2 on GitHub to install. At the same time, huggingface.co provides the effect of llm-jp-3-172b-alpha2 install, users can directly use llm-jp-3-172b-alpha2 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
llm-jp-3-172b-alpha2 install url in huggingface.co: