I need to convince my friend, Phyllis, that she should train a custom LLM for her Fortune 500 company using the MosaicML Platform.
Please write an email that explains why MosaicML's emphasis on cutting edge methodology, data privacy, and efficiency are so important.
End the email with a friendly inquiry about Phyllis's family.
MPT-30B-Instruct
:
Phyllis -
I know you're always looking for ways to improve efficiency at work and I think I've found a solution that you'll be very interested in.
LLMs are only as good as the data they're trained on and the training process can be very time consuming and resource intensive. MosaicML uses cutting edge methodology to efficiently train high quality models without compromising data privacy.
They also provide tools to easily connect to and use the model in your daily workflow.
I think you'd really enjoy speaking with their founder, we can set up a call if you're interested.
Also, I know it's been a tough year for your family, how are things?
Best,
Your Friend
How to Use
Note: This model requires that
trust_remote_code=True
be passed to the
from_pretrained
method. This is because we use a custom model architecture that is not yet part of the
transformers
package.
import transformers
model = transformers.AutoModelForCausalLM.from_pretrained(
'mosaicml/mpt-30b-instruct',
trust_remote_code=True
)
To use the optimized
triton implementation
of FlashAttention, you can load the model on GPU (
cuda:0
) with
attn_impl='triton'
and with
bfloat16
precision:
import torch
import transformers
name = 'mosaicml/mpt-30b-instruct'
config = transformers.AutoConfig.from_pretrained(name, trust_remote_code=True)
config.attn_config['attn_impl'] = 'triton'# change this to use triton-based FlashAttention
config.init_device = 'cuda:0'# For fast initialization directly on GPU!
model = transformers.AutoModelForCausalLM.from_pretrained(
name,
config=config,
torch_dtype=torch.bfloat16, # Load model weights in bfloat16
trust_remote_code=True
)
The model was trained initially on a sequence length of 2048. An additional pre-training phase was included for sequence length adaptation to 8192. However, ALiBi further enables users to increase the maximum sequence length during finetuning and/or inference. For example:
import transformers
name = 'mosaicml/mpt-30b-instruct'
config = transformers.AutoConfig.from_pretrained(name, trust_remote_code=True)
config.max_seq_len = 16384# (input + output) tokens can now be up to 16384
model = transformers.AutoModelForCausalLM.from_pretrained(
name,
config=config,
trust_remote_code=True
)
This model was trained with the MPT-30B tokenizer which is based on the
EleutherAI/gpt-neox-20b
tokenizer and includes additional padding and eos tokens.
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained('mosaicml/mpt-30b')
The model can then be used, for example, within a text-generation pipeline.
Note: when running Torch modules in lower precision, it is best practice to use the
torch.autocast context manager
.
from transformers import pipeline
with torch.autocast('cuda', dtype=torch.bfloat16):
inputs = tokenizer('Here is a recipe for vegan banana bread:\n', return_tensors="pt").to('cuda')
outputs = model.generate(**inputs, max_new_tokens=100)
print(tokenizer.batch_decode(outputs, skip_special_tokens=True))
# or using the HF pipeline
pipe = pipeline('text-generation', model=model, tokenizer=tokenizer, device='cuda:0')
with torch.autocast('cuda', dtype=torch.bfloat16):
print(
pipe('Here is a recipe for vegan banana bread:\n',
max_new_tokens=100,
do_sample=True,
use_cache=True))
Formatting
This model was trained on data formatted as follows:
defformat_prompt(instruction):
template = "Below is an instruction that describes a task. Write a response that appropriately completes the request.\n\n###Instruction\n{instruction}\n\n### Response\n"return template.format(instruction=instruction)
example = "Tell me a funny joke.\nDon't make it too funny though."
fmt_ex = format_prompt(instruction=example)
In the above example,
fmt_ex
is ready to be tokenized and sent through the model.
Model Description
The architecture is a modification of a standard decoder-only transformer.
The model has been modified from a standard transformer in the following ways:
This model was trained on 72 A100 40GB GPUs for 8 hours using the
MosaicML Platform
.
The model was trained with sharded data parallelism using
FSDP
and used the AdamW optimizer.
MPT-30B-Instruct can produce factually incorrect output, and should not be relied on to produce factually accurate information.
MPT-30B-Instruct was trained on various public datasets.
While great efforts have been taken to clean the pretraining data, it is possible that this model could generate lewd, biased or otherwise offensive outputs.
Acknowledgements
This model was finetuned by Sam Havens, Alex Trott, and the MosaicML NLP team
The license on this model does not constitute legal advice. We are not responsible for the actions of third parties who use this model. Please consult an attorney before using this model for commercial purposes.
Citation
Please cite this model using the following format:
@online{MosaicML2023Introducing,
author = {MosaicML NLP Team},
title = {Introducing MPT-30B: Raising the bar
for open-source foundation models},
year = {2023},
url = {www.mosaicml.com/blog/mpt-30b},
note = {Accessed: 2023-06-22},
urldate = {2023-06-22}
}
Runs of michaelfeil ct2fast-mpt-30b-instruct on huggingface.co
17
Total runs
0
24-hour runs
3
3-day runs
5
7-day runs
11
30-day runs
More Information About ct2fast-mpt-30b-instruct huggingface.co Model
ct2fast-mpt-30b-instruct huggingface.co is an AI model on huggingface.co that provides ct2fast-mpt-30b-instruct's model effect (), which can be used instantly with this michaelfeil ct2fast-mpt-30b-instruct model. huggingface.co supports a free trial of the ct2fast-mpt-30b-instruct model, and also provides paid use of the ct2fast-mpt-30b-instruct. Support call ct2fast-mpt-30b-instruct model through api, including Node.js, Python, http.
ct2fast-mpt-30b-instruct huggingface.co is an online trial and call api platform, which integrates ct2fast-mpt-30b-instruct's modeling effects, including api services, and provides a free online trial of ct2fast-mpt-30b-instruct, you can try ct2fast-mpt-30b-instruct online for free by clicking the link below.
michaelfeil ct2fast-mpt-30b-instruct online free url in huggingface.co:
ct2fast-mpt-30b-instruct is an open source model from GitHub that offers a free installation service, and any user can find ct2fast-mpt-30b-instruct on GitHub to install. At the same time, huggingface.co provides the effect of ct2fast-mpt-30b-instruct install, users can directly use ct2fast-mpt-30b-instruct installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
ct2fast-mpt-30b-instruct install url in huggingface.co: