Pre-trained language model with identical parameters to
gpt2-medium
, but with additional language modeling heads ("exits") connected to different layers of the model.
These 12 additional heads (in layers 2, 4, 6, 8, 10, 12, 14, 16, 18, 20, 22, 24) were trained on the English portion of
CC-100
while keeping the original pre-trained model parameters frozen.
The model can be used for the
Autocontrastive Decoding
text generation approach described in
Gera et al. 2023
, for
early-exiting
approaches, or for other algorithms that consider the next-token predictions of different model layers.
Usage
Harnessing the additional language modeling heads requires loading the model using the
auto-contrastive-generation library
(
pip install autocontrastive-gen
).
In a nutshell, the user creates a
MultiExitConfiguration
that determines model behavior at training and inference, and then loads the model using the dedicated
AutoMultiExitModel
class. After that, the model can be used with the
transformers
API like any other model. See the
GitHub
for detailed usage instructions.
For example, the code below initializes the model to use
Autocontrastive Decoding
, and then performs text generation in this chosen setting:
from transformers import AutoTokenizer
from autocontrastive_gen.modeling.configuration import MultiExitConfiguration
from autocontrastive_gen.modeling.auto_model import AutoMultiExitModel
# initialize a pre-trained multi-exit model to use auto-contrast between layer 24 and layer 12
multi_exit_config = MultiExitConfiguration(use_original_head=False,
contrast_layer_indices=(24, 12))
model = AutoMultiExitModel.from_pretrained("IBM/gpt2-medium-multiexit", multi_exit_config=multi_exit_config)
# perform text generation as usual
tokenizer = AutoTokenizer.from_pretrained("IBM/gpt2-medium-multiexit")
prompt = tokenizer("humpty dumpty sat on", return_tensors='pt')
generated_ids = model.generate(**prompt, max_new_tokens=15)
print(tokenizer.batch_decode(generated_ids))
@inproceedings{gera2023autocontrastive,
title={The Benefits of Bad Advice: Autocontrastive Decoding across Model Layers},
author={Gera, Ariel and Friedman, Roni and Arviv, Ofir and Gunasekara, Chulaka and Sznajder, Benjamin and Slonim, Noam and Shnarch, Eyal},
booktitle={Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)},
month={july},
address={Toronto, Canada},
year={2023}
}
Runs of ibm gpt2-medium-multiexit on huggingface.co
109
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About gpt2-medium-multiexit huggingface.co Model
gpt2-medium-multiexit huggingface.co is an AI model on huggingface.co that provides gpt2-medium-multiexit's model effect (), which can be used instantly with this ibm gpt2-medium-multiexit model. huggingface.co supports a free trial of the gpt2-medium-multiexit model, and also provides paid use of the gpt2-medium-multiexit. Support call gpt2-medium-multiexit model through api, including Node.js, Python, http.
gpt2-medium-multiexit huggingface.co is an online trial and call api platform, which integrates gpt2-medium-multiexit's modeling effects, including api services, and provides a free online trial of gpt2-medium-multiexit, you can try gpt2-medium-multiexit online for free by clicking the link below.
ibm gpt2-medium-multiexit online free url in huggingface.co:
gpt2-medium-multiexit is an open source model from GitHub that offers a free installation service, and any user can find gpt2-medium-multiexit on GitHub to install. At the same time, huggingface.co provides the effect of gpt2-medium-multiexit install, users can directly use gpt2-medium-multiexit installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
gpt2-medium-multiexit install url in huggingface.co: