This model is the fine-tuned version of
EleutherAI/gpt-j-6B
on the
GLUE MNLI dataset
.
MNLI dataset consists of pairs of sentences, a
premise
and a
hypothesis
.
The task is to predict the relation between the premise and the hypothesis, which can be:
entailment
: hypothesis follows from the premise,
contradiction
: hypothesis contradicts the premise,
neutral
: hypothesis and premise are unrelated.
We finetune the model as a Causal Language Model (CLM): given a sequence of tokens, the task is to predict the next token.
To achieve this, we create a stylised prompt string, following the approach of
T5 paper
.
mnli hypothesis: Your contributions were of no help with our students' education. premise: Your contribution helped make it possible for us to provide our students with a quality education. target: contradiction <|endoftext|>
Model description
GPT-J 6B is a transformer model trained using Ben Wang's
Mesh Transformer JAX
. "GPT-J" refers to the class of model, while "6B" represents the number of trainable parameters.
*
Each layer consists of one feedforward block and one self attention block.
†
Although the embedding matrix has a size of 50400, only 50257 entries are used by the GPT-2 tokenizer.
The model consists of 28 layers with a model dimension of 4096, and a feedforward dimension of 16384. The model
dimension is split into 16 heads, each with a dimension of 256. Rotary Position Embedding (RoPE) is applied to 64
dimensions of each head. The model is trained with a tokenization vocabulary of 50257, using the same set of BPEs as
GPT-2/GPT-3.
Fine tuning is done using the
train
split of the GLUE MNLI dataset and the performance is measured using the
validation_mismatched
split.
validation_mismatched
means validation examples are not derived from the same sources as those in the training set and therefore not closely resembling any of the examples seen at training time.
Data splits for the mnli dataset are the following
train
validation_matched
validation_mismatched
392702
9815
9832
Fine-tuning procedure
Fine tuned on a Graphcore IPU-POD64 using
popxl
.
Prompt sentences are tokenized and packed together to form 1024 token sequences, following
HF packing algorithm
. No padding is used.
The packing process works in groups of 1000 examples and discards any remainder from each group that isn't a whole sequence.
For the 392,702 training examples this gives a total of 17,762 sequences per epoch.
Since the model is trained to predict the next token, labels are simply the input sequence shifted by one token.
Given the training format, no extra care is needed to account for different sequences: the model does not need to know which sentence a token belongs to.
training steps: 300. Each epoch consists of ceil(17,762/128) steps, hence 300 steps are approximately 2 epochs.
Performance
The resulting model matches SOTA performance with 82.5% accuracy.
Total number of examples 9832
Number with badly formed result 0
Number with incorrect result 1725
Number with correct result 8107
[82.5%]
example 0 = {'prompt_text': "mnli hypothesis: Your contributions were of no help with our students' education. premise: Your contribution helped make it possible for us to provide our students with a quality education. target:", 'class_label': 'contradiction'}
result = {'generated_text': ' contradiction'}
First 10 generated_text and expected class_label results:
0: 'contradiction' contradiction
1: 'contradiction' contradiction
2: 'entailment' entailment
3: 'contradiction' contradiction
4: 'entailment' entailment
5: 'entailment' entailment
6: 'contradiction' contradiction
7: 'contradiction' contradiction
8: 'entailment' neutral
9: 'contradiction' contradiction
How to use
The model can be easily loaded using AutoModelForCausalLM.
You can use the pipeline API for text generation.
from transformers import pipeline, AutoModelForCausalLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained('EleutherAI/gpt-j-6B')
hf_model = AutoModelForCausalLM.from_pretrained("Graphcore/gptj-mnli", pad_token_id=tokenizer.eos_token_id)
generator = pipeline('text-generation', model=hf_model, tokenizer=tokenizer)
prompt = "mnli hypothesis: Your contributions were of no help with our students' education." \
"premise: Your contribution helped make it possible for us to provide our students with a quality education. target:"
out = generator(prompt, return_full_text=False, max_new_tokens=5, top_k=1)
# [{'generated_text': ' contradiction'}]
You can create prompt-like inputs starting from GLUE MNLI dataset using functions provided in the
data_utils.py
script.
from datasets import load_dataset
from data_utils import form_text, split_text
dataset = load_dataset('glue', 'mnli', split='validation_mismatched')
dataset = dataset.map(
form_text, remove_columns=['hypothesis', 'premise','label', 'idx'])
# dataset[0] {'text': "mnli hypothesis: Your contributions were of no help with our students' education. premise: Your contribution helped make it possible for us to provide our students with a quality education. target: contradiction<|endoftext|>"}
dataset = dataset.map(split_text, remove_columns=['text'])
# dataset[0] {'prompt_text': "mnli hypothesis: Your contributions were of no help with our students' education. premise: Your contribution helped make it possible for us to provide our students with a quality education. target:",# 'class_label': 'contradiction'}
Runs of Graphcore gptj-mnli on huggingface.co
0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About gptj-mnli huggingface.co Model
gptj-mnli huggingface.co is an AI model on huggingface.co that provides gptj-mnli's model effect (), which can be used instantly with this Graphcore gptj-mnli model. huggingface.co supports a free trial of the gptj-mnli model, and also provides paid use of the gptj-mnli. Support call gptj-mnli model through api, including Node.js, Python, http.
gptj-mnli huggingface.co is an online trial and call api platform, which integrates gptj-mnli's modeling effects, including api services, and provides a free online trial of gptj-mnli, you can try gptj-mnli online for free by clicking the link below.
Graphcore gptj-mnli online free url in huggingface.co:
gptj-mnli is an open source model from GitHub that offers a free installation service, and any user can find gptj-mnli on GitHub to install. At the same time, huggingface.co provides the effect of gptj-mnli install, users can directly use gptj-mnli installed effect in huggingface.co for debugging and trial. It also supports api for free installation.