A more detailed description of the model can be found in an article we published on the KBLab blog
here
and for the updated model
here
.
Update
: We have released updated versions of the model since the initial release. The original model described in the blog post is
v1.0
. The current version is
v2.0
. The newer versions are trained on longer paragraphs, and have a longer max sequence length.
v2.0
is trained with a stronger teacher model and is the current default.
from sentence_transformers import SentenceTransformer
sentences = ["Det här är en exempelmening", "Varje exempel blir konverterad"]
model = SentenceTransformer('KBLab/sentence-bert-swedish-cased')
embeddings = model.encode(sentences)
print(embeddings)
Loading an older model version (Sentence-Transformers)
Currently, the easiest way to load an older model version is to clone the model repository and load it from disk. For example, to clone the
v1.0
model:
Then you can load the model by pointing to the local folder where you cloned the model:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("path_to_model_folder/sentence-bert-swedish-cased")
Usage (HuggingFace Transformers)
Without
sentence-transformers
, you can use the model like this: First, you pass your input through the transformer model, then you have to apply the right pooling-operation on-top of the contextualized word embeddings.
from transformers import AutoTokenizer, AutoModel
import torch
#Mean Pooling - Take attention mask into account for correct averagingdefmean_pooling(model_output, attention_mask):
token_embeddings = model_output[0] #First element of model_output contains all token embeddings
input_mask_expanded = attention_mask.unsqueeze(-1).expand(token_embeddings.size()).float()
return torch.sum(token_embeddings * input_mask_expanded, 1) / torch.clamp(input_mask_expanded.sum(1), min=1e-9)
# Sentences we want sentence embeddings for
sentences = ['Det här är en exempelmening', 'Varje exempel blir konverterad']
# Load model from HuggingFace Hub# To load an older version, e.g. v1.0, add the argument revision="v1.0"
tokenizer = AutoTokenizer.from_pretrained('KBLab/sentence-bert-swedish-cased')
model = AutoModel.from_pretrained('KBLab/sentence-bert-swedish-cased')
# Tokenize sentences
encoded_input = tokenizer(sentences, padding=True, truncation=True, return_tensors='pt')
# Compute token embeddingswith torch.no_grad():
model_output = model(**encoded_input)
# Perform pooling. In this case, max pooling.
sentence_embeddings = mean_pooling(model_output, encoded_input['attention_mask'])
print("Sentence embeddings:")
print(sentence_embeddings)
Loading an older model (Hugginfface Transformers)
To load an older model specify the version tag with the
revision
arg. For example, to load the
v1.0
model, use the following code:
The model was evaluated on
SweParaphrase v1.0
and
SweParaphrase v2.0
. This test set is part of
SuperLim
-- a Swedish evaluation suite for natural langage understanding tasks. We calculated Pearson and Spearman correlation between predicted model similarity scores and the human similarity score labels. Results from
SweParaphrase v1.0
are displayed below.
Model version
Pearson
Spearman
v1.0
0.9183
0.9114
v1.1
0.9183
0.9114
v2.0
0.9283
0.9130
The following code snippet can be used to reproduce the above results:
In general,
v1.1
correlates the most with human assessment of text similarity on SweParaphrase v2.0. Below, we present zero-shot evaluation results on all data splits. They display the model's performance out of the box, without any fine-tuning.
Model version
Data split
Pearson
Spearman
v1.0
train
0.8355
0.8256
v1.1
train
0.8383
0.8302
v2.0
train
0.8209
0.8059
v1.0
dev
0.8682
0.8774
v1.1
dev
0.8739
0.8833
v2.0
dev
0.8638
0.8668
v1.0
test
0.8356
0.8476
v1.1
test
0.8393
0.8550
v2.0
test
0.8232
0.8213
SweFAQ v2.0
When it comes to retrieval tasks,
v2.0
performs the best by quite a substantial margin. It is better at matching the correct answer to a question compared to v1.1 and v1.0.
An article with more details on data and v1.0 of the model can be found on the
KBLab blog
.
Around 14.6 million sentences from English-Swedish parallel corpuses were used to train the model. Data was sourced from the
Open Parallel Corpus
(OPUS) and downloaded via the python package
opustools
. Datasets used were: JW300, Europarl, DGT-TM, EMEA, ELITR-ECA, TED2020, Tatoeba and OpenSubtitles.
The model was trained with the parameters:
DataLoader
:
torch.utils.data.dataloader.DataLoader
of length 180513 with parameters:
@misc{rekathati2021introducing,
author = {Rekathati, Faton},
title = {The KBLab Blog: Introducing a Swedish Sentence Transformer},
url = {https://kb-labb.github.io/posts/2021-08-23-a-swedish-sentence-transformer/},
year = {2021}
}
Acknowledgements
We gratefully acknowledge the HPC RIVR consortium (
www.hpc-rivr.si
) and EuroHPC JU (
eurohpc-ju.europa.eu/
) for funding this research by providing computing resources of the HPC system Vega at the Institute of Information Science (
www.izum.si
).
Runs of KBLab sentence-bert-swedish-cased on huggingface.co
62.8K
Total runs
1.3K
24-hour runs
1.8K
3-day runs
14.8K
7-day runs
36.1K
30-day runs
More Information About sentence-bert-swedish-cased huggingface.co Model
More sentence-bert-swedish-cased license Visit here:
sentence-bert-swedish-cased huggingface.co is an AI model on huggingface.co that provides sentence-bert-swedish-cased's model effect (), which can be used instantly with this KBLab sentence-bert-swedish-cased model. huggingface.co supports a free trial of the sentence-bert-swedish-cased model, and also provides paid use of the sentence-bert-swedish-cased. Support call sentence-bert-swedish-cased model through api, including Node.js, Python, http.
sentence-bert-swedish-cased huggingface.co is an online trial and call api platform, which integrates sentence-bert-swedish-cased's modeling effects, including api services, and provides a free online trial of sentence-bert-swedish-cased, you can try sentence-bert-swedish-cased online for free by clicking the link below.
KBLab sentence-bert-swedish-cased online free url in huggingface.co:
sentence-bert-swedish-cased is an open source model from GitHub that offers a free installation service, and any user can find sentence-bert-swedish-cased on GitHub to install. At the same time, huggingface.co provides the effect of sentence-bert-swedish-cased install, users can directly use sentence-bert-swedish-cased installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
sentence-bert-swedish-cased install url in huggingface.co: