This model is a
RoBERTa base model
that was further trained using a masked language modeling task on a compendium of english scientific textual examples from the life sciences using the
BioLang dataset
. It was then fine-tuned for token classification on the SourceData
sd-nlp
dataset with the
PANELIZATION
task to perform 'parsing' or 'segmentation' of figure legends into fragments corresponding to sub-panels.
Figures are usually composite representations of results obtained with heterogeneous experimental approaches and systems. Breaking figures into panels allows identifying more coherent descriptions of individual scientific experiments.
Intended uses & limitations
How to use
The intended use of this model is for 'parsing' figure legends into sub-fragments corresponding to individual panels as used in SourceData annotations (
https://sourcedata.embo.org
).
To have a quick check of the model:
from transformers import pipeline, RobertaTokenizerFast, RobertaForTokenClassification
example = """Fig 4. a, Volume density of early (Avi) and late (Avd) autophagic vacuoles.a, Volume density of early (Avi) and late (Avd) autophagic vacuoles from four independent cultures. Examples of Avi and Avd are shown in b and c, respectively. Bars represent 0.4����m. d, Labelling density of cathepsin-D as estimated in two independent experiments. e, Labelling density of LAMP-1."""
tokenizer = RobertaTokenizerFast.from_pretrained('roberta-base', max_len=512)
model = RobertaForTokenClassification.from_pretrained('EMBO/sd-panelization')
ner = pipeline('ner', model, tokenizer=tokenizer)
res = ner(example)
for r in res: print(r['word'], r['entity'])
Limitations and bias
The model must be used with the
roberta-base
tokenizer.
Training data
The model was trained for token classification using the
EMBO/sd-nlp PANELIZATION
dataset which includes manually annotated examples.
Training procedure
The training was run on an NVIDIA DGX Station with 4XTesla V100 GPUs.
sd-panelization huggingface.co is an AI model on huggingface.co that provides sd-panelization's model effect (), which can be used instantly with this EMBO sd-panelization model. huggingface.co supports a free trial of the sd-panelization model, and also provides paid use of the sd-panelization. Support call sd-panelization model through api, including Node.js, Python, http.
sd-panelization huggingface.co is an online trial and call api platform, which integrates sd-panelization's modeling effects, including api services, and provides a free online trial of sd-panelization, you can try sd-panelization online for free by clicking the link below.
EMBO sd-panelization online free url in huggingface.co:
sd-panelization is an open source model from GitHub that offers a free installation service, and any user can find sd-panelization on GitHub to install. At the same time, huggingface.co provides the effect of sd-panelization install, users can directly use sd-panelization installed effect in huggingface.co for debugging and trial. It also supports api for free installation.