This model is intended to be used for QA in the Vietnamese language so the valid set is Vietnamese only (but English works fine). The evaluation result below using 10% of the Vietnamese dataset.
MRCQuestionAnswering
using
XLM-RoBERTa
as a pre-trained language model. By default, XLM-RoBERTa will split word in to sub-words. But in my implementation, I re-combine sub-words representation (after encoded by BERT layer) into word representation using sum strategy.
Using pre-trained model
Hugging Face pipeline style (
NOT using sum features strategy
).
from transformers import pipeline
# model_checkpoint = "nguyenvulebinh/vi-mrc-large"
model_checkpoint = "nguyenvulebinh/vi-mrc-base"
nlp = pipeline('question-answering', model=model_checkpoint,
tokenizer=model_checkpoint)
QA_input = {
'question': "Bình là chuyên gia về gì ?",
'context': "Bình Nguyễn là một người đam mê với lĩnh vực xử lý ngôn ngữ tự nhiên . Anh nhận chứng chỉ Google Developer Expert năm 2020"
}
res = nlp(QA_input)
print('pipeline: {}'.format(res))
#{'score': 0.5782045125961304, 'start': 45, 'end': 68, 'answer': 'xử lý ngôn ngữ tự nhiên'}
from infer import tokenize_function, data_collator, extract_answer
from model.mrc_model import MRCQuestionAnswering
from transformers import AutoTokenizer
# model_checkpoint = "nguyenvulebinh/vi-mrc-large"
model_checkpoint = "nguyenvulebinh/vi-mrc-base"
tokenizer = AutoTokenizer.from_pretrained(model_checkpoint)
model = MRCQuestionAnswering.from_pretrained(model_checkpoint)
QA_input = {
'question': "Bình được công nhận với danh hiệu gì ?",
'context': "Bình Nguyễn là một người đam mê với lĩnh vực xử lý ngôn ngữ tự nhiên . Anh nhận chứng chỉ Google Developer Expert năm 2020"
}
inputs = [tokenize_function(*QA_input)]
inputs_ids = data_collator(inputs)
outputs = model(**inputs_ids)
answer = extract_answer(inputs, outputs, tokenizer)
print(answer)
# answer: Google Developer Expert. Score start: 0.9926977753639221, Score end: 0.9909810423851013
About
Built by Binh Nguyen
For more details, visit the project repository.
Runs of nguyenvulebinh vi-mrc-base on huggingface.co
28
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
-77
30-day runs
More Information About vi-mrc-base huggingface.co Model
vi-mrc-base huggingface.co is an AI model on huggingface.co that provides vi-mrc-base's model effect (), which can be used instantly with this nguyenvulebinh vi-mrc-base model. huggingface.co supports a free trial of the vi-mrc-base model, and also provides paid use of the vi-mrc-base. Support call vi-mrc-base model through api, including Node.js, Python, http.
vi-mrc-base huggingface.co is an online trial and call api platform, which integrates vi-mrc-base's modeling effects, including api services, and provides a free online trial of vi-mrc-base, you can try vi-mrc-base online for free by clicking the link below.
nguyenvulebinh vi-mrc-base online free url in huggingface.co:
vi-mrc-base is an open source model from GitHub that offers a free installation service, and any user can find vi-mrc-base on GitHub to install. At the same time, huggingface.co provides the effect of vi-mrc-base install, users can directly use vi-mrc-base installed effect in huggingface.co for debugging and trial. It also supports api for free installation.