Results on the MP-DocVQA dataset are reported in Table 2.
Training hyperparameters can be found in Table 8 of Appendix D.
How to use
Here is how to use this model to get the features of a given text in PyTorch:
import torch
from transformers import LayoutLMv3Processor, LayoutLMv3ForQuestionAnswering
processor = LayoutLMv3Processor.from_pretrained("rubentito/layoutlmv3-base-mpdocvqa", apply_ocr=False)
model = LayoutLMv3ForQuestionAnswering.from_pretrained("rubentito/layoutlmv3-base-mpdocvqa")
image = Image.open("example.jpg").convert("RGB")
question = "Is this a question?"
context = ["Example"]
boxes = [0, 0, 1000, 1000] # This is an example bounding box covering the whole image.
document_encoding = processor(image, question, context, boxes=boxes, return_tensors="pt")
outputs = model(**document_encoding)
# Get the answer
start_idx = torch.argmax(outputs.start_logits, axis=1)
end_idx = torch.argmax(outputs.end_logits, axis=1)
answers = self.processor.tokenizer.decode(input_tokens[start_idx: end_idx+1]).strip()
Metrics
Average Normalized Levenshtein Similarity (ANLS)
The standard metric for text-based VQA tasks (ST-VQA and DocVQA). It evaluates the method's reasoning capabilities while smoothly penalizes OCR recognition errors.
Check
Scene Text Visual Question Answering
for detailed information.
Answer Page Prediction Accuracy (APPA)
In the MP-DocVQA task, the models can provide the index of the page where the information required to answer the question is located. For this subtask accuracy is used to evaluate the predictions: i.e. if the predicted page is correct or not.
Check
Hierarchical multimodal transformers for Multi-Page DocVQA
for detailed information.
layoutlmv3-base-mpdocvqa huggingface.co is an AI model on huggingface.co that provides layoutlmv3-base-mpdocvqa's model effect (), which can be used instantly with this rubentito layoutlmv3-base-mpdocvqa model. huggingface.co supports a free trial of the layoutlmv3-base-mpdocvqa model, and also provides paid use of the layoutlmv3-base-mpdocvqa. Support call layoutlmv3-base-mpdocvqa model through api, including Node.js, Python, http.
layoutlmv3-base-mpdocvqa huggingface.co is an online trial and call api platform, which integrates layoutlmv3-base-mpdocvqa's modeling effects, including api services, and provides a free online trial of layoutlmv3-base-mpdocvqa, you can try layoutlmv3-base-mpdocvqa online for free by clicking the link below.
rubentito layoutlmv3-base-mpdocvqa online free url in huggingface.co:
layoutlmv3-base-mpdocvqa is an open source model from GitHub that offers a free installation service, and any user can find layoutlmv3-base-mpdocvqa on GitHub to install. At the same time, huggingface.co provides the effect of layoutlmv3-base-mpdocvqa install, users can directly use layoutlmv3-base-mpdocvqa installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
layoutlmv3-base-mpdocvqa install url in huggingface.co: