Yale-BIDS-Chen / medpmc-multi-fig-detection-vit

huggingface.co
Total runs: 0
24-hour runs: 0
7-day runs: 0
30-day runs: 0
Model's Last Updated: July 18 2026
image-classification

Introduction of medpmc-multi-fig-detection-vit

Model Details of medpmc-multi-fig-detection-vit

MedPMC Multi-Figure Detection Model: ViT

This repository provides the vision transformer-based multi-figure detection model used in the MedPMC data curation pipeline.

The model is a binary image classifier trained to predict whether a biomedical figure is a multi-panel / compound figure or a single-panel figure. It is intended for processing figures from biomedical literature, especially figures from PubMed Central (PMC) articles.

Task

The model performs binary image classification.

0: single-panel figure
1: multi-panel / compound figure
Usage
import torch
import timm
from PIL import Image
from torchvision import transforms

checkpoint_path = "model.pth.tar"
image_path = "example.jpg"

device = "cuda" if torch.cuda.is_available() else "cpu"

checkpoint = torch.load(checkpoint_path, map_location="cpu")
arch = checkpoint["arch"]
state_dict = checkpoint["state_dict"]

# Remove DataParallel/DDP prefix if present.
state_dict = {
    k.replace("module.", "", 1) if k.startswith("module.") else k: v
    for k, v in state_dict.items()
}

# Binary classifier.
model = timm.create_model(
    arch,
    pretrained=False,
    num_classes=2,
)

model.load_state_dict(state_dict, strict=True)
model = model.to(device)
model.eval()

preprocess = transforms.Compose([
    transforms.Resize((224, 224)),
    transforms.ToTensor(),
    transforms.Normalize(
        mean=(0.485, 0.456, 0.406),
        std=(0.229, 0.224, 0.225),
    ),
])

image = Image.open(image_path).convert("RGB")
inputs = preprocess(image).unsqueeze(0).to(device)

with torch.no_grad():
    logits = model(inputs)
    probs = torch.softmax(logits, dim=-1)
    pred = torch.argmax(probs, dim=-1).item()

print("Prediction:", pred)
print("Probabilities:", probs.cpu().tolist())

Example output:

Prediction: 1
Probabilities: [[0.08, 0.92]]

This means that the model predicts the input image as a multi-panel / compound figure.

Batch Inference
import torch
import timm
from PIL import Image
from pathlib import Path
from torchvision import transforms

checkpoint_path = "model.pth.tar"
image_dir = "sample"

device = "cuda" if torch.cuda.is_available() else "cpu"

checkpoint = torch.load(checkpoint_path, map_location="cpu")
arch = checkpoint["arch"]
state_dict = checkpoint["state_dict"]

state_dict = {
    k.replace("module.", "", 1) if k.startswith("module.") else k: v
    for k, v in state_dict.items()
}

model = timm.create_model(arch, pretrained=False, num_classes=2)
model.load_state_dict(state_dict, strict=True)
model = model.to(device)
model.eval()

preprocess = transforms.Compose([
    transforms.Resize((224, 224)),
    transforms.ToTensor(),
    transforms.Normalize(
        mean=(0.485, 0.456, 0.406),
        std=(0.229, 0.224, 0.225),
    ),
])

image_paths = sorted(
    list(Path(image_dir).glob("*.jpg")) +
    list(Path(image_dir).glob("*.jpeg")) +
    list(Path(image_dir).glob("*.png"))
)

for image_path in image_paths:
    image = Image.open(image_path).convert("RGB")
    inputs = preprocess(image).unsqueeze(0).to(device)

    with torch.no_grad():
        logits = model(inputs)
        probs = torch.softmax(logits, dim=-1)
        pred = torch.argmax(probs, dim=-1).item()

    print("Image:", image_path)
    print("Prediction:", pred)
    print("Probabilities:", probs.cpu().tolist())
License

The model is released for non-commercial research use under CC BY-NC-SA 4.0.

Citation

Citation information will be updated soon.

Runs of Yale-BIDS-Chen medpmc-multi-fig-detection-vit on huggingface.co

0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs

More Information About medpmc-multi-fig-detection-vit huggingface.co Model

More medpmc-multi-fig-detection-vit license Visit here:

https://choosealicense.com/licenses/cc-by-nc-sa-4.0

medpmc-multi-fig-detection-vit huggingface.co

medpmc-multi-fig-detection-vit huggingface.co is an AI model on huggingface.co that provides medpmc-multi-fig-detection-vit's model effect (), which can be used instantly with this Yale-BIDS-Chen medpmc-multi-fig-detection-vit model. huggingface.co supports a free trial of the medpmc-multi-fig-detection-vit model, and also provides paid use of the medpmc-multi-fig-detection-vit. Support call medpmc-multi-fig-detection-vit model through api, including Node.js, Python, http.

medpmc-multi-fig-detection-vit huggingface.co Url

https://huggingface.co/Yale-BIDS-Chen/medpmc-multi-fig-detection-vit

Yale-BIDS-Chen medpmc-multi-fig-detection-vit online free

medpmc-multi-fig-detection-vit huggingface.co is an online trial and call api platform, which integrates medpmc-multi-fig-detection-vit's modeling effects, including api services, and provides a free online trial of medpmc-multi-fig-detection-vit, you can try medpmc-multi-fig-detection-vit online for free by clicking the link below.

Yale-BIDS-Chen medpmc-multi-fig-detection-vit online free url in huggingface.co:

https://huggingface.co/Yale-BIDS-Chen/medpmc-multi-fig-detection-vit

medpmc-multi-fig-detection-vit install

medpmc-multi-fig-detection-vit is an open source model from GitHub that offers a free installation service, and any user can find medpmc-multi-fig-detection-vit on GitHub to install. At the same time, huggingface.co provides the effect of medpmc-multi-fig-detection-vit install, users can directly use medpmc-multi-fig-detection-vit installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

medpmc-multi-fig-detection-vit install url in huggingface.co:

https://huggingface.co/Yale-BIDS-Chen/medpmc-multi-fig-detection-vit

Url of medpmc-multi-fig-detection-vit

medpmc-multi-fig-detection-vit huggingface.co Url

Provider of medpmc-multi-fig-detection-vit huggingface.co

Yale-BIDS-Chen
ORGANIZATIONS

Other API from Yale-BIDS-Chen