dmusingu / cxr-vitl14-captioning-50ep

huggingface.co
Total runs: 0
24-hour runs: 0
7-day runs: 0
30-day runs: 0
Model's Last Updated: June 26 2026
image-feature-extraction

Introduction of cxr-vitl14-captioning-50ep

Model Details of cxr-vitl14-captioning-50ep

CXR ViT-L/14 — captioning 50ep

ViT-L/14 chest-radiograph vision encoder, pretrained from scratch with the objective: Global image -> full report captioning, extended to 50 epochs.

Part of a controlled study comparing pretraining objectives for chest-X-ray vision encoders (all share the same ViT-L/14 backbone, data, and budget).

Files
File What it is
encoder_final.pt Vision encoder — ViT-L/14 trunk only ( state_dict under key encoder_state , 294 tensors). Use this to extract frozen image features.
model_best.pt Full pretraining model — ViT-L/14 + GPT-2-style causal captioning decoder ( state_dict under key model_state ).
Architecture & training
  • ViT-L/14, 384x384 input -> 729 patch tokens (27x27 grid), trained from scratch in bf16. Text side (where present) uses the GPT-2 tokenizer (vocab 50257).
  • Training data: MIMIC-CXR (reports) and Chest ImaGenome (region-phrase / anatomy boxes).
  • Objective: Global image -> full report captioning, extended to 50 epochs.
Downstream results (frozen-feature transfer)

NIH AUC 0.645 (best of from-scratch); CheXpert 0.815; VQA BLEU-4 0.240; DiffVQA ROUGE-L 0.643.

Full per-task comparison across all 9 from-scratch encoders is in the project logbook.

Usage (vision encoder)
import torch
ckpt = torch.load("encoder_final.pt", map_location="cpu", weights_only=False)
vision_state = ckpt["encoder_state"]   # 294 tensors, ViT-L/14
# load into your ViT-L/14 implementation, then forward 384x384 images -> [B, 729, 1024]
Intended use & limitations
  • Research use only. Frozen feature extraction / fine-tuning for chest-X-ray tasks.
  • NOT a diagnostic device. No clinical or patient-facing use.
  • Trained only on adult frontal/lateral CXR distributions of the source datasets; may not generalize to other modalities, body regions, populations, or acquisition settings.
⚠️ License / Data Use Agreement

These weights are derived from MIMIC-CXR and Chest ImaGenome , which are distributed under the PhysioNet Credentialed Health Data Use Agreement . Redistribution and use of models derived from these data are subject to that DUA — you must hold the appropriate credentialed access and comply with its terms. Do not make this repository public or share access without confirming your DUA permits sharing derived model weights.

Citation

If you use this encoder, please cite the source datasets (MIMIC-CXR; Chest ImaGenome) and this project.

Runs of dmusingu cxr-vitl14-captioning-50ep on huggingface.co

0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs

More Information About cxr-vitl14-captioning-50ep huggingface.co Model

More cxr-vitl14-captioning-50ep license Visit here:

https://choosealicense.com/licenses/physionet-dua

cxr-vitl14-captioning-50ep huggingface.co

cxr-vitl14-captioning-50ep huggingface.co is an AI model on huggingface.co that provides cxr-vitl14-captioning-50ep's model effect (), which can be used instantly with this dmusingu cxr-vitl14-captioning-50ep model. huggingface.co supports a free trial of the cxr-vitl14-captioning-50ep model, and also provides paid use of the cxr-vitl14-captioning-50ep. Support call cxr-vitl14-captioning-50ep model through api, including Node.js, Python, http.

cxr-vitl14-captioning-50ep huggingface.co Url

https://huggingface.co/dmusingu/cxr-vitl14-captioning-50ep

dmusingu cxr-vitl14-captioning-50ep online free

cxr-vitl14-captioning-50ep huggingface.co is an online trial and call api platform, which integrates cxr-vitl14-captioning-50ep's modeling effects, including api services, and provides a free online trial of cxr-vitl14-captioning-50ep, you can try cxr-vitl14-captioning-50ep online for free by clicking the link below.

dmusingu cxr-vitl14-captioning-50ep online free url in huggingface.co:

https://huggingface.co/dmusingu/cxr-vitl14-captioning-50ep

cxr-vitl14-captioning-50ep install

cxr-vitl14-captioning-50ep is an open source model from GitHub that offers a free installation service, and any user can find cxr-vitl14-captioning-50ep on GitHub to install. At the same time, huggingface.co provides the effect of cxr-vitl14-captioning-50ep install, users can directly use cxr-vitl14-captioning-50ep installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

cxr-vitl14-captioning-50ep install url in huggingface.co:

https://huggingface.co/dmusingu/cxr-vitl14-captioning-50ep

Url of cxr-vitl14-captioning-50ep

cxr-vitl14-captioning-50ep huggingface.co Url

Provider of cxr-vitl14-captioning-50ep huggingface.co

dmusingu
ORGANIZATIONS

Other API from dmusingu

huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:June 06 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:June 06 2026