ViT-L/14 chest-radiograph vision encoder,
pretrained from scratch
with the objective:
Global image<->report contrastive alignment (SigLIP sigmoid loss).
Part of a controlled study comparing pretraining objectives for chest-X-ray vision encoders
(all share the same ViT-L/14 backbone, data, and budget).
Files
File
What it is
encoder_final.pt
Vision encoder
— ViT-L/14 trunk only (
state_dict
under key
encoder_state
, 294 tensors). Use this to extract frozen image features.
model_best.pt
Full pretraining model
— ViT-L/14 + bidirectional text encoder + projection heads + logit_scale (
state_dict
under key
model_state
).
Architecture & training
ViT-L/14, 384x384 input -> 729 patch tokens (27x27 grid), trained
from scratch
in bf16. Text side (where present) uses the GPT-2 tokenizer (vocab 50257).
Training data:
MIMIC-CXR (reports) and Chest ImaGenome (region-phrase / anatomy boxes).
Objective:
Global image<->report contrastive alignment (SigLIP sigmoid loss).
Downstream results (frozen-feature transfer)
Generative-PG mIoU 0.209 (best of from-scratch); CheXpert AUC 0.784; AD
[email protected]
0.050.
Full per-task comparison across all 9 from-scratch encoders is in the project logbook.
Usage (vision encoder)
import torch
ckpt = torch.load("encoder_final.pt", map_location="cpu", weights_only=False)
vision_state = ckpt["encoder_state"] # 294 tensors, ViT-L/14# load into your ViT-L/14 implementation, then forward 384x384 images -> [B, 729, 1024]
Intended use & limitations
Research use only.
Frozen feature extraction / fine-tuning for chest-X-ray tasks.
NOT a diagnostic device.
No clinical or patient-facing use.
Trained only on adult frontal/lateral CXR distributions of the source datasets; may not
generalize to other modalities, body regions, populations, or acquisition settings.
⚠️ License / Data Use Agreement
These weights are derived from
MIMIC-CXR
and
Chest ImaGenome
, which are distributed
under the
PhysioNet Credentialed Health Data Use Agreement
. Redistribution and use of
models derived from these data are subject to that DUA — you must hold the appropriate
credentialed access and comply with its terms. Do
not
make this repository public or
share access without confirming your DUA permits sharing derived model weights.
Citation
If you use this encoder, please cite the source datasets (MIMIC-CXR; Chest ImaGenome) and
this project.
Runs of dmusingu cxr-vitl14-siglip on huggingface.co
0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About cxr-vitl14-siglip huggingface.co Model
cxr-vitl14-siglip huggingface.co is an AI model on huggingface.co that provides cxr-vitl14-siglip's model effect (), which can be used instantly with this dmusingu cxr-vitl14-siglip model. huggingface.co supports a free trial of the cxr-vitl14-siglip model, and also provides paid use of the cxr-vitl14-siglip. Support call cxr-vitl14-siglip model through api, including Node.js, Python, http.
cxr-vitl14-siglip huggingface.co is an online trial and call api platform, which integrates cxr-vitl14-siglip's modeling effects, including api services, and provides a free online trial of cxr-vitl14-siglip, you can try cxr-vitl14-siglip online for free by clicking the link below.
dmusingu cxr-vitl14-siglip online free url in huggingface.co:
cxr-vitl14-siglip is an open source model from GitHub that offers a free installation service, and any user can find cxr-vitl14-siglip on GitHub to install. At the same time, huggingface.co provides the effect of cxr-vitl14-siglip install, users can directly use cxr-vitl14-siglip installed effect in huggingface.co for debugging and trial. It also supports api for free installation.