This SAE is trained on penultimate-layer activations of a DINOv2 ViT-B/14 model applied to insect images from
BIOSCAN-5M
. Its latents capture interpretable visual features that correspond to species-level morphological traits (e.g., wing venation, body coloration, antennal structure). These latents are used to steer a multimodal LLM (Qwen2.5-VL-72B) into generating natural-language trait annotations.
Architecture:
Base encoder: DINOv2 ViT-B/14 (frozen), activations from layer
-2
SAE input dimension (
d-vit
): 768
Expansion factor: 32 →
24,576 latent dimensions
Training data: patch-level activations from BIOSCAN-5M
Usage
Clone the
code repository
(which vendors the
saev
library), then load and run the SAE as follows:
import torch
import saev.nn
import saev.activations
from torchvision import datasets
from torch.utils.data import DataLoader
device = "cuda"if torch.cuda.is_available() else"cpu"# Build the image transform and DINOv2 ViT-B/14 backbone
img_transform = saev.activations.make_img_transform("dinov2", "sae.pt")
vit = saev.activations.make_vit("dinov2", "dinov2_vitb14")
# Wrap the ViT to record activations from layer 10 (penultimate), 256 patches
recorded_vit = saev.activations.RecordedVisionTransformer(
vit, n_patches=256, cls_token=True, layers=[10]
).to(device)
# Load the SAE checkpoint
sae = saev.nn.load("sae.pt").to(device)
sae.eval()
# --- Encode a batch of images ---# dataset: torchvision ImageFolder with images at 224x224
dataset = datasets.ImageFolder(root="/path/to/images/train")
defcollate_fn(batch):
images, labels = zip(*batch)
returnlist(images), torch.tensor(labels)
loader = DataLoader(dataset, batch_size=32, shuffle=False, collate_fn=collate_fn)
with torch.no_grad():
for images, labels in loader:
images_t = torch.stack(img_transform(images)).to(device)
# vit_acts: (batch, n_layers, n_patches+1, d_vit)
_, vit_acts = recorded_vit(images_t)
# Select layer 0 of the recorded layers, drop the CLS token
vit_acts = vit_acts[:, 0, 1:, :] # (batch, 256, 768)# SAE forward: returns (reconstruction, features, aux)
_, f_x, _ = sae(vit_acts) # f_x: (batch, 256, 24576)# Threshold activations to find active latents (default thresh=0.9)
active = (f_x > 0.9) # (batch, 256, 24576) bool
The active latent indices per patch identify which SAE dimensions fire on each image region. These are used downstream to find species-prominent latents and generate trait annotations via an MLLM. See
create_trait_dataset_mllm_sae.py
for the full pipeline.
Training Details
Training data:
BIOSCAN-5M insect images preprocessed into
ImageFolder
layout
Learning rate:
1e-3
Sparsity coefficient (alpha):
: 4e-4
Data patches:
patch-level (256 patches/image), unscaled mean and norm
Intended Use
Generating morphological trait annotations for organismal (insect) images
Interpretability research on vision foundation models via SAE latent analysis
Downstream fine-tuning of classifiers using trait-annotated data (e.g., with BioCLIP)
Citation
@inproceedings{
pahuja2026automatic,
title={Automatic Image-Level Morphological Trait Annotation for Organismal Images},
author={Vardaan Pahuja and Samuel Stevens and Alyson East and Sydne Record and Yu Su},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026},
url={https://openreview.net/forum?id=oFRbiaib5Q}
}
Acknowledgments
Supported by NSF CAREER #2443149, NSF OAC 2118240, and an Alfred P. Sloan Foundation Fellowship.
Computational resources provided by the Ohio Supercomputer Center.
SAE training infrastructure from
SAEV
.
Runs of osunlp sae-trait-annotation on huggingface.co
0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About sae-trait-annotation huggingface.co Model
sae-trait-annotation huggingface.co is an AI model on huggingface.co that provides sae-trait-annotation's model effect (), which can be used instantly with this osunlp sae-trait-annotation model. huggingface.co supports a free trial of the sae-trait-annotation model, and also provides paid use of the sae-trait-annotation. Support call sae-trait-annotation model through api, including Node.js, Python, http.
sae-trait-annotation huggingface.co is an online trial and call api platform, which integrates sae-trait-annotation's modeling effects, including api services, and provides a free online trial of sae-trait-annotation, you can try sae-trait-annotation online for free by clicking the link below.
osunlp sae-trait-annotation online free url in huggingface.co:
sae-trait-annotation is an open source model from GitHub that offers a free installation service, and any user can find sae-trait-annotation on GitHub to install. At the same time, huggingface.co provides the effect of sae-trait-annotation install, users can directly use sae-trait-annotation installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
sae-trait-annotation install url in huggingface.co: