MolmoPoint-8B is a fully-open VLM developed by the Allen Institute for AI (Ai2) that support image, video and multi-image understanding and grounding.
It has new pointing mechansim that improves image pointing, video pointing, and video tracking, see our technical report for details.
Note the huggingface MolmoPoint model does not support training, see our github repo for the training code.
We recommend running MolmoPoint with
logits_processor=model.build_logit_processor_from_inputs(model_inputs)
to enforce points tokens are generated in a valid way.
In MolmoPoint, instead of coordinates points will be generated as a series of special
tokens, decoding the tokens back into points requires some additional
metadata from the preprocessor.
The metadata is returned by the preprocessor using the
return_pointing_metadata
flag.
Then
model.extract_image_points
and
model.extract_video_points
do the decoding, they
return a list of ({image_id|timestamps}, object_id, pixel_x, pixel_y) output points.
Image Pointing Example:
from transformers import AutoProcessor, AutoModelForImageTextToText
import torch
import numpy as np
checkpoint_dir = "allenai/MolmoPoint-8B"# or path to a converted HF checkpoint
model = AutoModelForImageTextToText.from_pretrained(
checkpoint_dir,
trust_remote_code=True,
dtype="auto",
device_map="auto",
)
processor = AutoProcessor.from_pretrained(
checkpoint_dir,
trust_remote_code=True,
padding_side="left",
)
image_messages = [
{
"role": "user",
"content": [
{"type": "text", "text": "Point to the boats"},
{"type": "image", "image": "https://assets.thesparksite.com/uploads/sites/5550/2025/01/aerial-view-of-boats-yachts-water-bike-and-woode-2023-11-27-04-51-17-utc.jpg"},
{"type": "image", "image": "https://storage.googleapis.com/ai2-playground-molmo/promptTemplates/Stock_278013497.jpeg"},
]
}
]
inputs = processor.apply_chat_template(
image_messages,
tokenize=True,
add_generation_prompt=True,
return_tensors="pt",
return_dict=True,
padding=True,
return_pointing_metadata=True
)
metadata = inputs.pop("metadata")
inputs = {k: v.to("cuda") for k, v in inputs.items()}
with torch.inference_mode(), torch.autocast("cuda", dtype=torch.bfloat16):
output = model.generate(
**inputs,
logits_processor=model.build_logit_processor_from_inputs(inputs),
max_new_tokens=200
)
generated_tokens = output[:, inputs["input_ids"].size(1):]
generated_text = processor.post_process_image_text_to_text(generated_tokens, skip_special_tokens=False, clean_up_tokenization_spaces=False)[0]
points = model.extract_image_points(
generated_text,
metadata["token_pooling"],
metadata["subpatch_mapping"],
metadata["image_sizes"]
)
# points as a list of [object_id, image_num, x, y]# For multiple images, `image_num` is the index of the image the point is inprint(np.array(points))
Video Pointing Example:
video_path = "https://storage.googleapis.com/oe-training-public/demo_videos/many_penguins.mp4"
video_messages = [
{
"role": "user",
"content": [
dict(type="text", text="Point to the penguins"),
dict(type="video", video=video_path),
]
}
]
inputs = processor.apply_chat_template(
video_messages,
tokenize=True,
add_generation_prompt=True,
return_tensors="pt",
return_dict=True,
padding=True,
return_pointing_metadata=True
)
metadata = inputs.pop("metadata")
inputs = {k: v.to("cuda") for k, v in inputs.items()}
with torch.inference_mode(), torch.autocast("cuda", dtype=torch.bfloat16):
output = model.generate(
**inputs,
logits_processor=model.build_logit_processor_from_inputs(inputs),
max_new_tokens=200
)
generated_tokens = output[:, inputs['input_ids'].size(1):]
generated_text = processor.post_process_image_text_to_text(generated_tokens, skip_special_tokens=False, clean_up_tokenization_spaces=False)[0]
video_points = model.extract_video_points(
generated_text,
metadata["token_pooling"],
metadata["subpatch_mapping"],
metadata["timestamps"],
metadata["video_size"]
)
# points as a list of [object_id, image_num, x, y]# For tracking, object_id uniquely identifies objects that might appear multiple frames.print(np.array(video_points))
License and Use
This model is licensed under Apache 2.0. It is intended for research and educational use in accordance with Ai2’s Responsible Use Guidelines. This model is trained on third party datasets that are subject to academic and non-commercial research use only. Please review the sources to determine if this model is appropriate for your use case.
Runs of allenai MolmoPoint-8B on huggingface.co
3.9K
Total runs
0
24-hour runs
-202
3-day runs
-549
7-day runs
-788
30-day runs
More Information About MolmoPoint-8B huggingface.co Model
MolmoPoint-8B huggingface.co is an AI model on huggingface.co that provides MolmoPoint-8B's model effect (), which can be used instantly with this allenai MolmoPoint-8B model. huggingface.co supports a free trial of the MolmoPoint-8B model, and also provides paid use of the MolmoPoint-8B. Support call MolmoPoint-8B model through api, including Node.js, Python, http.
MolmoPoint-8B huggingface.co is an online trial and call api platform, which integrates MolmoPoint-8B's modeling effects, including api services, and provides a free online trial of MolmoPoint-8B, you can try MolmoPoint-8B online for free by clicking the link below.
allenai MolmoPoint-8B online free url in huggingface.co:
MolmoPoint-8B is an open source model from GitHub that offers a free installation service, and any user can find MolmoPoint-8B on GitHub to install. At the same time, huggingface.co provides the effect of MolmoPoint-8B install, users can directly use MolmoPoint-8B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.