kosmos-2-patch14-224 huggingface.co api & microsoft kosmos-2-patch14-224 github AI Model

Introduction of kosmos-2-patch14-224

Model Details of kosmos-2-patch14-224

Kosmos-2: Grounding Multimodal Large Language Models to the World

**[An image of a snowman warming himself by a fire.]**

This Hub repository contains a HuggingFace's transformers implementation of the original Kosmos-2 model from Microsoft.

How to Get Started with the Model

Use the code below to get started with the model.

import requests

from PIL import Image
from transformers import AutoProcessor, AutoModelForVision2Seq


model = AutoModelForVision2Seq.from_pretrained("microsoft/kosmos-2-patch14-224")
processor = AutoProcessor.from_pretrained("microsoft/kosmos-2-patch14-224")

prompt = "<grounding>An image of"

url = "https://huggingface.co/microsoft/kosmos-2-patch14-224/resolve/main/snowman.png"
image = Image.open(requests.get(url, stream=True).raw)

# The original Kosmos-2 demo saves the image first then reload it. For some images, this will give slightly different image input and change the generation outputs.
image.save("new_image.jpg")
image = Image.open("new_image.jpg")

inputs = processor(text=prompt, images=image, return_tensors="pt")

generated_ids = model.generate(
    pixel_values=inputs["pixel_values"],
    input_ids=inputs["input_ids"],
    attention_mask=inputs["attention_mask"],
    image_embeds=None,
    image_embeds_position_mask=inputs["image_embeds_position_mask"],
    use_cache=True,
    max_new_tokens=128,
)
generated_text = processor.batch_decode(generated_ids, skip_special_tokens=True)[0]

# Specify `cleanup_and_extract=False` in order to see the raw model generation.
processed_text = processor.post_process_generation(generated_text, cleanup_and_extract=False)

print(processed_text)
# `<grounding> An image of<phrase> a snowman</phrase><object><patch_index_0044><patch_index_0863></object> warming himself by<phrase> a fire</phrase><object><patch_index_0005><patch_index_0911></object>.`

# By default, the generated  text is cleanup and the entities are extracted.
processed_text, entities = processor.post_process_generation(generated_text)

print(processed_text)
# `An image of a snowman warming himself by a fire.`

print(entities)
# `[('a snowman', (12, 21), [(0.390625, 0.046875, 0.984375, 0.828125)]), ('a fire', (41, 47), [(0.171875, 0.015625, 0.484375, 0.890625)])]`

Tasks

This model is capable of performing different tasks through changing the prompts.

First, let's define a function to run a prompt.

Click to expand

import requests

from PIL import Image
from transformers import AutoProcessor, AutoModelForVision2Seq


model = AutoModelForVision2Seq.from_pretrained("microsoft/kosmos-2-patch14-224")
processor = AutoProcessor.from_pretrained("microsoft/kosmos-2-patch14-224")

url = "https://huggingface.co/microsoft/kosmos-2-patch14-224/resolve/main/snowman.png"
image = Image.open(requests.get(url, stream=True).raw)

def run_example(prompt):

    inputs = processor(text=prompt, images=image, return_tensors="pt")
    generated_ids = model.generate(
      pixel_values=inputs["pixel_values"],
      input_ids=inputs["input_ids"],
      attention_mask=inputs["attention_mask"],
      image_embeds=None,
      image_embeds_position_mask=inputs["image_embeds_position_mask"],
      use_cache=True,
      max_new_tokens=128,
    )
    generated_text = processor.batch_decode(generated_ids, skip_special_tokens=True)[0]
    _processed_text = processor.post_process_generation(generated_text, cleanup_and_extract=False)
    processed_text, entities = processor.post_process_generation(generated_text)

    print(processed_text)
    print(entities)
    print(_processed_text)

Here are the tasks Kosmos-2 could perform:

Click to expand

Multimodal Grounding

• Phrase Grounding

prompt = "<grounding><phrase> a snowman</phrase>"
run_example(prompt)

# a snowman is warming himself by the fire
# [('a snowman', (0, 9), [(0.390625, 0.046875, 0.984375, 0.828125)]), ('the fire', (32, 40), [(0.203125, 0.015625, 0.453125, 0.859375)])]

# <grounding><phrase> a snowman</phrase><object><patch_index_0044><patch_index_0863></object> is warming himself by<phrase> the fire</phrase><object><patch_index_0006><patch_index_0878></object>

• Referring Expression Comprehension

prompt = "<grounding><phrase> a snowman next to a fire</phrase>"
run_example(prompt)

# a snowman next to a fire
# [('a snowman next to a fire', (0, 24), [(0.390625, 0.046875, 0.984375, 0.828125)])]

# <grounding><phrase> a snowman next to a fire</phrase><object><patch_index_0044><patch_index_0863></object>

Multimodal Referring

• Referring expression generation

prompt = "<grounding><phrase> It</phrase><object><patch_index_0044><patch_index_0863></object> is"
run_example(prompt)

# It is snowman in a hat and scarf
# [('It', (0, 2), [(0.390625, 0.046875, 0.984375, 0.828125)])]

# <grounding><phrase> It</phrase><object><patch_index_0044><patch_index_0863></object> is snowman in a hat and scarf

Perception-Language Tasks

• Grounded VQA

prompt = "<grounding> Question: What is special about this image? Answer:"
run_example(prompt)

# Question: What is special about this image? Answer: The image features a snowman sitting by a campfire in the snow.
# [('a snowman', (71, 80), [(0.390625, 0.046875, 0.984375, 0.828125)]), ('a campfire', (92, 102), [(0.109375, 0.640625, 0.546875, 0.984375)])]

# <grounding> Question: What is special about this image? Answer: The image features<phrase> a snowman</phrase><object><patch_index_0044><patch_index_0863></object> sitting by<phrase> a campfire</phrase><object><patch_index_0643><patch_index_1009></object> in the snow.

• Grounded VQA with multimodal referring via bounding boxes

prompt = "<grounding> Question: Where is<phrase> the fire</phrase><object><patch_index_0005><patch_index_0911></object> next to? Answer:"
run_example(prompt)

# Question: Where is the fire next to? Answer: Near the snowman.
# [('the fire', (19, 27), [(0.171875, 0.015625, 0.484375, 0.890625)]), ('the snowman', (50, 61), [(0.390625, 0.046875, 0.984375, 0.828125)])]

# <grounding> Question: Where is<phrase> the fire</phrase><object><patch_index_0005><patch_index_0911></object> next to? Answer: Near<phrase> the snowman</phrase><object><patch_index_0044><patch_index_0863></object>.

Grounded Image captioning

• Brief

prompt = "<grounding> An image of"
run_example(prompt)

# An image of a snowman warming himself by a campfire.
# [('a snowman', (12, 21), [(0.390625, 0.046875, 0.984375, 0.828125)]), ('a campfire', (41, 51), [(0.109375, 0.640625, 0.546875, 0.984375)])]

# <grounding> An image of<phrase> a snowman</phrase><object><patch_index_0044><patch_index_0863></object> warming himself by<phrase> a campfire</phrase><object><patch_index_0643><patch_index_1009></object>.

• Detailed

prompt = "<grounding> Describe this image in detail:"
run_example(prompt)

# Describe this image in detail: The image features a snowman sitting by a campfire in the snow. He is wearing a hat, scarf, and gloves, with a pot nearby and a cup nearby. The snowman appears to be enjoying the warmth of the fire, and it appears to have a warm and cozy atmosphere.
# [('a campfire', (71, 81), [(0.171875, 0.015625, 0.484375, 0.984375)]), ('a hat', (109, 114), [(0.515625, 0.046875, 0.828125, 0.234375)]), ('scarf', (116, 121), [(0.515625, 0.234375, 0.890625, 0.578125)]), ('gloves', (127, 133), [(0.515625, 0.390625, 0.640625, 0.515625)]), ('a pot', (140, 145), [(0.078125, 0.609375, 0.265625, 0.859375)]), ('a cup', (157, 162), [(0.890625, 0.765625, 0.984375, 0.984375)])]

# <grounding> Describe this image in detail: The image features a snowman sitting by<phrase> a campfire</phrase><object><patch_index_0005><patch_index_1007></object> in the snow. He is wearing<phrase> a hat</phrase><object><patch_index_0048><patch_index_0250></object>,<phrase> scarf</phrase><object><patch_index_0240><patch_index_0604></object>, and<phrase> gloves</phrase><object><patch_index_0400><patch_index_0532></object>, with<phrase> a pot</phrase><object><patch_index_0610><patch_index_0872></object> nearby and<phrase> a cup</phrase><object><patch_index_0796><patch_index_1023></object> nearby. The snowman appears to be enjoying the warmth of the fire, and it appears to have a warm and cozy atmosphere.

Draw the bounding bboxes of the entities on the image

Once you have the entities , you can use the following helper function to draw their bounding bboxes on the image:

Click to expand

import cv2
import numpy as np
import os
import requests
import torch
import torchvision.transforms as T

from PIL import Image


def is_overlapping(rect1, rect2):
    x1, y1, x2, y2 = rect1
    x3, y3, x4, y4 = rect2
    return not (x2 < x3 or x1 > x4 or y2 < y3 or y1 > y4)


def draw_entity_boxes_on_image(image, entities, show=False, save_path=None):
    """_summary_
    Args:
        image (_type_): image or image path
        collect_entity_location (_type_): _description_
    """
    if isinstance(image, Image.Image):
        image_h = image.height
        image_w = image.width
        image = np.array(image)[:, :, [2, 1, 0]]
    elif isinstance(image, str):
        if os.path.exists(image):
            pil_img = Image.open(image).convert("RGB")
            image = np.array(pil_img)[:, :, [2, 1, 0]]
            image_h = pil_img.height
            image_w = pil_img.width
        else:
            raise ValueError(f"invaild image path, {image}")
    elif isinstance(image, torch.Tensor):
        image_tensor = image.cpu()
        reverse_norm_mean = torch.tensor([0.48145466, 0.4578275, 0.40821073])[:, None, None]
        reverse_norm_std = torch.tensor([0.26862954, 0.26130258, 0.27577711])[:, None, None]
        image_tensor = image_tensor * reverse_norm_std + reverse_norm_mean
        pil_img = T.ToPILImage()(image_tensor)
        image_h = pil_img.height
        image_w = pil_img.width
        image = np.array(pil_img)[:, :, [2, 1, 0]]
    else:
        raise ValueError(f"invaild image format, {type(image)} for {image}")

    if len(entities) == 0:
        return image

    new_image = image.copy()
    previous_bboxes = []
    # size of text
    text_size = 1
    # thickness of text
    text_line = 1  # int(max(1 * min(image_h, image_w) / 512, 1))
    box_line = 3
    (c_width, text_height), _ = cv2.getTextSize("F", cv2.FONT_HERSHEY_COMPLEX, text_size, text_line)
    base_height = int(text_height * 0.675)
    text_offset_original = text_height - base_height
    text_spaces = 3

    for entity_name, (start, end), bboxes in entities:
        for (x1_norm, y1_norm, x2_norm, y2_norm) in bboxes:
            orig_x1, orig_y1, orig_x2, orig_y2 = int(x1_norm * image_w), int(y1_norm * image_h), int(x2_norm * image_w), int(y2_norm * image_h)
            # draw bbox
            # random color
            color = tuple(np.random.randint(0, 255, size=3).tolist())
            new_image = cv2.rectangle(new_image, (orig_x1, orig_y1), (orig_x2, orig_y2), color, box_line)

            l_o, r_o = box_line // 2 + box_line % 2, box_line // 2 + box_line % 2 + 1

            x1 = orig_x1 - l_o
            y1 = orig_y1 - l_o

            if y1 < text_height + text_offset_original + 2 * text_spaces:
                y1 = orig_y1 + r_o + text_height + text_offset_original + 2 * text_spaces
                x1 = orig_x1 + r_o

            # add text background
            (text_width, text_height), _ = cv2.getTextSize(f"  {entity_name}", cv2.FONT_HERSHEY_COMPLEX, text_size, text_line)
            text_bg_x1, text_bg_y1, text_bg_x2, text_bg_y2 = x1, y1 - (text_height + text_offset_original + 2 * text_spaces), x1 + text_width, y1

            for prev_bbox in previous_bboxes:
                while is_overlapping((text_bg_x1, text_bg_y1, text_bg_x2, text_bg_y2), prev_bbox):
                    text_bg_y1 += (text_height + text_offset_original + 2 * text_spaces)
                    text_bg_y2 += (text_height + text_offset_original + 2 * text_spaces)
                    y1 += (text_height + text_offset_original + 2 * text_spaces)

                    if text_bg_y2 >= image_h:
                        text_bg_y1 = max(0, image_h - (text_height + text_offset_original + 2 * text_spaces))
                        text_bg_y2 = image_h
                        y1 = image_h
                        break

            alpha = 0.5
            for i in range(text_bg_y1, text_bg_y2):
                for j in range(text_bg_x1, text_bg_x2):
                    if i < image_h and j < image_w:
                        if j < text_bg_x1 + 1.35 * c_width:
                            # original color
                            bg_color = color
                        else:
                            # white
                            bg_color = [255, 255, 255]
                        new_image[i, j] = (alpha * new_image[i, j] + (1 - alpha) * np.array(bg_color)).astype(np.uint8)

            cv2.putText(
                new_image, f"  {entity_name}", (x1, y1 - text_offset_original - 1 * text_spaces), cv2.FONT_HERSHEY_COMPLEX, text_size, (0, 0, 0), text_line, cv2.LINE_AA
            )
            # previous_locations.append((x1, y1))
            previous_bboxes.append((text_bg_x1, text_bg_y1, text_bg_x2, text_bg_y2))

    pil_image = Image.fromarray(new_image[:, :, [2, 1, 0]])
    if save_path:
        pil_image.save(save_path)
    if show:
        pil_image.show()

    return new_image


# (The same image from the previous code example)
url = "https://huggingface.co/microsoft/kosmos-2-patch14-224/resolve/main/snowman.png"
image = Image.open(requests.get(url, stream=True).raw)

# From the previous code example
entities = [('a snowman', (12, 21), [(0.390625, 0.046875, 0.984375, 0.828125)]), ('a fire', (41, 47), [(0.171875, 0.015625, 0.484375, 0.890625)])]

# Draw the bounding bboxes
draw_entity_boxes_on_image(image, entities, show=True)

Here is the annotated image:

BibTex and citation info

@article{kosmos-2,
  title={Kosmos-2: Grounding Multimodal Large Language Models to the World},
  author={Zhiliang Peng and Wenhui Wang and Li Dong and Yaru Hao and Shaohan Huang and Shuming Ma and Furu Wei},
  journal={ArXiv},
  year={2023},
  volume={abs/2306}
}

@article{kosmos-1,
  title={Language Is Not All You Need: Aligning Perception with Language Models},
  author={Shaohan Huang and Li Dong and Wenhui Wang and Yaru Hao and Saksham Singhal and Shuming Ma and Tengchao Lv and Lei Cui and Owais Khan Mohammed and Qiang Liu and Kriti Aggarwal and Zewen Chi and Johan Bjorck and Vishrav Chaudhary and Subhojit Som and Xia Song and Furu Wei},
  journal={ArXiv},
  year={2023},
  volume={abs/2302.14045}
}

@article{metalm,
  title={Language Models are General-Purpose Interfaces},
  author={Yaru Hao and Haoyu Song and Li Dong and Shaohan Huang and Zewen Chi and Wenhui Wang and Shuming Ma and Furu Wei},
  journal={ArXiv},
  year={2022},
  volume={abs/2206.06336}
}

Runs of microsoft kosmos-2-patch14-224 on huggingface.co

162.2K

Total runs

24-hour runs

3-day runs

944

7-day runs

-5.0K

30-day runs

More Information About kosmos-2-patch14-224 huggingface.co Model

More kosmos-2-patch14-224 license Visit here:

https://choosealicense.com/licenses/mit

kosmos-2-patch14-224 huggingface.co

kosmos-2-patch14-224 huggingface.co is an AI model on huggingface.co that provides kosmos-2-patch14-224's model effect (), which can be used instantly with this microsoft kosmos-2-patch14-224 model. huggingface.co supports a free trial of the kosmos-2-patch14-224 model, and also provides paid use of the kosmos-2-patch14-224. Support call kosmos-2-patch14-224 model through api, including Node.js, Python, http.

kosmos-2-patch14-224 huggingface.co Url

https://huggingface.co/microsoft/kosmos-2-patch14-224

microsoft kosmos-2-patch14-224 online free

kosmos-2-patch14-224 huggingface.co is an online trial and call api platform, which integrates kosmos-2-patch14-224's modeling effects, including api services, and provides a free online trial of kosmos-2-patch14-224, you can try kosmos-2-patch14-224 online for free by clicking the link below.

microsoft kosmos-2-patch14-224 online free url in huggingface.co:

https://huggingface.co/microsoft/kosmos-2-patch14-224

kosmos-2-patch14-224 install

kosmos-2-patch14-224 is an open source model from GitHub that offers a free installation service, and any user can find kosmos-2-patch14-224 on GitHub to install. At the same time, huggingface.co provides the effect of kosmos-2-patch14-224 install, users can directly use kosmos-2-patch14-224 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

kosmos-2-patch14-224 install url in huggingface.co:

https://huggingface.co/microsoft/kosmos-2-patch14-224

huggingface.co

microsoft/table-transformer-detection

Total runs: 3.6M

Run Growth: 256.3K

Growth Rate: 7.11%

Updated:September 06 2023

huggingface.co

microsoft/deberta-v3-base

Total runs: 2.4M

Run Growth: 60.8K

Growth Rate: 2.54%

Updated:September 22 2022

huggingface.co

microsoft/TRELLIS-image-large

Total runs: 2.4M

Run Growth: -851.4K

Growth Rate: -35.81%

Updated:December 07 2024

huggingface.co

microsoft/BiomedNLP-BiomedBERT-base-uncased-abstract

Total runs: 1.7M

Run Growth: 773.5K

Growth Rate: 45.07%

Updated:November 07 2023

huggingface.co

microsoft/table-transformer-structure-recognition-v1.1-all

Total runs: 1.6M

Run Growth: 1.4M

Growth Rate: 86.30%

Updated:November 19 2023

huggingface.co

microsoft/Phi-3.5-vision-instruct

Total runs: 1.6M

Run Growth: 296.4K

Growth Rate: 19.46%

Updated:December 11 2025

huggingface.co

microsoft/Phi-4-mini-instruct

Total runs: 1.5M

Run Growth: 850.6K

Growth Rate: 57.15%

Updated:December 11 2025

huggingface.co

microsoft/mdeberta-v3-base

Total runs: 1.4M

Run Growth: 211.3K

Growth Rate: 15.18%

Updated:April 06 2023

huggingface.co

microsoft/table-transformer-structure-recognition

Total runs: 1.3M

Run Growth: 54.2K

Growth Rate: 4.26%

Updated:September 06 2023

huggingface.co

microsoft/VibeVoice-Realtime-0.5B

Total runs: 1.2M

Run Growth: 861.9K

Growth Rate: 70.32%

Updated:December 12 2025

huggingface.co

microsoft/deberta-v3-large

Total runs: 963.2K

Run Growth: -102.8K

Growth Rate: -10.95%

Updated:March 19 2023

huggingface.co

microsoft/tapex-base-finetuned-wikisql

Total runs: 911.5K

Run Growth: 13.4K

Growth Rate: 1.47%

Updated:January 25 2023

huggingface.co

microsoft/BiomedCLIP-PubMedBERT_256-vit_base_patch16_224

Total runs: 872.5K

Run Growth: 254.0K

Growth Rate: 29.26%

Updated:January 15 2025

huggingface.co

microsoft/Florence-2-large

Total runs: 841.8K

Run Growth: -337.6K

Growth Rate: -38.76%

Updated:August 05 2025

huggingface.co

microsoft/VibeVoice-ASR

Total runs: 737.4K

Run Growth: 179.7K

Growth Rate: 24.37%

Updated:January 27 2026

huggingface.co

microsoft/Phi-3-mini-4k-instruct

Total runs: 734.7K

Run Growth: 3.2K

Growth Rate: 0.43%

Updated:December 11 2025

huggingface.co

microsoft/Phi-3.5-mini-instruct

Total runs: 724.4K

Run Growth: -184.2K

Growth Rate: -25.06%

Updated:December 11 2025

huggingface.co

microsoft/trocr-base-printed

Total runs: 700.0K

Run Growth: -438.6K

Growth Rate: -61.73%

Updated:May 28 2024

huggingface.co

microsoft/Florence-2-base

Total runs: 690.9K

Run Growth: -139.5K

Growth Rate: -20.77%

Updated:August 05 2025

huggingface.co

microsoft/Phi-tiny-MoE-instruct

Total runs: 679.8K

Run Growth: 136.1K

Growth Rate: 20.26%

Updated:December 11 2025

huggingface.co

microsoft/layoutlmv3-base

Total runs: 656.6K

Run Growth: -92.9K

Growth Rate: -14.82%

Updated:April 10 2024

huggingface.co

microsoft/phi-2

Total runs: 622.7K

Run Growth: -1.0M

Growth Rate: -153.26%

Updated:December 08 2025

huggingface.co

microsoft/deberta-v3-small

Total runs: 615.8K

Run Growth: -551.8K

Growth Rate: -91.29%

Updated:September 26 2022

huggingface.co

microsoft/layoutlmv2-base-uncased

Total runs: 611.6K

Run Growth: -58.1K

Growth Rate: -9.87%

Updated:September 16 2022

huggingface.co

microsoft/deberta-large-mnli

Total runs: 602.2K

Run Growth: -100.6K

Growth Rate: -16.81%

Updated:May 22 2021

huggingface.co

microsoft/wavlm-base-plus

Total runs: 589.2K

Run Growth: 40.4K

Growth Rate: 6.82%

Updated:December 23 2021

huggingface.co

microsoft/llmlingua-2-xlm-roberta-large-meetingbank

Total runs: 584.5K

Run Growth: 330.0K

Growth Rate: 62.26%

Updated:January 08 2025

huggingface.co

microsoft/wavlm-large

Total runs: 574.6K

Run Growth: 213.5K

Growth Rate: 37.09%

Updated:February 03 2022

huggingface.co

microsoft/resnet-18

Total runs: 537.9K

Run Growth: 432.8K

Growth Rate: 81.51%

Updated:April 08 2024

huggingface.co

microsoft/phi-4

Total runs: 493.3K

Run Growth: -256.3K

Growth Rate: -52.90%

Updated:November 25 2025

huggingface.co

microsoft/deberta-xlarge-mnli

Total runs: 412.9K

Run Growth: -205.6K

Growth Rate: -49.35%

Updated:June 27 2022

huggingface.co

microsoft/swinv2-tiny-patch4-window16-256

Total runs: 404.6K

Run Growth: 5.4K

Growth Rate: 1.40%

Updated:December 10 2022

huggingface.co

microsoft/unispeech-sat-large-sv

Total runs: 398.2K

Run Growth: 397.1K

Growth Rate: 99.71%

Updated:December 18 2021

huggingface.co

microsoft/Phi-4-multimodal-instruct

Total runs: 375.2K

Run Growth: 66.7K

Growth Rate: 18.24%

Updated:December 11 2025

huggingface.co

microsoft/markuplm-base

Total runs: 361.2K

Run Growth: 143.5K

Growth Rate: 40.25%

Updated:December 15 2022

huggingface.co

microsoft/graphcodebert-base

Total runs: 334.3K

Run Growth: 99.7K

Growth Rate: 29.59%

Updated:September 27 2022

huggingface.co

microsoft/resnet-50

Total runs: 311.8K

Run Growth: 8.1K

Growth Rate: 2.65%

Updated:February 14 2024

huggingface.co

microsoft/VibeVoice-ASR-HF

Total runs: 309.0K

Run Growth: 71.7K

Growth Rate: 23.19%

Updated:March 09 2026

huggingface.co

microsoft/codebert-base

Total runs: 301.6K

Run Growth: 34.1K

Growth Rate: 11.25%

Updated:February 12 2022

huggingface.co

microsoft/BiomedNLP-BiomedBERT-base-uncased-abstract-fulltext

Total runs: 301.3K

Run Growth: 87.8K

Growth Rate: 27.79%

Updated:November 07 2023

huggingface.co

microsoft/DialoGPT-medium

Total runs: 275.1K

Run Growth: -52.2K

Growth Rate: -19.03%

Updated:February 29 2024

huggingface.co

microsoft/Phi-3-mini-128k-instruct

Total runs: 245.4K

Run Growth: 9.9K

Growth Rate: 3.89%

Updated:December 11 2025

huggingface.co

microsoft/beit-base-patch16-224

Total runs: 241.3K

Run Growth: 203.6K

Growth Rate: 88.76%

Updated:April 21 2024

huggingface.co

microsoft/xclip-base-patch32

Total runs: 234.8K

Run Growth: 127.7K

Growth Rate: 54.59%

Updated:February 04 2024

huggingface.co

microsoft/BiomedVLP-CXR-BERT-specialized

Total runs: 219.2K

Run Growth: 60.1K

Growth Rate: 35.12%

Updated:July 25 2025

huggingface.co

microsoft/VibeVoice-1.5B

Total runs: 215.3K

Run Growth: 147.4K

Growth Rate: 70.18%

Updated:January 22 2026

huggingface.co

microsoft/layoutlm-base-uncased

Total runs: 187.7K

Run Growth: 46.7K

Growth Rate: 25.67%

Updated:April 16 2024

huggingface.co

microsoft/speecht5_tts

Total runs: 184.4K

Run Growth: -3.1K

Growth Rate: -1.65%

Updated:November 08 2023

huggingface.co

microsoft/trocr-large-handwritten

Total runs: 170.9K

Run Growth: -145.0K

Growth Rate: -76.07%

Updated:May 28 2024

huggingface.co

microsoft/harrier-oss-v1-0.6b

Total runs: 167.1K

Run Growth: 167.1K

Growth Rate: 100.00%

Updated:March 30 2026

huggingface.co

microsoft/trocr-base-handwritten

Total runs: 151.6K

Run Growth: -15.2K

Growth Rate: -9.95%

Updated:February 11 2025

huggingface.co

microsoft/deberta-v2-xlarge

Total runs: 144.4K

Run Growth: -38.2K

Growth Rate: -26.11%

Updated:September 26 2022

huggingface.co

microsoft/llmlingua-2-bert-base-multilingual-cased-meetingbank

Total runs: 142.2K

Run Growth: 4.8K

Growth Rate: 3.76%

Updated:January 08 2025

huggingface.co

microsoft/unixcoder-base

Total runs: 139.9K

Run Growth: -46.7K

Growth Rate: -33.08%

Updated:July 31 2024

huggingface.co

microsoft/Phi-3-vision-128k-instruct

Total runs: 138.5K

Run Growth: 35.1K

Growth Rate: 27.12%

Updated:December 11 2025

huggingface.co

microsoft/wavlm-base-plus-sd

Total runs: 134.8K

Run Growth: 21.3K

Growth Rate: 14.65%

Updated:March 25 2022

huggingface.co

microsoft/Phi-mini-MoE-instruct

Total runs: 133.2K

Run Growth: 30.7K

Growth Rate: 23.60%

Updated:December 11 2025

huggingface.co

microsoft/wavlm-base-plus-sv

Total runs: 133.2K

Run Growth: -42.0K

Growth Rate: -32.00%

Updated:March 25 2022

huggingface.co

microsoft/infoxlm-large

Total runs: 130.3K

Run Growth: 9.6K

Growth Rate: 7.62%

Updated:August 04 2021

huggingface.co

microsoft/biogpt

Total runs: 129.9K

Run Growth: -150.9K

Growth Rate: -114.85%

Updated:February 03 2023

huggingface.co

microsoft/trocr-large-printed

Total runs: 128.3K

Run Growth: -587.5K

Growth Rate: -394.86%

Updated:May 28 2024

huggingface.co

microsoft/deberta-base

Total runs: 120.5K

Run Growth: -213.6K

Growth Rate: -172.71%

Updated:September 26 2022

huggingface.co

microsoft/mpnet-base

Total runs: 109.5K

Run Growth: -39.3K

Growth Rate: -36.02%

Updated:February 29 2024

huggingface.co

microsoft/Phi-3.5-MoE-instruct

Total runs: 107.9K

Run Growth: 10.3K

Growth Rate: 9.85%

Updated:December 11 2025

huggingface.co

microsoft/swin-base-patch4-window12-384-in22k

Total runs: 96.0K

Run Growth: 74.3K

Growth Rate: 85.61%

Updated:May 17 2022

huggingface.co

microsoft/speecht5_asr

Total runs: 89.4K

Run Growth: -36.9K

Growth Rate: -40.55%

Updated:March 23 2023

huggingface.co

microsoft/prophetnet-large-uncased

Total runs: 89.2K

Run Growth: -45.7K

Growth Rate: -51.53%

Updated:April 27 2023

huggingface.co

microsoft/harrier-oss-v1-270m

Total runs: 88.5K

Run Growth: 88.5K

Growth Rate: 100.00%

Updated:March 30 2026

huggingface.co

microsoft/kosmos-2.5

Total runs: 84.9K

Run Growth: -39.8K

Growth Rate: -47.29%

Updated:August 28 2025

huggingface.co

microsoft/codebert-base-mlm

Total runs: 84.4K

Run Growth: 72.3K

Growth Rate: 85.93%

Updated:January 09 2023

huggingface.co

microsoft/speecht5_hifigan

Total runs: 75.9K

Run Growth: -18.7K

Growth Rate: -25.28%

Updated:February 02 2023

huggingface.co

microsoft/layoutlmv3-large

Total runs: 73.6K

Run Growth: 1.9K

Growth Rate: 2.68%

Updated:September 16 2022

huggingface.co

microsoft/udop-large

Total runs: 72.2K

Run Growth: -30.7K

Growth Rate: -42.63%

Updated:December 03 2025

huggingface.co

microsoft/Phi-4-reasoning-vision-15B

Total runs: 72.2K

Run Growth: 45.8K

Growth Rate: 63.45%

Updated:March 19 2026

huggingface.co

microsoft/phi-1_5

Total runs: 69.5K

Run Growth: -18.5K

Growth Rate: -23.62%

Updated:November 25 2025

huggingface.co

microsoft/DialoGPT-small

Total runs: 68.5K

Run Growth: 12.3K

Growth Rate: 17.92%

Updated:February 29 2024

huggingface.co

microsoft/bitnet-b1.58-2B-4T-gguf

Total runs: 67.2K

Run Growth: 32.9K

Growth Rate: 48.99%

Updated:December 18 2025

huggingface.co

microsoft/Phi-3-mini-4k-instruct-gguf

Total runs: 61.8K

Run Growth: 5.7K

Growth Rate: 9.05%

Updated:December 11 2025

huggingface.co

microsoft/TRELLIS-text-xlarge

Total runs: 60.4K

Run Growth: 0

Growth Rate: 0.00%

Updated:March 24 2025

huggingface.co

microsoft/deberta-base-mnli

Total runs: 60.1K

Run Growth: -55.4K

Growth Rate: -93.48%

Updated:December 09 2021

huggingface.co

microsoft/Phi-3-medium-4k-instruct

Total runs: 56.3K

Run Growth: 44.4K

Growth Rate: 79.06%

Updated:December 11 2025

huggingface.co

microsoft/trocr-small-handwritten

Total runs: 54.6K

Run Growth: 39.9K

Growth Rate: 78.11%

Updated:May 28 2024

huggingface.co

microsoft/deberta-v3-xsmall

Total runs: 53.1K

Run Growth: 13.9K

Growth Rate: 26.65%

Updated:September 26 2022

huggingface.co

microsoft/rad-dino

Total runs: 52.5K

Run Growth: 15.3K

Growth Rate: 32.26%

Updated:October 09 2025

huggingface.co

microsoft/harrier-oss-v1-27b

Total runs: 51.7K

Run Growth: 51.7K

Growth Rate: 100.00%

Updated:March 30 2026

huggingface.co

microsoft/Florence-2-base-ft

Total runs: 50.8K

Run Growth: 11.9K

Growth Rate: 23.79%

Updated:August 05 2025

huggingface.co

microsoft/wavlm-base

Total runs: 46.7K

Run Growth: -11.0K

Growth Rate: -26.03%

Updated:December 23 2021

huggingface.co

microsoft/wavlm-base-sv

Total runs: 42.6K

Run Growth: 11.9K

Growth Rate: 27.41%

Updated:March 25 2022

huggingface.co

microsoft/swin-base-patch4-window7-224

Total runs: 42.3K

Run Growth: 9.8K

Growth Rate: 23.40%

Updated:September 11 2023

huggingface.co

microsoft/Phi-4-mini-reasoning

Total runs: 39.0K

Run Growth: 18.9K

Growth Rate: 48.35%

Updated:December 11 2025

huggingface.co

microsoft/BiomedVLP-CXR-BERT-general

Total runs: 36.4K

Run Growth: 12.4K

Growth Rate: 37.20%

Updated:July 25 2025

huggingface.co

microsoft/MiniLM-L12-H384-uncased

Total runs: 34.7K

Run Growth: 7.6K

Growth Rate: 23.13%

Updated:May 20 2021

huggingface.co

microsoft/Florence-2-large-ft

Total runs: 32.8K

Run Growth: -818

Growth Rate: -2.49%

Updated:August 05 2025

huggingface.co

microsoft/infoxlm-base

Total runs: 32.7K

Run Growth: -9.0K

Growth Rate: -28.30%

Updated:August 04 2021

huggingface.co

microsoft/Multilingual-MiniLM-L12-H384

Total runs: 31.2K

Run Growth: -25.8K

Growth Rate: -75.33%

Updated:August 10 2022

huggingface.co

microsoft/BiomedParse

Total runs: 31.0K

Run Growth: -8.1K

Growth Rate: -26.25%

Updated:October 11 2025

huggingface.co

microsoft/trocr-small-printed

Total runs: 30.1K

Run Growth: -7.5K

Growth Rate: -24.82%

Updated:May 28 2024

huggingface.co

microsoft/swin-tiny-patch4-window7-224

Total runs: 27.3K

Run Growth: -9.9K

Growth Rate: -35.15%

Updated:September 16 2023

huggingface.co

microsoft/rad-dino-maira-2

Total runs: 24.8K

Run Growth: 16.7K

Growth Rate: 68.42%

Updated:August 22 2024

huggingface.co

microsoft/BiomedNLP-BiomedBERT-large-uncased-abstract

Total runs: 24.2K

Run Growth: 21.5K

Growth Rate: 87.29%

Updated:November 07 2023

microsoft / kosmos-2-patch14-224

Introduction of kosmos-2-patch14-224

Model Details of kosmos-2-patch14-224

Kosmos-2: Grounding Multimodal Large Language Models to the World

How to Get Started with the Model

Tasks

Multimodal Grounding

• Phrase Grounding

• Referring Expression Comprehension

Multimodal Referring

• Referring expression generation

Perception-Language Tasks

• Grounded VQA

• Grounded VQA with multimodal referring via bounding boxes

Grounded Image captioning

• Brief

• Detailed

Draw the bounding bboxes of the entities on the image

BibTex and citation info

Runs of microsoft kosmos-2-patch14-224 on huggingface.co

More Information About kosmos-2-patch14-224 huggingface.co Model

More kosmos-2-patch14-224 license Visit here:

kosmos-2-patch14-224 huggingface.co

kosmos-2-patch14-224 huggingface.co Url

microsoft kosmos-2-patch14-224 online free

microsoft kosmos-2-patch14-224 online free url in huggingface.co:

kosmos-2-patch14-224 install

kosmos-2-patch14-224 install url in huggingface.co:

Url of kosmos-2-patch14-224

kosmos-2-patch14-224 huggingface.co Url

Provider of kosmos-2-patch14-224 huggingface.co

Other API from microsoft