cjwbw / pix2struct

Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding

replicate.com
Total runs: 6.1K
24-hour runs: 0
7-day runs: 0
30-day runs: 0
Github
Model's Last Updated: März 29 2023

Introduction of pix2struct

Model Details of pix2struct

Readme

Pix2Struct

Models included in this demo:

model_urls = {
    "textcaps": "google/pix2struct-textcaps-large", # Finetuned on TextCaps
    "screen2words": "google/pix2struct-screen2words-large", # Finetuned on Screen2Words
    "widgetcaption": "google/pix2struct-widget-captioning-large", # Finetuned on Widget Captioning (captioning a UI component on a screen)
    "infographics": "google/pix2struct-infographics-vqa-large", # Infographics
    "docvqa": "google/pix2struct-docvqa-large", # Visual question answering
    "ai2d": "google/pix2struct-ai2d-large", # Scienfic diagram
}

model_image

Pix2Struct is an image encoder - text decoder model that is trained on image-text pairs for various tasks, including image captionning and visual question answering. The full list of available models can be found on the Table 1 of the paper:

Table 1 - paper

The abstract of the model states that:

Visually-situated language is ubiquitous—sources range from textbooks with diagrams to web pages with images and tables, to mobile apps with buttons and forms. Perhaps due to this diversity, previous work has typically relied on domainspecific recipes with limited sharing of the underlying data, model architectures, and objectives. We present Pix2Struct, a pretrained image-to-text model for purely visual language understanding, which can be finetuned on tasks containing visually-situated language. Pix2Struct is pretrained by learning to parse masked screenshots of web pages into simplified HTML. The web, with its richness of visual elements cleanly reflected in the HTML structure, provides a large source of pretraining data well suited to the diversity of downstream tasks. Intuitively, this objective subsumes common pretraining signals such as OCR, language modeling, image captioning. In addition to the novel pretraining strategy, we introduce a variable-resolution input representation and a more flexible integration of language and vision inputs, where language prompts such as questions are rendered directly on top of the input image. For the first time, we show that a single pretrained model can achieve state-of-the-art results in six out of nine tasks across four domains: documents, illustrations, user interfaces, and natural images.

Contribution

This model was originally contributed by Kenton Lee, Mandar Joshi et al. and added to the Hugging Face ecosystem by Younes Belkada .

Citation

If you want to cite this work, please consider citing the original paper:

@misc{https://doi.org/10.48550/arxiv.2210.03347,
  doi = {10.48550/ARXIV.2210.03347},

  url = {https://arxiv.org/abs/2210.03347},

  author = {Lee, Kenton and Joshi, Mandar and Turc, Iulia and Hu, Hexiang and Liu, Fangyu and Eisenschlos, Julian and Khandelwal, Urvashi and Shaw, Peter and Chang, Ming-Wei and Toutanova, Kristina},

  keywords = {Computation and Language (cs.CL), Computer Vision and Pattern Recognition (cs.CV), FOS: Computer and information sciences, FOS: Computer and information sciences},

  title = {Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding},

  publisher = {arXiv},

  year = {2022},

  copyright = {Creative Commons Attribution 4.0 International}
}

Pricing of pix2struct replicate.com

Run time and cost

This model costs approximately $0.0021 to run on Replicate, or 476 runs per $1, but this varies depending on your inputs. It is also open source and you can run it on your own computer with Docker .

This model runs on Nvidia A100 (80GB) GPU hardware . Predictions typically complete within 2 seconds. The predict time for this model varies significantly based on the inputs.

Runs of cjwbw pix2struct on replicate.com

6.1K
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs

More Information About pix2struct replicate.com Model

pix2struct replicate.com

pix2struct replicate.com is an AI model on replicate.com that provides pix2struct's model effect (Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding), which can be used instantly with this cjwbw pix2struct model. replicate.com supports a free trial of the pix2struct model, and also provides paid use of the pix2struct. Support call pix2struct model through api, including Node.js, Python, http.

pix2struct replicate.com Url

https://replicate.com/cjwbw/pix2struct

cjwbw pix2struct online free

pix2struct replicate.com is an online trial and call api platform, which integrates pix2struct's modeling effects, including api services, and provides a free online trial of pix2struct, you can try pix2struct online for free by clicking the link below.

cjwbw pix2struct online free url in replicate.com:

https://replicate.com/cjwbw/pix2struct

pix2struct install

pix2struct is an open source model from GitHub that offers a free installation service, and any user can find pix2struct on GitHub to install. At the same time, replicate.com provides the effect of pix2struct install, users can directly use pix2struct installed effect in replicate.com for debugging and trial. It also supports api for free installation.

pix2struct install url in replicate.com:

https://replicate.com/cjwbw/pix2struct

pix2struct install url in github:

https://github.com/chenxwh/cog-pix2struct

Url of pix2struct

Provider of pix2struct replicate.com

Other API from cjwbw

replicate

Remove images background

Total runs: 8.3M
Run Growth: 0
Growth Rate: 0.00%
Updated:November 30 2022
replicate

openai/clip-vit-large-patch14 with Transformers

Total runs: 6.9M
Run Growth: 0
Growth Rate: 0.00%
Updated:September 22 2022
replicate

ZoeDepth: Combining relative and metric depth

Total runs: 4.5M
Run Growth: 0
Growth Rate: 0.00%
Updated:März 05 2023
replicate

Anime-themed text-to-image stable diffusion model

Total runs: 4.0M
Run Growth: 0
Growth Rate: 0.00%
Updated:März 20 2024
replicate

high-quality, highly detailed anime style stable-diffusion with better VAE

Total runs: 3.5M
Run Growth: 0
Growth Rate: 0.00%
Updated:Januar 15 2023
replicate

high-quality, highly detailed anime-style Stable Diffusion models

Total runs: 3.3M
Run Growth: 0
Growth Rate: 0.00%
Updated:Januar 23 2023
replicate

Real-ESRGAN: Real-World Blind Super-Resolution

Total runs: 2.2M
Run Growth: 0
Growth Rate: 0.00%
Updated:Februar 19 2023
replicate

powerful open-source visual language model

Total runs: 1.5M
Run Growth: 0
Growth Rate: 0.00%
Updated:November 30 2023
replicate

Dream Shaper stable diffusion

Total runs: 1.3M
Run Growth: 0
Growth Rate: 0.00%
Updated:März 12 2023
replicate

Stable Diffusion on Danbooru images

Total runs: 1.1M
Run Growth: 0
Growth Rate: 0.00%
Updated:Oktober 10 2022
replicate

Colorization using a Generative Color Prior for Natural Images

Total runs: 564.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:August 04 2022
replicate

Real-ESRGAN super-resolution model from ruDALL-E

Total runs: 483.5K
Run Growth: 0
Growth Rate: 0.00%
Updated:September 29 2022
replicate

Robust Monocular Depth Estimation

Total runs: 414.7K
Run Growth: 0
Growth Rate: 0.00%
Updated:Februar 15 2023
replicate

high-quality, highly detailed anime style stable-diffusion

Total runs: 353.7K
Run Growth: 0
Growth Rate: 0.00%
Updated:Dezember 20 2022
replicate

sd-v2 with diffusers, test version!

Total runs: 280.3K
Run Growth: 0
Growth Rate: 0.00%
Updated:Dezember 02 2022
replicate

a dreambooth model trained on a diverse set of analog photographs

Total runs: 234.4K
Run Growth: 0
Growth Rate: 0.00%
Updated:Januar 01 2023
replicate

Practicing Model Scaling for Photo-Realistic Image Restoration In the Wild. This version uses LLaVA-13b for captioning.

Total runs: 186.0K
Run Growth: 0
Growth Rate: 0.00%
Updated:Februar 24 2024
replicate

Demucs Music Source Separation

Total runs: 184.6K
Run Growth: 0
Growth Rate: 0.00%
Updated:Juli 02 2023
replicate

Advanced text-image comprehension and composition based on InternLM

Total runs: 164.3K
Run Growth: 0
Growth Rate: 0.00%
Updated:Oktober 02 2023
replicate

Multi-stage text-to-video generation

Total runs: 143.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:März 24 2023
replicate

Blind Face Restoration with Vector-Quantized Dictionary and Parallel Decoder

Total runs: 140.1K
Run Growth: 0
Growth Rate: 0.00%
Updated:August 04 2022
replicate

Stylized Audio-Driven Single Image Talking Face Animation

Total runs: 127.9K
Run Growth: 0
Growth Rate: 0.00%
Updated:Juni 01 2024
replicate

SeamlessM4T—Massively Multilingual & Multimodal Machine Translation

Total runs: 82.1K
Run Growth: 0
Growth Rate: 0.00%
Updated:September 14 2023
replicate

VideoCrafter2: Text-to-Video and Image-to-Video Generation and Editing

Total runs: 66.7K
Run Growth: 0
Growth Rate: 0.00%
Updated:Januar 31 2024
replicate

stable-diffusion with negative prompts, more scheduler

Total runs: 65.3K
Run Growth: 0
Growth Rate: 0.00%
Updated:Oktober 08 2022
replicate

Background removal model developed by BRIA.AI, trained on a carefully selected dataset and is available as an open-source model for non-commercial use.

Total runs: 55.1K
Run Growth: 0
Growth Rate: 0.00%
Updated:Februar 07 2024
replicate

with large-v2 checkpoint

Total runs: 54.5K
Run Growth: 0
Growth Rate: 0.00%
Updated:Dezember 16 2022
replicate

Unsupervised Night Image Enhancement

Total runs: 41.9K
Run Growth: 0
Growth Rate: 0.00%
Updated:August 14 2022
replicate

Text-to-Image Diffusion Models are Zero-Shot Video Generators

Total runs: 41.5K
Run Growth: 0
Growth Rate: 0.00%
Updated:April 08 2023
replicate

stable-diffusion with v1-5 checkpoint

Total runs: 35.5K
Run Growth: 0
Growth Rate: 0.00%
Updated:Oktober 26 2022
replicate

Tuning-Free Multi-Subject Image Generation with Localized Attention

Total runs: 34.1K
Run Growth: 0
Growth Rate: 0.00%
Updated:Mai 19 2023
replicate

high-quality highly detailed anime stylized latent diffusion model

Total runs: 31.8K
Run Growth: 0
Growth Rate: 0.00%
Updated:Februar 21 2023
replicate

mixed stable diffusion model

Total runs: 30.0K
Run Growth: 0
Growth Rate: 0.00%
Updated:März 05 2023
replicate

Portraits with stable-diffusion

Total runs: 24.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:Januar 25 2023
replicate

VQ-Diffusion for Text-to-Image Synthesis

Total runs: 20.7K
Run Growth: 0
Growth Rate: 0.00%
Updated:August 10 2022
replicate

Practicing Model Scaling for Photo-Realistic Image Restoration In the Wild. This is the SUPIR-v0Q model and does NOT use LLaVA-13b.

Total runs: 19.7K
Run Growth: 0
Growth Rate: 0.00%
Updated:Februar 24 2024
replicate

Image Manipulatinon with Diffusion Autoencoders

Total runs: 17.1K
Run Growth: 0
Growth Rate: 0.00%
Updated:August 04 2022
replicate

Generating Conditional 3D Implicit Functions

Total runs: 15.1K
Run Growth: 0
Growth Rate: 0.00%
Updated:Mai 20 2023
replicate

High-Quality Video Generation with Cascaded Latent Diffusion Models

Total runs: 13.9K
Run Growth: 0
Growth Rate: 0.00%
Updated:November 21 2023
replicate

Audio-Driven Synthesis of Photorealistic Portrait Animations

Total runs: 13.3K
Run Growth: 0
Growth Rate: 0.00%
Updated:April 01 2024
replicate

stable-diffusion models for high quality and detailed anime images

Total runs: 13.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:Januar 01 2023
replicate

Practicing Model Scaling for Photo-Realistic Image Restoration In the Wild. This is the SUPIR-v0F model and does NOT use LLaVA-13b.

Total runs: 13.0K
Run Growth: 0
Growth Rate: 0.00%
Updated:Februar 24 2024
replicate

Highly practical solution for robust monocular depth estimation by training on a combination of 1.5M labeled images and 62M+ unlabeled images

Total runs: 11.5K
Run Growth: 0
Growth Rate: 0.00%
Updated:Januar 24 2024
replicate

Zero-Shot Speech Editing and Text-to-Speech in the Wild

Total runs: 9.9K
Run Growth: 0
Growth Rate: 0.00%
Updated:März 15 2025
replicate

Pose-Invariant Hairstyle Transfer

Total runs: 9.6K
Run Growth: 0
Growth Rate: 0.00%
Updated:August 21 2022
replicate

Point-E: A System for Generating 3D Point Clouds from Complex Prompts

Total runs: 8.7K
Run Growth: 0
Growth Rate: 0.00%
Updated:Januar 17 2023
replicate

A linear estimator on top of clip to predict the aesthetic quality of pictures

Total runs: 8.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:November 18 2022
replicate

Decoding Micromotion in Low-dimensional Latent Spaces from StyleGAN

Total runs: 8.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:August 04 2022
replicate

fine-tuned Stable Diffusion model trained on the game art from Elden Ring

Total runs: 6.9K
Run Growth: 0
Growth Rate: 0.00%
Updated:November 03 2022
replicate

Zero-shot Image-to-Image Translation

Total runs: 6.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:Februar 12 2023
replicate

Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation

Total runs: 6.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:April 14 2024
replicate

Van Gough on Stable Diffusion via Dreambooth

Total runs: 5.5K
Run Growth: 0
Growth Rate: 0.00%
Updated:November 08 2022
replicate

Finte-tuned Stable Diffusion on high quality 3D images with a futuristic Sci-Fi theme

Total runs: 5.5K
Run Growth: 0
Growth Rate: 0.00%
Updated:Januar 08 2023
replicate

face alignment using stylegan-encoding

Total runs: 4.8K
Run Growth: 0
Growth Rate: 0.00%
Updated:Mai 27 2022
replicate

Clip-Guided Diffusion Model for Image Generation

Total runs: 4.5K
Run Growth: 0
Growth Rate: 0.00%
Updated:März 12 2022
replicate

Efficient Pretraining of Text-to-Image Models

Total runs: 4.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:September 16 2023
replicate

Separate Anything You Describe

Total runs: 4.1K
Run Growth: 0
Growth Rate: 0.00%
Updated:Oktober 20 2023
replicate

Inpainting using Denoising Diffusion Probabilistic Models

Total runs: 4.0K
Run Growth: 0
Growth Rate: 0.00%
Updated:August 17 2022
replicate

Learning Adapters towards Controllable for Text-to-Image Diffusion Models

Total runs: 3.9K
Run Growth: 0
Growth Rate: 0.00%
Updated:Februar 18 2023
replicate

End-to-End Document Image Enhancement Transformer

Total runs: 3.9K
Run Growth: 0
Growth Rate: 0.00%
Updated:August 30 2022
replicate

dreambooth trained on a very diverse dataset ranging from photographs to paintings

Total runs: 3.7K
Run Growth: 0
Growth Rate: 0.00%
Updated:Dezember 09 2022
replicate

Disco Diffusion style on Stable Diffusion via Dreambooth

Total runs: 3.5K
Run Growth: 0
Growth Rate: 0.00%
Updated:November 08 2022
replicate

Efficient Diffusion Model for Image Super-resolution by Residual Shifting

Total runs: 3.0K
Run Growth: 0
Growth Rate: 0.00%
Updated:August 02 2023
replicate

Real-Time High-Resolution Background Matting

Total runs: 2.7K
Run Growth: 0
Growth Rate: 0.00%
Updated:November 18 2022
replicate

Prompt-to-prompt image editing with cross-attention control

Total runs: 2.3K
Run Growth: 0
Growth Rate: 0.00%
Updated:Oktober 30 2022
replicate

Training-free Controllable Text-to-Video Generation

Total runs: 2.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:Mai 28 2023
replicate

lightweight text-to-speech (TTS) model, trained on 10.5K hours of audio data

Total runs: 2.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:April 16 2024
replicate

A Visual Language Model for GUI Agents

Total runs: 2.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:Februar 05 2024
replicate

herge_style on Stable Diffusion via Dreambooth

Total runs: 2.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:November 08 2022
replicate

Controlling Vision-Language Models for Universal Image Restoration

Total runs: 2.1K
Run Growth: 0
Growth Rate: 0.00%
Updated:Oktober 13 2023
replicate

Consistent Diffusion Features for Consistent Video Editing

Total runs: 2.0K
Run Growth: 0
Growth Rate: 0.00%
Updated:Januar 23 2024
replicate

Finetuned Stable-diffusion from Gerry Anderson Supermarionation

Total runs: 1.9K
Run Growth: 0
Growth Rate: 0.00%
Updated:März 04 2023
replicate

text-to-image generation

Total runs: 1.8K
Run Growth: 0
Growth Rate: 0.00%
Updated:August 10 2022
replicate

Diffusion Models as Text Painters

Total runs: 1.7K
Run Growth: 0
Growth Rate: 0.00%
Updated:Juni 04 2023
replicate

Open-source Distilled Stable Diffusion 100% speedup

Total runs: 1.7K
Run Growth: 0
Growth Rate: 0.00%
Updated:Dezember 13 2023
replicate

Hyper-SD: Trajectory Segmented Consistency Model for Efficient Image Synthesis

Total runs: 1.4K
Run Growth: 0
Growth Rate: 0.00%
Updated:April 27 2024
replicate

High-quality multilingual text-to-speech library

Total runs: 1.4K
Run Growth: 0
Growth Rate: 0.00%
Updated:März 03 2024
replicate

Panoptic Scene Graph Generation

Total runs: 1.3K
Run Growth: 0
Growth Rate: 0.00%
Updated:August 13 2022
replicate

text-to-video generation model

Total runs: 1.2K
Run Growth: 0
Growth Rate: 0.00%
Updated:November 26 2023