Model type:
Masked Autoregressive Text-to-Image Generation Model
Model size:
645M
Model precision:
torch.float16 (FP16)
Model resolution:
1024x1024
Model Description:
This is a model that can be used to generate and modify images based on text prompts. It is a
Masked Autoregressive (MAR)
diffusion model that uses a pretrained text encoder (
Phi-2
) and one VAE image tokenizer (
SDXL-VAE
).
import torch
from diffnext.pipelines import NOVAPipeline
model_id = "BAAI/nova-d48w1024-sdxl1024"
model_args = {"torch_dtype": torch.float16, "trust_remote_code": True}
pipe = NOVAPipeline.from_pretrained(model_id, **model_args)
pipe = pipe.to("cuda")
prompt = "a shiba inu wearing a beret and black turtleneck."
image = pipe(prompt).images[0]
image.save("shiba_inu.jpg")
Uses
Direct Use
The model is intended for research purposes only. Possible research areas and tasks include
Research on generative models.
Applications in educational or creative tools.
Generation of artworks and use in design and other artistic processes.
Probing and understanding the limitations and biases of generative models.
Safe deployment of models which have the potential to generate harmful content.
Excluded uses are described below.
Out-of-Scope Use
The model was not trained to be factual or true representations of people or events, and therefore using the model to generate such content is out-of-scope for the abilities of this model.
Misuse and Malicious Use
Using the model to generate content that is cruel to individuals is a misuse of this model. This includes, but is not limited to:
Mis- and disinformation.
Representations of egregious violence and gore.
Impersonating individuals without their consent.
Sexual content without consent of the people who might see it.
Sharing of copyrighted or licensed material in violation of its terms of use.
Intentionally promoting or propagating discriminatory content or harmful stereotypes.
Sharing content that is an alteration of copyrighted or licensed material in violation of its terms of use.
Generating demeaning, dehumanizing, or otherwise harmful representations of people or their environments, cultures, religions, etc.
Limitations and Bias
Limitations
The autoencoding part of the model is lossy.
The model cannot render complex legible text.
The model does not achieve perfect photorealism.
The fingers, .etc in general may not be generated properly.
The model was trained on a subset of the web datasets
LAION-5B
and
COYO-700M
, which contains adult, violent and sexual content.
Bias
While the capabilities of image generation models are impressive, they can also reinforce or exacerbate social biases.
Runs of BAAI nova-d48w1024-sdxl1024 on huggingface.co
2.5K
Total runs
-354
24-hour runs
-210
3-day runs
765
7-day runs
1.6K
30-day runs
More Information About nova-d48w1024-sdxl1024 huggingface.co Model
nova-d48w1024-sdxl1024 huggingface.co is an AI model on huggingface.co that provides nova-d48w1024-sdxl1024's model effect (), which can be used instantly with this BAAI nova-d48w1024-sdxl1024 model. huggingface.co supports a free trial of the nova-d48w1024-sdxl1024 model, and also provides paid use of the nova-d48w1024-sdxl1024. Support call nova-d48w1024-sdxl1024 model through api, including Node.js, Python, http.
nova-d48w1024-sdxl1024 huggingface.co is an online trial and call api platform, which integrates nova-d48w1024-sdxl1024's modeling effects, including api services, and provides a free online trial of nova-d48w1024-sdxl1024, you can try nova-d48w1024-sdxl1024 online for free by clicking the link below.
BAAI nova-d48w1024-sdxl1024 online free url in huggingface.co:
nova-d48w1024-sdxl1024 is an open source model from GitHub that offers a free installation service, and any user can find nova-d48w1024-sdxl1024 on GitHub to install. At the same time, huggingface.co provides the effect of nova-d48w1024-sdxl1024 install, users can directly use nova-d48w1024-sdxl1024 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
nova-d48w1024-sdxl1024 install url in huggingface.co: