Introducing RealFormer - A new approach to Photorealism over Supersampling
Introducing RealFormer, a novel image-to-image transformer model designed for enhancing photorealism in images, particularly focused on transforming synthetic images to more realistic ones.
Model Details
Model Description
RealFormer is an innovative Vision Transformer (ViT) based architecture that combines elements of Linear Attention (approximation attention) with Swin Transformers and adaptive instance normalization (AdaIN) for style transfer. It's designed to transform images,specifically targeted at the video game and animation industry, potentially enhancing their photorealism or applying style transfer.
Integration into game engines for real-time graphics enhancement - AdaIN layers are finetunable for video-game-specific usecases. In this implementation, the models have been pretrained on a variety of video for
super-sampling
,
photorealistic style transfer
and
reverse photorealism
.
Pre-processing step in computer vision pipelines to improve input image quality - Decoder layers can be frozen for task-specific usecases.
Photo editing software for synthesized image enhancement
Out-of-Scope Use
This model is not recommended for:
Generating or manipulating images in ways that could be deceptive or harmful
Tasks requiring perfect preservation of specific image details, as the transformation process may alter some artifacts of the image
Medical or forensic image analysis where any alteration could lead to misinterpretation. Remember, this is a model, not a classification or detection model.
Bias, Risks, and Limitations
The model may introduce biases present in the training data, potentially altering images in ways that reflect these biases.
There's a risk of over-smoothing or losing fine details in the image transformation process.
The model's performance may vary significantly depending on the input image characteristics and how similar they are to the training data.
As with any image manipulation tool, there's a potential for misuse in creating deceptive or altered images.
How to Get Started with the Model
Use the code below to get started with the model.
# Instantiate the model
model = ViTImage2Image(img_size=512, patch_size=16, emb_dim=768, num_heads=16, num_layers=8, hidden_dim=3072)
# Move model to GPU if available
device = torch.device("cuda"if torch.cuda.is_available() else"cpu")
model = model.to(device)
# Load an image
input_image = load_image('path_to_your_image.png')
input_image = input_image.to(device)
# Perform inferencewith torch.no_grad():
output = model(input_image, input_image) # Using input as both content and style for this example# Visualize or save the output
visualize_tensor(output, "Output Image")
Training Details
Training Data
The model was trained on
Pre-Training Dataset
and then the decoder layers were frozen to finetune it on the
Calibration Dataset for Grand Theft Auto V
. The former includes over 400,000 frames of footage from video games such as WatchDogs 2, Grand Theft Auto V, CyberPunk, several Hollywood films and high-defintion photos. The latter comprises of ~25,000 high-definition semantic segmentation map - rendered frame pairs captured from Grand Theft Auto V in-game and a UNet based Semantic Segmentation Model.
Training Procedure
Optimizer: Adam
Learning rate: 0.001
Batch size: 8
Steps per epoch: 3,125
Number of epochs: 100
Total number of steps: 312,500
Loss function: Combined L1 loss, Perpetual Loss, Style Transfer Loss, Total Variation loss
Preprocessing
Images and their corresponding style semantic maps were resized to fit the input-output window dimensions (512 x 512). Bit depth has been recorrected to 24bit (3 channel) for images with depth greater than 24bit.
Training Hyperparameters
Precision:fp32
Embedded dimensions: 768
Hidden dimensions: 3072
Attention Type: Linear Attention
Number of attention heads: 16
Number of attention layers: 8
Number of transformer encoder layers (feed-forward): 8
Number of transformer decoder layers (feed-forward): 8
RealFormer huggingface.co is an AI model on huggingface.co that provides RealFormer's model effect (), which can be used instantly with this aoxo RealFormer model. huggingface.co supports a free trial of the RealFormer model, and also provides paid use of the RealFormer. Support call RealFormer model through api, including Node.js, Python, http.
RealFormer huggingface.co is an online trial and call api platform, which integrates RealFormer's modeling effects, including api services, and provides a free online trial of RealFormer, you can try RealFormer online for free by clicking the link below.
aoxo RealFormer online free url in huggingface.co:
RealFormer is an open source model from GitHub that offers a free installation service, and any user can find RealFormer on GitHub to install. At the same time, huggingface.co provides the effect of RealFormer install, users can directly use RealFormer installed effect in huggingface.co for debugging and trial. It also supports api for free installation.