aoxo / RealFormer

huggingface.co
Total runs: 0
24-hour runs: 0
7-day runs: 0
30-day runs: 0
Model's Last Updated: January 15 2026
image-to-image

Introduction of RealFormer

Model Details of RealFormer

Introducing RealFormer - A new approach to Photorealism over Supersampling

Introducing RealFormer, a novel image-to-image transformer model designed for enhancing photorealism in images, particularly focused on transforming synthetic images to more realistic ones.

Model Details
Model Description

RealFormer is an innovative Vision Transformer (ViT) based architecture that combines elements of Linear Attention (approximation attention) with Swin Transformers and adaptive instance normalization (AdaIN) for style transfer. It's designed to transform images,specifically targeted at the video game and animation industry, potentially enhancing their photorealism or applying style transfer.

  • Developed by: Alosh Denny
  • Funded by [optional]: EmelinLabs
  • Shared by [optional]: EmelinLabs
  • Model type: Image-to-Image Transformer
  • Language(s) (NLP): None (Pre-trained Generative Image Model)
  • License: Apache-2.0
  • Finetuned from model [optional]: Novel; Pre-trained (not finetuned)
Model Sources [optional]
Uses
Direct Use

RealFormer is designed for image-to-image translation tasks. It can be used directly for:

  • Enhancing photorealism in synthetic images (e.g., transforming video game graphics to more realistic images)
  • Style transfer between rendered frames and post-processed frames
  • To be incorporated in pipeline with DLSS
Downstream Use

Potential downstream uses could include:

  • Integration into game engines for real-time graphics enhancement - AdaIN layers are finetunable for video-game-specific usecases. In this implementation, the models have been pretrained on a variety of video for super-sampling , photorealistic style transfer and reverse photorealism .
  • Pre-processing step in computer vision pipelines to improve input image quality - Decoder layers can be frozen for task-specific usecases.
  • Photo editing software for synthesized image enhancement
Out-of-Scope Use

This model is not recommended for:

  • Generating or manipulating images in ways that could be deceptive or harmful
  • Tasks requiring perfect preservation of specific image details, as the transformation process may alter some artifacts of the image
  • Medical or forensic image analysis where any alteration could lead to misinterpretation. Remember, this is a model, not a classification or detection model.
Bias, Risks, and Limitations
  • The model may introduce biases present in the training data, potentially altering images in ways that reflect these biases.
  • There's a risk of over-smoothing or losing fine details in the image transformation process.
  • The model's performance may vary significantly depending on the input image characteristics and how similar they are to the training data.
  • As with any image manipulation tool, there's a potential for misuse in creating deceptive or altered images.
How to Get Started with the Model

Use the code below to get started with the model.

# Instantiate the model
model = ViTImage2Image(img_size=512, patch_size=16, emb_dim=768, num_heads=16, num_layers=8, hidden_dim=3072)

# Move model to GPU if available
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
model = model.to(device)

# Load an image
input_image = load_image('path_to_your_image.png')
input_image = input_image.to(device)

# Perform inference
with torch.no_grad():
    output = model(input_image, input_image)  # Using input as both content and style for this example

# Visualize or save the output
visualize_tensor(output, "Output Image")
Training Details
Training Data

The model was trained on Pre-Training Dataset and then the decoder layers were frozen to finetune it on the Calibration Dataset for Grand Theft Auto V . The former includes over 400,000 frames of footage from video games such as WatchDogs 2, Grand Theft Auto V, CyberPunk, several Hollywood films and high-defintion photos. The latter comprises of ~25,000 high-definition semantic segmentation map - rendered frame pairs captured from Grand Theft Auto V in-game and a UNet based Semantic Segmentation Model.

Training Procedure
  • Optimizer: Adam
  • Learning rate: 0.001
  • Batch size: 8
  • Steps per epoch: 3,125
  • Number of epochs: 100
  • Total number of steps: 312,500
  • Loss function: Combined L1 loss, Perpetual Loss, Style Transfer Loss, Total Variation loss
Preprocessing

Images and their corresponding style semantic maps were resized to fit the input-output window dimensions (512 x 512). Bit depth has been recorrected to 24bit (3 channel) for images with depth greater than 24bit.

Training Hyperparameters
  • Precision:fp32
  • Embedded dimensions: 768
  • Hidden dimensions: 3072
  • Attention Type: Linear Attention
  • Number of attention heads: 16
  • Number of attention layers: 8
  • Number of transformer encoder layers (feed-forward): 8
  • Number of transformer decoder layers (feed-forward): 8
  • Activation function: ReLU
  • Patch Size: 8
  • Swin Window Size: 7
  • Swin Shift Size: 2
Speeds, Sizes, Times [optional]

[More Information Needed]

Evaluation
Testing Data, Factors & Metrics
Testing Data

[More Information Needed]

Factors

[More Information Needed]

Metrics

[More Information Needed]

Results

[More Information Needed]

Summary
Model Examination [optional]

[More Information Needed]

Environmental Impact

Carbon emissions can be estimated using the Machine Learning Impact calculator presented in Lacoste et al. (2019) .

  • Hardware Type: [More Information Needed]
  • Hours used: [More Information Needed]
  • Cloud Provider: [More Information Needed]
  • Compute Region: [More Information Needed]
  • Carbon Emitted: [More Information Needed]
Technical Specifications [optional]
Model Architecture and Objective

[More Information Needed]

Compute Infrastructure

[More Information Needed]

Hardware

[More Information Needed]

Software

[More Information Needed]

Citation [optional]

BibTeX:

[More Information Needed]

APA:

[More Information Needed]

Glossary [optional]

[More Information Needed]

More Information [optional]

[More Information Needed]

Model Card Authors [optional]

[More Information Needed]

Model Card Contact

[More Information Needed]

Runs of aoxo RealFormer on huggingface.co

0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs

More Information About RealFormer huggingface.co Model

More RealFormer license Visit here:

https://choosealicense.com/licenses/apache-2.0

RealFormer huggingface.co

RealFormer huggingface.co is an AI model on huggingface.co that provides RealFormer's model effect (), which can be used instantly with this aoxo RealFormer model. huggingface.co supports a free trial of the RealFormer model, and also provides paid use of the RealFormer. Support call RealFormer model through api, including Node.js, Python, http.

RealFormer huggingface.co Url

https://huggingface.co/aoxo/RealFormer

aoxo RealFormer online free

RealFormer huggingface.co is an online trial and call api platform, which integrates RealFormer's modeling effects, including api services, and provides a free online trial of RealFormer, you can try RealFormer online for free by clicking the link below.

aoxo RealFormer online free url in huggingface.co:

https://huggingface.co/aoxo/RealFormer

RealFormer install

RealFormer is an open source model from GitHub that offers a free installation service, and any user can find RealFormer on GitHub to install. At the same time, huggingface.co provides the effect of RealFormer install, users can directly use RealFormer installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

RealFormer install url in huggingface.co:

https://huggingface.co/aoxo/RealFormer

Url of RealFormer

RealFormer huggingface.co Url

Provider of RealFormer huggingface.co

aoxo
ORGANIZATIONS

Other API from aoxo

huggingface.co

Total runs: 15
Run Growth: -6
Growth Rate: -50.00%
Updated:December 10 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:December 14 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:April 26 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:December 19 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:January 12 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:September 28 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:October 21 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:November 28 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:March 19 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:November 17 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:December 04 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:November 25 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:December 21 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:December 21 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:December 09 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:November 18 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:October 30 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:December 17 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:November 26 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:September 02 2026