A tiny text-to-3D diffusion model. You give it a prompt, it gives you a
32x32x32 occupancy grid. It is 40,050,892 trained parameters and it trained
in under ten hours on a single A100.
This is the PixelModel line stepping out of 2D. v1 through v4 were tiny latent
diffusion for images; VoxelModel keeps the same argument (low resolution as a
deliberate aesthetic, not a compromise) and applies it to shape.
model.png
is
still here and still holds every weight, because that gimmick ports unchanged to
3D: it only ever cared about floats.
The one honest result
It works for bulky objects and it fails for thin ones.
Chairs, benches, tables, mushrooms, trees, cars and boats come back legible. A
sword comes back as 0.2% occupancy, which is a handful of stray voxels, and a key
comes back as a featureless bar.
That is not a training failure. A sword does not survive voxelization at 32^3 in
the
training data
either: the real Objaverse sword we voxelized for reference
sits at 2.9% occupancy and reads as a cross at best. The model reproduced what it
was shown. If you want swords, the answer is resolution, not more steps.
So: 32^3 is enough, for a specific and checkable definition of enough. That was
the question v1 existed to answer.
No VAE
VoxelModel operates directly on voxels, so there is no autoencoder in the stack.
Total footprint is 40.05M trained plus 37.8M frozen CLIP ViT-B/32, which is
77.9M. PixelModel v4 was 161M for comparable trained size, because it carries an
83.7M frozen
sd-vae-ft-mse
. This one is tiny without an asterisk.
Architecture
VoxelDiT
is a DiT with 3D patchify. A 32^3 grid at patch 4 gives 512 tokens,
dim 384, depth 12, heads 6. Conditioning is cross-attention on frozen CLIP
ViT-B/32 text features (40 tokens), plus adaLN-zero from the pooled embedding and
the timestep. Positional encoding is a 3D sincos over the 8x8x8 patch grid.
Objective is rectified flow with
target = x1 - x0
and logit-normal timestep
sampling,
t = sigmoid(randn)
.
Training
meshes
28,415 train, 1,024 held out
steps
90,000 at batch 256
hardware
one A100 80GB, 2.57 steps/s, 41.4 GiB
wall clock
9.7 hours
optimizer
AdamW 2e-4, betas (0.9, 0.99), grad clip 1.0
schedule
cosine to 1e-6, 500 step warmup
EMA
0.9999
CFG dropout
10%
Held-out loss bottomed at
0.018737 at step 64,000
and drifted up to 0.019314
by step 90,000. That is 3.08% of overfitting, and it means the run was about 2.8
hours longer than it needed to be. 90,000 steps over 28,415 meshes is 811 epochs
with no augmentation, which is the obvious thing to fix next. Yaw rotations plus
a mirror would give 8x for free on a cubic grid.
The shipped weights are the step 90,000 EMA, not the step 64,000 best, because
checkpoints were written to a fixed path and overwrote each other. Noted, and
worth 3.08% to you if you retrain.
Data
Meshes are
dylanebert/objaverse-lowpoly-obj
, captions are
tiange/Cap3D
,
joined on Objaverse UID for 58,374 captioned meshes. Voxelization centers on the
vertex mean, scales by max absolute coordinate, and rasterizes surfaces at pitch
2/31 with no fill, since filling solidifies closed volumes and pushes 4% of the
set past 90% occupancy.
caption_filter.py
then throws away half of it. Multi-object meshes are the main
defect in this corpus: normalization scales the whole scene bounding box into the
grid, so a mesh holding several objects shrinks each one into a cluster of blobs.
The caption gives this away before you load anything, so the filter rejects scene
and variant-set wording, a top level
and
joining two noun phrases, and captions
over 12 words. Colour pairs survive on purpose, so "a purple and green mushroom"
is kept and "a white cube and a white robot" is not.
Measured over all 58,374 captions, 29,439 survive, which is 50.4%.
Worth knowing if you build on this:
occupancy does not detect the defect.
Survivors and rejects have almost the same occupancy distribution, medians 4.3%
against 5.0% over a 900 mesh sample. No threshold would have caught it.
Writes a JSON of base64
np.packbits
grids. To load the weights from the PNG
instead of the safetensors:
from png_codec import load_model_png
model = load_model_png("model.png", "model_png.json", device="cuda")
Round-trip error against the safetensors is 9.7e-04 max absolute, which is fp16
quantization and nothing else.
What v1 does not do
Colour.
These meshes carry no vertex colours and no
.mtl
files, so the
model generates grey geometry only. Colour needs a GLB-with-materials source and
its own voxelization path.
A real eval metric.
FID does not transfer to 3D. Held-out rectified flow loss
is what the table above reports, and it is a training diagnostic, not a quality
measure. Rendered-FID with CLIP score on 2D renders, or IoU against a held-out
split, are the candidates. This is unresolved and we would rather say so than
publish a number we do not believe.
Files
file
what
model.safetensors
EMA weights, fp32
model.png
the same weights as a 6329x6329 PNG
model_png.json
parameter manifest for the PNG
voxel_dit.py
the model
train.py
training loop
sample.py
inference
caption_filter.py
the single-object filter
build_data.py
mesh download and voxelization
png_codec.py
PNG encode, decode and verify
Runs of bench-labs VoxelModel-v1 on huggingface.co
66
Total runs
-6
24-hour runs
-10
3-day runs
-10
7-day runs
-2.4K
30-day runs
More Information About VoxelModel-v1 huggingface.co Model
VoxelModel-v1 huggingface.co is an AI model on huggingface.co that provides VoxelModel-v1's model effect (), which can be used instantly with this bench-labs VoxelModel-v1 model. huggingface.co supports a free trial of the VoxelModel-v1 model, and also provides paid use of the VoxelModel-v1. Support call VoxelModel-v1 model through api, including Node.js, Python, http.
VoxelModel-v1 huggingface.co is an online trial and call api platform, which integrates VoxelModel-v1's modeling effects, including api services, and provides a free online trial of VoxelModel-v1, you can try VoxelModel-v1 online for free by clicking the link below.
bench-labs VoxelModel-v1 online free url in huggingface.co:
VoxelModel-v1 is an open source model from GitHub that offers a free installation service, and any user can find VoxelModel-v1 on GitHub to install. At the same time, huggingface.co provides the effect of VoxelModel-v1 install, users can directly use VoxelModel-v1 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.