GeoSpatial Prior: Synthetic 3D → Geometric Substrate Training
Incomplete Documentation
Claude provided an early document and it's full of problems that will need smoothing and refining.
This will not behave exactly as Claude says it will and there will be multiple refactors and compromises along the way.
Claude below
Abstract
A system for teaching geometric spatial reasoning to neural networks by rendering deterministic 3D scenes where every spatial relationship — position, occlusion, depth, lighting direction, scale — maps directly to known simplex coordinates. Rather than inferring geometric structure from 2D pixel statistics, we construct ground-truth spatial labels from a sectorized 5×5×5 perspective volume and use those labels to pretrain both a geometric classifier and a geometric CLIP variant. The result is a transferable spatial reasoning backbone that can be finetuned into any vision model, providing compositional understanding that current models lack.
1. Core Concept: The 5×5×5 Perspective Volume
1.1 Sectorized Space
The viewing frustum is divided into a 5×5×5 grid of sectors:
X axis
(horizontal): 5 columns, left to right
Y axis
(vertical): 5 rows, bottom to top
Z axis
(depth): 5 layers, near to far
This produces
125 sectors
, each representing a unique spatial address
(x, y, z)
where
x, y, z ∈ {0, 1, 2, 3, 4}
.
1.2 Perspective Scaling
Critical insight: sectors are not uniform cubes in world space. The perspective projection means:
Near sectors (z=0)
: Small world-space volume, large screen-space coverage. A coffee cup fills a near sector.
Far sectors (z=4)
: Enormous world-space volume, small screen-space coverage. A football stadium fits in a far sector.
Each sector's world-space dimensions scale with depth:
This means the same 5×5×5 grid naturally encodes both tabletop scenes (near sectors) and landscape vistas (far sectors) in a single unified coordinate system.
1.3 Sector Labels
Every object placement generates a deterministic label vector:
Label
Type
Description
sector_xyz
(int, int, int)
Primary sector address
sector_coverage
list[(int,int,int)]
All sectors the object spans
depth_order
int
Front-to-back ordering among all objects
occlusion_pct
float
Percentage of object occluded by nearer objects
occluded_by
list[int]
IDs of occluding objects
screen_bbox
(x1, y1, x2, y2)
Normalized screen-space bounding box
relative_scale
float
Object's screen size relative to its true size
lighting_sector
(int, int, int)
Sector of dominant light source
shadow_direction
(float, float)
Normalized shadow vector on ground plane
viewing_angle
(float, float, float)
Object's rotation relative to camera
1.4 Simplex Coordinate Mapping
The pentachoron's 5 vertices map to spatial dimensions:
Vertex
Spatial Meaning
v₀
Horizontal position (x)
v₁
Vertical position (y)
v₂
Depth / distance (z)
v₃
Scale / size relationship
v₄
Viewpoint / rotation encoding
An object at sector (2, 3, 1) with moderate scale and frontal view maps to a specific barycentric coordinate on the simplex. This mapping is
defined
, not learned — the training teaches the network to predict these coordinates from pixels.
2. Rendering Pipeline
2.1 Requirements
Speed
: Must generate millions of training pairs. Target: 100+ scenes/second on a single GPU.
Where
L_sector
includes a geometric component that penalizes predictions that violate simplex constraints (e.g., predicted sector coordinates must form valid configurations on the pentachoron).
Training data
: 1M–10M rendered scenes with exact labels.
3.2 Stage 2: Geometric CLIP
Take the pretrained spatial backbone from Stage 1 and use it as the vision encoder for a CLIP-style contrastive model.
Vision encoder
: Stage 1 backbone (frozen or lightly finetuned)
Text encoder
: Standard transformer, initialized from existing CLIP text encoder
Contrastive target
: Align image embeddings with text descriptions that include spatial language
Training captions are generated from scene labels:
defscene_to_caption(labels):
"""Generate spatial text from ground-truth labels."""
parts = []
for obj in labels["objects"]:
# Position
x, y, z = obj["sector_xyz"]
h_pos = ["far left", "left", "center", "right", "far right"][x]
v_pos = ["bottom", "lower", "middle", "upper", "top"][y]
d_pos = ["very close", "near", "middle distance", "far", "very far"][z]
parts.append(f"a {obj['type']} at {h_pos}{v_pos}, {d_pos}")
# Occlusionif obj["occlusion_pct"] > 0.2:
occluder = labels["objects"][obj["occluded_by"][0]]["type"]
parts.append(f"partially behind the {occluder}")
# Scale contextif z >= 3:
parts.append(f"appearing small in the distance")
# Spatial relationsfor rel in labels["scene_graph"]:
parts.append(f"the {rel['obj_a']} is {rel['relation']} the {rel['obj_b']}")
return", ".join(parts)
Key insight
: The geometric CLIP doesn't just learn "cat" ↔ image-of-cat. It learns "cat at upper-right, far distance, partially behind a tree" ↔ specific simplex configuration. Spatial prepositions become geometric operations, not statistical associations.
3.3 Stage 3: Transfer to Real Images
The pretrained geometric backbone transfers to real-world tasks:
Direct finetuning
: Replace the geometric CLIP's vision encoder in SD1.5's conditioning path. Now "cup on top of book" activates specific simplex configurations that were grounded in actual 3D relationships.
Inverse embedding
: Given a real image, extract simplex coordinates that describe its spatial structure. These become geometric conditioning signals for diffusion models.
Hybrid
: Use the geometric backbone as an auxiliary encoder alongside standard CLIP. The geometric channel provides spatial structure; CLIP provides semantic content. The geo_prior blends them on the simplex.
4. Dataset Scaling Strategy
4.1 Procedural Generation Tiers
Tier
Scenes
Resolution
Objects
Purpose
Tier 1
1M
256×256
Primitives only
Fast pretraining of spatial reasoning
Tier 2
5M
512×512
Composites + varied lighting
Full spatial classifier training
Tier 3
10M
512×512
Low-poly meshes + textures
Geometric CLIP pretraining
Tier 4
1M
512×512
Complex scenes + real textures
Bridge to photorealism
4.2 Augmentation via Sector Permutation
Because the 5×5×5 grid is symmetric, we can generate 6× data for free by:
Horizontal flip: maps sector (x, y, z) → (4-x, y, z)
Text encoder: initialize from OpenAI CLIP text encoder
Contrastive loss with hard negatives (spatial near-misses)
Training
Caption generation from scene labels
Hard negative mining (swap spatial relations in text)
Spatial preposition evaluation benchmark
Transfer evaluation: zero-shot spatial classification on real images
Integration with existing pipeline
Replace SD1.5 CLIP encoder with geometric CLIP
Measure impact on compositional generation
Compare geo_prior behavior with geometric vs standard CLIP
Phase 4: Transfer + Real-World Bridge [LOWER PRIORITY]
Domain transfer
Finetune geometric backbone on COCO with spatial annotations
Evaluate on spatial reasoning benchmarks (CLEVR, SpatialBench)
Test compositional generation improvement in SD1.5
Inverse embedding pipeline
Given real image → extract simplex coordinates
Use as conditioning signal for diffusion
Compare with CLIP-only conditioning
Hybrid encoder
Dual-stream: geometric backbone + CLIP
Learnable fusion on simplex
Evaluate on attribute binding + spatial composition jointly
7. Key Hypotheses to Validate
Sector classification transfers to real images
: A model trained entirely on synthetic 3D scenes can identify spatial sectors in photographs with >60% accuracy.
Geometric CLIP improves compositional generation
: Replacing standard CLIP with geometric CLIP in the SD1.5 pipeline produces measurably better spatial composition (evaluated via CLEVR-style spatial accuracy metrics).
Simplex coordinates are a natural spatial language
: The 5-vertex pentachoron provides sufficient dimensionality to encode the spatial relationships humans use in language (in front of, behind, above, below, next to, far away, etc.).
Vertex entropy predicts spatial complexity
: Scenes with more objects produce lower vertex entropy (hard routing), while single-object scenes produce higher entropy (attribute binding). This pattern, observed in the triad study, should emerge naturally from synthetic training.
The 5×5×5 grid scales
: The same sectorization works for both close-up tabletop scenes and panoramic landscapes by leveraging perspective scaling of sector volumes.
8. Technical Notes
8.1 Why 5×5×5?
5 matches the pentachoron vertex count — direct 1:1 axis-to-vertex mapping
125 sectors is fine enough for meaningful spatial discrimination without being computationally prohibitive
The number 5 appears consistently across the geometric vocabulary work (k=4 simplex has 5 vertices, 5 edge dimensions, etc.)
8.2 Why Not Photorealistic Rendering?
Photorealism introduces texture/material confounds that obscure spatial signal
Simple rendering is 100-1000× faster, enabling much larger datasets
Transfer from simple→real is well-studied (sim2real in robotics)
The geometric prior should learn spatial structure independent of visual style — using simple rendering enforces this
Phase 4 progressively adds visual complexity once spatial reasoning is established
8.3 Relationship to Existing Work
CLEVR
: Similar synthetic scene approach but limited to 2D spatial relations. Our 5×5×5 grid adds depth, scale, and perspective.
NeRF / 3D Gaussians
: Reconstruct 3D from 2D. We go the opposite direction — start with known 3D, teach networks to infer it from 2D.
Spatial transformers
: Learn attention over spatial positions. Our approach provides explicit spatial supervision rather than hoping attention learns spatial structure.
Scene graphs
: Prior work on scene graph prediction from images. Our contribution is grounding scene graphs in simplex geometry rather than abstract relation classification.
geoflow-scene-classifier-proto huggingface.co is an AI model on huggingface.co that provides geoflow-scene-classifier-proto's model effect (), which can be used instantly with this AbstractPhil geoflow-scene-classifier-proto model. huggingface.co supports a free trial of the geoflow-scene-classifier-proto model, and also provides paid use of the geoflow-scene-classifier-proto. Support call geoflow-scene-classifier-proto model through api, including Node.js, Python, http.
geoflow-scene-classifier-proto huggingface.co is an online trial and call api platform, which integrates geoflow-scene-classifier-proto's modeling effects, including api services, and provides a free online trial of geoflow-scene-classifier-proto, you can try geoflow-scene-classifier-proto online for free by clicking the link below.
AbstractPhil geoflow-scene-classifier-proto online free url in huggingface.co:
geoflow-scene-classifier-proto is an open source model from GitHub that offers a free installation service, and any user can find geoflow-scene-classifier-proto on GitHub to install. At the same time, huggingface.co provides the effect of geoflow-scene-classifier-proto install, users can directly use geoflow-scene-classifier-proto installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
geoflow-scene-classifier-proto install url in huggingface.co: