microsoft/TRELLIS.2

▲ 207 stars today★ 11,766⑂ 1,427

Native and Compact Structured Latents for 3D Generation

About microsoft/TRELLIS.2

microsoft/TRELLIS.2 is an open-source project on GitHub, mainly written in Python. Native and Compact Structured Latents for 3D Generation It currently holds 11,766 stars and 1,427 forks with 158 open issues, and was last pushed on 2026-07-10 (repository created 2025-11-26).

Project Overview

Git Homed tracks it on the Today's Trending board, currently at rank #22 with 207 new stars today.

GitHub Repository Details

Repository microsoft/TRELLIS.2 · default branch main · size 676187 KB · watchers 82 · source: GitHub REST API and repository README

README

Native and Compact Structured Latents for 3D Generation

https://github.com/microsoft/TRELLIS.2/blob/HEAD/Paper https://github.com/microsoft/TRELLIS.2/blob/HEAD/Hugging Face https://github.com/microsoft/TRELLIS.2/blob/HEAD/Project Page https://github.com/microsoft/TRELLIS.2/blob/HEAD/License

https://github.com/user-attachments/assets/63b43a7e-acc7-4c81-a900-6da450527d8f

(Compressed version due to GitHub size limits. See the full-quality video on our project page!)

TRELLIS.2 is a state-of-the-art large 3D generative model (4B parameters) designed for high-fidelity image-to-3D generation. It leverages a novel "field-free" sparse voxel structure termed O-Voxel to reconstruct and generate arbitrary 3D assets with complex topologies, sharp features, and full PBR materials.

✨ Features

1. High Quality, Resolution & Efficiency

Our 4B-parameter model generates high-resolution fully textured assets with exceptional fidelity and efficiency using vanilla DiTs. It utilizes a Sparse 3D VAE with 16× spatial downsampling to encode assets into a compact latent space.

| Resolution | Total Time* | Breakdown (Shape + Mat) | | :--- | :--- | :--- | | 512³ | ~3s | 2s + 1s | | 1024³ | ~17s | 10s + 7s | | 1536³ | ~60s | 35s + 25s |

*Tested on NVIDIA H100 GPU.

2. Arbitrary Topology Handling

The O-Voxel representation breaks the limits of iso-surface fields. It robustly handles complex structures without lossy conversion:

3. Rich Texture Modeling

Beyond basic colors, TRELLIS.2 models arbitrary surface attributes including Base Color, Roughness, Metallic, and Opacity, enabling photorealistic rendering and transparency support.

4. Minimalist Processing

Data processing is streamlined for instant conversions that are fully rendering-free and optimization-free.

🗺️ Roadmap

🛠️ Installation

Prerequisites

Installation Steps

1. Clone the repo:
    git clone -b main https://github.com/microsoft/TRELLIS.2.git --recursive
    cd TRELLIS.2
    

2. Install the dependencies: Before running the following command there are somethings to note:

Create a new conda environment named trellis2 and install the dependencies:
    . ./setup.sh --new-env --basic --flash-attn --nvdiffrast --nvdiffrec --cumesh --o-voxel --flexgemm
    
The detailed usage of setup.sh can be found by running . ./setup.sh --help.
    Usage: setup.sh [OPTIONS]
    Options:
        -h, --help              Display this help message
        --new-env               Create a new conda environment
        --basic                 Install basic dependencies
        --flash-attn            Install flash-attention
        --cumesh                Install cumesh
        --o-voxel               Install o-voxel
        --flexgemm              Install flexgemm
        --nvdiffrast            Install nvdiffrast
        --nvdiffrec             Install nvdiffrec
    

📦 Pretrained Weights

The pretrained model TRELLIS.2-4B is available on Hugging Face. Please refer to the model card there for more details.

| Model | Parameters | Resolution | Link | | :--- | :--- | :--- | :--- | | TRELLIS.2-4B | 4 Billion | 512³ - 1536³ | Hugging Face |

🚀 Usage

1. Image to 3D Generation

Minimal Example

Here is an example of how to use the pretrained models for 3D asset generation.

import os
os.environ['OPENCV_IO_ENABLE_OPENEXR'] = '1'
os.environ["PYTORCH_CUDA_ALLOC_CONF"] = "expandable_segments:True"  # Can save GPU memory
import cv2
import imageio
from PIL import Image
import torch
from trellis2.pipelines import Trellis2ImageTo3DPipeline
from trellis2.utils import render_utils
from trellis2.renderers import EnvMap
import o_voxel

1. Setup Environment Map

envmap = EnvMap(torch.tensor( cv2.cvtColor(cv2.imread('assets/hdri/forest.exr', cv2.IMREAD_UNCHANGED), cv2.COLOR_BGR2RGB), dtype=torch.float32, device='cuda' ))

2. Load Pipeline

pipeline = Trellis2ImageTo3DPipeline.from_pretrained("microsoft/TRELLIS.2-4B") pipeline.cuda()

3. Load Image & Run

image = Image.open("assets/example_image/T.png") mesh = pipeline.run(image)[0] mesh.simplify(16777216) # nvdiffrast limit

4. Render Video

video = render_utils.make_pbr_vis_frames(render_utils.render_video(mesh, envmap=envmap)) imageio.mimsave("sample.mp4", video, fps=15)

5. Export to GLB

glb = o_voxel.postprocess.to_glb( vertices = mesh.vertices, faces = mesh.faces, attr_volume = mesh.attrs, coords = mesh.coords, attr_layout = mesh.layout, voxel_size = mesh.voxel_size, aabb = [[-0.5, -0.5, -0.5], [0.5, 0.5, 0.5]], decimation_target = 1000000, texture_size = 4096, remesh = True, remesh_band = 1, remesh_project = 0, verbose = True ) glb.export("sample.glb", extension_webp=True)

Upon execution, the script generates the following files:

Note: The .glb file is exported in OPAQUE mode by default. Although the alpha channel is preserved within the texture map, it is not active initially. To enable transparency, import the asset into your 3D software and manually connect the texture's alpha channel to the material's opacity or alpha input.

Web Demo

app.py provides a simple web demo for image to 3D asset generation. you can run the demo with the following command:

python app.py

Then, you can access the demo at the address shown in the terminal.

2. PBR Texture Generation

Please refer to the example_texturing.py for an example of how to generate PBR textures for a given 3D shape. Also, you can use the app_texturing.py to run a web demo for PBR texture generation.

🏋️ Training

We provide the full training codebase, enabling users to train TRELLIS.2 from scratch or fine-tune it on custom datasets.

1. Data Preparation

Before training, raw 3D assets must be converted into the O-Voxel representation. This process includes mesh conversion, compact structured latent generation, and metadata preparation.

📂 Please refer to data_toolkit/README.md for detailed instructions on data preprocessing and dataset organization.

2. Running Training

Training is managed through the train.py script, which accepts multiple command-line arguments to configure experiments:

SC-VAE Training

To train the shape SC-VAE, run:

python train.py \
  --config configs/scvae/shape_vae_next_dc_f16c32_fp16.json \
  --output_dir results/shape_vae_next_dc_f16c32_fp16 \
  --data_dir "{\"ObjaverseXL_sketchfab\": {\"base\": \"datasets/ObjaverseXL_sketchfab\", \"mesh_dump\": \"datasets/ObjaverseXL_sketchfab/mesh_dumps\", \"dual_grid\": \"datasets/ObjaverseXL_sketchfab/dual_grid_256\", \"asset_stats\": \"datasets/ObjaverseXL_sketchfab/asset_stats\"}}"

This command trains the shape SC-VAE on the Objaverse-XL dataset using the shape_vae_next_dc_f16c32_fp16.json configuration. Training outputs will be saved to results/shape_vae_next_dc_f16c32_fp16.

The dataset is specified as a JSON string, where each dataset entry includes:

To fine-tune the model at a higher resolution, use the shape_vae_next_dc_f16c32_fp16_ft_512.json configuration. Remember to update the finetune_ckpt field and adjust the dataset paths accordingly.

To train the texture SC-VAE, run:

python train.py \
  --config configs/scvae/tex_vae_next_dc_f16c32_fp16.json \
  --output_dir results/tex_vae_next_dc_f16c32_fp16 \
  --data_dir "{\"ObjaverseXL_sketchfab\": {\"base\": \"datasets/ObjaverseXL_sketchfab\", \"pbr_dump\": \"datasets/ObjaverseXL_sketchfab/pbr_dumps\", \"pbr_voxel\": \"datasets/ObjaverseXL_sketchfab/pbr_voxels_256\", \"asset_stats\": \"datasets/ObjaverseXL_sketchfab/asset_stats\"}}"

Flow Model Training

To train the sparse structure flow model, run:

python train.py \
  --config configs/gen/ss_flow_img_dit_1_3B_64_bf16.json \
  --output_dir results/ss_flow_img_dit_1_3B_64_bf16 \
  --data_dir "{\"ObjaverseXL_sketchfab\": {\"base\": \"datasets/ObjaverseXL_sketchfab\", \"ss_latent\": \"datasets/ObjaverseXL_sketchfab/ss_latents/ss_enc_conv3d_16l8_fp16_64\", \"render_cond\": \"datasets/ObjaverseXL_sketchfab/renders_cond\"}}"

This command trains the sparse-structure flow model on the Objaverse-XL dataset using the specified configuration file. Outputs are saved to results/ss_flow_img_dit_1_3B_64_bf16.

The dataset configuration includes:

The second- and third-stage flow models for shape and texture generation can be trained using the following configurations:

Example commands:

# Shape flow model
python train.py \
  --config configs/gen/slat_flow_img2shape_dit_1_3B_512_bf16.json \
  --output_dir results/slat_flow_img2shape_dit_1_3B_512_bf16 \
  --data_dir "{\"ObjaverseXL_sketchfab\": {\"base\": \"datasets/ObjaverseXL_sketchfab\", \"shape_latent\": \"datasets/ObjaverseXL_sketchfab/shape_latents/shape_enc_next_dc_f16c32_fp16_512\", \"render_cond\": \"datasets/ObjaverseXL_sketchfab/renders_cond\"}}"

Texture flow model

python train.py \ --config configs/gen/slat_flow_imgshape2tex_dit_1_3B_512_bf16.json \ --output_dir results/slat_flow_imgshape2tex_dit_1_3B_512_bf16 \ --data_dir "{\"ObjaverseXL_sketchfab\": {\"base\": \"datasets/ObjaverseXL_sketchfab\", \"shape_latent\": \"datasets/ObjaverseXL_sketchfab/shape_latents/shape_enc_next_dc_f16c32_fp16_512\", \"pbr_latent\": \"datasets/ObjaverseXL_sketchfab/pbr_latents/tex_enc_next_dc_f16c32_fp16_512\", \"render_cond\": \"datasets/ObjaverseXL_sketchfab/renders_cond\"}}"

Higher-resolution fine-tuning can be performed by updating the finetune_ckpt field in the following configuration files and adjusting the dataset paths accordingly:

🧩 Related Packages

TRELLIS.2 is built upon several specialized high-performance packages developed by our team:

Core library handling the logic for converting between textured meshes and the O-Voxel representation, ensuring instant bidirectional transformation. Efficient sparse convolution implementation based on Triton, enabling rapid processing of sparse voxel structures. CUDA-accelerated mesh utilities used for high-speed post-processing, remeshing, decimation, and UV-unwrapping.

⚖️ License

This model and code are released under the MIT License.

Please note that certain dependencies operate under separate license terms:

📚 Citation

If you find this model useful for your research, please cite our work:

@article{
    xiang2025trellis2,
    title={Native and Compact Structured Latents for 3D Generation},
    author={Xiang, Jianfeng and Chen, Xiaoxue and Xu, Sicheng and Wang, Ruicheng and Lv, Zelong and Deng, Yu and Zhu, Hongyuan and Dong, Yue and Zhao, Hao and Yuan, Nicholas Jing and Yang, Jiaolong},
    journal={Tech report},
    year={2025}
}

GitHub Stars & Activity

11,766Stars
1,427Forks
158Open issues
PythonLanguage

GitHub Popularity

GitHub stars11,766
Forks1,427
Open issues158
Primary languagePython
LicenseMIT
Stars gained today207
Created2025-11-26
Last pushed2026-07-10

Trending History

Daily boardrank #22 · ▲ 207 stars

Related GitHub Projects

1

huggingface / transformers

Python★ 167,167⑂ 34,796▲ 94 stars
→
2

pytorch / pytorch

Python★ 104,110⑂ 31,933▲ 81 stars
→
3

fastapi / fastapi

Python★ 102,971⑂ 10,018▲ 56 stars
→
4

datawhalechina / hello-agents

Python★ 82,446⑂ 10,226▲ 202 stars
→
5

hugohe3 / ppt-master

Python★ 59,304⑂ 4,672▲ 515 stars
→
6

debpalash / VoiceStudio

Python★ 57,097⑂ 6,404▲ 1,066 stars
→
7

sgl-project / sglang

Python★ 36,966⑂ 9,435▲ 47 stars
→
8

anthropics / knowledge-work-plugins

Python★ 28,734⑂ 3,272▲ 626 stars
→

More Trending Repositories