EDBT 2026 Demo / reviewers in the wild / expert
Zibo Zhao 0001
dblp:237/0093-1
· DBLP profile ↗
14ranked-venue papers
2as first author
14since 2021 · last 2026
0009-0005-4831-8418ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 1 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Unified multi-modality conditional latent diffusion model for point cloud generation
Yihang Yang, Zibo Zhao 0001, Fukun Yin, Wen Liu 0003, Yuhan Ding, Biao Jiang, Gang Yu 0002, Tao Chen 0003 |
Pattern Recognit. | 2 |
| 2025 | FlexiTex: Enhancing Texture Generation via Visual GuidanceabstractRecent texture generation methods achieve impressive results due to the powerful generative prior they leverage from large-scale text-to-image diffusion models. However, abstract textual prompts are limited in providing global textural or shape information, which results in the texture generation methods producing blurry or inconsistent patterns. To tackle this, we present FlexiTex, embedding rich information via visual guidance to generate a high-quality texture. The core of FlexiTex is the Visual Guidance Enhancement module, which incorporates more specific information from visual guidance to reduce ambiguity in the text prompt and preserve high-frequency details. To further enhance the visual guidance, we introduce a Direction-Aware Adaptation module that automatically designs direction prompts based on different camera poses, avoiding the Janus problem and maintaining semantically global consistency. Benefiting from the visual guidance, FlexiTex produces quantitatively and qualitatively sound results, demonstrating its potential to advance texture generation for real-world applications. Dadong Jiang, Xianghui Yang, Zibo Zhao 0001, Zeqiang Lai, Shaoxiong Yang, Chunchao Guo, Xiaobo Zhou 0003, Zhihui Ke |
AAAI | 3 |
| 2025 | Scaling Mesh Generation via Compressive TokenizationabstractWe propose a compressive yet effective mesh tokenization, Blocked and Patchified Tokenization (BPT), facilitating the generation of meshes exceeding 8k faces. BPT compresses mesh sequences by employing block-wise indexing and patch aggregation, reducing their length by approximately 75% compared to the vanilla coordinate sequences. This compression milestone unlocks the potential to utilize mesh data with significantly more faces, thereby enhancing detail richness and improving generation robustness. Empowered with the BPT, we have built a foundation mesh generative model training on scaled mesh data to support flexible control for point clouds and images. Our model demonstrates the capability to generate meshes with intricate details and accurate topology, achieving SoTA performance on mesh generation and reaching the level for direct product usage. Haohan Weng, Zibo Zhao 0001, Biwen Lei, Xianghui Yang, Jian Liu 0036, Zeqiang Lai, Zhuo Chen 0054, Jie Jiang 0015, Chunchao Guo, Tong Zhang 0015, Shenghua Gao, C. L. Philip Chen |
CVPR | 2 |
| 2025 | RomanTex: Decoupling 3D-Aware Rotary Positional Embedded Multi-Attention Network for Texture Synthesis
Mingxin Yang, Zibo Zhao 0001, Jie Jiang 0015, Chunchao Guo |
ICCV | 6 |
| 2025 | Unleashing Vecset Diffusion Model for Fast Shape Generationabstract3D shape generation has greatly flourished through the development of so-called "native" 3D diffusion, particularly through the Vecset Diffusion Model (VDM). While recent advancements have shown promising results in generating high-resolution 3D shapes, VDM still struggles with high-speed generation. Challenges exist because of difficulties not only in accelerating diffusion sampling but also VAE decoding in VDM, areas under-explored in previous works. To address these challenges, we present FlashVDM, a systematic framework for accelerating both VAE and DiT in VDM. For DiT, FlashVDM enables flexible diffusion sampling with as few as 5 inference steps and comparable quality, which is made possible by stabilizing consistency distillation with our newly introduced Progressive Flow Distillation. For VAE, we introduce a lightning vecset decoder equipped with Adaptive KV Selection, Hierarchical Volume Decoding, and Efficient Network Design. By exploiting the locality of the vecset and the sparsity of shape surface in the volume, our decoder drastically lowers FLOPs, minimizing the overall decoding overhead. We apply FlashVDM to Hunyuan3D-2 to obtain Hunyuan3D-2 Turbo. Through systematic evaluation, we show that our model significantly outperforms existing fast 3D generation methods, achieving comparable performance to the state-of-the-art while reducing inference time by over 45x for reconstruction and 32x for generation. Code and models are available at https://github.com/Tencent/FlashVDM. Zeqiang Lai, Zibo Zhao 0001, Fuyun Wang, Huiwen Shi, Xianghui Yang, Qingxiang Lin, Jie Jiang 0015, Chunchao Guo, Xiangyu Yue 0001 |
ICCV | 3 |
| 2025 | FreeMesh: Boosting Mesh Generation with Coordinates MergingabstractThe next-coordinate prediction paradigm has emerged as the de facto standard in current auto-regressive mesh generation methods. Despite their effectiveness, there is no efficient measurement for the various tokenizers that serialize meshes into sequences. In this paper, we introduce a new metric Per-Token-Mesh-Entropy (PTME) to evaluate the existing mesh tokenizers theoretically without any training. Building upon PTME, we propose a plug-and-play tokenization technique called coordinate merging. It further improves the compression ratios of existing tokenizers by rearranging and merging the most frequent patterns of coordinates. Through experiments on various tokenization methods like MeshXL, MeshAnything V2, and Edgerunner, we further validate the performance of our method. We hope that the proposed PTME and coordinate merging can enhance the existing mesh tokenizers and guide the further development of native mesh generation. Jian Liu 0036, Haohan Weng, Biwen Lei, Xianghui Yang, Zibo Zhao 0001, Zhuo Chen 0054, Song Guo 0001, Tao Han 0002, Chunchao Guo |
ICML | 5 |
| 2025 | ShapeGPT: 3D Shape Generation With a Unified Multi-Modal Language ModelabstractThe advent of large language models, which enable flexibility through instruction-driven approaches, has revolutionized many traditional generative tasks, but large models for 3D data, particularly in comprehensively handling 3D shapes with other modalities, are still under-explored. By achieving instruction-based shape generation, versatile multi-modal generative shape models can significantly benefit various fields, such as 3D virtual construction and network-aided design. In this article, we present ShapeGPT, a shape-included multi-modal framework to leverage strong pre-trained language models to address multiple shape-relevant tasks. Specifically, ShapeGPT employs a “word-sentence-paragraph” framework to discretize continuous shapes into shape words, further assembles these words into shape sentences, and integrates shape with instructional text for multi-modal paragraphs. To learn this shape-language model, we use a three-stage training scheme, including shape representation, multi-modal alignment, and instruction-based generation, to align shape-language codebooks and learn the intricate correlations among these modalities. Extensive experiments demonstrate that ShapeGPT achieves comparable performance across shape-relevant tasks, including text-to-shape, shape-to-text, shape completion, and shape editing. Fukun Yin, Xin Chen 0040, Chi Zhang 0007, Biao Jiang, Zibo Zhao 0001, Wen Liu 0003, Gang Yu 0002, Tao Chen 0003 |
IEEE Trans. Multim. | 5 |
| 2024 | RoomDesigner: Encoding Anchor-latents for Style-consistent and Shape-compatible Indoor Scene GenerationabstractIndoor scene generation aims at creating shape-compatible, style-consistent furniture arrangements within a spatially reasonable layout. However, most existing approaches primarily focus on generating plausible furniture layouts without incorporating specific details related to individual furniture. To address this limitation, we propose a two-stage model integrating shape priors into the indoor scene generation by encoding furniture as anchor latent representations. In the first stage, we employ discrete vector quantization to encode each piece of furniture as anchor-latents. Based on the anchor-latents representation, the shape and location information of furniture was characterized by a concatenation of location, size, orientation, class, and our anchor latent. In the second stage, we leverage a transformer model to predict indoor scenes configuration autoregressively. Thanks to the proposed anchor-latents representations, our generative model can synthesis furniture in diverse shapes and produce physically plausible arrangements with shape-compatible and style-consistent furniture. Furthermore, our method facilitates various human interaction applications, such as style-consistent scene completion, object mismatch correction, and controllable object-level editing. Experimental results on the 3D-Front dataset demonstrate that our approach can generate more consistent and compatible indoor scenes compared to existing methods, even without shape retrieval. Additionally, extensive ablation studies confirm the effectiveness of our design choices in the indoor scene generation model. Yiqun Zhao, Zibo Zhao 0001, Jing Li 0117, Sixun Dong, Shenghua Gao |
3DV | 2 |
| 2024 | Paint3D: Paint Anything 3D With Lighting-Less Texture Diffusion ModelsabstractThis paper presents Paint3D, a novel coarse-to-fine generative framework that is capable of producing high-resolution, lighting-less, and diverse 2K UV texture maps for untextured 3D meshes conditioned on text or image inputs. The key challenge addressed is generating high-quality textures without embedded illumination information, which allows the textures to be re-lighted or re-edited within modern graphics pipelines. To achieve this, our method first leverages a pre-trained depth-aware 2D diffusion model to generate view-conditional images and per-form multi-view texture fusion, producing an initial coarse texture map. However, as 2D models cannot fully repre-sent 3D shapes and disable lighting effects, the coarse texture map exhibits incomplete areas and illumination artifacts. To resolve this, we train separate UV Inpainting and UVHD diffusion models specialized for the shape-aware re-finement of incomplete areas and the removal of illumination artifacts. Through this coarse-to-fine process, Paint3D can produce high-quality 2K UV textures that maintain se-mantic consistency while being lighting-less, significantly advancing the state-of-the-art in texturing 3D objects. Xianfang Zeng, Xin Chen 0040, Zhongqi Qi, Wen Liu 0003, Zibo Zhao 0001, Zhibin Wang 0004, Yong Liu 0007, Gang Yu 0002 |
CVPR | 5 |
| 2024 | HeroMaker: Human-centric Video Editing with Motion PriorsabstractVideo generation and editing, particularly human-centric video editing, has seen a surge of interest in its potential to create immersive and dynamic content. A fundamental challenge is ensuring temporal coherence and visual harmony across frames, especially in handling large-scale human motion and maintaining consistency over long sequences. The previous methods, such as zero-shot text-to-video methods with diffusion model, struggle with flickering and length limitations. In contrast, methods employing Video-2D representations grapple with accurately capturing complex structural relationships in large-scale human motion. Simultaneously, some patterns on the human body appear intermittently throughout the video, posing a knotty problem in identifying visual correspondence. To address the above problems, we present HeroMaker. This human-centric video editing framework manipulates the person's appearance within the input video and achieves consistent results across frames. Specifically, we propose to learn the motion priors, which represent the correspondences between dual canonical fields and each video frame, by leveraging the body mesh-based human motion warping and neural deformation-based margin refinement in the video reconstruction framework to ensure the semantic correctness of canonical fields. HeroMaker performs human-centric video editing by manipulating the dual canonical fields and combining them with motion priors to synthesize temporally coherent and visually plausible results. Comprehensive experiments demonstrate that our approach surpasses existing methods regarding temporal consistency, visual quality, and semantic coherence. Zibo Zhao 0001, Yihao Zhi, Yiqun Zhao, Binbin Huang 0004, Ruoyu Wang 0014, Michael Xuan, Shenghua Gao |
ACM Multimedia | 2 |
| 2024 | TSP-Transformer: Task-Specific Prompts Boosted Transformer for Holistic Scene UnderstandingabstractHolistic scene understanding includes semantic segmentation, surface normal estimation, object boundary detection, depth estimation, etc. The key aspect of this problem is to learn representation effectively, as each subtask builds upon not only correlated but also distinct attributes. Inspired by visual-prompt tuning, we propose a Task-Specific Prompts Transformer, dubbed TSP-Transformer, for holistic scene understanding. It features a vanilla transformer in the early stage and tasks-specific prompts transformer encoder in the lateral stage, where tasks-specific prompts are augmented. By doing so, the transformer layer learns the generic information from the shared parts and is endowed with task-specific capacity. First, the tasks-specific prompts serve as induced priors for each task effectively. Moreover, the task-specific prompts can be seen as switches to favor task-specific representation learning for different tasks. Extensive experiments on NYUD-v2 and PASCAL-Context show that our method achieves state-of-the-art performance, validating the effectiveness of our method for holistic scene understanding. We also provide our code in the following link1. Jing Li 0117, Zibo Zhao 0001, Dongze Lian, Binbin Huang 0004, Shenghua Gao |
WACV | 3 |
| 2024 | Instruct Pix-to-3D: Instructional 3D object generation from a single image
Wen Liu 0003, Wanzhang Li, Zibo Zhao 0001, Fukun Yin, Xin Chen 0018, Lei Zhao 0035, Tao Chen 0003 |
Neurocomputing | 4 |
| 2023 | Michelangelo: Conditional 3D Shape Generation based on Shape-Image-Text Aligned Latent RepresentationabstractWe present a novel alignment-before-generation approach to tackle the challenging task of generating general 3D shapes based on 2D images or texts. Directly learning a conditional generative model from images or texts to 3D shapes is prone to producing inconsistent results with the conditions because 3D shapes have an additional dimension whose distribution significantly differs from that of 2D images and texts. To bridge the domain gap among the three modalities and facilitate multi-modal-conditioned 3D shape generation, we explore representing 3D shapes in a shape-image-text-aligned space. Our framework comprises two models: a Shape-Image-Text-Aligned Variational Auto-Encoder (SITA-VAE) and a conditional Aligned Shape Latent Diffusion Model (ASLDM). The former model encodes the 3D shapes into the shape latent space aligned to the image and text and reconstructs the fine-grained 3D neural fields corresponding to given shape embeddings via the transformer-based decoder. The latter model learns a probabilistic mapping function from the image or text space to the latent shape space. Our extensive experiments demonstrate that our proposed approach can generate higher-quality and more diverse 3D shapes that better semantically conform to the visual or textural conditional inputs, validating the effectiveness of the shape-image-text-aligned space for cross-modality 3D shape generation. Zibo Zhao 0001, Wen Liu 0003, Xin Chen 0040, Xianfang Zeng, Rui Wang 0099, Tao Chen 0003, Gang Yu 0002, Shenghua Gao |
NeurIPS | 1 |
| 2021 | Prior Based Human CompletionabstractWe study a very challenging task, human image completion, which tries to recover the human body part with a reasonable human shape from the corrupted region. Since each human body part is unique, it is infeasible to restore the missing part by borrowing textures from other visible regions. Thus, we propose two types of learned priors to compensate for the damaged region. One is a structure prior, it uses a human parsing map to represent the human body structure. The other is a structure-texture correlation prior. It learns a structure and a texture memory bank, which encodes the common body structures and texture patterns, respectively. With the aid of these memory banks, the model could utilize the visible pattern to query and fetch a similar structure and texture pattern to introduce additional reasonable structures and textures for the corrupted region. Besides, since multiple potential human shapes are underlying the corrupted region, we propose multi-scale structure discriminators to further restore a plausible topological structure. Experiments on various large-scale benchmarks demonstrate the effectiveness of our proposed method. Zibo Zhao 0001, Wen Liu 0003, Yanyu Xu 0001, Xianing Chen, Weixin Luo, Bohui Zhu, Tong Liu 0037, Binqiang Zhao, Shenghua Gao |
CVPR | 1 |