Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Zhibing Li

dblp:215/1997 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
5since 2021 · last 2025
0009-0002-4528-5495ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
4 papers
Visual content generation and editing · 63% Computational photography and imaging · 10% Image and video coding · 10%
Artificial intelligence
3 papers
3D vision · 36% Vision and language · 32% Generative modeling · 16%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Visual content generation and editing
3d content generation
2.432025
SS4D: Native 4D Generative Model via Structured Spacetime Latents · ACM Trans. Graph. 2025
Hi3DEval: Advancing 3D Generation Evaluation with Hierarchical Validity · NeurIPS 2025
HyperDreamer: Hyper-Realistic 3D Content Generation and Editing from a Single Image · SIGGRAPH Asia 2023
Computer vision › 3D vision
3d content generation
0.912025
Hi3DEval: Advancing 3D Generation Evaluation with Hierarchical Validity · NeurIPS 2025
Computer vision › 3D vision › 3d reconstruction
multi-view reconstruction
0.912025
IDArb: Intrinsic Decomposition for Arbitrary Number of Input Views and Illuminations · ICLR 2025
Visual content generation and editing › 3d content generation
4d content generation
0.912025
SS4D: Native 4D Generative Model via Structured Spacetime Latents · ACM Trans. Graph. 2025
Computational photography and imaging
intrinsic image decomposition
0.912025
IDArb: Intrinsic Decomposition for Arbitrary Number of Input Views and Illuminations · ICLR 2025
Image and video coding
quality assessment
0.912025
Hi3DEval: Advancing 3D Generation Evaluation with Hierarchical Validity · NeurIPS 2025
Machine learning › Trustworthy machine learning › AI safety
human-aligned evaluation
0.812024
GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation · CVPR 2024
Computer vision › Vision and language › vision-language model
multimodal large language model
0.812024
GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation · CVPR 2024
Machine learning › Generative modeling › diffusion model › 3d shape generation
text-to-3d generation
0.812024
GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation · CVPR 2024
Computer vision › Vision and language
vision-language model
0.812024
GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation · CVPR 2024
Visual content generation and editing
3d content editing
0.712023
HyperDreamer: Hyper-Realistic 3D Content Generation and Editing from a Single Image · SIGGRAPH Asia 2023
Visual content generation and editing › 3d content generation
image-to-3d generation
0.712023
HyperDreamer: Hyper-Realistic 3D Content Generation and Editing from a Single Image · SIGGRAPH Asia 2023
Visual content generation and editing › material editing
texture editing
0.712023
HyperDreamer: Hyper-Realistic 3D Content Generation and Editing from a Single Image · SIGGRAPH Asia 2023
Rendering › appearance acquisition › material acquisition
material estimation
0.212023
HyperDreamer: Hyper-Realistic 3D Content Generation and Editing from a Single Image · SIGGRAPH Asia 2023

Methods — techniques the papers use, named apart from their topics

diffusion model · 3.3video-based representation · 1.7multi-agent annotation · 1.7hybrid 3d representation · 1.7cross-view attention · 1.7temporal downsampling · 0.9spacetime latent representation · 0.9factorized 4d convolution · 0.9pairwise comparison · 0.8elo rating · 0.8semantic segmentation · 0.7
YearPublicationVenuePosition
2025 IDArb: Intrinsic Decomposition for Arbitrary Number of Input Views and Illuminations
abstract
Capturing geometric and material information from images remains a fundamental challenge in computer vision and graphics. Traditional optimization-based methods often require hours of computational time to reconstruct geometry, material properties, and environmental lighting from dense multi-view inputs, while still struggling with inherent ambiguities between lighting and material. On the other hand, learning-based approaches leverage rich material priors from existing 3D object datasets but face challenges with maintaining multi-view consistency. In this paper, we introduce IDArb, a diffusion-based model designed to perform intrinsic decomposition on an arbitrary number of images under varying illuminations. Our method achieves highly accurate and multi-view consistent estimation on surface normals and material properties. This is made possible through a novel cross-view, cross-domain attention module and an illumination-augmented, view-adaptive training strategy. Additionally, we introduce ARB-Objaverse, a new dataset that provides large-scale multi-view intrinsic data and renderings under diverse lighting conditions, supporting robust training. Extensive experiments demonstrate that IDArb outperforms state-of-the-art methods both qualitatively and quantitatively. Moreover, our approach facilitates a range of downstream tasks, including single-image relighting, photometric stereo, and 3D reconstruction, highlighting its broad applicability in realistic 3D content creation. Project website: https://lizb6626.github.io/IDArb/.
Zhibing Li, Jing Tan 0002, Mengchen Zhang 0001, Jiaqi Wang 0003, Dahua Lin
ICLR1
2025 Hi3DEval: Advancing 3D Generation Evaluation with Hierarchical Validity
abstract
Despite rapid advances in 3D content generation, quality assessment for the generated 3D assets remains challenging.Existing methods mainly rely on image-based metrics and operate solely at the object level, limiting their ability to capture spatial Despite rapid advances in 3D content generation, quality assessment for the generated 3D assets remains challenging.Existing methods mainly rely on image-based metrics and operate solely at the object level, limiting their ability to capture spatial coherence, material authenticity, and high-fidelity local details.1) To address these challenges, we introduce Hi3DEval, a hierarchical evaluation framework tailored for 3D generative content. It combines both object-level and part-level evaluation, enabling holistic assessments across multiple dimensions as well as fine-grained quality analysis. Additionally, we extend texture evaluation beyond aesthetic appearance by explicitly assessing material realism, focusing on attributes such as albedo, saturation, and metallicness. 2) To support this framework, we construct Hi3DBench, a large-scale dataset comprising diverse 3D assets and high-quality annotations, accompanied by a reliable multi-agent annotation pipeline.We further propose a 3D-aware automated scoring system based on hybrid 3D representations. Specifically, we leverage video-based representations for object-level and material-subject evaluations to enhance modeling of spatio-temporal consistency and employ pretrained 3D features for part-level perception.Extensive experiments demonstrate that our approach outperforms existing image-based metrics in modeling 3D characteristics and achieves superior alignment with human preference, providing a scalable alternative to manual evaluations.
Long Zhuo, Ziyang Chu, Zhibing Li, Liang Pan, Dahua Lin, Ziwei Liu 0002
NeurIPS5
2025 SS4D: Native 4D Generative Model via Structured Spacetime Latents
abstract
We present SS4D, a native 4D generative model that synthesizes dynamic 3D objects directly from monocular video. Unlike prior approaches that construct 4D representations by optimizing over 3D or video generative models, we train a generator directly on 4D data, achieving high fidelity, temporal coherence, and structural consistency. At the core of our method is a compressed set of structured spacetime latents. Specifically, (1) To address the scarcity of 4D training data, we build on a pre-trained single-image-to-3D model, preserving strong spatial consistency. (2) Temporal consistency is enforced by introducing dedicated temporal layers that reason across frames. (3) To support efficient training and inference over long video sequences, we compress the latent sequence along the temporal axis using factorized 4D convolutions and temporal downsampling blocks. In addition, we employ a carefully designed training strategy to enhance robustness against occlusion and motion blur, leading to high-quality generation. Extensive experiments show that SS4D produces spatio-temporally consistent 4D objects with superior quality and efficiency, significantly outperforming state-of-the-art methods on both synthetic and real-world datasets.
Zhibing Li, Mengchen Zhang 0001, Jing Tan 0002, Jiaqi Wang 0003, Dahua Lin
ACM Trans. Graph.1
2024 GPT-4V(ision) is a Human-Aligned Evaluator for Text-to-3D Generation
abstract
Despite recent advances in text-to-3D generative methods, there is a notable absence of reliable evaluation metrics. Existing metrics usually focus on a single criterion each, such as how well the asset aligned with the input text. These metrics lack the flexibility to generalize to different evaluation criteria and might not align well with human preferences. Conducting user preference studies is an alternative that offers both adaptability and human-aligned results. User studies, however, can be very ex-pensive to scale. This paper presents an automatic, ver-satile, and human-aligned evaluation metric for text-to-3D generative models. To this end, we first develop a prompt generator using GPT-4V to generate evaluating prompts, which serve as input to compare text-to-3D models. We further design a method instructing GPT-4V to compare two 3D assets according to user-defined crite-ria. Finally, we use these pairwise comparison results to assign these models Elo ratings. Experimental results suggest our metric strongly aligns with human preference across different evaluation criteria. Our code is available at https://github.com/3DTopia/GPTEval3D.
Guandao Yang, Zhibing Li, Kai Zhang 0045, Ziwei Liu 0002, Leonidas J. Guibas, Dahua Lin, Gordon Wetzstein
CVPR3
2023 HyperDreamer: Hyper-Realistic 3D Content Generation and Editing from a Single Image
abstract
3D content creation from a single image is a long-standing yet highly desirable task. Recent advances introduce 2D diffusion priors, yielding reasonable results. However, existing methods are not hyper-realistic enough for post-generation usage, as users cannot view, render and edit the resulting 3D content from a full range. To address these challenges, we introduce HyperDreamer with several key designs and appealing properties: 1) Full-range viewable: 360° mesh modeling with high-resolution textures enables the creation of visually compelling 3D models from a full range of observation points. 2) Full-range renderable: Fine-grained semantic segmentation and data-driven priors are incorporated as guidance to learn reasonable albedo, roughness, and specular properties of the materials, enabling semantic-aware arbitrary material estimation. 3) Full-range editable: For a generated model or their own data, users can interactively select any region via a few clicks and efficiently edit the texture with text-based guidance. Extensive experiments demonstrate the effectiveness of HyperDreamer in modeling region-aware materials with high-resolution textures and enabling user-friendly editing. We believe that HyperDreamer holds promise for advancing 3D content creation and finding applications in various domains.
Zhibing Li, Shuai Yang 0001, Pan Zhang 0001, Xingang Pan, Jiaqi Wang 0003, Dahua Lin, Ziwei Liu 0002
SIGGRAPH Asia2