Zike Wu

dblp:331/1483 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Generative modeling · 68% 3D vision · 23% Representation and self-supervised learning · 7%
Computer graphics and multimedia
2 papers
Visual content generation and editing · 43% Geometric modeling and processing · 43% Rendering · 14%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%

Topics — the 20 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.732024
Diffusion Time-step Curriculum for One Image to 3D Generation · CVPR 2024
Consistent3D: Towards Consistent High-Fidelity Text-to-3D Generation with Deterministic Sampling Prior · CVPR 2024
MVGamba: Unify 3D Content Generation as State Space Sequence Modeling · NeurIPS 2024
Machine learning › Generative modeling › diffusion model
score distillation sampling
1.522024
Diffusion Time-step Curriculum for One Image to 3D Generation · CVPR 2024
Consistent3D: Towards Consistent High-Fidelity Text-to-3D Generation with Deterministic Sampling Prior · CVPR 2024
Computer vision › 3D vision
3d generation
0.812024
Diffusion Time-step Curriculum for One Image to 3D Generation · CVPR 2024
Machine learning › Generative modeling › diffusion model
deterministic sampler
0.812024
Consistent3D: Towards Consistent High-Fidelity Text-to-3D Generation with Deterministic Sampling Prior · CVPR 2024
Computer vision › 3D vision › 3d generation
image-to-3d generation
0.812024
Diffusion Time-step Curriculum for One Image to 3D Generation · CVPR 2024
Machine learning › Generative modeling › diffusion model › diffusion sampling
ODE-based sampling
0.812024
Consistent3D: Towards Consistent High-Fidelity Text-to-3D Generation with Deterministic Sampling Prior · CVPR 2024
Machine learning › Generative modeling › diffusion model › 3d shape generation
text-to-3d generation
0.812024
Consistent3D: Towards Consistent High-Fidelity Text-to-3D Generation with Deterministic Sampling Prior · CVPR 2024
Visual content generation and editing
3d content editing
0.812024
View-Consistent 3D Editing with Gaussian Splatting · ECCV (35) 2024
Visual content generation and editing
3d content generation
0.812024
MVGamba: Unify 3D Content Generation as State Space Sequence Modeling · NeurIPS 2024
Geometric modeling and processing
3d reconstruction
0.812024
MVGamba: Unify 3D Content Generation as State Space Sequence Modeling · NeurIPS 2024
Geometric modeling and processing
3d scene representation
0.812024
View-Consistent 3D Editing with Gaussian Splatting · ECCV (35) 2024
Rendering
gaussian splatting
0.812024
View-Consistent 3D Editing with Gaussian Splatting · ECCV (35) 2024
Visual content generation and editing › 3d scene editing
multi-view consistent 3d editing
0.812024
View-Consistent 3D Editing with Gaussian Splatting · ECCV (35) 2024
Geometric modeling and processing › 3d reconstruction
multi-view reconstruction
0.812024
MVGamba: Unify 3D Content Generation as State Space Sequence Modeling · NeurIPS 2024
Machine learning › Representation and self-supervised learning › representation learning
invariant representation learning
0.612022
Invariant Representation Learning for Multimedia Recommendation · ACM Multimedia 2022
Recommender systems
multimodal recommendation
0.612022
Invariant Representation Learning for Multimedia Recommendation · ACM Multimedia 2022
Computer vision › 3D vision
3d reconstruction
0.212024
Diffusion Time-step Curriculum for One Image to 3D Generation · CVPR 2024
Machine learning › Generative modeling › diffusion model › 3d-aware diffusion
multi-view diffusion
0.212024
MVGamba: Unify 3D Content Generation as State Space Sequence Modeling · NeurIPS 2024
Computer vision › 3D vision
neural radiance field
0.212024
Diffusion Time-step Curriculum for One Image to 3D Generation · CVPR 2024
Machine learning › Trustworthy machine learning › robustness › spurious correlation
spurious correlation mitigation
0.212022
Invariant Representation Learning for Multimedia Recommendation · ACM Multimedia 2022

Methods — techniques the papers use, named apart from their topics

state space model · 1.5score distillation sampling · 1.53d gaussian splatting · 1.5invariant learning · 1.1causal inference · 1.1ordinary differential equation · 0.8gaussian splatting · 0.8diffusion time-step curriculum · 0.8diffusion model · 0.8consistency distillation · 0.8
YearPublicationVenuePosition
2024 Consistent3D: Towards Consistent High-Fidelity Text-to-3D Generation with Deterministic Sampling Prior
abstract
Score distillation sampling (SDS) and its variants have greatly boosted the development of text-to-3D generation, but are vulnerable to geometry collapse and poor textures yet. To solve this issue, we first deeply analyze the SDS and find that its distillation sampling process indeed corresponds to the trajectory sampling of a stochastic differential equation (SDE): SDS samples along an SDE trajectory to yield a less noisy sample which then serves as a guidance to optimize a 3D model. However, the randomness in SDE sampling often leads to a diverse and unpredictable sample which is not always less noisy, and thus is not a consistently correct guidance, explaining the vulnerability of SDS. Since for any SDE, there always exists an ordinary differential equation (ODE) whose trajectory sampling can deterministically and consistently converge to the desired target point as the SDE, we propose a novel and effective “Consistent3D” method that explores the ODE deterministic sampling prior for text-to-3D generation. Specifically, at each training iteration, given a rendered image by a 3D model, we first estimate its desired 3D score function by a pre-trained 2D diffusion model, and build an ODE for trajectory sampling. Next, we design a consistency distillation sampling loss which samples along the ODE trajectory to generate two adjacent samples and uses the less noisy sample to guide another more noisy one for distilling the deterministic prior into the 3D model. Experimental results show the efficacy of our Consistent3D in generating high-fidelity and diverse 3D objects and large-scale scenes, as shown in Fig. 1. The codes are available at https://github.com/sail-sg/Consistent3D.
Zike Wu, Pan Zhou 0002, Xuanyu Yi, Xiaoding Yuan, Hanwang Zhang
CVPR1
2024 Diffusion Time-step Curriculum for One Image to 3D Generation
abstract
Score distillation sampling (SDS) has been widely adopted to overcome the absence of unseen views in reconstructing 3D objects from a single image. It leverages pretrained 2D diffusion models as teacher to guide the reconstruction of student 3D models. Despite their remarkable success, SDS-based methods often encounter geometric artifacts and texture saturation. We find out the crux is the overlooked indiscriminate treatment of diffusion time-steps during optimization: it unreasonably treats the student-teacher knowledge distillation to be equal at all time-steps and thus entangles coarse-grained and fine-grained modeling. Therefore, we propose the Diffusion Time-step Curriculum one-image-to-3D pipeline (DTC123), which involves both the teacher and student models collaborating with the time-step curriculum in a coarse-to-fine manner. Extensive experiments on NeRF4, RealFusion15, GSO and Level50 benchmark demonstrate that DTC123 can produce multiview consistent, high-quality, and diverse 3D assets. Codes and more generation demos will be released in https://github.com/yxymessi/DTC123.
Xuanyu Yi, Zike Wu, Qingshan Xu 0001, Pan Zhou 0002, Joo-Hwee Lim, Hanwang Zhang
CVPR2
2024 View-Consistent 3D Editing with Gaussian Splatting
Xuanyu Yi, Zike Wu, Na Zhao 0004, Long Chen 0016, Hanwang Zhang
ECCV (35)3
2024 MVGamba: Unify 3D Content Generation as State Space Sequence Modeling
abstract
Recent 3D large reconstruction models (LRMs) can generate high-quality 3D content in sub-seconds by integrating multi-view diffusion models with scalable multi-view reconstructors. Current works further leverage 3D Gaussian Splatting as 3D representation for improved visual quality and rendering efficiency. However, we observe that existing Gaussian reconstruction models often suffer from multi-view inconsistency and blurred textures. We attribute this to the compromise of multi-view information propagation in favor of adopting powerful yet computationally intensive architectures (\eg, Transformers). To address this issue, we introduce MVGamba, a general and lightweight Gaussian reconstruction model featuring a multi-view Gaussian reconstructor based on the RNN-like State Space Model (SSM). Our Gaussian reconstructor propagates causal context containing multi-view information for cross-view self-refinement while generating a long sequence of Gaussians for fine-detail modeling with linear complexity. With off-the-shelf multi-view diffusion models integrated, MVGamba unifies 3D generation tasks from a single image, sparse images, or text prompts. Extensive experiments demonstrate that MVGamba outperforms state-of-the-art baselines in all 3D content generation scenarios with approximately only $0.1\times$ of the model size. The codes are available at \url{https://github.com/SkyworkAI/MVGamba}.
Xuanyu Yi, Zike Wu, Qiuhong Shen, Qingshan Xu 0001, Pan Zhou 0002, Joo-Hwee Lim, Shuicheng Yan, Xinchao Wang, Hanwang Zhang
NeurIPS2
2022 Invariant Representation Learning for Multimedia Recommendation
abstract
Multimedia recommendation forms a personalized ranking task with multimedia content representations which are mostly extracted via generic encoders. However, the generic representations introduce spurious correlations --- the meaningless correlation from the recommendation perspective. For example, suppose a user bought two dresses on the same model, this co-occurrence would produce a correlation between the model and purchases, but the correlation is spurious from the view of fashion recommendation. Existing work alleviates this issue by customizing preference-aware representations, requiring high-cost analysis and design.
Xiaoyu Du 0002, Zike Wu, Fuli Feng, Xiangnan He 0001, Jinhui Tang 0001
ACM Multimedia2