Zhangyu Lai

dblp:377/5699 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
2 papers
Image and video processing · 53% Visual content generation and editing · 47%
Artificial intelligence
2 papers
Generative modeling · 57% 3D vision · 43%

Topics — the 8 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.012026
AnomalyPainter: Vision-Language-Diffusion Synergy for Realistic and Diverse Unseen Industrial Anomaly Synthesis · AAAI 2026
Machine learning › Generative modeling › diffusion model
latent diffusion model
1.012026
AnomalyPainter: Vision-Language-Diffusion Synergy for Realistic and Diverse Unseen Industrial Anomaly Synthesis · AAAI 2026
Image and video processing › pattern detection
anomaly detection
1.012026
AnomalyPainter: Vision-Language-Diffusion Synergy for Realistic and Diverse Unseen Industrial Anomaly Synthesis · AAAI 2026
Visual content generation and editing
image generation
1.012026
AnomalyPainter: Vision-Language-Diffusion Synergy for Realistic and Diverse Unseen Industrial Anomaly Synthesis · AAAI 2026
Image and video processing › pattern detection › anomaly detection
industrial anomaly detection
1.012026
AnomalyPainter: Vision-Language-Diffusion Synergy for Realistic and Diverse Unseen Industrial Anomaly Synthesis · AAAI 2026
Computer vision › 3D vision › 3d scene modeling › scene representation
3d gaussian representation
0.812024
Director3D: Real-world Camera Trajectory and 3D Scene Generation from Text · NeurIPS 2024
Computer vision › 3D vision
3d scene reconstruction
0.812024
Director3D: Real-world Camera Trajectory and 3D Scene Generation from Text · NeurIPS 2024
Visual content generation and editing › 3d content generation
text-to-3d generation
0.812024
Director3D: Real-world Camera Trajectory and 3D Scene Generation from Text · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

vision-language large model · 2.0texture-guided generation · 2.0controlnet · 2.0score distillation sampling · 1.5multi-view latent diffusion · 1.5diffusion transformer · 1.5
YearPublicationVenuePosition
2026 AnomalyPainter: Vision-Language-Diffusion Synergy for Realistic and Diverse Unseen Industrial Anomaly Synthesis
abstract
Visual anomaly detection is limited by the lack of sufficient anomaly data. While existing anomaly synthesis methods have made remarkable progress, achieving both realism and diversity in synthesis remains a major obstacle. To address this, we propose AnomalyPainter, a novel framework that breaks the diversity-realism trade-off dilemma through synergizing Vision Language Large Model (VLLM), Latent Diffusion Model (LDM), and our newly introduced texture library Tex-9K. Tex-9K is a professional texture library containing 75 categories and 8792 texture assets crafted for diverse anomaly synthesis. Leveraging VLLM's general knowledge, reasonable anomaly text descriptions are generated for each industrial object and matched with relevant diverse textures from Tex-9K. These textures then guide the LDM via ControlNet to paint on normal images. Furthermore, we introduce Texture-Aware Latent Init to stabilize the natural-image-trained ControlNet for industrial images. Extensive experiments show that AnomalyPainter outperforms existing methods in realism, diversity, and generalization, achieving superior downstream performance.
Zhangyu Lai, Jianghang Lin, Yansong Qu, Ming Li 0010, Liujuan Cao
AAAI1
2024 Director3D: Real-world Camera Trajectory and 3D Scene Generation from Text
abstract
Recent advancements in 3D generation have leveraged synthetic datasets with ground truth 3D assets and predefined camera trajectories. However, the potential of adopting real-world datasets, which can produce significantly more realistic 3D scenes, remains largely unexplored. In this work, we delve into the key challenge of the complex and scene-specific camera trajectories found in real-world captures. We introduce Director3D, a robust open-world text-to-3D generation framework, designed to generate both real-world 3D scenes and adaptive camera trajectories. To achieve this, (1) we first utilize a Trajectory Diffusion Transformer, acting as the \emph{Cinematographer}, to model the distribution of camera trajectories based on textual descriptions. Next, a Gaussian-driven Multi-view Latent Diffusion Model serves as the \emph{Decorator}, modeling the image sequence distribution given the camera trajectories and texts. This model, fine-tuned from a 2D diffusion model, directly generates pixel-aligned 3D Gaussians as an immediate 3D scene representation for consistent denoising. Lastly, the 3D Gaussians are further refined by a novel SDS++ loss as the \emph{Detailer}, which incorporates the prior of the 2D diffusion model. Extensive experiments demonstrate that Director3D outperforms existing methods, offering superior performance in real-world 3D generation.
Zhangyu Lai, Linning Xu, Yansong Qu, Liujuan Cao, Shengchuan Zhang, Bo Dai 0002, Rongrong Ji
NeurIPS2