Yuyang Yin

dblp:354/8945 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2026
0009-0005-7443-6194ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Generative modeling · 58% 3D vision · 25% Vision and language · 9%
Computer graphics and multimedia
4 papers
Image and video processing · 36% Visual content generation and editing · 35% Virtual and augmented reality · 24%
Databases, data mining, and information retrieval
1 paper
Knowledge graphs · 100%

Topics — the 22 heaviest of 23, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
2.332025
ClassDiffusion: More Aligned Personalization Tuning with Explicit Class Guidance · ICLR 2025
Diffusion4D: Fast Spatial-temporal Consistent 4D generation via Video Diffusion Models · NeurIPS 2024
CLE Diffusion: Controllable Light Enhancement Diffusion Model · ACM Multimedia 2023
Computer vision › 3D vision › 3d scene modeling › scene representation
3d gaussian representation
0.912025
CLIP-GS: Unifying Vision-Language Representation with 3D Gaussian Splatting · ICCV 2025
Computer vision › 3D vision
3d reconstruction
0.912025
Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions · NeurIPS 2025
Computer vision › 3D vision › geometric deep learning
3d representation learning
0.912025
CLIP-GS: Unifying Vision-Language Representation with 3D Gaussian Splatting · ICCV 2025
Machine learning › Generative modeling › diffusion model › controllable generation
compositional generation
0.912025
ClassDiffusion: More Aligned Personalization Tuning with Explicit Class Guidance · ICLR 2025
Machine learning › Generative modeling › video generation
controllable video generation
0.912025
Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions · NeurIPS 2025
Machine learning › Representation and self-supervised learning
multimodal representation learning
0.912025
CLIP-GS: Unifying Vision-Language Representation with 3D Gaussian Splatting · ICCV 2025
Machine learning › Generative modeling › diffusion model › text-to-image generation
personalized text-to-image generation
0.912025
ClassDiffusion: More Aligned Personalization Tuning with Explicit Class Guidance · ICLR 2025
Machine learning › Generative modeling
video generation
0.912025
Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions · NeurIPS 2025
Knowledge graphs
entity linking
0.912025
Reverse Region-to-Entity Annotation for Pixel-Level Visual Entity Linking · AAAI 2025
Virtual and augmented reality › immersive content generation › XR content creation
immersive scene generation
0.912025
TiP4GEN: Text to Immersive Panorama 4D Scene Generation · ACM Multimedia 2025
Machine learning › Generative modeling › diffusion model
video diffusion model
0.812024
Diffusion4D: Fast Spatial-temporal Consistent 4D generation via Video Diffusion Models · NeurIPS 2024
Visual content generation and editing › 3d content generation
4d content generation
0.812024
Diffusion4D: Fast Spatial-temporal Consistent 4D generation via Video Diffusion Models · NeurIPS 2024
Machine learning › Generative modeling › diffusion model
conditional diffusion model
0.712023
CLE Diffusion: Controllable Light Enhancement Diffusion Model · ACM Multimedia 2023
Image and video processing
image enhancement
0.712023
CLE Diffusion: Controllable Light Enhancement Diffusion Model · ACM Multimedia 2023
Image and video processing › image enhancement
low-light image enhancement
0.712023
CLE Diffusion: Controllable Light Enhancement Diffusion Model · ACM Multimedia 2023
Computer vision › 3D vision › neural rendering
3d gaussian splatting
0.312025
TiP4GEN: Text to Immersive Panorama 4D Scene Generation · ACM Multimedia 2025
Computer vision › 3D vision
3d scene reconstruction
0.312025
TiP4GEN: Text to Immersive Panorama 4D Scene Generation · ACM Multimedia 2025
Computer vision › Vision and language
region-level understanding
0.312025
Reverse Region-to-Entity Annotation for Pixel-Level Visual Entity Linking · AAAI 2025
Visual content generation and editing › video generation
text-to-video generation
0.312025
TiP4GEN: Text to Immersive Panorama 4D Scene Generation · ACM Multimedia 2025
Visual content generation and editing
video generation
0.312025
Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions · NeurIPS 2025
Rendering
gaussian splatting
0.212024
Diffusion4D: Fast Spatial-temporal Consistent 4D generation via Video Diffusion Models · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

visual semantic tokenization · 1.7stereo navigation image processing · 1.7reverse annotation · 1.7region-interacted attention · 1.7multimodal conditioning · 1.7transformer · 0.9semantic preservation loss · 0.9fine-tuning · 0.9diffusion model · 0.9depth alignment · 0.9cross-attention · 0.9contrastive loss · 0.93d gaussian splatting · 0.9score distillation sampling · 0.8classifier-free guidance · 0.8segment anything model · 0.7illumination embedding · 0.7
YearPublicationVenuePosition
2026 Template-guided interpretable reasoning with execution feedback for LLM-based program repair
Sichong Hao, Yuyang Yin
Inf. Softw. Technol.4
2025 Reverse Region-to-Entity Annotation for Pixel-Level Visual Entity Linking
abstract
Visual Entity Linking (VEL) is a crucial task for achieving fine-grained visual understanding, matching objects within images (visual mentions) to entities in a knowledge base. Previous VEL tasks rely on textual inputs, but writing queries for complex scenes can be challenging. Visual inputs like clicks or bounding boxes offer a more convenient alternative. Therefore, we propose a new task, Pixel-Level Visual Entity Linking (PL-VEL), which uses pixel masks from visual inputs to refer to objects, supplementing reference methods for VEL. To facilitate research on this task, we have constructed the MaskOVEN-Wiki dataset through an entirely automatic reverse region-entity annotation framework. This dataset contains over 5 million annotations aligning pixel-level regions with entity-level labels, which will advance visual understanding towards fine-grained. Moreover, as pixel masks correspond to semantic regions in an image, we enhance previous patch-interacted attention with region-interacted attention by a visual semantic tokenization approach. Manual evaluation results indicate that the reverse annotation framework achieved a 94.8% annotation success rate. Experimental results show that models trained on this dataset improved accuracy by 18 points compared to zero-shot models. Additionally, the semantic tokenization method achieved a 5-point accuracy improvement over the trained baseline.
Zhengfei Xu, Sijia Zhao, Yanchao Hao, Yuyang Yin
AAAI6
2025 CLIP-GS: Unifying Vision-Language Representation with 3D Gaussian Splatting
abstract
Recent works in 3D multimodal learning have made remarkable progress. However, typically 3D multimodal models are only capable of handling point clouds. Compared to the emerging 3D representation technique, 3D Gaussian Splatting (3DGS), the spatially sparse point cloud cannot depict the texture information of 3D objects, resulting in inferior reconstruction capabilities. This limitation constrains the potential of point cloud-based 3D multimodal representation learning. In this paper, we present CLIP-GS, a novel multimodal representation learning framework grounded in 3DGS. We introduce the GS Tokenizer to generate serialized gaussian tokens, which are then processed through transformer layers pre-initialized with weights from point cloud models, resulting in the 3DGS embeddings. CLIP-GS leverages contrastive loss between 3DGS and the visual-text embeddings of CLIP, and we introduce an image voting loss to guide the directionality and convergence of gradient optimization. Furthermore, we develop an efficient way to generate triplets of 3DGS, images, and text, facilitating CLIP-GS in learning unified multimodal representations. Leveraging the well-aligned multimodal representations, CLIP-GS demonstrates versatility and outperforms point cloud-based models on various 3D tasks, including multimodal retrieval, zero-shot, and few-shot classification.
Siyu Jiao, Haoye Dong, Yuyang Yin, Zequn Jie, Yinlong Qian, Yao Zhao 0001, Humphrey Shi, Yunchao Wei
ICCV3
2025 ClassDiffusion: More Aligned Personalization Tuning with Explicit Class Guidance
abstract
Recent text-to-image customization works have proven successful in generating images of given concepts by fine-tuning diffusion models on a few examples. However, tuning-based methods inherently tend to overfit the concepts, resulting in failure to create the concept under multiple conditions (*e.g.*, headphone is missing when generating "a <sks>`dog wearing a headphone"). Interestingly, we notice that the base model before fine-tuning exhibits the capability to compose the base concept with other elements (*e.g.*, "a dog wearing a headphone"), implying that the compositional ability only disappears after personalization tuning. We observe a semantic shift in the customized concept after fine-tuning, indicating that the personalized concept is not aligned with the original concept, and further show through theoretical analyses that this semantic shift leads to increased difficulty in sampling the joint conditional probability distribution, resulting in the loss of the compositional ability. Inspired by this finding, we present **ClassDiffusion**, a technique that leverages a **semantic preservation loss** to explicitly regulate the concept space when learning a new concept. Although simple, this approach effectively prevents semantic drift during the fine-tuning process of the target concepts. Extensive qualitative and quantitative experiments demonstrate that the use of semantic preservation loss effectively improves the compositional abilities of fine-tuning models. Lastly, we also extend our ClassDiffusion to personalized video generation, demonstrating its flexibility.
Jiannan Huang 0002, Jun Hao Liew, Hanshu Yan, Yuyang Yin, Yao Zhao 0001, Humphrey Shi, Yunchao Wei
ICLR4
2025 TiP4GEN: Text to Immersive Panorama 4D Scene Generation
abstract
With the rapid advancement and widespread adoption of VR/AR technologies, there is a growing demand for the creation of high-quality, immersive dynamic scenes. However, existing generation works predominantly concentrate on the creation of static scenes or narrow perspective-view dynamic scenes, falling short of delivering a truly 360-degree immersive experience from any viewpoint. In this paper, we introduce TiP4GEN, an advanced text-to-dynamic panorama scene generation framework that enables fine-grained content control and synthesizes motion-rich, geometry-consistent panoramic 4D scenes. TiP4GEN integrates panorama video generation and dynamic scene reconstruction to create 360-degree immersive virtual environments. For video generation, we introduce a Dual-branch Generation Model consisting of a panorama branch and a perspective branch, responsible for global and local view generation, respectively. A bidirectional cross-attention mechanism facilitates comprehensive information exchange between the branches. For scene reconstruction, we propose a Geometry-aligned Reconstruction Model based on 3D Gaussian Splatting. By aligning spatial-temporal point clouds using metric depth maps and initializing scene cameras with estimated poses, our method ensures geometric consistency and temporal coherence for the reconstructed scenes. Extensive experiments demonstrate the effectiveness of our proposed designs and the superiority of TiP4GEN in generating visually compelling and motion-coherent dynamic panoramic scenes.
Hanwen Liang, Dejia Xu, Yuyang Yin, Konstantinos N. Plataniotis, Yao Zhao 0001, Yunchao Wei
ACM Multimedia4
2025 Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions
abstract
The synthesis of realistic Martian landscape videos, essential for mission rehearsal and robotic simulation, presents unique challenges. These primarily stem from the scarcity of high-quality Martian data and the significant domain gap relative to terrestrial imagery.To address these challenges, we introduce a holistic solution comprising two main components: 1) a data curation framework, Multimodal Mars Synthesis (M3arsSynth), which processes stereo navigation images to render high-fidelity 3D video sequences. 2) a video-based Martian terrain generator (MarsGen), that utilizes multimodal conditioning data to accurately synthesize novel, 3D-consistent frames. Our data are sourced from NASA’s Planetary Data System (PDS), covering diverse Martian terrains and dates, enabling the production of physics-accurate 3D surface models at metric-scale resolution. During inference, MarsGen is conditioned on an initial image frame and can be guided by specified camera trajectories or textual prompts to generate new environments.Experimental results demonstrate that our solution surpasses video synthesis approaches trained on terrestrial data, achieving superior visual quality and 3D structural consistency.
Zhiwen Fan, Wenyan Cong, Xinhang Liu, Yuyang Yin, Matthew Foutter, Panwang Pan, Chenyu You, Yue Wang 0041, Zhangyang Wang, Yao Zhao 0001, Marco Pavone 0001, Yunchao Wei
NeurIPS5
2024 Diffusion4D: Fast Spatial-temporal Consistent 4D generation via Video Diffusion Models
abstract
The availability of large-scale multimodal datasets and advancements in diffusion models have significantly accelerated progress in 4D content generation. Most prior approaches rely on multiple images or video diffusion models, utilizing score distillation sampling for optimization or generating pseudo novel views for direct supervision. However, these methods are hindered by slow optimization speeds and multi-view inconsistency issues. Spatial and temporal consistency in 4D geometry has been extensively explored respectively in 3D-aware diffusion models and traditional monocular video diffusion models. Building on this foundation, we propose a strategy to migrate the temporal consistency in video diffusion models to the spatial-temporal consistency required for 4D generation. Specifically, we present a novel framework, \textbf{Diffusion4D}, for efficient and scalable 4D content generation. Leveraging a meticulously curated dynamic 3D dataset, we develop a 4D-aware video diffusion model capable of synthesizing orbital views of dynamic 3D assets. To control the dynamic strength of these assets, we introduce a 3D-to-4D motion magnitude metric as guidance. Additionally, we propose a novel motion magnitude reconstruction loss and 3D-aware classifier-free guidance to refine the learning and generation of motion dynamics. After obtaining orbital views of the 4D asset, we perform explicit 4D construction with Gaussian splatting in a coarse-to-fine manner. Extensive experiments demonstrate that our method surpasses prior state-of-the-art techniques in terms of generation efficiency and 4D geometry consistency across various prompt modalities.
Hanwen Liang, Yuyang Yin, Dejia Xu, Hanxue Liang, Zhangyang Wang, Konstantinos N. Plataniotis, Yao Zhao 0001, Yunchao Wei
NeurIPS2
2023 CLE Diffusion: Controllable Light Enhancement Diffusion Model
abstract
Low light enhancement has gained increasing importance with the rapid development of visual creation and editing. However, most existing enhancement algorithms are designed to homogeneously increase the brightness of images to a pre-defined extent, limiting the user experience. To address this issue, we propose Controllable Light Enhancement Diffusion Model, dubbed CLE Diffusion, a novel diffusion framework to provide users with rich controllability.Built with a conditional diffusion model, we introduce an illumination embedding to let users control their desired brightness level. Additionally, we incorporate the Segment-Anything Model (SAM) to enable user-friendly region controllability, where users can click on objects to specify the regions they wish to enhance. Extensive experiments demonstrate that CLE Diffusion achieves competitive performance regarding quantitative metrics, qualitative results, and versatile controllability. Project page: https://yuyangyin.github.io/CLEDiffusion
Yuyang Yin, Dejia Xu, Chuangchuang Tan, Ping Liu 0004, Yao Zhao 0001, Yunchao Wei
ACM Multimedia1