Hanyang Kong

dblp:241/9666 · DBLP profile ↗
← Back
5ranked-venue papers
5as first author
5since 2021 · last 2025
0000-0002-5895-5112ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Generative modeling · 61% 3D vision · 37% Representation and self-supervised learning · 2%
Computer graphics and multimedia
4 papers
Rendering · 100%

Topics — the 17 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
2.332025
Generative Sparse-View Gaussian Splatting · CVPR 2025
DreamDrone: Text-to-Image Diffusion Models Are Zero-Shot Perpetual View Generators · ECCV (13) 2024
Priority-Centric Human Motion Generation in Discrete Latent Space · ICCV 2023
Rendering
gaussian splatting
1.722025
Rogsplat: Robust Gaussian Splatting Via Generative Priors · ICCV 2025
Efficient Gaussian Splatting for Monocular Dynamic Scene Rendering via Sparse Time-Variant Attribute Modeling · AAAI 2025
Rendering
novel view synthesis
1.622025
Generative Sparse-View Gaussian Splatting · CVPR 2025
DreamDrone: Text-to-Image Diffusion Models Are Zero-Shot Perpetual View Generators · ECCV (13) 2024
Computer vision › 3D vision › neural rendering
3d gaussian splatting
0.912025
Generative Sparse-View Gaussian Splatting · CVPR 2025
Computer vision › 3D vision
3d reconstruction
0.912025
Generative Sparse-View Gaussian Splatting · CVPR 2025
Computer vision › 3D vision
3d scene reconstruction
0.912025
Efficient Gaussian Splatting for Monocular Dynamic Scene Rendering via Sparse Time-Variant Attribute Modeling · AAAI 2025
Machine learning › Generative modeling
generative prior
0.912025
Rogsplat: Robust Gaussian Splatting Via Generative Priors · ICCV 2025
Machine learning › Generative modeling › diffusion model
image diffusion model
0.912025
Generative Sparse-View Gaussian Splatting · CVPR 2025
Computer vision › 3D vision › 3d reconstruction › dynamic 3d reconstruction
monocular dynamic scene reconstruction
0.912025
Efficient Gaussian Splatting for Monocular Dynamic Scene Rendering via Sparse Time-Variant Attribute Modeling · AAAI 2025
Rendering › temporal rendering
dynamic scene rendering
0.912025
Efficient Gaussian Splatting for Monocular Dynamic Scene Rendering via Sparse Time-Variant Attribute Modeling · AAAI 2025
Rendering
neural rendering
0.912025
Efficient Gaussian Splatting for Monocular Dynamic Scene Rendering via Sparse Time-Variant Attribute Modeling · AAAI 2025
Rendering › novel view synthesis
sparse-view novel view synthesis
0.912025
Generative Sparse-View Gaussian Splatting · CVPR 2025
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.812024
DreamDrone: Text-to-Image Diffusion Models Are Zero-Shot Perpetual View Generators · ECCV (13) 2024
Machine learning › Generative modeling › diffusion model
discrete diffusion model
0.712023
Priority-Centric Human Motion Generation in Discrete Latent Space · ICCV 2023
Machine learning › Generative modeling › diffusion model › human motion generation
text-to-motion generation
0.712023
Priority-Centric Human Motion Generation in Discrete Latent Space · ICCV 2023
Computer vision › 3D vision › 3d scene reconstruction
dynamic scene reconstruction
0.312025
Generative Sparse-View Gaussian Splatting · CVPR 2025
Machine learning › Representation and self-supervised learning › representation learning › discrete representation learning
vector-quantized representation
0.212023
Priority-Centric Human Motion Generation in Discrete Latent Space · ICCV 2023

Methods — techniques the papers use, named apart from their topics

sparse anchor-grid representation · 1.7kernel representation · 1.7generative prior · 1.7gaussian splatting · 1.7diffusion model · 1.7MLP · 1.7zero-shot view generation · 1.5transformer · 0.7noise schedule · 0.7VQ-VAE · 0.7
YearPublicationVenuePosition
2025 Efficient Gaussian Splatting for Monocular Dynamic Scene Rendering via Sparse Time-Variant Attribute Modeling
abstract
Rendering dynamic scenes from monocular videos is a crucial yet challenging task. The recent deformable Gaussian Splatting has emerged as a robust solution to represent real-world dynamic scenes. However, it often leads to heavily redundant Gaussians, attempting to fit every training view at various time steps, leading to slower rendering speeds. Additionally, the attributes of Gaussians in static areas are time-invariant, making it unnecessary to model every Gaussian, which can cause jittering in static regions. In practice, the primary bottleneck in rendering speed for dynamic scenes is the number of Gaussians. In response, we introduce Efficient Dynamic Gaussian Splatting (EDGS), which represents dynamic scenes via sparse time-variant attribute modeling. Our approach formulates dynamic scenes using a sparse anchor-grid representation, with the motion flow of dense Gaussians calculated via a classical kernel representation. Furthermore, we propose an unsupervised strategy to efficiently filter out anchors corresponding to static areas. Only anchors associated with deformable objects are input into MLPs to query time-variant attributes. Experiments on two real-world datasets demonstrate that our EDGS significantly improves the rendering speed with superior rendering quality compared to previous state-of-the-art methods.
Hanyang Kong, Xingyi Yang, Xinchao Wang
AAAI1
2025 Generative Sparse-View Gaussian Splatting
abstract
Novel view synthesis from limited observations remains a significant challenge due to the lack of information in under-sampled regions, often resulting in noticeable artifacts. We introduce Generative Sparse-View Gaussian Splatting (GSGS), a general pipeline designed to enhance the rendering quality of 3D/4D Gaussian Splatting (GS) when training views are sparse. Our method generates unseen views using generative models, specifically leveraging pre-trained image diffusion models to iteratively refine view consistency and hallucinate additional images at pseudo views. This approach improves 3D/4D scene reconstruction by explicitly enforcing semantic correspondences during the generation of unseen views, thereby enhancing geometric consistency-unlike purely generative methods that often fail to maintain view consistency. Extensive evaluations on various 3D/4D datasets—including Blender, LLFF, Mip-NeRF360, and Neural 3D Video-Demonstrate that our GS-GS outperforms existing state-of-the-art methods in rendering quality without sacrificing efficiency.
Hanyang Kong, Xingyi Yang, Xinchao Wang
CVPR1
2025 Rogsplat: Robust Gaussian Splatting Via Generative Priors
Hanyang Kong, Xingyi Yang, Xinchao Wang
ICCV1
2024 DreamDrone: Text-to-Image Diffusion Models Are Zero-Shot Perpetual View Generators
Hanyang Kong, Dongze Lian, Michael Bi Mi, Xinchao Wang
ECCV (13)1
2023 Priority-Centric Human Motion Generation in Discrete Latent Space
abstract
Text-to-motion generation is a formidable task, aiming to produce human motions that align with the input text while also adhering to human capabilities and physical laws. While there have been advancements in diffusion models, their application in discrete spaces remains underexplored. Current methods often overlook the varying significance of different motions, treating them uniformly. It is essential to recognize that not all motions hold the same relevance to a particular textual description. Some motions, being more salient and informative, should be given precedence during generation. In response, we introduce a Priority-Centric Motion Discrete Diffusion Model (M2DM), which utilizes a Transformer-based VQ-VAE to derive a concise, discrete motion representation, incorporating a global self-attention mechanism and a regularization term to counteract code collapse. We also present a motion discrete diffusion model that employs an innovative noise schedule, determined by the significance of each motion token within the entire motion sequence. This approach retains the most salient motions during the reverse diffusion process, leading to more semantically rich and varied motions. Additionally, we formulate two strategies to gauge the importance of motion tokens, drawing from both textual and visual indicators. Comprehensive experiments on the HumanML3D and KIT-ML datasets confirm that our model surpasses existing techniques in fidelity and diversity, particularly for intricate textual descriptions.
Hanyang Kong, Kehong Gong, Dongze Lian, Michael Bi Mi, Xinchao Wang
ICCV1