VLDB 2026 Research / reviewers in the wild / expert
Qi Sun 0005
dblp:05/4187-5
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2025
0009-0004-5204-765XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EG4D: Explicit Generation of 4D Object without Score DistillationabstractIn recent years, the increasing demand for dynamic 3D assets in design and gaming applications has given rise to powerful generative pipelines capable of synthesizing high-quality 4D objects.
Previous methods generally rely on score distillation sampling (SDS) algorithm to infer the unseen views and motion of 4D objects, thus leading to unsatisfactory results with defects like over-saturation and Janus problem.
Therefore, inspired by recent progress of video diffusion models, we propose to optimize a 4D representation by explicitly generating multi-view videos from one input image.
However, it is far from trivial to handle practical challenges faced by such a pipeline, including dramatic temporal inconsistency, inter-frame geometry and texture diversity, and semantic defects brought by video generation results.
To address these issues, we propose EG4D, a novel multi-stage framework that generates high-quality and consistent 4D assets without score distillation.
Specifically, collaborative techniques and solutions are developed, including an attention injection strategy to synthesize temporal-consistent multi-view videos, a robust and efficient dynamic reconstruction method based on Gaussian Splatting, and a refinement stage with diffusion prior for semantic restoration.
The qualitative comparisons and quantitative results demonstrate that our framework outperforms the baselines in generation quality by a considerable margin. Qi Sun 0005, Zhiyang Guo, Ziyu Wan, Jing Nathan Yan, Shengming Yin, Wengang Zhou 0001, Jing Liao 0001, Houqiang Li |
ICLR | 1 |
| 2025 | Animus3D: Text-driven 3D Animation via Motion Score DistillationabstractWe present Animus3D , a text-driven 3D animation framework that generates motion field given a static 3D asset and text prompt. Previous methods mostly leverage the vanilla Score Distillation Sampling (SDS) objective to distill motion from pretrained text-to-video diffusion, leading to animations with minimal movement or noticeable jitter. To address this, our approach introduces a novel SDS alternative, Motion Score Distillation (MSD). Specifically, we introduce a LoRA-enhanced video diffusion model that defines a static source distribution rather than pure noise as in SDS, while another inversion-based noise estimation technique ensures appearance preservation when guiding motion. To further improve motion fidelity, we incorporate explicit temporal and spatial regularization terms that mitigate geometric distortions across time and space. Additionally, we propose a motion refinement module to upscale the temporal resolution and enhance fine-grained details, overcoming the fixed-resolution constraints of the underlying video model. Extensive experiments demonstrate that Animus3D successfully animates static 3D assets from diverse text prompts, generating significantly more substantial and detailed motion than state-of-the-art baselines while maintaining high visual integrity. Code will be released upon acceptance. Qi Sun 0005, Can Wang 0007, Jiaxiang Shang, Wensen Feng, Jing Liao 0001 |
SIGGRAPH Asia | 1 |
| 2025 | DISA: Disentangled Dual-Branch Framework for Affordance-Aware Human InsertionabstractAffordance-aware human insertion is a controllable human synthesis task aimed at seamlessly integrating a person into a scene while aligning human pose with contextual scene affordance and preserving human visual identity. Previous methods, typically reliant on a general framework of inpainting that injects all conditional information into a single branch, often struggle with the complexities of real-world contexts and the nuanced attributes of human figures. To this end, we present a novel Disentangled dual-branch framework for Affordance-aware human insertion task (DISA) , which focuses on both scene context comprehension and precise person attribute extraction. Specifically, our dual-branch design facilitates diffusion models to ensure disentangled and precise manipulations: one branch utilizes an additional network for deep scene context comprehension and control, while the other branch employs a parallel encoder to extract the feature of the reference person and injects this information through cross-attention mechanism. Furthermore, to comprehensively evaluate affordance-aware human insertion task, we introduce a new metric to assess the preservation of visual identity. We conduct a broad variety of evaluation experiments and validate the diversity and robustness of our method in different settings and downstream applications. Both qualitative and quantitative experimental analysis demonstrates that our approach outperforms previous methods in terms of image quality, pose accuracy, and visual identity preservation. Xuanqing Cao, Wengang Zhou 0001, Qi Sun 0005, Weilun Wang, Li Li 0040, Houqiang Li |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | LayoutEnc: Leveraging Enhanced Layout Representations for Transformer-based Complex Scene SynthesisabstractIn complex scene synthesis, the effective representation of layouts is paramount. This paper introduces LayoutEnc, an advanced approach specifically designed to enhance layout representation by improving interpretability, robustness, and expressiveness, thereby facilitating more efficient image transformation. Distinct from conventional approaches that homogenize layout and image data, LayoutEnc distinctively processes various data modalities, enhancing the fidelity and interpretability of the layout representation. We apply stochastic noise injection to image tokens to align training and inference conditions, thereby fortifying the robustness of the layout representation. Additionally, LayoutEnc employs a two-stage multi-scale guidance learning strategy, to meticulously extract and refine semantic and textural features from training images. This enriched layout representation is then adeptly integrated into a transformer-based image generation framework, facilitating controlled and nuanced scene synthesis. Experimental results on the COCO-stuff and Visual Genome datasets demonstrate that LayoutEnc outperforms prior works in metrics such as FID and Scene-FID scores. The code and demo are available on https://github.com/qsun1/LayoutEnc . Qi Sun 0005, Min Wang 0019, Li Li 0040, Wengang Zhou 0001, Houqiang Li |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | FOREST2SEQ: Revitalizing Order Prior for Sequential Indoor Scene Synthesis
Qi Sun 0005, Hang Zhou 0007, Wengang Zhou 0001, Li Li 0040, Houqiang Li |
ECCV (25) | 1 |
| 2024 | Recurrent Generic Contour-Based Instance Segmentation With Progressive LearningabstractContour-based instance segmentation has been actively studied, thanks to its flexibility and elegance in processing visual objects within complex backgrounds. In this work, we propose a novel deep network architecture,i.e., PolySnake, for generic contour-based instance segmentation. Motivated by the classic Snake algorithm, the proposed PolySnake achieves superior and robust segmentation performance with an iterative and progressive contour refinement strategy. Technically, PolySnake introduces a recurrent update operator to estimate the object contour iteratively. It maintains a single estimate of the contour that is progressively deformed toward the object boundary. At each iteration, PolySnake builds a semantic-rich representation for the current contour and feeds it to the recurrent operator for further contour adjustment. Through the iterative refinements, the contour progressively converges to a stable status that tightly encloses the object instance. Beyond the scope of general instance segmentation, extensive experiments are conducted to validate the effectiveness and generalizability of our PolySnake in two additional specific task scenarios, including scene text detection and lane detection. The results demonstrate that the proposed PolySnake outperforms the existing advanced methods on several multiple prevalent benchmarks across the three tasks. The codes and pre-trained models are available at https://github.com/fh2019ustc/PolySnake. Hao Feng 0009, Keyi Zhou, Wengang Zhou 0001, Yufei Yin, Jiajun Deng, Qi Sun 0005, Houqiang Li |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | S3IM: Stochastic Structural SIMilarity and Its Unreasonable Effectiveness for Neural FieldsabstractRecently, Neural Radiance Field (NeRF) has shown great success in rendering novel-view images of a given scene by learning an implicit representation with only posed RGB images. NeRF and relevant neural field methods (e.g., neural surface representation) typically optimize a point-wise loss and make point-wise predictions, where one data point corresponds to one pixel. Unfortunately, this line of research failed to use the collective supervision of distant pixels, although it is known that pixels in an image or scene can provide rich structural information. To the best of our knowledge, we are the first to design a nonlocal multiplex training paradigm for NeRF and relevant neural field methods via a novel Stochastic Structural SIMilarity (S3IM) loss that processes multiple data points as a whole set instead of process multiple inputs independently. Our extensive experiments demonstrate the unreasonable effectiveness of S3IM in improving NeRF and neural surface representation for nearly free. The improvements of quality metrics can be particularly significant for those relatively difficult tasks: e.g., the test MSE loss unexpectedly drops by 90% for TensoRF and DVGO over eight novel view synthesis tasks; a 198% F-score gain and a 64% Chamfer L1distance reduction for NeuS over eight surface reconstruction tasks. Moreover, S3IM is consistently robust even with sparse inputs, corrupted images, and dynamic scenes. Zeke Xie, Xindi Yang, Qi Sun 0005, Yixiang Jiang, Haoran Wang 0004, Yunfeng Cai, Mingming Sun 0001 |
ICCV | 4 |