VLDB 2026 Research / reviewers in the wild / expert
Shoulong Zhang
dblp:294/0179
· DBLP profile ↗
10ranked-venue papers
4as first author
10since 2021 · last 2025
0000-0003-1626-4054ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Cafe-Talk: Generating 3D Talking Face Animation with Multimodal Coarse- and Fine-grained ControlabstractSpeech-driven 3D talking face method should offer both accurate lip synchronization and controllable expressions. Previous methods solely adopt discrete emotion labels to globally control expressions throughout sequences while limiting flexible fine-grained facial control within the spatiotemporal domain. We propose a diffusion-transformer-based 3D talking face generation model, Cafe-Talk, which simultaneously incorporates coarse- and fine-grained multimodal control conditions. Nevertheless, the entanglement of multiple conditions challenges achieving satisfying performance. To disentangle speech audio and fine-grained conditions, we employ a two-stage training pipeline. Specifically, Cafe-Talk is initially trained using only speech audio and coarse-grained conditions. Then, a proposed fine-grained control adapter gradually adds fine-grained instructions represented by action units (AUs), preventing unfavorable speech-lip synchronization. To disentangle coarse- and fine-grained conditions, we design a swap-label training mechanism, which enables the dominance of the fine-grained conditions. We also devise a mask-based CFG technique to regulate the occurrence and intensity of fine-grained control. In addition, a text-based detector is introduced with text-AU alignment to enable natural language user input and further support multimodal control. Extensive experimental results prove that Cafe-Talk achieves state-of-the-art lip synchronization and expressiveness performance and receives wide acceptance in fine-grained control in user studies. Hejia Chen, Haoxian Zhang, Shoulong Zhang, Sisi Zhuang, Yuan Zhang 0020, Pengfei Wan 0001, Di Zhang 0026, Shuai Li 0001 |
ICLR | 3 |
| 2025 | The Effect of Unexpected Visual Stimuli on Short-Term Memory in Immersive Experience
Shoulong Zhang, Yutian Xiao, Xuejing Lu, Shuai Li 0001 |
ICXR | 1 |
| 2025 | Phys4DRT: Physics-based 4D Generation for Real-Time Interaction with Time-Frequency Supervision
Yuntian Xiao, Shoulong Zhang, Jiahao Cui 0001, Shuai Li 0001 |
ACM Multimedia | 2 |
| 2025 | Reactffusion: Physical Contact-guided Diffusion Model for Reaction Generation
Shoulong Zhang, Shuai Li 0001 |
ACM Multimedia | 2 |
| 2024 | SIE-DepthNet: Semantic-Guided Monocular Depth Estimation for Dynamic Environment
Zilong Song, Yang Gao 0032, Sijia Dai, Shuai Li 0001, Aimin Hao, Shoulong Zhang |
ICXR | 6 |
| 2024 | Conditional room layout generation based on graph neural networks
Zhihan Yao, Jiahao Cui 0001, Shoulong Zhang, Shuai Li 0001, Aimin Hao |
Comput. Graph. | 4 |
| 2023 | Propose-and-Complete: Auto-regressive Semantic Group Generation for Personalized Scene Synthesis
Shoulong Zhang, Shuai Li 0001, Xinwei Huang, Wenchong Xu, Aimin Hao, Hong Qin 0001 |
BMVC | 1 |
| 2022 | Distribution-motivated 3D Style Characterization Based on Latent Feature Decomposition
Xinwei Huang, Shuai Li 0001, Shoulong Zhang, Aimin Hao, Hong Qin 0001 |
Comput. Aided Des. | 3 |
| 2021 | Point Cloud Semantic Scene Completion from RGB-D ImagesabstractIn this paper, we devise a novel semantic completion network, called point cloud semantic scene completion network (PCSSC-Net), for indoor scenes solely based on point clouds. Existing point cloud completion networks still suffer from their inability of fully recovering complex structures and contents from global geometric descriptions neglecting semantic hints. To extract and infer comprehensive information from partial input, we design a patch-based contextual encoder to hierarchically learn point-level, patch-level, and scene-level geometric and contextual semantic information with a divide-and-conquer strategy. Consider that the scene semantics afford a high-level clue of constituting geometry for an indoor scene environment, we articulate a semantics-guided completion decoder where semantics could help cluster isolated points in the latent space and infer complicated scene geometry. Given the fact that real-world scans tend to be incomplete as ground truth, we choose to synthesize scene dataset with RGB-D images and annotate complete point clouds as ground truth for the supervised training purpose. Extensive experiments validate that our new method achieves the state-of-the-art performance, in contrast with the current methods applied to our dataset. Shoulong Zhang, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
AAAI | 1 |
| 2021 | Knowledge-inspired 3D Scene Graph Prediction in Point CloudabstractPrior knowledge integration helps identify semantic entities and their relationships in a graphical representation, however, its meaningful abstraction and intervention remain elusive. This paper advocates a knowledge-inspired 3D scene graph prediction method solely based on point clouds. At the mathematical modeling level, we formulate the task as two sub-problems: knowledge learning and scene graph prediction with learned prior knowledge. Unlike conventional methods that learn knowledge embedding and regular patterns from encoded visual information, we propose to suppress the misunderstandings caused by appearance similarities and other perceptual confusion. At the network design level, we devise a graph auto-encoder to automatically extract class-dependent representations and topological patterns from the one-hot class labels and their intrinsic graphical structures, so that the prior knowledge can avoid perceptual errors and noises. We further devise a scene graph prediction model to predict credible relationship triplets by incorporating the related prototype knowledge with perceptual information. Comprehensive experiments confirm that, our method can successfully learn representative knowledge embedding, and the obtained prior knowledge can effectively enhance the accuracy of relationship predictions. Our thorough evaluations indicate the new method can achieve the state-of-the-art performance compared with other scene graph prediction methods. Shoulong Zhang, Shuai Li 0001, Aimin Hao, Hong Qin 0001 |
NeurIPS | 1 |