VLDB 2026 Research / reviewers in the wild / expert
Caoyuan Ma
dblp:223/7396
· DBLP profile ↗
8ranked-venue papers
2as first author
8since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Subjective Camera 1.0: Bridging Human Cognition and Visual Reconstruction Through Sequence-Aware Sketch-Guided DiffusionabstractWe introduce the concept of a subjective camera to reconstruct meaningful moments that physical cameras fail to capture. We propose Subjective Camera 1.0, a framework for reconstructing real-world scenes from readily accessible subjective readouts, i.e., textual descriptions and progressively drawn rough sketches. Built on optimization-based alignment of diffusion models, our approach avoids large-scale paired training data and mitigates generalization issues. To address the challenge of integrating multiple abstract concepts in real-world scenarios, we design a Sequence-Aware Sketch-Guided Diffusion framework with three loss terms for concept-wise sequential optimization, following the natural order of subjective readouts. Experiments on two datasets demonstrate that our method achieves state-of-the-art performance in image quality as well as spatial and semantic alignment with target scenes. User studies with 40 participants further confirm that our approach is consistently preferred. Our project page is at: subjective-camera.github.io Dongfang Sun, Caoyuan Ma, Shiqin Wang, Zheng Wang 0007, Zhixiang Wang 0001 |
ICCV | 3 |
| 2025 | You Think, You ACT: the New Task of Arbitrary Text to Motion Generation
Runqi Wang, Caoyuan Ma, Hanrui Xu, Zheng Wang 0007 |
ICCV | 2 |
| 2025 | Leader is Guided: Interactive Motion Generation via Lead-Follow Paradigm and Trajectory GuidanceabstractGenerating interactive motion from texts has garnered significant attention in recent years. While text inputs offer greater flexibility, in many practical applications, there is a need to controllably impose strict constraints on the motion range or trajectory of virtual characters. However, existing trajectory-based methods are designed for single-actor scenarios and lack support for interactivity in interactive motions. Moreover, text-only methods struggle to accurately convey user-intended trajectories. The distribution shift between training and inference often leads to trajectory deviation and physical interpenetration. To address the questions mentioned, we introduce two key concepts: (1) Lead-Follow Paradigm: Inspired by role allocation in partner dancing, we decompose complex interactive motion tasks into a Lead-Follow paradigm. The leader's path is optimized first, and the follower's motion is subsequently adjusted for coherence and alignment. (2) Trajectory Guidance: We highlight the pivotal role of 3D trajectory guidance in interactive motion generation and accurately reflect user intentions. Through 3D trajectory control, we can more controllably generate the desired motion while avoiding physical interpenetration. In addition, we further investigate the refinement of motion scopes for interactive agents and propose an effective optimization strategy to enhance motion coherence and controllability. Experimental results show that the proposed approach, by more effectively using trajectory, outperforms existing methods in both realism and accuracy. Runqi Wang, Caoyuan Ma, Jian Zhao 0013, Hanrui Xu, Dongfang Sun, Zheng Wang 0007, Xuelong Li 0001 |
ACM Multimedia | 2 |
| 2025 | Behave Your Motion: Habit-preserved Cross-category Animal Motion TransferabstractAnimal motion embodies species-specific behavioral habits, making the transfer of motion across categories a critical yet complex task for applications in animation and virtual reality. Existing motion transfer methods, primarily focused on human motion, emphasize skeletal alignment (motion retargeting) or stylistic consistency (motion style transfer), often neglecting the preservation of distinct habitual behaviors in animals. To bridge this gap, we propose a novel habit-preserved motion transfer framework for cross-category animal motion. Built upon a generative framework, our model introduces a habit-preservation module with category-specific habit encoder, allowing it to learn motion priors that capture distinctive habitual characteristics. Furthermore, we integrate a large language model (LLM) to facilitate the motion transfer to previously unobserved species. To evaluate the effectiveness of our approach, we introduce the DeformingThings4D-skl dataset, a quadruped dataset with skeletal bindings, and conduct extensive experiments and quantitative analyses, which validate the superiority of our proposed model. Zhimin Zhang 0008, Bi'an Du, Caoyuan Ma, Zheng Wang 0007, Wei Hu 0003 |
ACM Multimedia | 3 |
| 2024 | Learning a Dynamic Neural Human via Poses Guided Dual Spaces FeatureabstractLearning human representations from video is becoming increasingly important in various applications. However, due to the limited information in videos and the complexity of human deformation, existing methods cannot faithfully reconstruct the image representation of humans, including clothing folds and light and shadow. Our method is built upon a deformation-based approach, which uses pose-guided joint learning to derive human representations in both canonical space and observation space, thereby enhancing the model’s performance in human details. We conducted several experiments on publicly available datasets using our approach, achieving highly realistic reconstruction results that are difficult to distinguish from real frames. Our approach also showed improved overall evaluation metrics for video frames that were not visible in the original view angle. Caoyuan Ma, Runqi Wang, Wu Liu 0005, Ziqiao Zhou, Zheng Wang 0007 |
AVSS | 1 |
| 2024 | HumanNeRF-SE: A Simple yet Effective Approach to Animate HumanNeRF with Diverse PosesabstractWe present HumanNeRF-SE, a simple yet effective method that synthesizes diverse novel pose images with sim-ple input. Previous HumanNeRF works require a large number of optimizable parameters to fit the human images. Instead, we reload these approaches by combining explicit and implicit human representations to design both general-ized rigid deformation and specific non-rigid deformation. Our key insight is that explicit shape can reduce the sam-pling points used to fit implicit representation, and frozen blending weights from SMPL constructing a generalized rigid deformation can effectively avoid overfitting and im-prove pose generalization performance. Our architecture involving both explicit and implicit representation is sim-ple yet effective. Experiments demonstrate our model can synthesize images under arbitrary poses with few-shot input and increase the speed of synthesizing images by 15 times through a reduction in computational complexity without using any existing acceleration modules. Compared to the state-of-the-art HumanNeRF studies, HumanNeRF-SE achieves better performance with fewer learnable parame-ters and less training time. Caoyuan Ma, Yu-Lun Liu 0001, Zhixiang Wang 0001, Wu Liu 0005, Xinchen Liu, Zheng Wang 0007 |
CVPR | 1 |
| 2024 | Causal Deciphering and Inpainting in Spatio-Temporal Dynamics via Diffusion ModelabstractSpatio-temporal (ST) prediction has garnered a De facto attention in earth sciences, such as meteorological prediction, human mobility perception. However, the scarcity of data coupled with the high expenses involved in sensor deployment results in notable data imbalances. Furthermore, models that are excessively customized and devoid of causal connections further undermine the generalizability and interpretability. To this end, we establish a causal framework for ST predictions, termed CaPaint, which targets to identify causal regions in data and endow model with causal reasoning ability in a two-stage process. Going beyond this process, we utilize the back-door adjustment to specifically address the sub-regions identified as non-causal in the upstream phase. Specifically, we employ a novel image inpainting technique. By using a fine-tuned unconditional Diffusion Probabilistic Model (DDPM) as the generative prior, we in-fill the masks defined as environmental parts, offering the possibility of reliable extrapolation for potential data distributions. CaPaint overcomes the high complexity dilemma of optimal ST causal discovery models by reducing the data generation complexity from exponential to quasi-linear levels. Extensive experiments conducted on five real-world ST benchmarks demonstrate that integrating the CaPaint concept allows models to achieve improvements ranging from 4.3% to 77.3%. Moreover, compared to traditional mainstream ST augmenters, CaPaint underscores the potential of diffusion models in ST enhancement, offering a novel paradigm for this field. Our project is available at https://anonymous.4open.science/r/12345-DFCC. Yifan Duan, Jian Zhao 0006, pengcheng, Junyuan Mao, Hao Wu 0098, Jingyu Xu 0002, Shilong Wang 0002, Caoyuan Ma, Kai Wang 0036, Kun Wang 0056, Xuelong Li 0001 |
NeurIPS | 8 |
| 2022 | Improving Stack Overflow question title generation with copying enhanced CodeBERT model and bi-modal information
Fengji Zhang, Xiao Yu 0008, Jacky W. Keung, Zhiwen Xie, Zhen Yang 0022, Caoyuan Ma, Zhimin Zhang 0008 |
Inf. Softw. Technol. | 7 |