VLDB 2026 Research / reviewers in the wild / expert
Jiajun Su
dblp:210/2464
· DBLP profile ↗
14ranked-venue papers
2as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Text-to-video person re-identification benchmark: Dataset and dual-modal contextual alignment
Jiajun Su, Simin Zhan, Pudu Liu, Jianqing Zhu, Huanqiang Zeng |
Neurocomputing | 1 |
| 2026 | UDD: Unsupervised denoising diffusion for noisy multi-focus image fusion
Pudu Liu, Jiajun Su, Simin Zhan, Jianqing Zhu, Huanqiang Zeng |
Neural Networks | 3 |
| 2025 | SAT-HMR: Real-Time Multi-Person 3D Mesh Estimation via Scale-Adaptive TokensabstractWe propose SAT-HMR, a one-stage framework for real-time multi-person 3D human mesh estimation from a single RGB image. While current one-stage methods, which follow a DETR-style pipeline, achieve state-of-the-art (SOTA) performance with high-resolution inputs, we observe that this particularly benefits the estimation of individuals in smaller scales of the image (e.g., those of young age or far from the camera), but at the cost of significantly increased computation overhead. To address this, we introduce scale-adaptive tokens that are dynamically adjusted based on the relative scale of each individual in the image within the DETR framework. Specifically, individuals in smaller scales are processed at higher resolutions, larger ones at lower resolutions, and background regions are further distilled. These scale-adaptive tokens more efficiently encode the image features, facilitating subsequent decoding to regress the human mesh, while allowing the model to allocate computational resources more effectively and focus on more challenging cases. Experiments show that our method preserves the accuracy benefits of high-resolution processing while substantially reducing computational cost, achieving real-time inference with performance comparable to SOTA methods. Chi Su, Xiaoxuan Ma 0001, Jiajun Su, Yizhou Wang 0001 |
CVPR | 3 |
| 2025 | Learning Whole-Body Control for Small-Sized Quadruped Robots with a Flexible SpineabstractImproving the adaptability of small-sized quadruped robots has been a longstanding challenge in robotics. However, the weak whole-body coordination in existing small-sized quadruped robots limits their locomotion in many environments. In this work, we propose a teacher-student online learning framework for agile whole-body control of small-sized quadruped robots with a flexible spine. We first select a simple and effective gait pattern, the diagonal symmetrical sequence, using a dynamics model. Based on the reference motions provided by the gait pattern and combined with privileged information, we train a teacher policy to generate high-quality motion data. After setting the state space to match the actual robot’s state space, we initialize the robot’s initial state using the teacher data and train a student policy. Finally, we deploy the student policy on the SQuRo-Lite, a small-sized quadruped robot with a flexible spine, demonstrating that our approach can achieve stable yet dynamic locomotion for walking and turning. In the variable-spacing slalom experiment, the robot is able to flexibly adjust the motion patterns of its spine and legs based on commands, enabling dynamic changes in its turning radius. This further validates that our approach can achieve agile whole-body control for small-sized quadruped robots. This work helps broaden the application scenarios of small quadruped robots. Dixuan Jiang, Guanglu Jia, Changwen Dong, Jiajun Su |
IROS | 4 |
| 2025 | VMarker-Pro: Probabilistic 3D Human Mesh Estimation From Virtual MarkersabstractMonocular 3D human mesh estimation faces challenges due to depth ambiguity and the complexity of mapping images to complex parameter spaces. Recent methods propose to use 3D poses as a proxy representation, which often lose crucial body shape information, leading to mediocre performance. Conversely, advanced motion capture systems, though accurate, are impractical for markerless wild images. Addressing these limitations, we introduce an innovative intermediate representation as virtual markers, which are learned from large-scale mocap data, mimicking the effects of physical markers. Building upon virtual markers, we propose VMarker, which detects virtual markers from wild images, and the intact mesh with realistic shapes can be obtained by simply interpolation from these markers. To address occlusions that obscure 3D virtual marker estimation, we further enhance our method with VMarker-Pro, a probabilistic framework that models the distribution of 3D virtual marker positions using diffusion models, enabling the generation of multiple plausible meshes aligned with images for robust 3D mesh estimation. Our approaches surpass existing methods on three benchmark datasets, particularly demonstrating significant improvements on the SURREAL dataset, which features diverse body shapes. Additionally, VMarker-Pro excels in accurately modeling data distributions, significantly enhancing performance in occluded scenarios. Xiaoxuan Ma 0001, Jiajun Su, Yuan Xu 0022, Wentao Zhu 0004, Chunyu Wang 0001, Yizhou Wang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | ScoreHypo: Probabilistic Human Mesh Estimation with Hypothesis ScoringabstractMonocular 3D human mesh estimation is an ill-posed problem, characterized by inherent ambiguity and occlusion. While recent probabilistic methods propose generating multiple solutions, little attention is paid to obtaining high-quality estimates from them. To address this limitation, we introduce ScoreHypo, a versatile framework by first leveraging our novel HypoNet to generate multiple hy-potheses, followed by employing a meticulously designed scorer, ScoreNet, to evaluate and select high-quality esti-mates. ScoreHypo formulates the estimation process as a re-verse denoising process, where HypoNet produces a diverse set of plausible estimates that effectively align with the im-age cues. Subsequently, ScoreNet is employed to rigorously evaluate and rank these estimates based on their quality and finally identify superior ones. Experimental results demon-strate that HypoNet outperforms existing state-of-the-art probabilistic methods as a multi-hypothesis mesh estimator. Moreover, the estimates selected by ScoreNet significantly outperform random generation or simple averaging. Notably, the trained ScoreNet exhibits generalizability, as it can effectively score existing methods and significantly reduce their errors by more than 15%. Code and models are available at ht tps: / /xy02- 05. gi thub. io/ScoreHypo. Yuan Xu 0022, Xiaoxuan Ma 0001, Jiajun Su, Wentao Zhu 0004, Yu Qiao 0001, Yizhou Wang 0001 |
CVPR | 3 |
| 2024 | A Hybrid Method to Interest-informed and Mobility-aware Mobile Service Migration in Edge ComputingabstractMobile edge computing(MEC) is an innovative technology that deploys computing resources around the demand side to provide near-request and responsiveness-guaranteed computing and storage services. A major attention paid by related works in this direction is mobility, where mobile traces of both edge users and servers are analyzed and exploited for accommodating offloading and migration requests for computation resources in a highly dynamic MEC environment. Our research in this work suggests that information of user interests, in terms of points of interest (POI), can be exploited in conjunction with mobility as well and proposes a hybrid method for for interest-informed and mobility-aware service migration path selection(HIMS). It synthesizes a trajectory prediction model and user interests prediction one for selecting target servers and reliable service migration paths. Experimental results demonstrate that our approach outperforms traditional methods across multiple performance metrics, especially those with sole input of mobility. Mengxuan Dai, Yunni Xia, Xu Wang 0024, Xingli Zhong, Hui Liu 0003, Qinglan Peng, Xiaoning Sun, Jiajun Su |
ICWS | 10 |
| 2024 | Modality-Consistent Attention for Visible-Infrared Vehicle Re-IdentificationabstractVisible-infrared vehicle re-identification (VIVR) seeks to match vehicle images of the same identity taken by cameras of different modalities. The noticeable disparity between visible and infrared modalities leads to attention deviations, causing deep models to incorrectly focus on different local regions of vehicles in visible and infrared images. We observed that the spatial distributions of distinguishing local regions, such as logos, front windows, and wheels, exhibit similarity in average images obtained from both visible and infrared images. Based on this, we propose a modality-consistent attention (MCA) approach for VIVR. Unlike image-level attention, our MCA is identity-level attention that holistically emphasizes the distinguishing regions of a vehicle identity across multiple images captured from various viewpoints. Furthermore, we constrain the differences between the identity-level spatial attention masks resulting from visible and infrared modalities. This approach helps deep networks focus consistently on learning the distinguishing local characteristics of vehicles across different modalities and viewpoints. Our experiments on RGBN300 and MSVR310 datasets demonstrate that our approach achieves state-of-the-art performance. Jiajun Su, Jianqing Zhu, Liu Liu 0014, Huanqiang Zeng |
IEEE Signal Process. Lett. | 2 |
| 2023 | A Multi-Agent Deep Reinforcement Learning-Based Approach to Mobility-Aware Caching
Shiyun Shao, Yong Ma 0005, Yunni Xia, Jiajun Su, Lingmeng Liu, Kaiwei Chen, Qinglan Peng |
CollaborateCom (2) | 5 |
| 2023 | 3D Human Mesh Estimation from Virtual MarkersabstractInspired by the success of volumetric 3D pose estimation, some recent human mesh estimators propose to estimate 3D skeletons as intermediate representations, from which, the dense 3D meshes are regressed by exploiting the mesh topology. However, body shape information is lost in extracting skeletons, leading to mediocre performance. The advanced motion capture systems solve the problem by placing dense physical markers on the body surface, which allows to extract realistic meshes from their non-rigid motions. However, they cannot be applied to wild images without markers. In this work, we present an intermediate representation, named virtual markers, which learns 64 landmark keypoints on the body surface based on the large-scale mocap data in a generative style, mimicking the effects of physical markers. The virtual markers can be accurately detected from wild images and can reconstruct the intact meshes with realistic shapes by simple interpolation. Our approach outperforms the state-of-the-art methods on three datasets. In particular, it surpasses the existing methods by a notable margin on the SURREAL dataset, which has diverse body shapes. Code is available at https://github.com/ShirleyMaxx/VirtualMarker Xiaoxuan Ma 0001, Jiajun Su, Chunyu Wang 0001, Wentao Zhu 0004, Yizhou Wang 0001 |
CVPR | 2 |
| 2023 | ChimpACT: A Longitudinal Dataset for Understanding Chimpanzee BehaviorsabstractUnderstanding the behavior of non-human primates is crucial for improving animal welfare, modeling social behavior, and gaining insights into distinctively human and phylogenetically shared behaviors. However, the lack of datasets on non-human primate behavior hinders in-depth exploration of primate social interactions, posing challenges to research on our closest living relatives. To address these limitations, we present ChimpACT, a comprehensive dataset for quantifying the longitudinal behavior and social relations of chimpanzees within a social group. Spanning from 2015 to 2018, ChimpACT features videos of a group of over 20 chimpanzees residing at the Leipzig Zoo, Germany, with a particular focus on documenting the developmental trajectory of one young male, Azibo. ChimpACT is both comprehensive and challenging, consisting of 163 videos with a cumulative 160,500 frames, each richly annotated with detection, identification, pose estimation, and fine-grained spatiotemporal behavior labels. We benchmark representative methods of three tracks on ChimpACT: (i) tracking and identification, (ii) pose estimation, and (iii) spatiotemporal action detection of the chimpanzees. Our experiments reveal that ChimpACT offers ample opportunities for both devising new methods and adapting existing ones to solve fundamental computer vision tasks applied to chimpanzee groups, such as detection, pose estimation, and behavior analysis, ultimately deepening our comprehension of communication and sociality in non-human primates. Xiaoxuan Ma 0001, Stephan P. Kaufhold, Jiajun Su, Wentao Zhu 0004, Jack Terwilliger, Andres Meza 0001, Yixin Zhu 0001, Federico Rossano, Yizhou Wang 0001 |
NeurIPS | 3 |
| 2022 | VirtualPose: Learning Generalizable 3D Human Pose Models from Virtual Data
Jiajun Su, Chunyu Wang 0001, Xiaoxuan Ma 0001, Wenjun Zeng 0001, Yizhou Wang 0001 |
ECCV (6) | 1 |
| 2021 | Context Modeling in 3D Human Pose Estimation: A Unified PerspectiveabstractEstimating 3D human pose from a single image suffers from severe ambiguity since multiple 3D joint configurations may have the same 2D projection. The state-of-the-art methods often rely on context modeling methods such as pictorial structure model (PSM) or graph neural network (GNN) to reduce ambiguity. However, there is no study that rigorously compares them side by side. So we first present a general formula for context modeling in which both PSM and GNN are its special cases. By comparing the two methods, we found that the end-to-end training scheme in GNN and the limb length constraints in PSM are two complementary factors to improve results. To combine their advantages, we propose ContextPose based on attention mechanism that allows enforcing soft limb length constraints in a deep network. The approach effectively reduces the chance of getting absurd 3D pose estimates with incorrect limb lengths and achieves state-of-the-art results on two benchmark datasets. More importantly, the introduction of limb length constraints into deep networks enables the approach to achieve much better generalization performance. Xiaoxuan Ma 0001, Jiajun Su, Chunyu Wang 0001, Hai Ci, Yizhou Wang 0001 |
CVPR | 2 |
| 2018 | Attentive Generative Adversarial Network for Raindrop Removal From a Single ImageabstractRaindrops adhered to a glass window or camera lens can severely hamper the visibility of a background scene and degrade an image considerably. In this paper, we address the problem by visually removing raindrops, and thus transforming a raindrop degraded image into a clean one. The problem is intractable, since first the regions occluded by raindrops are not given. Second, the information about the background scene of the occluded regions is completely lost for most part. To resolve the problem, we apply an attentive generative network using adversarial training. Our main idea is to inject visual attention into both the generative and discriminative networks. During the training, our visual attention learns about raindrop regions and their surroundings. Hence, by injecting this information, the generative network will pay more attention to the raindrop regions and the surrounding structures, and the discriminative network will be able to assess the local consistency of the restored regions. This injection of visual attention to both generative and discriminative networks is the main contribution of this paper. Our experiments show the effectiveness of our approach, which outperforms the state of the art methods quantitatively and qualitatively. Rui Qian 0003, Robby T. Tan, Wenhan Yang, Jiajun Su, Jiaying Liu 0001 |
CVPR | 4 |