VLDB 2026 Research / reviewers in the wild / expert
Zhongyu Jiang
dblp:172/4606
· DBLP profile ↗
27ranked-venue papers
4as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 4 first-author · 21 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SAMURAI: Motion-Aware Memory for Training-Free Visual Object Tracking With SAM 2abstractThe Segment Anything Model 2 (SAM 2) has demonstrated exceptional performance in object segmentation tasks but encounters challenges in visual object tracking, particularly in handling crowded scenes with fast-moving or self-occluding objects. Additionally, its fixed-window memory mechanism indiscriminately retains past frames, leading to error accumulation. This issue results in incorrect memory retention during occlusions, causing the model to condition future predictions on unreliable features and leading to identity switches or drift in crowded scenes. This paper introduces SAMURAI, an enhanced adaptation of SAM 2 that integrates temporal motion cues with a novel motion-aware memory selection strategy. SAMURAI effectively predicts object motion and refines mask selection, achieving robust and precise tracking without requiring retraining or fine-tuning. It demonstrates strong training-free performance across multiple VOT benchmark datasets, underscoring its generalization capability. SAMURAI achieves state-of-the-art performance on LaSOText, GOT-10k, and TrackingNet, while also delivering competitive results on LaSOT, VOT2020-ST, VOT2022-ST, and VOS benchmarks such as SA-V. These results highlight SAMURAI's robustness in complex tracking scenarios and its potential for real-world applications in dynamic environments with an optimized memory selection mechanism. Code and results are available at https://github.com/yangchris11/samurai. Cheng-Yeng Yang, Hsiang-Wei Huang, Wenhao Chai, Zhongyu Jiang, Jenq-Neng Hwang |
IEEE Trans. Image Process. | 4 |
| 2025 | The Role of Deductive and Inductive Reasoning in Large Language ModelsabstractChengkun Cai, Xu Zhao, Haoliang Liu, Zhongyu Jiang, Tianfang Zhang, Zongkai Wu, Jenq-Neng Hwang, Lei Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Chengkun Cai, Haoliang Liu, Zhongyu Jiang, Tianfang Zhang, Zongkai Wu, Jenq-Neng Hwang, Lei Li 0050 |
ACL (1) | 4 |
| 2025 | Human Motion Instruction TuningabstractThis paper presents LLaMo (Large Language and Human Motion Assistant), a multimodal framework for human motion instruction tuning. In contrast to conventional instruction-tuning approaches that convert non-linguistic inputs, such as video or motion sequences, into language tokens, LLaMo retains motion in its native form for instruction tuning. This method preserves motion-specific details that are often diminished in tokenization, thereby improving the model’s ability to interpret complex human behaviors. By processing both video and motion data alongside textual inputs, LLaMo enables a flexible, human-centric analysis. Experimental evaluations across high-complexity domains, including human behaviors and professional activities, indicate that LLaMo effectively captures domain-specific knowledge, enhancing comprehension and prediction in motion-intensive scenarios. We hope LLaMo offers a foundation for future multimodal AI systems with broad applications, from sports analytics to behavioral prediction. Lei Li 0050, Sen Jia 0003, Zhongyu Jiang, Feng Zhou 0007, Ju Dai, Tianfang Zhang, Zongkai Wu, Jenq-Neng Hwang |
CVPR | 4 |
| 2025 | MambaMOT: State-Space Model as Motion Predictor for Multi-Object TrackingabstractIn the field of multi-object tracking (MOT), traditional methods often rely on the Kalman filter for motion prediction, leveraging its strengths in linear motion scenarios. However, the inherent limitations of these methods become evident when confronted with complex, nonlinear motions and occlusions prevalent in dynamic environments like sports and dance. This paper explores the possibilities of replacing the Kalman filter with a learning-based motion model that effectively enhances tracking accuracy and adaptability beyond the constraints of Kalman filter-based tracker. In this paper, our proposed method MambaMOT and MambaMOT+, demonstrate advanced performance on challenging MOT datasets such as DanceTrack and SportsMOT, showcasing their ability to handle intricate, nonlinear motion patterns and frequent occlusions more effectively than traditional methods. Hsiang-Wei Huang, Cheng-Yen Yang, Wenhao Chai, Zhongyu Jiang, Jenq-Neng Hwang |
ICASSP | 4 |
| 2024 | RT-Pose: A 4D Radar Tensor-Based 3D Human Pose Estimation and Localization Benchmark
Yuan-Hao Ho, Jen-Hao Cheng, Sheng-Yao Kuan, Zhongyu Jiang, Wenhao Chai, Hsiang-Wei Huang, Chih-Lung Lin, Jenq-Neng Hwang |
ECCV (63) | 4 |
| 2024 | 2D Human Pose Estimation Calibration and Keypoint Visibility ClassificationabstractThe confidence scores of 2D pose estimation are widely utilized in various fields, including multi-view 3D human pose estimation, skeleton-based human tracking, human action recognition, human re-identification, etc. Despite widespread use, confidence scores from 2D pose estimation methods are unreliable in indicating the accuracy of estimation results, particularly in occlusion situations, i.e., keypoints with high confidence scores may have low accuracy and vice versa. To address this issue, we propose a new 2D human pose estimation calibration method in this paper. Our method not only enhances the accuracy of 2D pose estimation but also aligns the confidence scores with the quality and visibility of keypoints. We achieve 77.6 mAP in the COCO val dataset, compared with 76.5 mAP of the original HRNet. For key-point visibility prediction, we can reach 89.4% accuracy, 87.6% precision, and 97.0% recall in the COCO val dataset. Zhongyu Jiang, Haorui Ji, Cheng-Yen Yang, Jenq-Neng Hwang |
ICASSP | 1 |
| 2024 | A Density-Guided Temporal Attention Transformer for Indiscernible Object Counting in Underwater VideosabstractDense object counting or crowd counting has come a long way thanks to the recent development in the vision community. However, indiscernible object counting, which aims to count the number of targets that are blended with respect to their surroundings, has been a challenge. Image-based object counting datasets have been the mainstream of the current publicly available datasets. Therefore, we propose a large-scale dataset called YoutubeFish-35, which contains a total of 35 sequences of high-definition videos with high frame-per-second and more than 159,000 annotated center points across a selected variety of scenes. For bench-marking purposes, we select three mainstream methods for dense object counting and carefully evaluate them on the newly collected dataset. We propose TransVidCount, a new strong baseline that combines density and regression branches along the temporal domain in a unified framework and can effectively tackle indiscernible object counting with state-of-the-art performance on YoutubeFish-35 dataset. Cheng-Yen Yang, Hsiang-Wei Huang, Zhongyu Jiang, Farron Wallace, Jenq-Neng Hwang |
ICASSP | 3 |
| 2024 | MPM: A Unified 2D-3D Human Pose Representation via Masked Pose Modeling
Zhenyu Zhang 0030, Wenhao Chai, Zhongyu Jiang, Tian Ye 0001, Mingli Song, Jenq-Neng Hwang, Gaoang Wang |
PRCV (11) | 3 |
| 2024 | Back to Optimization: Diffusion-based Zero-Shot 3D Human Pose EstimationabstractLearning-based methods have dominated the 3D human pose estimation (HPE) tasks with significantly better performance in most benchmarks than traditional optimization-based methods. Nonetheless, 3D HPE in the wild is still the biggest challenge for learning-based models, whether with 2D-3D lifting, image-to-3D, or diffusion-based methods, since the trained networks implicitly learn camera intrinsic parameters and domain-based 3D human pose distributions and estimate poses by statistical average. On the other hand, the optimization-based methods estimate results case-by-case, which can predict more diverse and sophisticated human poses in the wild. By combining the advantages of optimization-based and learning-based methods, we propose the Zero-shot Diffusion-based Optimization (ZeDO) pipeline for 3D HPE to solve the problem of cross-domain and in-the-wild 3D HPE. Our multi-hypothesis ZeDO achieves state-of-the-art (SOTA) performance on Human3.6M, with minMPJPE 51.4mm, without training with any 2D-3D or image-3D pairs. Moreover, our single-hypothesis ZeDO achieves SOTA performance on 3DPW dataset with PA-MPJPE 40.3mm on cross-dataset evaluation, which even outperforms learning-based methods trained on 3DPW. Our code is available here: https://github.com/ipl-uw/ZeDO-Release. Zhongyu Jiang, Zhuoran Zhou, Lei Li 0050, Wenhao Chai, Cheng-Yen Yang, Jenq-Neng Hwang |
WACV | 1 |
| 2023 | Global Adaptation meets Local Generalization: Unsupervised Domain Adaptation for 3D Human Pose EstimationabstractWhen applying a pre-trained 2D-to-3D human pose lifting model to a target unseen dataset, large performance degradation is commonly encountered due to domain shift issues. We observe that the degradation is caused by two factors: 1) the large distribution gap over global positions of poses between the source and target datasets due to variant camera parameters and settings, and 2) the deficient diversity of local structures of poses in training. To this end, we combine global adaptation and local generalization in PoseDA, a simple yet effective framework of unsupervised domain adaptation for 3D human pose estimation. Specifically, global adaptation aims to align global positions of poses from the source domain to the target domain with a proposed global position alignment (GPA) module. And local generalization is designed to enhance the diversity of 2D-3D pose mapping with a local pose augmentation (LPA) module. These modules bring significant performance improvement without introducing additional learnable parameters. In addition, we propose local pose augmentation (LPA) to enhance the diversity of 3D poses following an adversarial training scheme consisting of 1) a augmentation generator that generates the parameters of pre-defined pose transformations and 2) an anchor discriminator to ensure the reality and quality of the augmented data. Our approach can be applicable to almost all 2D-3D lifting models. PoseDA achieves 61.3 mm of MPJPE on MPI-INF-3DHP under a cross-dataset evaluation setup, improving upon the previous state-of-the-art method by 10.2%. Wenhao Chai, Zhongyu Jiang, Jenq-Neng Hwang, Gaoang Wang |
ICCV | 2 |
| 2023 | Multi-Object Tracking by Iteratively Associating Detections with Uniform Appearance for Trawl-Based Fishing Bycatch MonitoringabstractThe aim of in-trawl catch monitoring for use in fishing operations is to detect, track and classify fish targets in real-time from video footage. Information gathered could be used to release unwanted bycatch in real-time. However, traditional multi-object tracking (MOT) methods have limitations, as they are developed for tracking vehicles or pedestrians with linear motions and diverse appearances, which are different from the scenarios such as livestock monitoring. Therefore, we propose a novel MOT method, built upon an existing observation-centric tracking algorithm, by adopting a new iterative association step to significantly boost the performance of tracking targets with a uniform appearance. The iterative association module is an extendable component that can be merged into most existing tracking methods. Our method offers improved performance in tracking targets with uniform appearance and outperforms state-of-the-art techniques on our underwater fish datasets as well as the MOT17 dataset, without increasing latency nor sacrificing accuracy as measured by HOTA, MOTA, and IDF1 performance metrics. Cheng-Yen Yang, Yu Shyang Tan, Melanie J. Underwood, Charlotte Bodie, Zhongyu Jiang, Steve George, Karl Warr, Jenq-Neng Hwang, Emma Jones |
ICIP | 5 |
| 2023 | CameraPose: Weakly-Supervised Monocular 3D Human Pose Estimation by Leveraging In-the-wild 2D AnnotationsabstractTo improve the generalization of 3D human pose estimators, many existing deep learning based models focus on adding different augmentations to training poses. However, data augmentation techniques are limited to the "seen" pose combinations and hard to infer poses with rare "unseen" joint positions. To address this problem, we present CameraPose, a weakly-supervised framework for 3D human pose estimation from a single image, which can not only be applied on 2D-3D pose pairs but also on 2D alone annotations. By adding a camera parameter branch, any in-the-wild 2D annotations can be fed into our pipeline to boost the training diversity and the 3D poses can be implicitly learned by reprojecting back to 2D. Moreover, CameraPose introduces a refinement network module with confidence-guided loss to further improve the quality of noisy 2D keypoints extracted by 2D pose estimators. Experimental results demonstrate that the CameraPose brings in clear improvements on cross-scenario datasets. Notably, it outperforms the baseline method by 3mm on the most challenging dataset 3DPW. In addition, by combining our proposed refinement network module with existing 3D pose estimators, their performance can be improved in cross-scenario evaluation. Cheng-Yen Yang, Jiajia Luo, Yuyin Sun, Nan Qiao 0009, Ke Zhang 0028, Zhongyu Jiang, Jenq-Neng Hwang, Cheng-Hao Kuo |
WACV | 7 |
| 2022 | Unsupervised Domain Adaptation Learning for Hierarchical Infant Pose Recognition with Synthetic DataabstractThe Alberta Infant Motor Scale (AIMS) is a well-known assessment scheme that evaluates the gross motor development of infants by recording the number of specific poses achieved. With the aid of the image-based pose recognition model, the AIMS evaluation procedure can be shortened and automated, providing early diagnosis or indicator of potential developmental disorder. Due to limited public infant-related datasets, many works use the SMIL-based method to generate synthetic infant images for training. However, this domain mismatch between real and synthetic training samples often leads to performance degradation during inference. In this paper, we present a CNN-based model which takes any infant image as input and predicts the coarse and fine-level pose labels. The model consists of an image branch and a pose branch, which respectively generates the coarse-level logits facilitated by the unsupervised domain adaptation and the 3D keypoints using the HRNet with SMPLify optimization. Then the outputs of these branches will be sent into the hierarchical pose recognition module to estimate the fine-level pose labels. We also collect and label a new AIMS dataset, which co—tains 750 real and 4000 synthetic infants images with AIMS pose labels. Our experimental results show that the proposed method can significantly align the distribution of synthetic and real-world datasets, thus achieving accurate performance on fine-grained infant pose recognition. Cheng-Yen Yang, Zhongyu Jiang, Shih-Yu Gu, Jenq-Neng Hwang, Jang-Hee Yoo |
ICME | 2 |
| 2022 | Unsupervised universal hierarchical multi-person 3D pose estimation for natural scenes
Renshu Gu, Zhongyu Jiang, Gaoang Wang, Kevin McQuade, Jenq-Neng Hwang |
Multim. Tools Appl. | 2 |
| 2022 | Text recognition in natural scenes based on deep learning
Zhongyu Jiang, Liang He 0012 |
Multim. Tools Appl. | 2 |
| 2021 | Hierarchical Pose Classification for Infant Action Analysis and Mental Development AssessmentabstractBased on Alberta Infant Motor Scale (AIMS), a questionnaire that tracks an infant’s motor function, an infant’s mental development can be evaluated by recording poses a baby can achieve. Therefore, it is meaningful to propose a systematic image-based pose classifier to classify infant actions based on AIMS to provide early diagnosis of a potential develop-mental disorder such as Autism. This paper presents a hierarchical pose classifier, given a baby image frame that com-bines the benefits of 3D human pose estimation and scene context information. Due to privacy policies, we cannot collect enough real infant images/videos for experiments. In-stead, we generate synthetic baby images with the help of the Skinned Multi-Infant Linear (SMIL) model. Images are first fed into a ResNet-50 for coarse-level pose classification. A stacked hourglass CNN and a hierarchical 3D pose estimation scheme are used for 2D/3D pose estimation. Finally, an innovative Hierarchical Infant Pose Classifier (HIPC) takes the estimated 3D keypoints and coarse-level pose classification confidence scores to give the fine-level baby pose classification results. Our experimental results show that our hierarchical pose classifier achieves accurate and stable performance on infant pose recognition. Jianxiong Zhou, Zhongyu Jiang, Jang-Hee Yoo, Jenq-Neng Hwang |
ICASSP | 2 |
| 2021 | ROD2021 Challenge: A Summary for Radar Object Detection Challenge for Autonomous Driving ApplicationsabstractThe Radar Object Detection 2021 (ROD2021) Challenge, held in the ACM International Conference on Multimedia Retrieval (ICMR) 2021, has been introduced to detect and classify objects purely using an FMCW radar for autonomous driving applications. As a robust sensor to all-weather conditions, radar has rich information hidden in the radio frequencies, which can potentially achieve object detection and classification. This insight will provide a new object perception solution for an autonomous vehicle even in adverse driving scenarios. The ROD2021 Challenge is the first public benchmark focusing on this topic, which attracts great attention and participation. There are more than 260 participants among 37 teams from more than 10 countries with different academic and industrial affiliations, contributing about 300 submissions in the first phase and 400 submissions in the second phase. The final performance is evaluated by average precision (AP). Results add strong value and a better understanding of the radar object detection task for the autonomous vehicle community. Yizhou Wang 0005, Jenq-Neng Hwang, Gaoang Wang, Hui Liu 0011, Kwang-Ju Kim, Hung-Min Hsu, Jiarui Cai, Haotian Zhang 0005, Zhongyu Jiang, Renshu Gu |
ICMR | 9 |
| 2021 | RODNet: Radar Object Detection using Cross-Modal SupervisionabstractRadar is usually more robust than the camera in severe driving scenarios, e.g., weak/strong lighting and bad weather. However, unlike RGB images captured by a camera, the semantic information from the radar signals is noticeably difficult to extract. In this paper, we propose a deep radar object detection network (RODNet), to effectively detect objects purely from the carefully processed radar frequency data in the format of range-azimuth frequency heatmaps (RAMaps). Three different 3D autoencoder based architectures are introduced to predict object confidence distribution from each snippet of the input RAMaps. The final detection results are then calculated using our post-processing method, called location-based non-maximum suppression (L-NMS). Instead of using burdensome human-labeled ground truth, we train the RODNet using the annotations generated automatically by a novel 3D localization method using a camera-radar fusion (CRF) strategy. To train and evaluate our method, we build a new dataset - CRUW, containing synchronized videos and RAMaps in various driving scenarios. After intensive experiments, our RODNet shows favorable object detection performance without the presence of the camera. Yizhou Wang 0005, Zhongyu Jiang, Jenq-Neng Hwang, Guanbin Xing, Hui Liu 0011 |
WACV | 2 |
| 2021 | Low-light image enhancement based on Retinex decomposition and adaptive gamma correctionabstractAbstract Low‐light images suffer from poor visibility and noise. In this paper, a low‐light image enhancement method based on Retinex decomposition is proposed. A pyramid network is first utilized to extract multi‐scale features to improve the quality of Retinex decomposition. Then the decomposed illumination is refined via an adaptive Gamma correction network to handle non‐uniform illumination, while the decomposed reflectance is refined with a lightweight network. Finally, the enhanced image is obtained by element‐wise multiplication between the refined illumination and reflectance components. Quantitative and qualitative experiments demonstrate the superiority of our method over state‐of‐the‐art image enhancement methods. Jing-Yu Yang 0002, Huanjing Yue, Zhongyu Jiang, Kun Li 0001 |
IET Image Process. | 4 |
| 2021 | Deep edge map guided depth super resolution
Zhongyu Jiang, Huanjing Yue, Yukun Lai, Jing-Yu Yang 0002, Yonghong Hou, Chunping Hou |
Signal Process. Image Commun. | 1 |
| 2021 | Deep noise estimation and removal for real-world noisy images
Huanjing Yue, Zhongyu Jiang, Shengdi Zhou, Jing-Yu Yang 0002, Yonghong Hou, Chunping Hou |
Signal Process. Image Commun. | 2 |
| 2021 | Reference guided image super-resolution via efficient dense warping and adaptive fusion
Huanjing Yue, Zhongyu Jiang, Jing-Yu Yang 0002, Chunping Hou |
Signal Process. Image Commun. | 3 |
| 2020 | Multi-Person Hierarchical 3D Pose Estimation in Natural VideosabstractDespite the increasing need of analyzing human poses on the street and in the wild, multi-person 3D pose estimation using monocular static or moving camera in real-world scenarios remains a challenge, either requiring large-scale training data or high computation complexity due to the high degrees of freedom in 3D human poses. We propose a novel scheme to effectively track and hierarchically estimate 3D human poses in natural videos in an efficient fashion. Without the need of using labelled 3D training data, we formulate torso estimation as a Perspective-N-Point (PNP) problem, and limb pose estimation as an optimization problem, and hierarchically structure the high dimensional poses to efficiently address the challenge. Experiments show good performance and high efficiency of multi-person 3D pose estimation on real-world videos, including street scenarios and various human daily activities from fixed and moving cameras, resulting in great new opportunities to understand and predict human behaviors. Renshu Gu, Gaoang Wang, Zhongyu Jiang, Jenq-Neng Hwang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Light Weight Stereo Matching via Deep Extraction and Integration of Low and High Level InformationabstractDeep convolutional neural networks (CNN) have demonstrated remarkable progress in stereo matching recently. However, disparity estimation in the ill-posed regions is still difficult. In addition, CNN based stereo matching methods often have impractical computational complexity and memory consumption. To address these problems we propose an end-to-end light weight CNN architecture to effectively learn and integrate low and high level information. To achieve this, a novel enhancement block built upon group convolution and dilated-convolution is proposed. Compared with state-of-the-art methods, the proposed method achieved competitive performance with the least number of network parameters on the Flyingthings3d and KITTI datasets. Yonghong Hou, Pichao Wang, Zhongyu Jiang, Wanqing Li 0001 |
ICME | 4 |
| 2019 | System Reliability Allocation and Optimization Based on Generalized Birnbaum Importance MeasureabstractImportance measure can be used to identify the most vulnerable components with respect to system functionality or failure. Traditional importance measures may not be effective to evaluate the contribution of an individual component if its reliability value does not fall in the full range between 0 and 1. Based on the cost-reliability relation, this paper proposes a generalized Birnbaum importance measure (GBIM) to quantify the contribution of individual components to system reliability improvement by considering reliability range, manufacturing complexity, and technology feasibility. Since GBIM possesses several unique features in terms of guiding system reliability optimization, in this paper, we further develop a GBIM-based genetic algorithm to solve a type of optimal reliability allocation problem. The numerical studies show that both the computational efficiency and the near global optimality based on GBIM outperforms the methods using the traditional importance measures. Shubin Si, Mingli Liu, Zhongyu Jiang, Tongdan Jin, Zhiqiang Cai 0003 |
IEEE Trans. Reliab. | 3 |
| 2018 | Depth Super-Resolution From RGB-D Pairs With Transform and Spatial Domain RegularizationabstractThis paper proposes a depth super-resolution method with both transform and spatial domain regularization. In the transform domain regularization, nonlocal correlations are exploited via an auto-regressive model, where each patch is further sparsified with a locally-trained transform to consider intra-patch correlations. In the spatial domain regularization, we propose a multi-directional total variation (MTV) prior to characterize the geometrical structures spatially orientated at arbitrary directions in depth maps. To achieve adaptive regularization, the MTV is weighted for each directional finite difference considering local characteristics of RGB-D data. We develop an accelerated proximal gradient algorithm to solve the proposed model. Quantitative and qualitative evaluations compared with state-of-the-art methods demonstrate that the proposed method achieves superior depth super-resolution performance for various configurations of magnification factors and datasets. Zhongyu Jiang, Yonghong Hou, Huanjing Yue, Jing-Yu Yang 0002, Chunping Hou |
IEEE Trans. Image Process. | 1 |
| 2015 | Rectangle fitting via quadratic programmingabstractThis paper investigates rectangle fitting via optimization approaches. We summarize two basic requirements for rectangular fitting, leading to a basic model that are non-convex and difficult to attack. To avoid potential trapping of local minima, we extend the basic model with centroid and orientation constraints into a quadratic programming. To achieve reliable fitting from noisy points, slack variables are introduced to soften hard constraints. The scalability to problem size are further addressed by careful selecting only a small fraction of slack variables. Results on clean dataset, noisy dataset, and practical data show that our method is able to reliably fit rectangles for various kinds of data. Jing-Yu Yang 0002, Zhongyu Jiang |
MMSP | 2 |