EDBT 2026 Demo / reviewers in the wild / expert
Jingzhe Ma
dblp:172/0859
· DBLP profile ↗
12ranked-venue papers
4as first author
11since 2021 · last 2026
0000-0002-7090-6989ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Security and privacy · 3 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SAFE: A Semantic-Appearance Full-body Editor for pedestrian video anonymization
Jingzhe Ma, Chao Fan 0001, Dingqiang Ye, Jinfeng Yang, Dongyang Jin, Fuad Mire Hassan, Shiqi Yu 0001 |
Pattern Recognit. | 1 |
| 2025 | On Denoising Walking Videos for Gait RecognitionabstractTo capture individual gait patterns, excluding identity-irrelevant cues in walking videos, such as clothing texture and color, remains a persistent challenge for vision-based gait recognition. Traditional silhouette- and pose-based methods, though theoretically effective at removing such distractions, often fall short of high accuracy due to their sparse and less informative inputs. Emerging end-to-end methods address this by directly denoising RGB videos using human priors. Building on this trend, we propose DenoisingGait, a novel gait denoising method. Inspired by the philosophy that "what I cannot create, I do not understand", we turn to generative diffusion models, uncovering how they partially filter out irrelevant factors for gait understanding. Additionally, we introduce a geometry-driven Feature Matching module, which, combined with background removal via human silhouettes, condenses the multi-channel diffusion features at each foreground pixel into a two-channel direction vector. Specifically, the proposed within- and cross-frame matching respectively capture the local vectorized structures of gait appearance and motion, producing a novel flow-like gait representation termed Gait Feature Field, which further reduces residual noise in diffusion features. Experiments on the CCPG, CASIA-B*, and SUSTech1K datasets demonstrate that DenoisingGait achieves a new SoTA performance in most cases for both within- and cross-domain evaluations. Code is available at https://github.com/ShiqiYu/OpenGait. Dongyang Jin, Chao Fan 0001, Jingzhe Ma, Jingkai Zhou, Shiqi Yu 0001 |
CVPR | 3 |
| 2025 | Human Identification at a Distance: Challenges, Methods and Results on the Competition HID 2025abstractHuman identification at a distance (HID) faces challenges due to the difficulty of acquiring traditional biometric modalities like face and fingerprints. Gait recognition offers a viable solution since it can be captured at a distance. To promote progress in gait recognition and provide a fair evaluation platform, the International Competition on Human Identification at a Distance (HID) has been organized annually since 2020. Since 2023, the competition has adopted the challenging SUSTech-Competition dataset, which includes significant variations in clothing, carried objects, and view angles. No training data is provided, requiring participants to train their models using external datasets. Each year, the competition applies a different random seed to generate distinct evaluation splits, reducing the risk of overfitting and ensuring fair evaluation of cross-domain generalization. Although the previous two competitions (HID 2023 and HID 2024) already utilized this dataset, HID 2025 aimed explicitly to explore whether algorithmic improvements could surpass the accuracy limits observed previously. Despite these heightened challenges, participants again demonstrated significant advancements, with the highest accuracy reaching 94.2%, setting a new benchmark for this dataset. We also analyze key technical trends and outline potential directions for future research on gait recognition. Jingzhe Ma, Jianlong Yu, Zunxiao Xu, Xue Cheng, Zepeng Wang 0002, Kazuki Osamura, Rujie Liu, Narishige Abe, Shunli Zhang 0005, Haojun Xie, Weiming Wu, Wenxiong Kang, Qingshuo Gao, Jiaming Xiong, Xianye Ben, Lei Chen 0095, Lichen Song, Junjian Cui, Haijun Xiong, Junhao Lu, Bin Feng 0001, Baoquan Zhao, Ke Xu 0001, Yongzhen Huang, Liang Wang 0001, Manuel J. Marín-Jiménez, Md. Atiqur Rahman Ahad, Shiqi Yu 0001 |
IJCB | 1 |
| 2025 | KDDA-balance: Knowledge-driven domain adaptation with correlated gait information for elderly balance assessment
Zhaoyang Ge, Shujie Huang, Huiqing Cheng, Jingzhe Ma, Zhuang Tong, Mingliang Xu 0001 |
Knowl. Based Syst. | 4 |
| 2025 | OpenGait: A Comprehensive Benchmark Study for Gait Recognition Toward Better PracticalityabstractGait recognition, a rapidly advancing vision technology for person identification from a distance, has made significant strides in indoor settings. However, evidence suggests that existing methods often yield unsatisfactory results when applied to newly released real-world gait datasets. Furthermore, conclusions drawn from indoor gait datasets may not easily generalize to outdoor ones. Therefore, the primary goal of this paper is to present a comprehensive benchmark study aimed at improving practicality rather than solely focusing on enhancing performance. To this end, we developed OpenGait, a flexible and efficient gait recognition platform. Using OpenGait, we conducted in-depth ablation experiments to revisit recent developments in gait recognition. Surprisingly, we detected some imperfect parts of some prior methods and thereby uncovered several critical yet previously neglected insights. These findings led us to develop three structurally simple yet empirically powerful and practically robust baseline models: DeepGaitV2, SkeletonGait, and SkeletonGait++, which represent the appearance-based, model-based, and multi-modal methodologies for gait pattern description, respectively. In addition to achieving state-of-the-art performance, our careful exploration provides new perspectives on the modeling experience of deep gait models and the representational capacity of typical gait modalities. In the end, we discuss the key trends and challenges in current gait recognition, aiming to inspire further advancements towards better practicality. Chao Fan 0001, Saihui Hou, Chuanfu Shen, Jingzhe Ma, Dongyang Jin, Yongzhen Huang, Shiqi Yu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Adaptive Dual-Axis Style-Based Recalibration Network With Class-Wise Statistics Loss for Imbalanced Medical Image ClassificationabstractSalient and small lesions (e.g., microaneurysms on fundus) both play significant roles in real-world disease diagnosis under medical image examinations. Although deep neural networks (DNNs) have achieved promising medical image classification performance, they often have limitations in capturing both salient and small lesion information, restricting performance improvement in imbalanced medical image classification. Recently, with the advent of DNN-based style transfer in medical image generation, the roles of clinical styles have attracted great interest, as they are crucial indicators of lesions. Motivated by this observation, we propose a novel Adaptive Dual-Axis Style-based Recalibration (ADSR) module, leveraging the potential of clinical styles to guide DNNs in effectively learning salient and small lesion information from a dual-axis perspective. ADSR first emphasizes salient lesion information via global style-based adaptation, then captures small lesion information with pixel-wise style-based fusion. We construct an ADSR-Net for imbalanced medical image classification by stacking multiple ADSR modules. Additionally, DNNs typically adopt cross-entropy loss for parameter optimization, which ignores the impacts of class-wise predicted probability distributions. To address this, we introduce a new Class-wise Statistics Loss (CWS) combined with CE to further boost imbalanced medical image classification results. Extensive experiments on five imbalanced medical image datasets demonstrate not only the superiority of ADSR-Net and CWS over state-of-the-art (SOTA) methods but also their improved confidence calibration results. For example, ADSR-Net with the proposed loss significantly outperforms CABNet50 by 21.39% and 27.82% in F1 and B-ACC while reducing 3.31% and 4.57% in ECE and BS on ISIC2018. Xiaoqing Zhang 0001, Zunjie Xiao, Jingzhe Ma, Jilu Zhao, Shuai Zhang 0029, Runzhi Li, Yi Pan 0001, Jiang Liu 0001 |
IEEE Trans. Image Process. | 3 |
| 2024 | SkeletonGait: Gait Recognition Using Skeleton MapsabstractThe choice of the representations is essential for deep gait recognition methods. The binary silhouettes and skeletal coordinates are two dominant representations in recent literature, achieving remarkable advances in many scenarios. However, inherent challenges remain, in which silhouettes are not always guaranteed in unconstrained scenes, and structural cues have not been fully utilized from skeletons. In this paper, we introduce a novel skeletal gait representation named skeleton map, together with SkeletonGait, a skeleton-based method to exploit structural information from human skeleton maps. Specifically, the skeleton map represents the coordinates of human joints as a heatmap with Gaussian approximation, exhibiting a silhouette-like image devoid of exact body structure. Beyond achieving state-of-the-art performances over five popular gait datasets, more importantly, SkeletonGait uncovers novel insights about how important structural features are in describing gait and when they play a role. Furthermore, we propose a multi-branch architecture, named SkeletonGait++, to make use of complementary features from both skeletons and silhouettes. Experiments indicate that SkeletonGait++ outperforms existing state-of-the-art methods by a significant margin in various scenarios. For instance, it achieves an impressive rank-1 accuracy of over 85% on the challenging GREW dataset. The source code is available at https://github.com/ShiqiYu/OpenGait. Chao Fan 0001, Jingzhe Ma, Dongyang Jin, Chuanfu Shen, Shiqi Yu 0001 |
AAAI | 2 |
| 2024 | BigGait: Learning Gait Representation You Want by Large Vision ModelsabstractGait recognition stands as one of the most pivotal remote identification technologies and progressively expands across research and industry communities. However, existing gait recognition methods heavily rely on task-specific upstream driven by supervised learning to provide explicit gait representations like silhouette sequences, which in-evitably introduce expensive annotation costs and poten-tial error accumulation. Escaping from this trend, this work explores effective gait representations based on the all-purpose knowledge produced by task-agnostic Large Vision Models (LVMs) and proposes a simple yet efficient gait framework, termed B igGait. Specifically, the Gait Repre-sentation Extractor (GRE) within BigGait draws upon design principles from established gait representations, effectively transforming all-purpose knowledge into implicit gait representations without requiring third-party supervision signals. Experiments on CCPG, CAISA-B* and SUSTechlK indicate that BigGait significantly outperforms the previous methods in both within-domain and cross-domain tasks in most cases, and provides a more practical paradigm for learning the next-generation gait representation. Fi-nally, we delve into prospective challenges and promising directions in LVMs-based gait recognition, aiming to in-spire future work in this emerging topic. The source code is available at https://github.com/ShiqiYu/OpenGait. Dingqiang Ye, Chao Fan 0001, Jingzhe Ma, Xiaoming Liu 0002, Shiqi Yu 0001 |
CVPR | 3 |
| 2024 | Learned HDR Image Compression for Perceptually Optimal Storage and Display
Peibei Cao, Jingzhe Ma, Yu-Chieh Yuan, Zhiyong Xie, Haiqing Bai, Kede Ma |
ECCV (49) | 3 |
| 2024 | Passersby-Anonymizer: Safeguard the Privacy of Passersby in Social VideosabstractIn the current era of pervasive short video content, the exposure of passersby’s data frequently raises privacy concerns. Traditional anonymization techniques for passersby, like blurring and mosaicing, are often used before uploading such videos. However, these methods tend to degrade the informational richness of the visual content, markedly reducing the quality of the anonymized videos. Recent advancements of diffusion models have paved the way for text-guided image and video synthesis, yet applying these models to the anonymization of passersby poses three main challenges: i) bridging the domain gap between specific passersby data and high-quality image/video datasets that are used for pre-training diffusion models, ii) ensuring temporal consistency in the anonymized videos, and iii) preserving the integrity of video subjects’ content while exclusively anonymizing passersby-related information. To address these challenges, we propose the Passersby-Anonymizer, a novel diffusion-based framework for anonymizing identity-specific attributes in video content. At its core, our model introduces a spatial content adapter (SCA) to adapt to the visual patterns of passersby image datasets. We introduce a Temporal Content Stabilizer (TCS) to maintain the temporal consistency of the anonymized videos. Furthermore, we design a mask-aware training strategy that specifically targets the anonymization of the mask region while preserving the integrity of other contents. Our experimental evaluations demonstrate that our model effectively addresses the challenge of anonymizing passersby without compromising the informational integrity of the social videos. The source code is available at https://github.com/HappyDeepLearning/Passersby-Anonymizer. Jingzhe Ma, Haoyu Luo, Zixu Huang, Dongyang Jin, Johann A. Briffa, Norman Poh, Shiqi Yu 0001 |
IJCB | 1 |
| 2024 | An Identity-Preserved Framework for Human Motion TransferabstractHuman motion transfer (HMT) aims to generate a video clip for the target subject by imitating the source subject’s motion. Although previous methods have achieved good results in synthesizing good-quality videos, they lose sight of individualized motion information from the source and target motions, which is significant for the realism of the motion in the generated video. To address this problem, we propose a novel identity-preserved HMT network, termedIDPres. This network is a skeleton-based approach that uniquely incorporates the target’s individualized motion and skeleton information to augment identity representations. This integration significantly enhances the realism of movements in the generated videos. Our method focuses on the fine-grained disentanglement and synthesis of motion. To improve the representation learning capability in latent space and facilitate the training ofIDPres, we introduce three training schemes. These schemes enableIDPresto concurrently disentangle different representations and accurately control them, ensuring the synthesis of ideal motions. To evaluate the proportion of individualized motion information in the generated video, we are the first to introduce a new quantitative metric called Identity Score (ID-Score), motivated by the success of gait recognition methods in capturing identity information. Moreover, we collect an identity-motion paired dataset,Dancer101, consisting of solo-dance videos of 101 subjects from the public domain, providing a benchmark to prompt the development of HMT methods. Extensive experiments demonstrate that the proposedIDPresmethod surpasses existing state-of-the-art techniques in terms of reconstruction accuracy, realistic motion, and identity preservation. Jingzhe Ma, Xiaoqing Zhang 0001, Shiqi Yu 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2020 | An Ensemble Deep Learning Architecture for Multilabel Classification on TI-RADSabstractIn recent years, thyroid nodule is being one of the most common nodular lesions. Ultrasonography is widely used in the clinical diagnosis of thyroid nodules. With the development of artificial intelligence, there emerge great progress in medical image diagnosis. Clinically, physician diagnose malignant or benign by many pathological features. In this work, we propose an ensemble architecture to resolve multi-label problem, which integrate three methods to extract features on thyroid nodule for ultrasound images. They are EfficientNet, feature engineering and feature pyramid network. We consider five kinds of pathological features quantified by Thyroid Imaging Report and Data System (TI-RADS). In the experiments, we use two datasets. One includes 587 original ultrasound images on thyroid nodules collected from the local health physical center of a 3A hospital. The other is a public dataset in the MICCAI 2020 competition. It contains 3644 ultrasound images of thyroid images. We use receiver operating characteristic curve (ROC) to evaluate the model, the area under curve (AUC) on every kind of features. The experimental results show that the proposed method can effectively assist doctors in diagnosing ultrasound images of thyroid nodules. Xueli Duan, Shaobo Duan, Pei Jiang 0008, Runzhi Li, Jingzhe Ma, Hongling Zhao, Honghua Dai 0001 |
BIBM | 6 |