Dongyang Jin

dblp:256/4998 · DBLP profile ↗
← Back
11ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Security and privacy · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SCALAR: Scale-wise Controllable Visual Autoregressive Learning
abstract
Controllable image synthesis, which enables fine-grained control over generated outputs, has emerged as a key focus in visual generative modeling. However, controllable generation remains challenging for Visual Autoregressive (VAR) models due to their hierarchical, next-scale prediction style. Existing VAR-based methods often suffer from inefficient control encoding and disruptive injection mechanisms that compromise both fidelity and efficiency. In this work, we present SCALAR, a controllable generation method based on VAR, incorporating a Scale-wise Conditional Decoding mechanism. SCALAR leverages a pretrained image encoder to extract semantic control signal encodings, which are projected into scale-specific representations and injected into the corresponding layers of the VAR backbone. This design provides persistent and structurally aligned guidance throughout the generation process. Building on SCALAR, we develop SCALAR-Uni, a unified extension that aligns multiple control modalities into a shared latent space, supporting flexible multi-conditional guidance in a single model. Extensive experiments show that SCALAR achieves superior generation quality and control precision across various tasks.
Ryan Xu, Dongyang Jin, Yancheng Bai, Rui Lan, Xu Duan, Xiangxiang Chu
AAAI2
2026 SAFE: A Semantic-Appearance Full-body Editor for pedestrian video anonymization
Jingzhe Ma, Chao Fan 0001, Dingqiang Ye, Jinfeng Yang, Dongyang Jin, Fuad Mire Hassan, Shiqi Yu 0001
Pattern Recognit.5
2025 Exploring More from Multiple Gait Modalities for Human Identification
abstract
The gait, as a kind of soft biometric characteristic, can reflect the distinct walking patterns of individuals at a distance, exhibiting a promising technique for unrestrained human identification. With largely excluding gait-unrelated cues hidden in RGB videos, the silhouette and skeleton, though visually compact, have acted as two of the most prevailing gait modalities for a long time. Recently, several attempts have been made to introduce more informative data forms like human parsing and optical flow images to capture gait characteristics, along with multi-branch architectures. However, due to the inconsistency within model designs and experiment settings, we argue that a comprehensive and fair comparative study among these popular gait modalities, involving the representational capacity and fusion strategy exploration, is still lacking. From the perspectives of fine vs. coarse-grained shape and whole vs. pixel-wise motion modeling, this work presents an in-depth investigation of three popular gait representations, i.e., silhouette, human parsing, and optical flow, with various fusion evaluations, and experimentally exposes their similarities and differences. Based on the obtained insights, we further develop a C²Fusion strategy, consequently building our new framework MultiGait++. C²Fusion preserves commonalities while highlighting differences to enrich the learning of gait features. To verify our findings and conclusions, extensive experiments on Gait3D, GREW, CCPG, and SUSTech1K are conducted.
Dongyang Jin, Chao Fan 0001, Shiqi Yu 0001
AAAI1
2025 On Denoising Walking Videos for Gait Recognition
abstract
To capture individual gait patterns, excluding identity-irrelevant cues in walking videos, such as clothing texture and color, remains a persistent challenge for vision-based gait recognition. Traditional silhouette- and pose-based methods, though theoretically effective at removing such distractions, often fall short of high accuracy due to their sparse and less informative inputs. Emerging end-to-end methods address this by directly denoising RGB videos using human priors. Building on this trend, we propose DenoisingGait, a novel gait denoising method. Inspired by the philosophy that "what I cannot create, I do not understand", we turn to generative diffusion models, uncovering how they partially filter out irrelevant factors for gait understanding. Additionally, we introduce a geometry-driven Feature Matching module, which, combined with background removal via human silhouettes, condenses the multi-channel diffusion features at each foreground pixel into a two-channel direction vector. Specifically, the proposed within- and cross-frame matching respectively capture the local vectorized structures of gait appearance and motion, producing a novel flow-like gait representation termed Gait Feature Field, which further reduces residual noise in diffusion features. Experiments on the CCPG, CASIA-B*, and SUSTech1K datasets demonstrate that DenoisingGait achieves a new SoTA performance in most cases for both within- and cross-domain evaluations. Code is available at https://github.com/ShiqiYu/OpenGait.
Dongyang Jin, Chao Fan 0001, Jingzhe Ma, Jingkai Zhou, Shiqi Yu 0001
CVPR1
2025 Pose as Clinical Prior: Learning Dual Representations for Scoliosis Screening
Zirui Zhou, Zizhao Peng, Dongyang Jin, Chao Fan 0001, Fengwei An, Shiqi Yu 0001
MICCAI (13)3
2025 OpenGait: A Comprehensive Benchmark Study for Gait Recognition Toward Better Practicality
abstract
Gait recognition, a rapidly advancing vision technology for person identification from a distance, has made significant strides in indoor settings. However, evidence suggests that existing methods often yield unsatisfactory results when applied to newly released real-world gait datasets. Furthermore, conclusions drawn from indoor gait datasets may not easily generalize to outdoor ones. Therefore, the primary goal of this paper is to present a comprehensive benchmark study aimed at improving practicality rather than solely focusing on enhancing performance. To this end, we developed OpenGait, a flexible and efficient gait recognition platform. Using OpenGait, we conducted in-depth ablation experiments to revisit recent developments in gait recognition. Surprisingly, we detected some imperfect parts of some prior methods and thereby uncovered several critical yet previously neglected insights. These findings led us to develop three structurally simple yet empirically powerful and practically robust baseline models: DeepGaitV2, SkeletonGait, and SkeletonGait++, which represent the appearance-based, model-based, and multi-modal methodologies for gait pattern description, respectively. In addition to achieving state-of-the-art performance, our careful exploration provides new perspectives on the modeling experience of deep gait models and the representational capacity of typical gait modalities. In the end, we discuss the key trends and challenges in current gait recognition, aiming to inspire further advancements towards better practicality.
Chao Fan 0001, Saihui Hou, Chuanfu Shen, Jingzhe Ma, Dongyang Jin, Yongzhen Huang, Shiqi Yu 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 SkeletonGait: Gait Recognition Using Skeleton Maps
abstract
The choice of the representations is essential for deep gait recognition methods. The binary silhouettes and skeletal coordinates are two dominant representations in recent literature, achieving remarkable advances in many scenarios. However, inherent challenges remain, in which silhouettes are not always guaranteed in unconstrained scenes, and structural cues have not been fully utilized from skeletons. In this paper, we introduce a novel skeletal gait representation named skeleton map, together with SkeletonGait, a skeleton-based method to exploit structural information from human skeleton maps. Specifically, the skeleton map represents the coordinates of human joints as a heatmap with Gaussian approximation, exhibiting a silhouette-like image devoid of exact body structure. Beyond achieving state-of-the-art performances over five popular gait datasets, more importantly, SkeletonGait uncovers novel insights about how important structural features are in describing gait and when they play a role. Furthermore, we propose a multi-branch architecture, named SkeletonGait++, to make use of complementary features from both skeletons and silhouettes. Experiments indicate that SkeletonGait++ outperforms existing state-of-the-art methods by a significant margin in various scenarios. For instance, it achieves an impressive rank-1 accuracy of over 85% on the challenging GREW dataset. The source code is available at https://github.com/ShiqiYu/OpenGait.
Chao Fan 0001, Jingzhe Ma, Dongyang Jin, Chuanfu Shen, Shiqi Yu 0001
AAAI3
2024 Passersby-Anonymizer: Safeguard the Privacy of Passersby in Social Videos
abstract
In the current era of pervasive short video content, the exposure of passersby’s data frequently raises privacy concerns. Traditional anonymization techniques for passersby, like blurring and mosaicing, are often used before uploading such videos. However, these methods tend to degrade the informational richness of the visual content, markedly reducing the quality of the anonymized videos. Recent advancements of diffusion models have paved the way for text-guided image and video synthesis, yet applying these models to the anonymization of passersby poses three main challenges: i) bridging the domain gap between specific passersby data and high-quality image/video datasets that are used for pre-training diffusion models, ii) ensuring temporal consistency in the anonymized videos, and iii) preserving the integrity of video subjects’ content while exclusively anonymizing passersby-related information. To address these challenges, we propose the Passersby-Anonymizer, a novel diffusion-based framework for anonymizing identity-specific attributes in video content. At its core, our model introduces a spatial content adapter (SCA) to adapt to the visual patterns of passersby image datasets. We introduce a Temporal Content Stabilizer (TCS) to maintain the temporal consistency of the anonymized videos. Furthermore, we design a mask-aware training strategy that specifically targets the anonymization of the mask region while preserving the integrity of other contents. Our experimental evaluations demonstrate that our model effectively addresses the challenge of anonymizing passersby without compromising the informational integrity of the social videos. The source code is available at https://github.com/HappyDeepLearning/Passersby-Anonymizer.
Jingzhe Ma, Haoyu Luo, Zixu Huang, Dongyang Jin, Johann A. Briffa, Norman Poh, Shiqi Yu 0001
IJCB4
2024 Predefined-Time Consensus for Second-Order Nonlinear Multiagent Systems via Sliding Mode Technique
abstract
This paper investigates the predefined-time consensus for a class of second-order nonlinear multi-agent systems via sliding mode technique. Fuzzy logic systems are utilized to estimate unknown continuous functions. Sliding mode control is employed to address the estimation errors generated by the estimation process and terms caused by unmodeled dynamics. By incorporating hyperbolic tangent functions into the protocol and utilizing their properties, singularity issues are avoided. A predefined-time fuzzy protocol is designed to ensure that the consensus error reaches a small neighborhood near zero within the predefined time. The effectiveness of the proposed approach is demonstrated through a simulation example.
Dongyang Jin, Zhengrong Xiang
IEEE Trans. Fuzzy Syst.1
2024 Adaptive Event-Triggered Fixed-Time Fault-Tolerant Consensus Control for a Class of Multiagent Systems
abstract
This article investigates an adaptive event-triggered fixed-time fault-tolerant consensus control for a category of multiagent systems (MASs). First, radial basis function (RBF) neural networks (NNs) are applied to handle unknown nonlinear terms. Furthermore, the backstepping technique is employed to construct the fixed-time event-triggered consensus control scheme by utilizing the command filter technique. The proposed scheme can guarantee the boundedness of all signals in the closed-loop system and the fixed-time convergence of consensus error. Additionally, a simulation example is presented to verify the validity of the presented results.
Dongyang Jin, Zhengrong Xiang
IEEE Trans. Syst. Man Cybern. Syst.1
2022 Relative Pose Estimation for Light Field Cameras Based on LF-Point-LF-Point Correspondence Model
abstract
In this paper, we propose a relative pose estimation algorithm for micro-lens array (MLA)-based conventional light field (LF) cameras. First, by employing the matched LF-point pairs, we establish the LF-point-LF-point correspondence model to represent the correlation between LF features of the same 3D scene point in a pair of LFs. Then, we employ the proposed correspondence model to estimate the relative camera pose, which includes a linear solution and a non-linear optimization on manifold. Unlike prior related algorithms, which estimated relative poses based on the recovered depths of scene points, we adopt the estimated disparities to avoid the inaccuracy in recovering depths due to the ultra-small baseline between sub-aperture images of LF cameras. Experimental results on both simulated and real scene data have demonstrated the effectiveness of the proposed algorithm compared with classical as well as state-of-art relative pose estimation algorithms.
Saiping Zhang, Dongyang Jin, Yuchao Dai, Fuzheng Yang 0001
IEEE Trans. Image Process.2