Yalong Jiang

dblp:210/2829 · DBLP profile ↗
← Back
18ranked-venue papers
9as first author
13since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Uncertainty-guided vertex-parameter bidirectional refinement for hand pose and shape estimation
Shuiping Gou, Yalong Jiang, Yu Sha, Yingping Li
Neurocomputing4
2026 From Abstract Events to Grounded Cues: Cue-Guided Vision-Language Anomaly Detection
abstract
Video anomaly detection (VAD) must be reliable under large appearance variation and weak supervision, yet provide explanations grounded in human-interpretable evidence. Vision-Language Models (VLMs) often suffer from brittle direct visual-text alignment which is unstable across domains, while deep models are accurate but have limited interpretability. We address this by introducing event-related but more concrete cues as intermediate representations, making the mapping from frames to evidence more stable than directly mapping frames to event labels. Building on this idea, we propose a two-stage cooperative framework: a VLM discovers cues and generates cue-guided pseudo frame-level labels, and a Symbolic Learning Model (SLM) learns from them to produce segment-level cue presence estimates and anomaly scores via cross-modal matching. In inference, cue estimates, SLM scores and video segments are combined by VLM to output final decisions with cue-based rationales. Experiments show competitive performance. Code is available athttps://github.com/AllenYLJiang/Atoms-to-Events-Categorical-Evidence-Composition-for-Video-Anomaly-Detection.
Yalong Jiang, Lian Huai, Yuyu Liu, Xingqun Jiang
IEEE Signal Process. Lett.1
2025 Local Patterns Generalize Better for Novel Anomalies
abstract
Video anomaly detection (VAD) aims to identify novel actions or events which are unseen during training. Existing mainstream VAD techniques typically focus on the global patterns with redundant details and struggle to generalize to unseen samples. In this paper, we propose a framework that identifies the local patterns which generalize to novel samples and models the dynamics of local patterns. The capability of extracting spatial local patterns is achieved through a two-stage process involving image-text alignment and cross-modality attention. Generalizable representations are built by focusing on semantically relevant components which can be recombined to capture the essence of novel anomalies, reducing unnecessary visual data variances. To enhance local patterns with temporal clues, we propose a State Machine Module (SMM) that utilizes earlier high-resolution textual tokens to guide the generation of precise captions for subsequent low-resolution observations. Furthermore, temporal motion estimation complements spatial local patterns to detect anomalies characterized by novel spatial distributions or distinctive dynamics. Extensive experiments on popular benchmark datasets demonstrate the achievement of state-of-the-art performance. Code is available at https://github.com/AllenYLJiang/Local-Patterns-Generalize-Better/.
Yalong Jiang
ICLR1
2025 MGFA-Unet: A Lightweight Crack Segmentation Network for Crowdsourcing
abstract
Deep learning techniques have demonstrated considerable potential in the field of crack detection. Especially in the emerging smart city infrastructure applications, it is very widespread, among which the crack detection of crowdsourcing mobile terminals is an innovative urban infrastructure detection model. To address the challenges of achieving precise crack segmentation and overcoming deployment constraints on mobile platforms, we propose a U-Net model that utilizes multiscale global information fusion attention. Multi-Scale Global Information Fusion Attention U-Net introduces a global attention module in the encoder, enabling the model to effectively capture critical crack features and maintain focused attention on target regions. In the meantime, the decoder incorporates a multiscale dilated fusion attention module, leveraging various dilation rates and combining dual attention mechanisms. This design effectively captures both detailed information on crack defects and extensive contextual features, facilitating efficient global context modeling. Our crowdsourcing-based mobile detection framework enables real-time, large-scale crack monitoring with unprecedented coverage density. Evaluations on four datasets show state-of-the-art accuracy (IoU: 76.02%, Dice: 86.33%) and efficiency (${2.85M}$parameters), showcasing its viability for ubiquitous intelligent infrastructure monitoring.
Nan Jiang 0013, Tongtong Zhou, Zhixiang Qian, Lihong Tong, Yalong Jiang
ICPADS5
2025 CLIP-TNseg: A Multi-Modal Hybrid Framework for Thyroid Nodule Segmentation in Ultrasound Images
abstract
Thyroid nodule segmentation in ultrasound images is crucial for accurate diagnosis and treatment planning. However, existing methods struggle with segmentation accuracy, interpretability, and generalization. This letter proposes CLIP-TNseg, a novel framework that integrates a multimodal large model with a neural network architecture to address these challenges. We innovatively divide visual features into coarse-grained and fine-grained components, leveraging textual integration with coarse-grained features for enhanced semantic understanding. Specifically, the Coarse-grained Branch extracts high-level semantic features from a frozen CLIP model, while the Fine-grained Branch refines spatial details using U-Net-style residual blocks. Extensive experiments on the newly collected PKTN dataset and other public datasets demonstrate the competitive performance of CLIP-TNseg. Additional ablation experiments confirm the critical contribution of textual inputs, particularly highlighting the effectiveness of our carefully designed textual prompts compared to fixed or absent textual information.
Boxiong Wei, Yalong Jiang, Liquan Mao, Qi Zhao 0037
IEEE Signal Process. Lett.3
2024 VLAVAD: Vision-Language Models Assisted Unsupervised Video Anomaly Detection
Changkang Li, Yalong Jiang
BMVC2
2024 Progressive prediction: Video anomaly detection via multi-grained prediction
abstract
Abstract Video Anomaly Detection (VAD) has been an active research field for several decades. However, most existing approaches merely extract a single type of feature from videos and define a single paradigm to indicate the extent of abnormalities. A coarse‐to‐fine three‐level prediction is built by integrating different levels of spatio‐temporal representations, better highlighting the difference between normal and abnormal behaviors. First, an object‐level trajectory prediction is proposed to model human historical position using a graph transformer network. Subsequently, skeleton‐level prediction is achieved by incorporating the positional information from the trajectory prediction. More importantly, based on the predicted skeleton, a skeleton‐guided pixel‐level region prediction is performed. A novel Skeleton Conditioned Generative Adversarial Network (SCGAN) is designed to explore the correlation between skeleton‐level and pixel‐level motion prediction. Benefiting from SCGAN, the prediction of human regions is contributed by both coarse‐grained and fine‐grained motion features. This three‐level prediction, namely Progressive Prediction Video Anomaly Detection (P 3 VAD), enlarges the prediction error on irregular motion patterns. Besides, a pixel‐level analysis method is proposed to achieve Background‐bias Elimination (BE) and denoise the predicted region. Experimental results validate the effectiveness of P 3 VAD on the four benchmark datasets (ShanghaiTech, CUHK Avenue, IITB‐Corridor, and ADOC).
Xianlin Zeng, Yalong Jiang, Yufeng Wang 0004, Wenrui Ding
IET Image Process.2
2024 Reasonable Anomaly Detection Based on Long-Term Sequence Modeling
abstract
Video anomaly detection is a challenging task due to the unpredictable nature of abnormal actions, sophisticated semantics and a lack in training data. The visual representations of most existing approaches are limited by short-term sequences which cannot provide necessary clues for achieving reasonable detections. In this paper, we propose to comprehensively represent the motion patterns in human actions by learning from long-term sequences. Firstly, a Stacked State Machine (SSM) model with distinctive basis functions is proposed to represent the temporal dependencies which are consistent across long-term observations. Secondly, the dependencies are leveraged in filtering out problematic motion estimations which are influenced by short-term observation noises, plausible motion parameters are obtained in this way. Finally, SSM model predicts future states based on past ones, the divergence between the predictions with inherent normal patterns and observed ones determines anomalies which violate normal motion patterns. To address the challenges in drone-based surveillance, a dataset which is more diversified than existing ones is built. Extensive experiments are carried out to evaluate the proposed approach on the dataset and existing ones. Improvements over state-of-the-art methods can be observed. The proposed dataset will be made publicly available. Code is available athttps://github.com/AllenYLJiang/Anomaly-Detection-in-Sequences.
Yalong Jiang, Changkang Li, Wenrui Ding, Jinzhi Xiang, Zheru Chi
IEEE Trans. Circuits Syst. Video Technol.1
2024 Multi-Person Pose Tracking With Sparse Key-Point Flow Estimation and Hierarchical Graph Distance Minimization
abstract
In this paper, we propose a novel framework for multi-person pose estimation and tracking on challenging scenarios. In view of occlusions and motion blurs which hinder the performance of pose tracking, we proposed to model humans as graphs and perform pose estimation and tracking by concentrating on the visible parts of human bodies which are informative about complete skeletons under incomplete observations. Specifically, the proposed framework involves three parts: (i) A Sparse Key-point Flow Estimating Module (SKFEM) and a Hierarchical Graph Distance Minimizing Module (HGMM) for estimating pixel-level and human-level motion, respectively; (ii) Pixel-level appearance consistency and human-level structural consistency are combined in measuring the visibility scores of body joints. The scores guide the pose estimator to predict complete skeletons by observing high-visibility parts, under the assumption that visible and invisible parts are inherently correlated in human part graphs. The pose estimator is iteratively fine-tuned to achieve this capability; (iii) Multiple historical frames are combined to benefit tracking which is implemented using HGMM. The proposed approach not only achieves state-of-the-art performance on PoseTrack datasets but also contributes to significant improvements in other tasks such as human-related anomaly detection.
Yalong Jiang, Wenrui Ding, Zheru Chi
IEEE Trans. Image Process.1
2023 A Physically Explainable Framework for Human-Related Anomaly Detection
abstract
Due to the complexity in understanding human behaviors under limited observations and insufficient training data, video anomaly detection is challenging. Most of existing approaches solely rely on visual clues and suffer from noisy observations which easily lead to unreasonable predictions. In this paper, we introduce physically explainable dynamics to enhance visual representations. Firstly, a Physical Intuition (PI) module is proposed to be combined with a Visual Representation (VR) module in estimating the forces applied on subjects. Secondly, a hierarchical structure is proposed to facilitate PI module in achieving physically plausible descriptions of human movements while maintaining the consistency with visual representations. Thirdly, a novel anomaly score is proposed considering the distributions of forces. Extensive experimental results on five benchmark datasets show that state-of-the-art performance can be achieved by the proposed framework with a strong robustness.
Yalong Jiang, Huining Li, Changkang Li
ICASSP1
2023 A Hierarchical Spatio-Temporal Graph Convolutional Neural Network for Anomaly Detection in Videos
abstract
Deep learning models have been widely used for anomaly detection in surveillance videos. Typical models are equipped with the capability to reconstruct normal videos and evaluate the reconstruction errors on anomalous videos to indicate the extent of abnormalities. However, existing approaches suffer from two disadvantages. Firstly, they can only encode the movements of each identity independently, without considering the interactions among identities which may also indicate anomalies. Secondly, they leverage inflexible models whose structures are fixed under different scenes, this configuration disables the understanding of scenes. In this paper, we propose a Hierarchical Spatio-Temporal Graph Convolutional Neural Network (HSTGCNN) to address these problems, the HSTGCNN is composed of multiple branches that correspond to different levels of graph representations. High-level graph representations encode the trajectories of people and the interactions among multiple identities while low-level graph representations encode the local body postures of each person. Furthermore, we propose to weightedly combine multiple branches that are better at different scenes. An improvement over single-level graph representations is achieved in this way. An understanding of scenes is achieved and serves anomaly detection. High-level graph representations are assigned higher weights to encode moving speed and directions of people in low-resolution videos while low-level graph representations are assigned higher weights to encode human skeletons in high-resolution videos. Experimental results show that the proposed HSTGCNN significantly outperforms current state-of-the-art models on four benchmark datasets (UCSD Pedestrian, ShanghaiTech, CUHK Avenue and IITB-Corridor) by using much less learnable parameters.
Xianlin Zeng, Yalong Jiang, Wenrui Ding, Yafeng Hao, Zifeng Qiu
IEEE Trans. Circuits Syst. Video Technol.2
2022 Towards Accurate Binary Neural Networks via Modeling Contextual Dependencies
Xingrun Xing, Yangguang Li 0001, Wei Li 0022, Wenrui Ding, Yalong Jiang, Yufeng Wang 0004, Chunlei Liu 0001, Xianglong Liu 0001
ECCV (11)5
2022 Binary Dense Predictors for Human Pose Estimation Based on Dynamic Thresholds and Filtering
abstract
Binary neural networks (BNNs) contribute a lot to the efficiency of image classification models. However, in dense predication tasks such as human pose estimation, predictions in different locations are coupled and rely on the extraction of features across entire images. As a result, more robust and adaptive binarization is required to bridge the performance gap between binarized and full precision models. We propose two approaches to conduct image-aware and pixel-aware dynamic binarization in a model for human pose estimation. Firstly, a simplified dynamic thresholding is leveraged in the backbone to determine unique binarization thresholds for each image. Secondly, in the decoder, we decouple binarization for each pixel according to the activations surrounding the pixel. Dynamic filtering modules are proposed to determine a different binarization strategy for each pixel. Compared with the strong baselines, the proposed framework improves 5.2% and 3.6% mAP on the COCO test-dev benchmark for ResNet-18/34 architectures respectively.
Xingrun Xing, Yalong Jiang, Baochang Zhang 0001, Wenrui Ding, Huan Peng
ICASSP2
2020 Cam-Net: Compressed Attentive Multi-Granularity Network For Dynamic Scene Classification
abstract
Dynamic scene classification on portable platforms is extremely challenging due to the contradiction between model complexity and computing resources. To resolve this long-standing dilemma, we propose the compressed attentive multi-granularity network (CAM-Net) in a two-step manner. First, we present a novel AM-Net based on multi-granularity attention units to boost the performance of the full-precision model. It captures and enhances both coarse and fine target-related information. Then, we introduce an efficient binary approximation to AM-Net to improve computing efficiency, leading to CAM-Net. Particularly, a grouping guidance approach is adopted to guide the reconstruction of full-precision weights from binary ones. With this guidance, CAM-Net can significantly reduce memory usage as well as CPU consumption, yet only cause a slight decline in accuracy. Extensive experiments have been conducted on three benchmark datasets, i.e., Maryland, YUPENN ++ and ActivityNet, demonstrating the effectiveness and superiority of the proposed method on scene classification.
Wenrui Ding, Yanjun Zhu, Yuanjun Huang, Yalong Jiang, Baochang Zhang 0001
ICIP5
2019 A CNN Model for Semantic Person Part Segmentation With Capacity Optimization
abstract
In this paper, a deep learning model with an optimal capacity is proposed to improve the performance of person part segmentation. Previous efforts in optimizing the capacity of a CNN model suffer from a lack of large datasets as well as the over-dependence on a single-modality CNN which is not effective in learning. We make several efforts in addressing these problems. Firstly, other datasets are utilized to train a CNN module for pre-processing image data and a segmentation performance improvement is achieved without a time-consuming annotation process. Secondly, we propose a novel way of integrating two complementary modules to enrich the feature representations for more reliable inferences. Thirdly, the factors to determine the capacity of a CNN model are studied and two novel methods are proposed to adjust (optimize) the capacity of a CNN to match it to the complexity of a task. The over-fitting and under-fitting problems are eased by using our methods. Experimental results show that our model outperforms the state-of-the-art deep learning models with a better generalization ability and a lower computational complexity.
Yalong Jiang, Zheru Chi
IEEE Trans. Image Process.1
2018 Person Part Segmentation based on Weak Supervision
Yalong Jiang, Zheru Chi
BMVC1
2018 A Novel Structure of Convolutional Layers with a Higher Performance-Complexity Ratio for Semantic Segmentation
abstract
In this paper, we study an important factor that determines the capacity of a CNN model and propose a novel structure of convolutional layers with a higher performance-complexity ratio. Firstly, the relationship of the model capacity and the number of parameters versus segmentation performance is explored. Secondly, a mechanism is proposed to optimize the structure of a CNN model for a specific task. The mechanism also provides better convergence than current state-of-the-art methods for factorizing convolutional layers, such as MobileNet. Thirdly, we propose a measure based on the mutual information between hidden activations and inputs/outputs to compute the capacity of a CNN model. This measure is highly correlated with segmentation performance. Experimental results on the segmentation of the PASCAL Person Parts Dataset show that the linear dependency among convolutional kernels is an important factor determining the capacity of a CNN model. It is also demonstrated that our approach can successfully adjust the model capacity to best match to the complexity of a dataset. The optimized CNN model achieves the similar performance to Deeplab-V2 on the segmentation task with 100 × less parameters, resulting in a significantly improved performance-complexity ratio.
Yalong Jiang, Zheru Chi
ICARCV1
2017 A scale-invariant framework for image classification with deep learning
abstract
In this paper, we propose a scale-invariant framework based on Convolutional Neural Networks (CNNs). The network exhibits robustness to scale and resolution variations in data. Previous efforts in achieving scale invariance were made on either integrating several variant-specific CNNs or data augmentation. However, these methods did not solve the fundamental problem that CNNs develop different feature representations for the variants of the same image. The topology proposed by this paper develops a uniform representation for each of the variants of the same image. The uniformity is acquired by concatenating scale-variant and scale-invariant features to enlarge the feature space so that the case when input images are of diverse variations but from the same class can be distinguished from another case when images are of different classes. Higher-order decision boundaries lead to the success of the framework. Experimental results on a challenging dataset substantiates that our framework performs better than traditional frameworks with the same number of free parameters. Our proposed framework can also achieve a higher training efficiency.
Yalong Jiang, Zheru Chi
SMC1