EDBT 2026 Demo / reviewers in the wild / expert
Yang Liu 0246
dblp:51/3710-246
· DBLP profile ↗
60ranked-venue papers
13as first author
58since 2021 · last 2026
0000-0002-1312-0146ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 8 first-author · 30 since 2021Artificial intelligence and machine learning · 28 · 5 first-author · 26 since 2021Computer networks · 5 · 5 since 2021Systems, architecture and hardware · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Projecting to Consensus: Communication-Efficient Collaborative Learning Across Heterogeneous Networks
Jing Liu 0050, Yao Du 0001, Yang Liu 0246, Zehua Wang 0001, Peng Sun 0007, Victor C. M. Leung |
ICC | 3 |
| 2026 | MeritFL: Self-Regulating Federated Learning via Merit-Gated Communication
Zhengliang Guo, Kun Yang 0010, Linxiao Gong, Yang Liu 0246, Jing Liu 0050 |
ISCAS | 4 |
| 2026 | Cross-modal attention fusion of RGB and skeleton for multimodal-driven video anomaly detection
Boan Chen, Weide Liu, Jinmei Liu, Baoquan Zhao, Yang Liu 0246 |
Pattern Recognit. | 6 |
| 2026 | Multimodal human video generation with uncertainty-aware pose guidance
Kun Yang 0010, Yuanyuan Meng, Juncen Guo, Yanda Meng, Songwen Pei, Jing Liu 0050, Yang Liu 0246 |
Pattern Recognit. | 9 |
| 2026 | Condition-Dependent Causal Discovery: A Polynomial Chaos Framework for Systems With Parametric UncertaintyabstractIdentifying causal structures within industrial processes is vital for effective optimization and control, yet it faces significant challenges due to fluctuating operational conditions. Conventional causal discovery techniques are constrained by their reliance on static relationship assumptions, rendering them inadequate for real-world scenarios where causal dynamics shift with parameters such as temperature and pressure. Furthermore, current uncertainty-aware methods typically focus solely on epistemic uncertainty arising from limited data, neglecting the functional dependency of causal strengths on measurable system parameters. To address this, we propose physics-informed polynomial chaos theory for causal discovery (PIPCT-CD), a novel framework that explicitly models causal edges as polynomial functions of operating parameters. By introducing a novel polynomial chaos expansion-based conditional independence test with a robust score-based learning strategy, PIPCT-CD accurately detects these dynamic interactions and quantifies associated uncertainties. Comprehensive validation on industrial-scale process networks demonstrates that PIPCT-CD achieves an F1-score of 0.800 on an electrical distribution system and 0.632 on a chemical refinery process, outperforming established baseline methods. Weide Liu, Yang Liu 0246 |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2026 | Channel-Independence for Traffic Forecasting: A Cascaded Spatio-Temporal MLP FrameworkabstractThe criticality of efficient traffic forecasting in Intelligent Transportation System (ITS) has garnered significant academic attention. This study addresses the prevalent issue of distribution shift in real-world datasets, which often degrades performance, and explores the effectiveness of the channel-independence (CI), a technique recently proposed to mitigate this issue. While Spatio-Temporal Graph Neural Networks (STGNNs) are noted for their flexibility to represent road structures, their designs typically lack the capability to integrate CI without disrupting the spatial relationships, potentially limiting the performance. We present a novel approach that successfully integrates CI into spatial-temporal forecasting by incorporating distinct temporal, spatial, and predefined graph structure information within each channel. Moreover, STGNNs frequently emphasize intricate designs, which result in increased computational demands while offering only marginal improvements in accuracy. This paper presents ST-MLP, a streamlined spatio-temporal model constructed exclusively from cascaded Multi-Layer Perceptron (MLP) modules and linear layers. Experimental results indicate that ST-MLP outperforms numerous existing STGNNs in both accuracy and computational efficiency. Our findings advocate for further investigation into more streamlined and effective neural network architectures within spatial-temporal forecasting research. Zepu Wang, Yuqi Nie, Yang Liu 0246, John M. Mulvey, H. Vincent Poor, Azzedine Boukerche, Nam H. Nguyen, Peng Sun 0007 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2026 | STNMamba: Mamba-Based Spatial-Temporal Normality Learning for Video Anomaly DetectionabstractVideo anomaly detection (VAD) has been extensively researched due to its potential for intelligent video systems. However, most existing methods based on CNNs and transformers still suffer from substantial computational burdens and have room for improvement in learning spatial-temporal normality. Recently, Mamba has shown great potential for modeling long-range dependencies with linear complexity, providing an effective solution to the above dilemma. To this end, we propose a lightweight and effective Mamba-based network named STNMamba, which incorporates carefully designed Mamba modules to enhance the learning of spatial-temporal normality. Firstly, we develop a dual-encoder architecture, where the spatial encoder equipped with Multi-Scale Vision Space State Blocks (MS-VSSB) extracts multi-scale appearance features, and the temporal encoder employs Channel-Aware Vision Space State Blocks (CA-VSSB) to capture significant motion patterns. Secondly, a Spatial-Temporal Interaction Module (STIM) is introduced to integrate spatial and temporal information across multiple levels, enabling effective modeling of intrinsic spatial-temporal consistency. Within this module, the Spatial-Temporal Fusion Block (STFB) is proposed to fuse the spatial and temporal features into a unified feature space, and the memory bank is utilized to store spatial-temporal prototypes of normal patterns, restricting the model's ability to represent anomalies. Extensive experiments on three benchmark datasets demonstrate that our STNMamba achieves competitive performance with fewer parameters and lower computational costs than existing methods. Zhangxun Li, Mengyang Zhao 0002, Yang Liu 0246, Jiamu Sheng, Xinhua Zeng, Tian Wang 0002, Kewei Wu, Yu-Gang Jiang 0001 |
IEEE Trans. Multim. | 4 |
| 2026 | Privacy-Preserving Video Anomaly Detection: A SurveyabstractThe video anomaly detection (VAD) aims to automatically analyze spatiotemporal patterns in surveillance videos collected from open spaces to detect anomalous events that may cause harm, such as fighting, stealing, and car accidents. However, vision-based surveillance systems such as closed-circuit television (CCTV) often capture personally identifiable information. The lack of transparency and interpretability in video transmission and usage raises public concerns about privacy and ethics, limiting the real-world application of VAD. Recently, researchers have focused on privacy concerns in VAD by conducting systematic studies from various perspectives, including data, features, and systems, making privacy-preserving VAD (P2VAD) a hotspot in the AI community. However, the current research in P2VAD is fragmented, and prior reviews have mostly focused on methods using RGB sequences, overlooking privacy leakage and appearance bias considerations. To address this gap, this article is the first to systematically review the progress of P2VAD, defining its scope and providing an intuitive taxonomy. We outline the basic assumptions, learning frameworks, and optimization objectives of various approaches, analyzing their strengths, weaknesses, and potential correlations. In addition, we provide open access to research resources such as benchmark datasets and available code. Finally, we discuss key challenges and future opportunities from the perspectives of AI development and P2VAD deployment, aiming to the guide future work in the field. Yang Liu 0246, Siao Liu, Xiaoguang Zhu, Hao Yang 0055, Juncen Guo, Liangyu Teng, Dingkang Yang, Yan Wang 0068, Jing Liu 0050 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2026 | A New Semi-Supervised Video Anomaly Detection Baseline in Lack of Anomalous SamplesabstractVideo anomaly detection (VAD) has been widely studied for its important applications in multimedia community. Recently, many Weakly Supervised VAD (WS-VAD) methods have been proposed, which tend to treat VAD as a classification task through multiple instance learning and result in the need to collect sufficient anomaly classes and samples to be used for training a classifier. However, anomaly events tend to be open-set and rare in real-world applications, so we often have difficulty collecting all anomaly classes and enough sample anomalies, which is a difficult situation for WS-VAD to cope with. To this end, we consider to treat VAD as an out-of-distribution detection task rather than a classification task and propose a simple but effective semi-supervised baseline method. First, we leverage the powerful zero-shot capability of large visual language models to generate summary text descriptions for videos and extract visual features as intermediates for subsequent use. Next, we use a text encoder to extract language features and combine them with visual features to obtain robust multimodal features. Finally, we introduce an out-of-distribution detection method that learns the center of normality in multimodal space from normal and unlabeled samples, while deviating abnormal samples from the center to cope with the scarcity of abnormal samples. To implement our baseline method, we also provide a new semi-supervised dataset by reorganizing an existing benchmark, which is the first available dataset in the VAD community that provides trimmed videos consisting of complete abnormal events. Experiments demonstrate that our method performs more robustly when fewer anomaly classes and anomaly samples collected. Mengyang Zhao 0002, Haiyang Yu 0004, Teng Fu 0001, Yang Liu 0246, Wei Zhou 0021, Bin Li 0015, Xiangyang Xue 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | SVR-YOLO11: Real-Time Animal Detection for Situational Awareness in SAR OperationsabstractReal-time animal detection in search-and-rescue (SAR) operations represents a critical challenge for situational awareness systems, particularly when deploying lightweight solutions on resource-constrained edge computing platforms. Current detection methods suffer from computational bottlenecks that compromise either accuracy or real-time performance, limiting their effectiveness in time-critical rescue scenarios. This paper presents SVR-YOLO11, an enhanced detection network specifically optimized for real-time animal identification in SAR applications. The proposed architecture introduces three key in-novations: (1) a Slim-Neck design paradigm that preserves inter-channel connections while reducing computational complexity, (2) integration of Large Separable Kernel Attention (LSKA) modules with dynamic upsampling to improve detection accuracy for morphologically similar animal species across diverse natural backgrounds, and (3) direct deployment capability on edge devices without requiring complex model conversion processes. Experimental validation on the Animals. v4i dataset demonstrates that SVR-YOLO11 achieves superior performance with 68.3% mAP50-95 while maintaining only 2.843M parameters and 6.5 GFLOPs computational cost. Real-time testing on multiple edge computing platforms, including Raspberry Pi and NVIDIA Jetson devices, confirms the network's practical applicability for field deployment. The system's effectiveness is further validated through a comprehensive head-mounted display demonstration using Unity Engine simulations that replicate dynamic SAR environments. These results establish SVR-YOLO11 as a ro-bust solution for real-time animal detection in distributed edge computing scenarios, directly supporting enhanced situational awareness in critical rescue operations. Zhongyi He, Yang Liu 0246, Hao Yang 0055, Peng Sun 0007, Boan Chen |
DS-RT | 2 |
| 2025 | Uncertainty-Aware Crime Prediction With Spatial Temporal Multivariate Graph Neural NetworksabstractCrime prediction (CP) plays a pivotal role in urban analytics, contributing significantly to personal safety and societal stability. Unlike conventional time series forecasting, CP faces unique difficulties due to the inherent sparsity of crime incidents, particularly within small spatial regions and limited time windows. This sparsity, coupled with the non-Gaussian distribution of crime data—characterized by an excess of zero events and over-dispersion—presents a critical challenge for the signal processing community. In this regard, we propose a novel framework, Spatial-Temporal Multivariate Zero-Inflated Negative Binomial Graph Neural Networks (STMGNN-ZINB), which integrates diffusion and convolutional graph networks to capture spatial, temporal, and multivariate dependencies. By leveraging a Zero-Inflated Negative Binomial distribution, the STMGNN-ZINB effectively models the over-dispersed and zero-heavy nature of crime data, significantly improving both prediction accuracy and confidence interval estimation. Experimental results on real-world datasets demonstrate that our STMGNN-ZINB outperforms state-of-the-art CP methods, offering a robust tool for crime early warning and explicable insights into urban crime dynamics. Zepu Wang, Huajie Yang, Weimin Lyu, Yang Liu 0246, Peng Sun 0007, Sharath Chandra Guntuku |
ICASSP | 5 |
| 2025 | Towards Advanced Emotional Care: Embodied Emotional Care System for Humanoid RobotsabstractIn modern healthcare, emotional well-being is critical to patient recovery and overall outcomes. However, limited availability of trained professionals and time constraints often hinder the delivery of consistent emotional support. To address this gap, we propose the Embodied Emotional Care System (EECS), a comprehensive humanoid robotic framework designed to deliver personalized emotional care through an integrated, multi-layered architecture. EECS analyzes dynamic facial expressions and real-time vocal inputs to extract the patient’s emotional state and semantic information, constructs context-aware prompts processed by an LLM for reasoning, and ultimately generates empathetic dialogues synchronized with human-like facial expressions and natural body movements to address diverse emotional support needs. Experimental results show that deploying EECS on a humanoid robot significantly boosts patient engagement through real-time multimodal interaction, delivering deeper emotional support and a more human-like therapeutic experience. Furthermore, it bridges gaps in professional emotional support resources, offering a feasible pathway to improve overall healthcare quality. Yang Chang, Aoxing Li, Yuxuan Lin 0001, Lizheng Liu, Yang Liu 0246, Jing Liu 0050, Yan Wang 0068, Zhongxue Gan 0001 |
ICME | 6 |
| 2025 | Identify, Isolate, and Purge: Mitigating Hallucinations in LVLMs via Self-Evolving DistillationabstractLarge Vision-Language Models (LVLMs) have demonstrated remarkable advancements in numerous areas such as multimedia. However, hallucination issues significantly limit their credibility and application potential. Existing mitigation methods typically rely on external tools or the comparison of multi-round inference, which significantly increase inference time. In this paper, we propose SElf-Evolving Distillation (SEED), which identifies hallucinations within the inner knowledge of LVLMs, isolates and purges them, and then distills the purified knowledge back into the model, enabling self-evolution. Furthermore, we identified that traditional distillation methods are prone to inducing void spaces in the output space of LVLMs. To address this issue, we propose a Mode-Seeking Evolving approach, which performs distillation to capture the dominant modes of the purified knowledge distribution, thereby avoiding the chaotic results that could emerge from void spaces. Moreover, we introduce a Hallucination Elimination Adapter, which corrects the dark knowledge of the original model by learning purified knowledge. Extensive experiments on multiple benchmarks validate the superiority of our SEED, demonstrating substantial improvements in mitigating hallucinations for representative LVLM models such as LLaVA-1.5 and InternVL2. Remarkably, the F1 score of LLaVA-1.5 on the hallucination evaluation metric POPE-Random improved from 81.3 to 88.3. Xiu Su, Yang Liu 0246, Shan You, Chang Xu 0002 |
ACM Multimedia | 5 |
| 2025 | DanmakuTPPBench: A Multi-modal Benchmark for Temporal Point Process Modeling and UnderstandingabstractWe introduce DanmakuTPPBench, a comprehensive benchmark designed to advance multi-modal Temporal Point Process (TPP) modeling in the era of Large Language Models (LLMs). While TPPs have been widely studied for modeling temporal event sequences, existing datasets are predominantly unimodal, hindering progress in models that require joint reasoning over temporal, textual, and visual information. To address this gap, DanmakuTPPBench comprises two complementary components:(1) DanmakuTPP-Events, a novel dataset derived from the Bilibili video platform, where user-generated bullet comments (Danmaku) naturally form multi-modal events annotated with precise timestamps, rich textual content, and corresponding video frames;(2) DanmakuTPP-QA, a challenging question-answering dataset constructed via a novel multi-agent pipeline powered by state-of-the-art LLMs and multi-modal LLMs (MLLMs), targeting complex temporal-textual-visual reasoning. We conduct extensive evaluations using both classical TPP models and recent MLLMs, revealing significant performance gaps and limitations in current methods’ ability to model multi-modal event dynamics. Our benchmark establishes strong baselines and calls for further integration of TPP modeling into the multi-modal language modeling landscape. Project page: https://github.com/FRENKIE-CHIANG/DanmakuTPPBench. Jichu Li, Yang Liu 0246, Dingkang Yang, Quyu Kong |
NeurIPS | 3 |
| 2025 | M2S2L: Mamba-based Multi-Scale Spatial-temporal Learning for Video Anomaly DetectionabstractVideo anomaly detection (VAD) is an essential task in the image processing community with prospects in video surveillance, which faces fundamental challenges in balancing detection accuracy with computational efficiency. As video content becomes increasingly complex with diverse behavioral patterns and contextual scenarios, traditional VAD approaches struggle to provide robust assessment for modern surveillance systems. Existing methods either lack comprehensive spatial-temporal modeling or require excessive computational resources for real-time applications. In this regard, we present a Mamba-based multi-scale spatial-temporal learning (M2S2L) framework in this paper. The proposed method employs hierarchical spatial encoders operating at multiple granularities and multi-temporal encoders capturing motion dynamics across different time scales. We also introduce a feature decomposition mechanism to enable task-specific optimization for appearance and motion reconstruction, facilitating more nuanced behavioral modeling and quality-aware anomaly assessment. Experiments on three benchmark datasets demonstrate that M2S2L framework achieves 98.5%, 92.1%, and 77.9% frame-level AUCs on UCSD Ped2, CUHK Avenue, and ShanghaiTech respectively, while maintaining efficiency with 20.1G FLOPs and 45 FPS inference speed, making it suitable for practical surveillance deployment. Yang Liu 0246, Boan Chen, Xiaoguang Zhu, Jing Liu 0050, Peng Sun 0007, Wei Zhou 0013 |
VCIP | 1 |
| 2025 | CNN-DAG-Editor: A Convolutional Neural Network offloading analyzer with Multi-Objective Dynamic Adaptive Resource Competitive Swarm Optimization
Bobo Ju, Yang Liu 0246, Jing Liu 0050, Peng Sun 0007 |
Comput. Networks | 2 |
| 2025 | Rethinking prediction-based video anomaly detection from local-global normality perspective
Mengyang Zhao 0002, Xinhua Zeng, Yang Liu 0246, Jing Liu 0050, Chengxin Pang |
Expert Syst. Appl. | 3 |
| 2025 | CRCL: Causal Representation Consistency Learning for Anomaly Detection in Surveillance VideosabstractVideo Anomaly Detection (VAD) remains a fundamental yet formidable task in the video understanding community, with promising applications in areas such as information forensics and public safety protection. Due to the rarity and diversity of anomalies, existing methods only use easily collected regular events to model the inherent normality of normal spatial-temporal patterns in an unsupervised manner. Although such methods have made significant progress benefiting from the development of deep learning, they attempt to model the statistical dependency between observable videos and semantic labels, which is a crude description of normality and lacks a systematic exploration of its underlying causal relationships. Previous studies have shown that existing unsupervised VAD models are incapable of label-independent data offsets (e.g., scene changes) in real-world scenarios and may fail to respond to light anomalies due to the overgeneralization of deep neural networks. Inspired by causality learning, we argue that there exist causal factors that can adequately generalize the prototypical patterns of regular events and present significant deviations when anomalous instances occur. In this regard, we propose Causal Representation Consistency Learning (CRCL) to implicitly mine potential scene-robust causal variable in unsupervised video normality learning. Specifically, building on the structural causal models, we propose scene-debiasing learning and causality-inspired normality learning to strip away entangled scene bias in deep representations and learn causal video normality, respectively. Extensive experiments on benchmarks validate the superiority of our method over conventional deep representation learning. Moreover, ablation studies and extension validation show that the CRCL can cope with label-independent biases in multi-scene settings and maintain stable performance with only limited training data available. Yang Liu 0246, Hongjin Wang, Zepu Wang, Xiaoguang Zhu, Jing Liu 0050, Peng Sun 0007, Jianwei Du, Victor C. M. Leung |
IEEE Trans. Image Process. | 1 |
| 2024 | De-Confounded Data-Free Knowledge Distillation for Handling Distribution ShiftsabstractData-Free Knowledge Distillation (DFKD) is a promising task to train high-performance small models to enhance actual deployment without relying on the original training data. Existing methods commonly avoid relying on private data by utilizing synthetic or sampled data. However, a long-overlooked issue is that the severe distribution shifts between their substitution and original data, which mani-fests as huge differences in the quality of images and class proportions. The harmful shifts are essentially the con-founder that significantly causes performance bottlenecks. To tackle the issue, this paper proposes a novel perspective with causal inference to disentangle the student models from the impact of such shifts. By designing a customized causal graph, we first reveal the causalities among the variables in the DFKD task. Subsequently, we propose a Knowledge Distillation Causal Intervention (KDCI) framework based on the backdoor adjustment to de-confound the confounder. KDCI can be flexibly combined with most existing state-of-the-art baselines. Experiments in combination with six representative DFKD methods demonstrate the effectiveness of our KDCI, which can obviously help existing methods under almost all settings, e.g., improving the base-line by up to 15.54% accuracy on the CIFAR-100 dataset. Dingkang Yang, Zhaoyu Chen 0001, Yang Liu 0246, Siao Liu, Lihua Zhang 0002, Lizhe Qi |
CVPR | 4 |
| 2024 | Towards Multimodal Sentiment Analysis Debiasing via Bias Purification
Dingkang Yang, Mingcheng Li, Dongling Xiao, Yang Liu 0246, Kun Yang 0010, Zhaoyu Chen 0001, Peng Zhai, Ke Li 0015, Lihua Zhang 0002 |
ECCV (58) | 4 |
| 2024 | Denoising Diffusion-Augmented Hybrid Video Anomaly Detection via Reconstructing Noised Frames
Kai Cheng 0001, Yaning Pan, Yang Liu 0246, Xinhua Zeng |
IJCAI | 3 |
| 2024 | Sampling to Distill: Knowledge Transfer from Open-World DataabstractData-Free Knowledge Distillation (DFKD) is a novel task that aims to train high-performance student models using only the pre-trained teacher network without original training data. Most of the existing DFKD methods rely heavily on additional generation modules to synthesize the substitution data resulting in high computational costs and ignoring the massive amounts of easily accessible, low-cost, unlabeled open-world data. Meanwhile, existing methods ignore the domain shift issue between the substitution data and the original data, resulting in knowledge from teachers not always trustworthy and structured knowledge from data becoming a crucial supplement. To tackle the issue, we propose a novel Open-world Data Sampling Distillation (ODSD) method for the DFKD task without the redundant generation process. First, we try to sample open-world data close to the original data's distribution by an adaptive sampling module and introduce a low-noise representation to alleviate the domain shift issue. Then, we build structured relationships of multiple data examples to exploit data knowledge through the student model itself and the teacher's structured representation. Extensive experiments on CIFAR-10, CIFAR-100, NYUv2, and ImageNet show that our ODSD method achieves state-of-the-art performance with lower FLOPs and parameters. Especially, we improve 1.50%-9.59% accuracy on the ImageNet dataset and avoid training the separate generator for each class. Zhaoyu Chen 0001, Jie Zhang 0107, Dingkang Yang, Zuhao Ge, Yang Liu 0246, Siao Liu, Yunquan Sun, Lizhe Qi |
ACM Multimedia | 6 |
| 2024 | Toward Robust Incomplete Multimodal Sentiment Analysis via Hierarchical Representation LearningabstractMultimodal Sentiment Analysis (MSA) is an important research area that aims to understand and recognize human sentiment through multiple modalities. The complementary information provided by multimodal fusion promotes better sentiment analysis compared to utilizing only a single modality. Nevertheless, in real-world applications, many unavoidable factors may lead to situations of uncertain modality missing, thus hindering the effectiveness of multimodal modeling and degrading the model’s performance. To this end, we propose a Hierarchical Representation Learning Framework (HRLF) for the MSA task under uncertain missing modalities. Specifically, we propose a fine-grained representation factorization module that sufficiently extracts valuable sentiment information by factorizing modality into sentiment-relevant and modality-specific representations through crossmodal translation and sentiment semantic reconstruction. Moreover, a hierarchical mutual information maximization mechanism is introduced to incrementally maximize the mutual information between multi-scale representations to align and reconstruct the high-level semantics in the representations. Ultimately, we propose a hierarchical adversarial learning mechanism that further aligns and adapts the latent distribution of sentiment-relevant representations to produce robust joint multimodal representations. Comprehensive experiments on three datasets demonstrate that HRLF significantly improves MSA performance under uncertain modality missing cases. Mingcheng Li, Dingkang Yang, Yang Liu 0246, Shunli Wang 0001, Jiawei Chen 0012, Shuaibing Wang, Jinjie Wei, Qingyao Xu, Xiaolu Hou, Ziyun Qian, Dongliang Kou, Lihua Zhang 0002 |
NeurIPS | 3 |
| 2024 | Memory-enhanced appearance-motion consistency framework for video anomaly detection
Zhiyuan Ning 0002, Zile Wang, Yang Liu 0246, Jing Liu 0050 |
Comput. Commun. | 3 |
| 2024 | Memory-enhanced spatial-temporal encoding framework for industrial anomaly detection system
Yang Liu 0246, Bobo Ju, Dingkang Yang, Liyuan Peng, Peng Sun 0007, Chengfang Li, Hao Yang 0055, Jing Liu 0050 |
Expert Syst. Appl. | 1 |
| 2024 | Normality learning reinforcement for anomaly detection in surveillance videos
Kai Cheng 0001, Xinhua Zeng, Yang Liu 0246, Yaning Pan, Xinzhe Li 0004 |
Knowl. Based Syst. | 3 |
| 2024 | DiffSkill: Improving Reinforcement Learning through diffusion-based skill denoiser for robotic manipulation
Siao Liu, Yang Liu 0246, Linqiang Hu, Ziqing Zhou, Zhile Zhao, Wei Li 0055, Zhongxue Gan 0001 |
Knowl. Based Syst. | 2 |
| 2024 | Decoding Silent Reading EEG Signals Using Adaptive Feature Graph Convolutional NetworkabstractDecoding silent reading Electroencephalography (EEG) signals is challenging because of its low signal-to-noise ratio. In addition, EEG signals are typically non-Euclidean structured, therefore merely using a two-dimensional matrix to represent the variation of sampling points of each channel in time cannot richly represent the spatial connection between channels. Furthermore, due to the individual differences in EEG signals, a fixed representation cannot adequately represent the temporal and spatial associations between channels in real time. In this letter, we use the feature matrix and its adaptive graph structure to represent each EEG signal. Then, we use them as inputs and propose a novel Adaptive Feature Graph Convolutional Network (AFGCN) to decode the silent reading EEG signals. We classify silent reading EEG signals under different tasks of 16 subjects from two publicly available datasets. The experimental results demonstrate that our proposed method achieves higher decoding accuracy than state-of-the-art EEG classification networks on both datasets. Among them, the highest classification accuracy for the four classes is 83.33%. The study could promote the application and development of BCI technology for silent reading EEG signal decoding. It can also provide an efficient and convenient communication method for patients with language impairment. Chengfang Li, Gaoyun Fang, Yang Liu 0246, Jing Liu 0050 |
IEEE Signal Process. Lett. | 3 |
| 2024 | AMP-Net: Appearance-Motion Prototype Network Assisted Automatic Video Anomaly Detection SystemabstractAs essential tools for industry safety protection, automatic video anomaly detection systems (AVADS) are designed to detect anomalous events of concern in surveillance videos. Existing VAD methods lack effective exploration of the prototypical appearance and motion features leading to poor performance in realistic scenarios. Specifically, they either misreport regular events as anomalies due to insufficient representation power, or lead to missed detections with over-power generalization. In this regard, we propose an appearance-motion prototype network (AMP-net) that uses external memories to record prototype features and augments the appearance-motion prototype with a spatial-temporal fusion. In addition, AMP-net sequentially fuses appearance features from deep to shallow to utilize multiscale spatial context. Additionally, we introduce temporal attention to capture important dynamics and enhance AMP-net for representing regular motion. The proposed method achieves a delicate balance of effective representation of normal events and limited generalization to anomalies. Experiments on three benchmark datasets demonstrate that our method can accurately detect anomalous events, achieving performance comparable to state-of-the-art methods with frame-level AUCs of 98.7%, 92.4%, and 78.8% on the UCSD Ped2, CUHK Avenue, and ShanghaiTech datasets. Moreover, we conducted a case study on the self-collected industrial dataset, and the results indicate that our AMP-net can cope with complex industrial scenarios and outperform existing methods. Yang Liu 0246, Jing Liu 0050, Kun Yang 0010, Bobo Ju, Siao Liu, Dingkang Yang, Peng Sun 0007 |
IEEE Trans. Ind. Informatics | 1 |
| 2023 | Surrogate-Assisted Evolution of Convolutional Neural Networks by Collaboratively Optimizing the Basic Blocks and TopologiesabstractConvolutional neural networks (CNNs) are prominent in many fields owing to their outstanding feature extraction abilities. Many excellent CNNs have been carefully designed by algorithm researchers; however, the design process is limited by the inherent knowledge of the researchers. Inspired by the existing successful block-based neural architecture search, we develop a collaboratively automatically evolutionary CNNs algorithm (CAE-CNN), which employs a surrogate-assisted genetic algorithm to search for a satisfactory CNN architecture by collaboratively optimizing the basic blocks in ResNet and DenseNet and their topologies. The encoding space of CAE-CNN consists of three basic units (pooling, ResNet, and DenseNet units) and a connection topology with a variable size. To effectively evolve the population, we design a double-module crossover operation and a multi-type mutation operation to collaboratively evolve the units and the topology. To address the problem of a rapidly increasing search space caused by the topological search, we use a random forest as the surrogate model to estimate the fitness of an individual to accelerate the search. CAE-CNN is a completely automatic algorithm in which a satisfactory CNN architecture can be obtained without any manual intervention. Experimental results show that CAE-CNN could archive competitive performance in terms of classification accuracy on four image classification datasets, and it consumes fewer computing resources than many algorithms. Jing Liu 0050, Yang Liu 0246 |
CEC | 3 |
| 2023 | Spatial-Temporal Graph Convolutional Network Boosted Flow-Frame Prediction For Video Anomaly DetectionabstractVideo Anomaly Detection (VAD) is a critical technology for intelligent surveillance systems and remains a challenging task in the signal processing community. An intuitive idea for VAD is to use a two-stream network to learn appearance and motion normality, respectively. However, existing approaches usually design a network architecture for the appearance stream with effort, then apply a similar architecture to the motion stream, ignoring the unique appearance and motion characteristics. In this paper, we propose STGCN-FFP, an unsupervised Spatial-Temporal Graph Convolutional Networks (STGCN) boosted Flow-Frame Prediction model. Specifically, we first design an STGCN-based memory module to extract and memorize normal patterns for optical flow, which is more suitable for learning motion normality. Then, we use a memory-augmented auto-encoder to model normal appearance patterns. Finally, the latent representation of two streams is fused to predict future frames, boosting the model to learn spatial-temporal normality. To our knowledge, STGCN-FFP is the first work applying STGCN to uniquely model the motion normality. Our method performs comparably to the state-of-the-art methods on three benchmarks. Kai Cheng 0001, Xinhua Zeng, Yang Liu 0246, Mengyang Zhao 0002, Chengxin Pang, Xing Hu 0006 |
ICASSP | 3 |
| 2023 | MSN-net: Multi-Scale Normality Network for Video Anomaly DetectionabstractExisting unsupervised video anomaly detection methods often suffer from performance degradation due to the overgeneralization of deep models. In this paper, we propose a simple yet effective Multi-Scale Normality network (MSN-net) that uses hierarchical memories to learn multi-level prototypical spatial-temporal patterns of normal events. Specifically, the hierarchical memory module interacts with the encoder through the reading and writing operations during the training phase, preserving multi-scale normality in three separate memory pools. Then, the decoder decodes the features rewritten by the memorized normality to predict future frames so that its ability to predict anomalies is diminished. Experimental results show that MSN-net performs comparably to the state-of-the-art methods, and extension analysis demonstrates the effectiveness of multi-scale normality learning. Yang Liu 0246, Dingkang Yang, Jing Liu 0050 |
ICASSP | 1 |
| 2023 | Adversarial Contrastive Distillation with Adaptive DenoisingabstractAdversarial Robustness Distillation (ARD) is a novel method to boost the robustness of small models. Unlike general adversarial training, its robust knowledge transfer can be less easily restricted by the model capacity. However, the teacher model that provides the robustness of knowledge does not always make correct predictions, interfering with the student’s robust performance. Besides, in the previous ARD methods, the robustness comes entirely from one-to-one imitation, ignoring the relationship between examples. To this end, we propose a novel structured ARD method called Contrastive Relationship DeNoise Distillation (CRDND). We design an adaptive compensation module to model the instability of the teacher. Moreover, we utilize the contrastive relationship to explore implicit robustness knowledge among multiple examples. Experimental results on multiple attack benchmarks show CRDND can transfer robust knowledge efficiently and achieves state-of-the-art performance. Zhaoyu Chen 0001, Dingkang Yang, Yang Liu 0246, Siao Liu, Lizhe Qi |
ICASSP | 4 |
| 2023 | Improving Generalization in Visual Reinforcement Learning via Conflict-aware Gradient Agreement AugmentationabstractLearning a policy with great generalization to unseen environments remains challenging but critical in visual reinforcement learning. Despite the success of augmentation combination in the supervised learning generalization, naively applying it to visual RL algorithms may damage the training efficiency, suffering from serve performance degradation. In this paper, we first conduct qualitative analysis and illuminate the main causes: (i) high-variance gradient magnitudes and (ii) gradient conflicts existed in various augmentation methods. To alleviate these issues, we propose a general policy gradient optimization framework, named Conflict-aware Gradient Agreement Augmentation (CG2A), and better integrate augmentation combination into visual RL algorithms to address the generalization bias. In particular, CG2A develops a Gradient Agreement Solver to adaptively balance the varying gradient magnitudes, and introduces a Soft Gradient Surgery strategy to alleviate the gradient conflicts. Extensive experiments demonstrate that CG2A significantly improves the generalization performance and sample efficiency of visual RL algorithms. Siao Liu, Zhaoyu Chen 0001, Yang Liu 0246, Dingkang Yang, Zhile Zhao, Ziqing Zhou, Xie Yi, Wei Li 0055, Zhongxue Gan 0001 |
ICCV | 3 |
| 2023 | AIDE: A Vision-Driven Multi-View, Multi-Modal, Multi-Tasking Dataset for Assistive Driving PerceptionabstractDriver distraction has become a significant cause of severe traffic accidents over the past decade. Despite the growing development of vision-driven driver monitoring systems, the lack of comprehensive perception datasets restricts road safety and traffic security. In this paper, we present an AssIstive Driving pErception dataset (AIDE) that considers context information both inside and outside the vehicle in naturalistic scenarios. AIDE facilitates holistic driver monitoring through three distinctive characteristics, including multi-view settings of driver and scene, multi-modal annotations of face, body, posture, and gesture, and four pragmatic task designs for driving understanding. To thoroughly explore AIDE, we provide experimental benchmarks on three kinds of baseline frameworks via extensive methods. Moreover, two fusion strategies are introduced to give new insights into learning effective multi-stream/modal representations. We also systematically investigate the importance and rationality of the key components in AIDE and benchmarks. The project link is https://github.com/ydk122024/AIDE. Dingkang Yang, Zhi Xu 0010, Shunli Wang 0001, Mingcheng Li, Yang Liu 0246, Kun Yang 0010, Zhaoyu Chen 0001, Yan Wang 0068, Jing Liu 0050, Peixuan Zhang, Peng Zhai, Lihua Zhang 0002 |
ICCV | 8 |
| 2023 | Spatio-Temporal Domain Awareness for Multi-Agent Collaborative PerceptionabstractMulti-agent collaborative perception as a potential application for vehicle-to-everything communication could significantly improve the perception performance of autonomous vehicles over single-agent perception. However, several challenges remain in achieving pragmatic information sharing in this emerging research. In this paper, we propose SCOPE, a novel collaborative perception frame-work that aggregates the spatio-temporal awareness characteristics across on-road agents in an end-to-end manner. Specifically, SCOPE has three distinct strengths: i) it considers effective semantic cues of the temporal context to enhance current representations of the target agent; ii) it aggregates perceptually critical spatial information from heterogeneous agents and overcomes localization errors via multi-scale feature interactions; iii) it integrates multi-source representations of the target agent based on their complementary contributions by an adaptive fusion paradigm. To thoroughly evaluate SCOPE, we consider both real-world and simulated scenarios of collaborative 3D object detection tasks on three datasets. Extensive experiments show the superiority of our approach and the necessity of the proposed components. The project link is https://ydk122024.github.io/SCOPE/. Kun Yang 0010, Dingkang Yang, Mingcheng Li, Yang Liu 0246, Jing Liu 0050, Hanqi Wang, Peng Sun 0007 |
ICCV | 5 |
| 2023 | Learning Causality-inspired Representation Consistency for Video Anomaly DetectionabstractVideo anomaly detection is an essential yet challenging task in the multimedia community, with promising applications in smart cities and secure communities. Existing methods attempt to learn abstract representations of regular events with statistical dependence to model the endogenous normality, which discriminates anomalies by measuring the deviations to the learned distribution. However, conventional representation learning is only a crude description of video normality and lacks an exploration of its underlying causality. The learned statistical dependence is unreliable for diverse regular events in the real world and may cause high false alarms due to over generalization. Inspired by causal representation learning, we think that there exists a causal variable capable of adequately representing the general patterns of regular events in which anomalies will present significant variations. Therefore, we design a causality-inspired representation consistency (CRC) framework to implicitly learn the unobservable causal variables of normality directly from available normal videos and detect abnormal events with the learned representation consistency. Extensive experiments show that the causality-inspired normality is robust to regular events with label-independent shifts, and the proposed CRC framework can quickly and accurately detect various complicated anomalies from real-world surveillance videos. Yang Liu 0246, Zhaoyang Xia, Mengyang Zhao 0002, Donglai Wei 0002, Siao Liu, Bobo Ju, Gaoyun Fang, Jing Liu 0050 |
ACM Multimedia | 1 |
| 2023 | DSDCLA: driving style detection via hybrid CNN-LSTM with multi-level attention fusion
Jing Liu 0050, Yang Liu 0246, Hanqi Wang |
Appl. Intell. | 2 |
| 2023 | A High-Reliability Edge-Side Mobile Terminal Shared Computing Architecture Based on Task Triple-Stage Full-Cycle MonitoringabstractEdge computing has emerged as a promising paradigm for addressing the challenges of latency, bandwidth, and energy consumption in the era of big data and the intelligent Internet of Things. However, the limited computing resources of edge devices and their vulnerability to failures pose significant challenges to the reliability and availability of edge computing systems. To this end, we propose a novel architecture for reliable edge computing that leverages the collective computing power of mobile edge devices in this article. Our architecture employs a task-oriented triple-stage monitoring mechanism to ensure system reliability. Moreover, we present a shared computing framework that allows edge devices to dynamically share their computing resources based on the current availability and workload. We evaluate the effectiveness of the proposed architecture with several computational tasks, including$\pi $calculation and video processing. The results show that our architecture achieves high reliability and availability while also improving the performance and energy efficiency of edge devices. Bobo Ju, Yang Liu 0246, Guixiang Gan, Zengwen Li, Linhua Jiang |
IEEE Internet Things J. | 2 |
| 2023 | Stochastic video normality network for abnormal event detection in surveillance videos
Yang Liu 0246, Dingkang Yang, Gaoyun Fang, Donglai Wei 0002, Mengyang Zhao 0002, Kai Cheng 0001, Jing Liu 0050 |
Knowl. Based Syst. | 1 |
| 2023 | Target and source modality co-reinforcement for emotion understanding from asynchronous multimodal sequences
Dingkang Yang, Yang Liu 0246, Can Huang 0002, Mingcheng Li, Kun Yang 0010, Yan Wang 0068, Peng Zhai, Lihua Zhang 0002 |
Knowl. Based Syst. | 2 |
| 2023 | Distributional and spatial-temporal robust representation learning for transportation activity recognition
Jing Liu 0050, Yang Liu 0246, Xiaoguang Zhu |
Pattern Recognit. | 2 |
| 2023 | Learning Graph Enhanced Spatial-Temporal Coherence for Video Anomaly DetectionabstractVideo Anomaly Detection (VAD) is a critical yet challenging task in the signal processing community. Since part abnormal events cannot be detected by analyzing spatial or temporal information alone, learning spatial-temporal coherence has been proven the key to effective VAD. To this end, we propose a Graph Enhanced Spatial-Temporal Attention (GESTA) to address unsupervised VAD by learning the spatial-temporal coherence of normal events. Firstly, we propose a Dynamic Graph Recurrent Neural Network (DGRNN) to extract the motion patterns. Then, we propose a Spatial-Temporal Attention Module (STAM) to better model spatial-temporal coherence by integrating the prototypical spatial and temporal information. Finally, the fused spatial-temporal features are fed into the decoder to predict future frames. In testing phase, the anomaly with irregular information will result in poor prediction results. Experiments on three benchmarks demonstrate that our GESTA performs comparably to the state-of-the-art methods, and extensive analysis proves the effectiveness of DGRNN and STAM. Kai Cheng 0001, Yang Liu 0246, Xinhua Zeng |
IEEE Signal Process. Lett. | 2 |
| 2023 | OSIN: Object-Centric Scene Inference Network for Unsupervised Video Anomaly DetectionabstractVideo Anomaly Detection (VAD) is an essential yet challenging task in the signal processing community, which aims to understand the spatial and temporal contextual interactions between objects and surrounding scenes to detect unexpected events in surveillance videos. However, existing unsupervised methods either use a single network to learn global prototype patterns without making a unique distinction between foreground objects and background scenes or try to strip objects from frames, ignoring that the essence of anomalies lies in unusual object-scene interactions. To this end, this letter proposes an Object-centric Scene Inference Network (OSIN) that uses a well-designed three-stream structure to learn both global scene normality and local object-specific normal patterns as well as explore the object-scene interactions using scene memory networks. Experimental results on three benchmark datasets demonstrate the effectiveness of the proposed OSIN model, which achieves frame-level AUCs of 91.7%, 79.6%, and 98.3% on the CUHK Avenue, ShanghaiTech, and UCSD Ped2 datasets, respectively. Yang Liu 0246, Zhengliang Guo, Jing Liu 0050, Chengfang Li |
IEEE Signal Process. Lett. | 1 |
| 2023 | Two-Stage Alignments Framework for Unsupervised Domain Adaptation on Time Series DataabstractUnsupervised Domain Adaptation (UDA) aims to free models from labeled information of target domain by minimizing the discrepancy of distributions between different domains. Most existing methods are designed to learn domain-invariant features either by domain discrimination or by matching lower-order moments. However, these methods are not robust due to the limited representation of statistical characteristics for non-Gaussian distributions and thus fail in domain matching. In addition, they often focus on matching distributions while not considering class decision boundaries between domains. To address these issues, we propose a novel Two-Stage Alignments Framework (TSAF) for UAD, which not only performs arbitrary-order moment matching to approximately characterize complex non-Gaussian distributions, but also utilizes domain-specific decision boundaries to align the probabilistic outputs of classifiers. Moreover, the reconstruction-based task is introduced to enhance the representation of the inherent characteristics for specific distribution. Extensive experiments on three real-world time series datasets demonstrate that: 1) our model evidently outperforms many state-of-the-art domain adaptation methods in cross-domain classification tasks; 2) TSAF can learn domain-invariant features efficiently. Xiaowei Xiang, Yang Liu 0246, Gaoyun Fang, Jing Liu 0050, Mengyang Zhao 0002 |
IEEE Signal Process. Lett. | 2 |
| 2022 | Emotion Recognition for Multiple Context Awareness
Dingkang Yang, Shunli Wang 0001, Yang Liu 0246, Peng Zhai, Liuzhen Su, Mingcheng Li, Lihua Zhang 0002 |
ECCV (37) | 4 |
| 2022 | Learning Task-Specific Representation for Video Anomaly Detection with Spatial-Temporal AttentionabstractThe automatic detection of abnormal events in surveillance videos with weak supervision has been formulated as a multiple instance learning task, which aims to localize the clips containing abnormal events temporally with the video-level labels. However, most existing methods rely on the features extracted by the pre-trained action recognition models, which are not discriminative enough for video anomaly detection. In this work, we propose a spatial-temporal attention mechanism to learn inter- and intra-correlations of video clips, and the boosted features are encouraged to be task-specific via the mutual cosine embedding loss. Experimental results on standard benchmarks demonstrate the effectiveness of the spatial-temporal attention, and our method achieves superior performance to the state-of-the-art methods. Yang Liu 0246, Jing Liu 0050, Xiaoguang Zhu, Donglai Wei 0002 |
ICASSP | 1 |
| 2022 | Look, Listen and Pay More Attention: Fusing Multi-Modal Information for Video Violence DetectionabstractViolence detection is an essential and challenging problem in the computer vision community. Most existing works focus on single modal data analysis, which is not effective when multi-modality is available. Therefore, we propose a two-stage multi-modal information fusion method for violence detection: 1) the first stage adopts multiple instance learning strategies to refine video-level hard labels into clip-level soft labels, and 2) the next stage uses multi-modal information fused attention module to achieve fusion, and supervised learning is carried out using the soft labels generated at the first stage. Extensive empirical evidence on the XD-Violence dataset shows that our method outperforms the state-of-the-art methods. Donglai Wei 0002, Chen-Geng Liu, Yang Liu 0246, Jing Liu 0050, Xiao-Guang Zhu, Xinhua Zeng |
ICASSP | 3 |
| 2022 | Learning Appearance-Motion Normality for Video Anomaly DetectionabstractVideo anomaly detection is a challenging task in the Computer vision community. Most single task-based methods do not consider the independence of unique spatial and temporal patterns, while two-stream structures lack the exploration of the correlations. In this paper, we propose spatial-temporal memories augmented two-stream auto-encoder framework, which learns the appearance normality and motion normal-ity independently and explores the correlations via adversar-ial learning. Specifically, we first design two proxy tasks to train the two-stream structure to extract appearance and motion features in isolation. Then, the prototypical features are recorded in the corresponding spatial and temporal memory pools. Finally, the encoding-decoding network performs ad-versariallearning with the discriminator to explore the corre-lations between spatial and temporal patterns. Experimental results show that our framework outperforms the state-of-the-art methods, achieving AUCs of 98.1% and 89.8% on UCSD Ped2 and CUHK Avenue datasets. Yang Liu 0246, Jing Liu 0050, Mengyang Zhao 0002, Dingkang Yang, Xiaoguang Zhu |
ICME | 1 |
| 2022 | MAR2MIX: A Novel Model for Dynamic Problem in Multi-agent Reinforcement Learning
Gaoyun Fang, Yang Liu 0246, Jing Liu 0050 |
ICONIP (4) | 2 |
| 2022 | Exploiting Spatial-temporal Correlations for Video Anomaly DetectionabstractVideo anomaly detection (VAD) remains a challenging task in the pattern recognition community due to the ambiguity and diversity of abnormal events. Existing deep learning-based VAD methods usually leverage proxy tasks to learn the normal patterns and discriminate the instances that deviate from such patterns as abnormal. However, most of them do not take full advantage of spatial-temporal correlations among video frames, which is critical for understanding normal patterns. In this paper, we address unsupervised VAD by learning the evolution regularity of appearance and motion in the long and short-term and exploit the spatial-temporal correlations among consecutive frames in normal videos more adequately. Specifically, we proposed to utilize the spatiotemporal long short-term memory (ST-LSTM) to extract and memorize spatial appearances and temporal variations in a unified memory cell. In addition, inspired by the generative adversarial network, we introduce a discriminator to perform adversarial learning with the ST-LSTM to enhance the learning capability. Experimental results on standard benchmarks demonstrate the effectiveness of spatial-temporal correlations for unsupervised VAD. Our method achieves competitive performance compared to the state-of-the-art methods with AUCs of 96.7%, 87.8%, and 73.1% on the UCSD Ped2, CUHK Avenue, and ShanghaiTech, respectively. Mengyang Zhao 0002, Yang Liu 0246, Jing Liu 0050, Xinhua Zeng |
ICPR | 2 |
| 2022 | Abnormal Event Detection with Self-guiding Multi-instance Ranking FrameworkabstractThe detection of abnormal events in surveillance videos with weak supervision is a challenging task, which tries to temporally find abnormal frames using readily accessible video-level labels. In this paper, we propose a self-guiding multi-instance ranking (SMR) framework, which has explored task-specific deep representations and considered the temporal correlations between video clips. Specifically, we apply a clustering algorithm to fine-tune the features extracted by the pre-trained 3D-convolutional-based models. Besides, the clustering module can generate clip-level labels for abnormal videos, and the pseudo-labels are in part used to supervise the training of the multi-instance regression. While implementing the regression module, we compare the effectiveness of various recurrent neural networks, and the results demonstrate the necessity of temporal correlations for weakly supervised video anomaly detection tasks. Experimental results on two standard benchmarks reveal that the SMR framework is comparable to the state-of-the-art approaches, with frame-level AUCs of 81.7% and 92.4% on the UCF-crime and UCSD Ped2 datasets respectively. Additionally, ablation studies and visualization results prove the effectiveness of the component, and our framework can accurately locate abnormal events. Yang Liu 0246, Jing Liu 0050 |
IJCNN | 1 |
| 2022 | Attention-Based Auto-Encoder Framework for Abnormal Driving DetectionabstractWith the popularity of smartphones, abnormal driving detection via smartphone sensors has been proposed in recent years. However, existing methods are insufficient in exploring feature extraction, so the practical value is limited due to the low accuracy. To address this problem, we propose an attention-based auto-encoder framework for abnormal driving detection that combines the advantages of bi-directional long short-term memory and self-attention. Specifically, these two modules are embedded in the auto-encoder for modeling latent vector and exploring the internal correlations of spatial-temporal features, respectively, so as to improve the capability of reconstructing driving time series using small and representative features. We conduct experiments on the real-world datasets, and the results show that the proposed framework achieves significant performance with recall and F1-score of 96.2% and 95.0%, superior to the other baselines. Jing Liu 0050, Yang Liu 0246, Donglai Wei 0002, Xinhua Zeng |
ISCAS | 2 |
| 2022 | Multi-level Attention Fusion for Multimodal Driving Maneuver RecognitionabstractSensor-based driving maneuver recognition (DMR) is a fundamental and challenging task in ubiquitous computing, which uses multimodal signals from embedded sensors such as accelerometers and gyroscopes to recognize driving maneuvers. However, the spatial-temporal features from neural networks are often treated equally, which may limit the performance of the model in predicting maneuvers. In this paper, we propose a novel hybrid neural network model based on multi-level attention fusion for multimodal DMR. The proposed model utilizes convolutional neural networks and gated recurrent unit to extract temporal-spatial features from multimodal sensing signals and propose the multi-level attention fusion to explore the significant patterns over local and global periods. In addition, We design three different levels of fusion (early, late, and full fusion) to explore the effects of different attention fusions on the model. Extensive experiments on the real-world dataset show that the proposed model achieves superior performance to the baseline methods, and multi-level attention fusion brings 6.17% gain to the F1-score. Jing Liu 0050, Yang Liu 0246, Chengwen Tian, Mengyang Zhao 0002, Xinhua Zeng |
ISCAS | 2 |
| 2022 | MSAF: Multimodal Supervise-Attention Enhanced Fusion for Video Anomaly DetectionabstractThe complementarity of multimodal signal is essential for video anomaly detection. However, existing methods either lack exploration to multimodal data or ignore the implicit alignment of multimodal features. In our work, we address this problem using a novel fusion method and propose a Multimodal Supervise-Attention enhanced Fusion (MSAF) framework under weak supervision. Our framework can be divided into two parts: 1) the multimodal labels refinement part refines video-level ground truth into pseudo clip-level labels for subsequent training, 2) the multimodal supervise-attention fusion network enhances features via implicitly aligning different information, then fusing them effectively to predict anomaly scores with the help of refined labels. We validate our framework on four challenging datasets: ShanghaiTech, UCF-Crime, LAD, and XD-Violence. Extensive experiments on the benchmarks demonstrate the effectiveness of our framework, which achieves comparable results on several benchmarks and outperforms current state-of-the-art methods on the XD-Violence audiovisual multimodal dataset. Donglai Wei 0002, Yang Liu 0246, Xiaoguang Zhu, Jing Liu 0050, Xinhua Zeng |
IEEE Signal Process. Lett. | 2 |
| 2022 | Contextual and Cross-Modal Interaction for Multi-Modal Speech Emotion RecognitionabstractSpeech emotion recognition combining linguistic content and audio signals in the dialog is a challenging task. Nevertheless, previous approaches have failed to explore emotion cues in contextual interactions and ignored the long-range dependencies between elements from different modalities. To tackle the above issues, this letter proposes a multimodal speech emotion recognition method using audio and text data. We first present a contextual transformer module to introduce contextual information via embedding the previous utterances between interlocutors, which enhances the emotion representation of the current utterance. Then, the proposed cross-modal transformer module focuses on the interactions between text and audio modalities, adaptively promoting the fusion from one modality to another. Furthermore, we construct associative topological relation over mini-batch and learn the association between deep fused features with graph convolutional network. Experimental results on the IEMOCAP and MELD datasets show that our method outperforms current state-of-the-art methods. Dingkang Yang, Yang Liu 0246, Lihua Zhang 0002 |
IEEE Signal Process. Lett. | 3 |
| 2022 | Adaptive Weighted Losses With Distribution Approximation for Efficient Consistency-Based Semi-Supervised LearningabstractRecent semi-supervised learning (SSL) algorithms such as FixMatch achieve state-of-the-art performance by exploiting consistency regularization and entropy minimization techniques. However, many consistency-based SSL algorithms extract pseudo-labels from unlabeled data through a fixed threshold and ignore the different learning progress of each category, which makes the easy-to-learn categories have more examples contributing to the loss, resulting in a class-imbalance problem and affecting training efficiency. In order to improve the training reliability, we propose adaptive weighted losses (AWL). Through the evaluation of the class-wise learning progress, the loss contribution of the pseudo-labeled data of each category is continuously and dynamically adjusted during the learning process, and the pseudo-label discrimination ability of the model can be steadily improved. Moreover, to improve the training efficiency, we propose a bidirectional distribution approximation (DA) method, which introduces the consistency information of the predictions under the threshold into the loss calculation, and significantly improves the model convergence speed. Through the combination of AWL and DA, our method surpasses the performance of other algorithms on multiple benchmarks with a faster convergence efficiency, especially in the case of labeled data extremely limited. For example, AWL&DA achieves 95.29% test accuracy on the CIFAR-10-40-labels experiment and 92.56% accuracy on a faster experiment setting with only$2^{18}$iterations. Yang Liu 0246 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Stack Multiple Shallow Autoencoders into a Strong One: A New Reconstruction-Based Method to Detect Anomaly
Hanqi Wang, Xing Hu 0006, Yang Liu 0246, Jing Liu 0050, Linhua Jiang |
ICONIP (1) | 5 |
| 2004 | Improved shot boundary detection method based on text edgesabstractShot boundary detection is a pre-requisite technique for video indexing and retrieval. To avoid the influence of flashlight on abrupt shot detection, many edge-based techniques are studied thoroughly. However, these techniques are still susceptible to miss and mistake detecting the abrupt changes. Our observation shows that one of the reasons for these errors is the existence of superimposed text which has rich edges and is ever presented in video frames. To provide a solution, we present a novel method that utilizes the edge type, text edge (edge in text area) or non-text-edge (edge in other text area), reducing erroneous detection with the appearance of video text. Compared to other edge-based detection techniques, experimental results show that our proposed method achieves preferable performance. Liuhong Liang, Yang Liu 0246, Xiangyang Xue 0001, Hong Lu 0001, Yap-Peng Tan |
ICARCV | 2 |
| 2004 | Effective video text detection using line featuresabstractText superimposed on video frames provides synoptic or supplemental information on video semantics. In this paper, we propose a novel method to detect superimposed text effectively. First, we detect edges by an improved Canny edge detector. Then, a line-feature vector graph is generated based on the edge map and the stroke information is extracted. Finally text regions are generated and filtered according to line features. Experimental results show that, without much increasing the computational cost, our proposed method could suppress the false alarms notably. Furthermore, our method can be easily customized to applications with different tradeoffs in recall and precision. Yang Liu 0246, Hong Lu 0001, Xiangyang Xue 0001, Yap-Peng Tan |
ICARCV | 1 |