EDBT 2026 Demo / reviewers in the wild / expert
Jing Liu 0050
dblp:72/2590-50
· DBLP profile ↗
42ranked-venue papers
7as first author
41since 2021 · last 2026
0000-0002-2819-0200ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 2 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 16 since 2021Computer networks · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Projecting to Consensus: Communication-Efficient Collaborative Learning Across Heterogeneous Networks
Jing Liu 0050, Yao Du 0001, Yang Liu 0246, Zehua Wang 0001, Peng Sun 0007, Victor C. M. Leung |
ICC | 1 |
| 2026 | Enhancing Collaborative Learning Efficiency via Control-Theoretic Merit Gating in Federated Networks
Jing Liu 0050, Gaoyun Fang, Liangyu Teng, Lang Qian, Bo Hu 0002, Peng Sun 0007 |
ICC | 1 |
| 2026 | MeritFL: Self-Regulating Federated Learning via Merit-Gated Communication
Zhengliang Guo, Kun Yang 0010, Linxiao Gong, Yang Liu 0246, Jing Liu 0050 |
ISCAS | 5 |
| 2026 | A Coordinated Optimization Framework for Intelligent Agents With Online Evolutive LearningabstractAs a prevalent field of study in machine learning, intelligent agents can perceive surroundings and make informed decisions. In many research areas such as autopilot systems, undersea explorations, and distributed robotics, researchers have traditionally employed unidirectional systems, which usually rely on perceptions from sensors to controllers, or end-to-end models, which generate actions directly from raw data. Nonetheless, in unidirectional systems, controller efficiency is intrinsically linked to sensor accuracy, which makes unidirectional systems lack a self-improving capability. Meanwhile, compared with functionally separated frameworks, end-to-end methods may have their own limitations in scalability, generality, interoperability, training costs, and so on. To fulfill this gap, we propose a new Coordinated Optimization Framework for Intelligent Agents (COIA). We introduce an inverted optimization channel from controllers to sensors in traditional functionally separated frameworks through communication between devices, enabling closed-loop online evolutive learning. To the best of our knowledge, this paper first presents a universal coordinated optimization framework among supervised learning and RL models, without human labels or intervention. Our method allows heterogeneous agents to autonomously adapt to some special situations in open environments, which forms a basis of networked Artificial General Intelligence (AGI). We design an experimental paradigm of COIA with concrete cases, which shows a significantly large performance margin over unidirectional and end-to-end models. The performance margin grows with task complexity. Lang Qian, Jiayue Jin, Peng Sun 0007, Jing Liu 0050, Bo Hu 0002, Azzedine Boukerche |
IEEE Internet Things J. | 4 |
| 2026 | Multimodal human video generation with uncertainty-aware pose guidance
Kun Yang 0010, Yuanyuan Meng, Juncen Guo, Yanda Meng, Songwen Pei, Jing Liu 0050, Yang Liu 0246 |
Pattern Recognit. | 8 |
| 2026 | Privacy-Preserving Video Anomaly Detection: A SurveyabstractThe video anomaly detection (VAD) aims to automatically analyze spatiotemporal patterns in surveillance videos collected from open spaces to detect anomalous events that may cause harm, such as fighting, stealing, and car accidents. However, vision-based surveillance systems such as closed-circuit television (CCTV) often capture personally identifiable information. The lack of transparency and interpretability in video transmission and usage raises public concerns about privacy and ethics, limiting the real-world application of VAD. Recently, researchers have focused on privacy concerns in VAD by conducting systematic studies from various perspectives, including data, features, and systems, making privacy-preserving VAD (P2VAD) a hotspot in the AI community. However, the current research in P2VAD is fragmented, and prior reviews have mostly focused on methods using RGB sequences, overlooking privacy leakage and appearance bias considerations. To address this gap, this article is the first to systematically review the progress of P2VAD, defining its scope and providing an intuitive taxonomy. We outline the basic assumptions, learning frameworks, and optimization objectives of various approaches, analyzing their strengths, weaknesses, and potential correlations. In addition, we provide open access to research resources such as benchmark datasets and available code. Finally, we discuss key challenges and future opportunities from the perspectives of AI development and P2VAD deployment, aiming to the guide future work in the field. Yang Liu 0246, Siao Liu, Xiaoguang Zhu, Hao Yang 0055, Juncen Guo, Liangyu Teng, Dingkang Yang, Yan Wang 0068, Jing Liu 0050 |
IEEE Trans. Neural Networks Learn. Syst. | 10 |
| 2025 | MF-AttnBiLSTM: Traffic Flow Prediction via Hybrid Signal Decomposition and Dual-Stream Temporal Attention LearningabstractAccurate traffic flow prediction is crucial for intelligent transportation systems supporting emerging applications such as autonomous driving and vehicle-infrastructure cooperation. However, existing methods often struggle to effectively disentangle the inherent trend, seasonal, and noise components within traffic flow data, thereby limiting prediction accuracy. To address this issue, we propose MF-AttnBiLSTM, a novel hybrid framework combining signal processing and temporal attention-based deep learning model through a decompose-then-predict strategy. Our approach first employs moving average to extract the trend component and discrete Fourier transform to isolate dominant seasonal patterns from the residuals. Subsequently, a dual-stream architecture utilizes multi-head self-attention-enhanced bidirectional LSTMs to independently model the temporal dynamics of the decomposed trend and seasonal components. The final prediction aggregates the outputs from both streams. Extensive experiments on PeMS04 and PeMS07 datasets demonstrate that MF-AttnBiLSTM significantly outperforms state-of-the-art baselines and exhibits robustness across varying traffic conditions. Ablation studies further confirm the efficacy of each component, particularly highlighting the significant contribution of the signal decomposition stage to overall performance improvement. Luyao Niu, Zepu Wang, Jing Liu 0050, Azzedine Boukerche, Peng Sun 0007 |
GLOBECOM | 3 |
| 2025 | Towards Advanced Emotional Care: Embodied Emotional Care System for Humanoid RobotsabstractIn modern healthcare, emotional well-being is critical to patient recovery and overall outcomes. However, limited availability of trained professionals and time constraints often hinder the delivery of consistent emotional support. To address this gap, we propose the Embodied Emotional Care System (EECS), a comprehensive humanoid robotic framework designed to deliver personalized emotional care through an integrated, multi-layered architecture. EECS analyzes dynamic facial expressions and real-time vocal inputs to extract the patient’s emotional state and semantic information, constructs context-aware prompts processed by an LLM for reasoning, and ultimately generates empathetic dialogues synchronized with human-like facial expressions and natural body movements to address diverse emotional support needs. Experimental results show that deploying EECS on a humanoid robot significantly boosts patient engagement through real-time multimodal interaction, delivering deeper emotional support and a more human-like therapeutic experience. Furthermore, it bridges gaps in professional emotional support resources, offering a feasible pathway to improve overall healthcare quality. Yang Chang, Aoxing Li, Yuxuan Lin 0001, Lizheng Liu, Yang Liu 0246, Jing Liu 0050, Yan Wang 0068, Zhongxue Gan 0001 |
ICME | 7 |
| 2025 | M2S2L: Mamba-based Multi-Scale Spatial-temporal Learning for Video Anomaly DetectionabstractVideo anomaly detection (VAD) is an essential task in the image processing community with prospects in video surveillance, which faces fundamental challenges in balancing detection accuracy with computational efficiency. As video content becomes increasingly complex with diverse behavioral patterns and contextual scenarios, traditional VAD approaches struggle to provide robust assessment for modern surveillance systems. Existing methods either lack comprehensive spatial-temporal modeling or require excessive computational resources for real-time applications. In this regard, we present a Mamba-based multi-scale spatial-temporal learning (M2S2L) framework in this paper. The proposed method employs hierarchical spatial encoders operating at multiple granularities and multi-temporal encoders capturing motion dynamics across different time scales. We also introduce a feature decomposition mechanism to enable task-specific optimization for appearance and motion reconstruction, facilitating more nuanced behavioral modeling and quality-aware anomaly assessment. Experiments on three benchmark datasets demonstrate that M2S2L framework achieves 98.5%, 92.1%, and 77.9% frame-level AUCs on UCSD Ped2, CUHK Avenue, and ShanghaiTech respectively, while maintaining efficiency with 20.1G FLOPs and 45 FPS inference speed, making it suitable for practical surveillance deployment. Yang Liu 0246, Boan Chen, Xiaoguang Zhu, Jing Liu 0050, Peng Sun 0007, Wei Zhou 0013 |
VCIP | 4 |
| 2025 | Domain generalization with semi-supervised learning for people-centric activity recognition
Jing Liu 0050, Xing Hu 0006 |
Sci. China Inf. Sci. | 1 |
| 2025 | CNN-DAG-Editor: A Convolutional Neural Network offloading analyzer with Multi-Objective Dynamic Adaptive Resource Competitive Swarm Optimization
Bobo Ju, Yang Liu 0246, Jing Liu 0050, Peng Sun 0007 |
Comput. Networks | 3 |
| 2025 | Rethinking prediction-based video anomaly detection from local-global normality perspective
Mengyang Zhao 0002, Xinhua Zeng, Yang Liu 0246, Jing Liu 0050, Chengxin Pang |
Expert Syst. Appl. | 4 |
| 2025 | CRCL: Causal Representation Consistency Learning for Anomaly Detection in Surveillance VideosabstractVideo Anomaly Detection (VAD) remains a fundamental yet formidable task in the video understanding community, with promising applications in areas such as information forensics and public safety protection. Due to the rarity and diversity of anomalies, existing methods only use easily collected regular events to model the inherent normality of normal spatial-temporal patterns in an unsupervised manner. Although such methods have made significant progress benefiting from the development of deep learning, they attempt to model the statistical dependency between observable videos and semantic labels, which is a crude description of normality and lacks a systematic exploration of its underlying causal relationships. Previous studies have shown that existing unsupervised VAD models are incapable of label-independent data offsets (e.g., scene changes) in real-world scenarios and may fail to respond to light anomalies due to the overgeneralization of deep neural networks. Inspired by causality learning, we argue that there exist causal factors that can adequately generalize the prototypical patterns of regular events and present significant deviations when anomalous instances occur. In this regard, we propose Causal Representation Consistency Learning (CRCL) to implicitly mine potential scene-robust causal variable in unsupervised video normality learning. Specifically, building on the structural causal models, we propose scene-debiasing learning and causality-inspired normality learning to strip away entangled scene bias in deep representations and learn causal video normality, respectively. Extensive experiments on benchmarks validate the superiority of our method over conventional deep representation learning. Moreover, ablation studies and extension validation show that the CRCL can cope with label-independent biases in multi-scene settings and maintain stable performance with only limited training data available. Yang Liu 0246, Hongjin Wang, Zepu Wang, Xiaoguang Zhu, Jing Liu 0050, Peng Sun 0007, Jianwei Du, Victor C. M. Leung |
IEEE Trans. Image Process. | 5 |
| 2024 | DyHGDAT: Dynamic Hypergraph Dual Attention Network for multi-agent trajectory predictionabstractModeling the interactions among agents based on their historical trajectories is key to precise multi-agent trajectory prediction. Hypergraph Convolutional Networks (HGCN) have become a proper choice for capturing high-order interactions among agents in this field. However, most existing works only consider static hypergraphs, and ignore that in a hypergraph, the power of influence varies between vertices (or hyperedges). Therefore, we propose DyHGDAT, a dynamic hypergraph dual attention network to capture the high-order interactions among agents, which not only models the evolution of hypergraph over time but also highlights the vertices and hyperedges with larger impacts. We apply DyHGDAT to a CVAE-based prediction system for predicting plausible trajectories. To validate the effectiveness of prediction, we evaluate our proposed method on two well-established trajectory prediction datasets: the ETH/UCY datasets and the Stanford Drone Dataset (SDD). The experimental results show that with DyHGDAT, the CVAE-based prediction system out-performs state-of-the-art methods by 12.5%/5.3% in ADE/FDE on ETH/UCY, and the improvement on SDD is 6.4%/7.4%. Weilong Lin, Xinhua Zeng, Chengxin Pang, Jing Teng, Jing Liu 0050 |
ICRA | 5 |
| 2024 | Memory-enhanced appearance-motion consistency framework for video anomaly detection
Zhiyuan Ning 0002, Zile Wang, Yang Liu 0246, Jing Liu 0050 |
Comput. Commun. | 4 |
| 2024 | Memory-enhanced spatial-temporal encoding framework for industrial anomaly detection system
Yang Liu 0246, Bobo Ju, Dingkang Yang, Liyuan Peng, Peng Sun 0007, Chengfang Li, Hao Yang 0055, Jing Liu 0050 |
Expert Syst. Appl. | 9 |
| 2024 | Decoding Silent Reading EEG Signals Using Adaptive Feature Graph Convolutional NetworkabstractDecoding silent reading Electroencephalography (EEG) signals is challenging because of its low signal-to-noise ratio. In addition, EEG signals are typically non-Euclidean structured, therefore merely using a two-dimensional matrix to represent the variation of sampling points of each channel in time cannot richly represent the spatial connection between channels. Furthermore, due to the individual differences in EEG signals, a fixed representation cannot adequately represent the temporal and spatial associations between channels in real time. In this letter, we use the feature matrix and its adaptive graph structure to represent each EEG signal. Then, we use them as inputs and propose a novel Adaptive Feature Graph Convolutional Network (AFGCN) to decode the silent reading EEG signals. We classify silent reading EEG signals under different tasks of 16 subjects from two publicly available datasets. The experimental results demonstrate that our proposed method achieves higher decoding accuracy than state-of-the-art EEG classification networks on both datasets. Among them, the highest classification accuracy for the four classes is 83.33%. The study could promote the application and development of BCI technology for silent reading EEG signal decoding. It can also provide an efficient and convenient communication method for patients with language impairment. Chengfang Li, Gaoyun Fang, Yang Liu 0246, Jing Liu 0050 |
IEEE Signal Process. Lett. | 4 |
| 2024 | AMP-Net: Appearance-Motion Prototype Network Assisted Automatic Video Anomaly Detection SystemabstractAs essential tools for industry safety protection, automatic video anomaly detection systems (AVADS) are designed to detect anomalous events of concern in surveillance videos. Existing VAD methods lack effective exploration of the prototypical appearance and motion features leading to poor performance in realistic scenarios. Specifically, they either misreport regular events as anomalies due to insufficient representation power, or lead to missed detections with over-power generalization. In this regard, we propose an appearance-motion prototype network (AMP-net) that uses external memories to record prototype features and augments the appearance-motion prototype with a spatial-temporal fusion. In addition, AMP-net sequentially fuses appearance features from deep to shallow to utilize multiscale spatial context. Additionally, we introduce temporal attention to capture important dynamics and enhance AMP-net for representing regular motion. The proposed method achieves a delicate balance of effective representation of normal events and limited generalization to anomalies. Experiments on three benchmark datasets demonstrate that our method can accurately detect anomalous events, achieving performance comparable to state-of-the-art methods with frame-level AUCs of 98.7%, 92.4%, and 78.8% on the UCSD Ped2, CUHK Avenue, and ShanghaiTech datasets. Moreover, we conducted a case study on the self-collected industrial dataset, and the results indicate that our AMP-net can cope with complex industrial scenarios and outperform existing methods. Yang Liu 0246, Jing Liu 0050, Kun Yang 0010, Bobo Ju, Siao Liu, Dingkang Yang, Peng Sun 0007 |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | MGR3Net: Multigranularity Region Relation Representation Network for Facial Expression Recognition in Affective RobotsabstractAutomatic facial expression recognition (FER) based on face images is essential for affective robots, which are designed for interactive companions and intelligent healthcare. Although existing DL-based FERs have made significant progress, an accurate FER model in robots is challenging due to the subtle differences in facial expressions across various scenarios. To address this issue, we propose a multigranularity region relation representation network (MGR3Net) to improve the robustness and generalization of FER via attention-guided global-local fusion. The MGR3Net is composed of three modules: multigranularity attention (MGA), holistic-regional feature extractor (HRFE), and hybrid feature fusion. In the MGA module, we first process each holistic cropped face image into three granularity of face regions from coarse to fine, which are four region-cropped faces,$2^{2}$face partitions, and$4^{2}$face partitions. Then, we propose the region attention relation cell to model the relationship between each region and the aggregated representation while preserving the spatial information of the local features. In the HRFE module, we align multigranularity features from the coarse space to the finer space and extract one holistic embedding and multiple region embeddings for each granularity. Finally, we use a hybrid-level fusion strategy to combine global-local features from the three granularities for final classification. Extensive experiments demonstrate that the MGR3Net outperforms the state-of-the-art methods evaluated on the in-the-lab datasets, in-the-wild datasets, and occlusion/pose-based sets. Yan Wang 0068, Shaoqi Yan, Wei Song 0007, Antonio Liotta, Jing Liu 0050, Dingkang Yang, Shuyong Gao |
IEEE Trans. Ind. Informatics | 5 |
| 2023 | Surrogate-Assisted Evolution of Convolutional Neural Networks by Collaboratively Optimizing the Basic Blocks and TopologiesabstractConvolutional neural networks (CNNs) are prominent in many fields owing to their outstanding feature extraction abilities. Many excellent CNNs have been carefully designed by algorithm researchers; however, the design process is limited by the inherent knowledge of the researchers. Inspired by the existing successful block-based neural architecture search, we develop a collaboratively automatically evolutionary CNNs algorithm (CAE-CNN), which employs a surrogate-assisted genetic algorithm to search for a satisfactory CNN architecture by collaboratively optimizing the basic blocks in ResNet and DenseNet and their topologies. The encoding space of CAE-CNN consists of three basic units (pooling, ResNet, and DenseNet units) and a connection topology with a variable size. To effectively evolve the population, we design a double-module crossover operation and a multi-type mutation operation to collaboratively evolve the units and the topology. To address the problem of a rapidly increasing search space caused by the topological search, we use a random forest as the surrogate model to estimate the fitness of an individual to accelerate the search. CAE-CNN is a completely automatic algorithm in which a satisfactory CNN architecture can be obtained without any manual intervention. Experimental results show that CAE-CNN could archive competitive performance in terms of classification accuracy on four image classification datasets, and it consumes fewer computing resources than many algorithms. Jing Liu 0050, Yang Liu 0246 |
CEC | 2 |
| 2023 | MSN-net: Multi-Scale Normality Network for Video Anomaly DetectionabstractExisting unsupervised video anomaly detection methods often suffer from performance degradation due to the overgeneralization of deep models. In this paper, we propose a simple yet effective Multi-Scale Normality network (MSN-net) that uses hierarchical memories to learn multi-level prototypical spatial-temporal patterns of normal events. Specifically, the hierarchical memory module interacts with the encoder through the reading and writing operations during the training phase, preserving multi-scale normality in three separate memory pools. Then, the decoder decodes the features rewritten by the memorized normality to predict future frames so that its ability to predict anomalies is diminished. Experimental results show that MSN-net performs comparably to the state-of-the-art methods, and extension analysis demonstrates the effectiveness of multi-scale normality learning. Yang Liu 0246, Dingkang Yang, Jing Liu 0050 |
ICASSP | 5 |
| 2023 | A Novel Efficient Multi-View Traffic-Related Object Detection FrameworkabstractWith the rapid development of intelligent transportation system applications, a tremendous amount of multi-view video data has emerged to enhance vehicle perception. However, performing video analytics efficiently by exploiting the spatial-temporal redundancy from video data remains challenging. Accordingly, we propose a novel traffic-related framework named CEVAS to achieve efficient object detection using multi-view video data. Briefly, a fine-grained input filtering policy is introduced to produce a reasonable region of interest from the captured images. Also, we design a sharing object manager to manage the information of objects with spatial redundancy and share their results with other vehicles. We further derive a content-aware model selection policy to select detection methods adaptively. Experimental results show that our framework significantly reduces response latency while achieving the same detection accuracy as the state-of-the-art methods. Kun Yang 0010, Jing Liu 0050, Dingkang Yang, Hanqi Wang, Peng Sun 0007 |
ICASSP | 2 |
| 2023 | AIDE: A Vision-Driven Multi-View, Multi-Modal, Multi-Tasking Dataset for Assistive Driving PerceptionabstractDriver distraction has become a significant cause of severe traffic accidents over the past decade. Despite the growing development of vision-driven driver monitoring systems, the lack of comprehensive perception datasets restricts road safety and traffic security. In this paper, we present an AssIstive Driving pErception dataset (AIDE) that considers context information both inside and outside the vehicle in naturalistic scenarios. AIDE facilitates holistic driver monitoring through three distinctive characteristics, including multi-view settings of driver and scene, multi-modal annotations of face, body, posture, and gesture, and four pragmatic task designs for driving understanding. To thoroughly explore AIDE, we provide experimental benchmarks on three kinds of baseline frameworks via extensive methods. Moreover, two fusion strategies are introduced to give new insights into learning effective multi-stream/modal representations. We also systematically investigate the importance and rationality of the key components in AIDE and benchmarks. The project link is https://github.com/ydk122024/AIDE. Dingkang Yang, Zhi Xu 0010, Shunli Wang 0001, Mingcheng Li, Yang Liu 0246, Kun Yang 0010, Zhaoyu Chen 0001, Yan Wang 0068, Jing Liu 0050, Peixuan Zhang, Peng Zhai, Lihua Zhang 0002 |
ICCV | 12 |
| 2023 | Spatio-Temporal Domain Awareness for Multi-Agent Collaborative PerceptionabstractMulti-agent collaborative perception as a potential application for vehicle-to-everything communication could significantly improve the perception performance of autonomous vehicles over single-agent perception. However, several challenges remain in achieving pragmatic information sharing in this emerging research. In this paper, we propose SCOPE, a novel collaborative perception frame-work that aggregates the spatio-temporal awareness characteristics across on-road agents in an end-to-end manner. Specifically, SCOPE has three distinct strengths: i) it considers effective semantic cues of the temporal context to enhance current representations of the target agent; ii) it aggregates perceptually critical spatial information from heterogeneous agents and overcomes localization errors via multi-scale feature interactions; iii) it integrates multi-source representations of the target agent based on their complementary contributions by an adaptive fusion paradigm. To thoroughly evaluate SCOPE, we consider both real-world and simulated scenarios of collaborative 3D object detection tasks on three datasets. Extensive experiments show the superiority of our approach and the necessity of the proposed components. The project link is https://ydk122024.github.io/SCOPE/. Kun Yang 0010, Dingkang Yang, Mingcheng Li, Yang Liu 0246, Jing Liu 0050, Hanqi Wang, Peng Sun 0007 |
ICCV | 6 |
| 2023 | Learning Causality-inspired Representation Consistency for Video Anomaly DetectionabstractVideo anomaly detection is an essential yet challenging task in the multimedia community, with promising applications in smart cities and secure communities. Existing methods attempt to learn abstract representations of regular events with statistical dependence to model the endogenous normality, which discriminates anomalies by measuring the deviations to the learned distribution. However, conventional representation learning is only a crude description of video normality and lacks an exploration of its underlying causality. The learned statistical dependence is unreliable for diverse regular events in the real world and may cause high false alarms due to over generalization. Inspired by causal representation learning, we think that there exists a causal variable capable of adequately representing the general patterns of regular events in which anomalies will present significant variations. Therefore, we design a causality-inspired representation consistency (CRC) framework to implicitly learn the unobservable causal variables of normality directly from available normal videos and detect abnormal events with the learned representation consistency. Extensive experiments show that the causality-inspired normality is robust to regular events with label-independent shifts, and the proposed CRC framework can quickly and accurately detect various complicated anomalies from real-world surveillance videos. Yang Liu 0246, Zhaoyang Xia, Mengyang Zhao 0002, Donglai Wei 0002, Siao Liu, Bobo Ju, Gaoyun Fang, Jing Liu 0050 |
ACM Multimedia | 9 |
| 2023 | How2comm: Communication-Efficient and Collaboration-Pragmatic Multi-Agent PerceptionabstractMulti-agent collaborative perception has recently received widespread attention as an emerging application in driving scenarios. Despite the advancements in previous efforts, challenges remain due to various noises in the perception procedure, including communication redundancy, transmission delay, and collaboration heterogeneity. To tackle these issues, we propose \textit{How2comm}, a collaborative perception framework that seeks a trade-off between perception performance and communication bandwidth. Our novelties lie in three aspects. First, we devise a mutual information-aware communication mechanism to maximally sustain the informative features shared by collaborators. The spatial-channel filtering is adopted to perform effective feature sparsification for efficient communication. Second, we present a flow-guided delay compensation strategy to predict future characteristics from collaborators and eliminate feature misalignment due to temporal asynchrony. Ultimately, a pragmatic collaboration transformer is introduced to integrate holistic spatial semantics and temporal context clues among agents. Our framework is thoroughly evaluated on several LiDAR-based collaborative detection datasets in real-world and simulated scenarios. Comprehensive experiments demonstrate the superiority of How2comm and the effectiveness of all its vital components. The code will be released at https://github.com/ydk122024/How2comm. Dingkang Yang, Kun Yang 0010, Jing Liu 0050, Zhi Xu 0010, Rongbin Yin, Peng Zhai, Lihua Zhang 0002 |
NeurIPS | 4 |
| 2023 | DSDCLA: driving style detection via hybrid CNN-LSTM with multi-level attention fusion
Jing Liu 0050, Yang Liu 0246, Hanqi Wang |
Appl. Intell. | 1 |
| 2023 | Stochastic video normality network for abnormal event detection in surveillance videos
Yang Liu 0246, Dingkang Yang, Gaoyun Fang, Donglai Wei 0002, Mengyang Zhao 0002, Kai Cheng 0001, Jing Liu 0050 |
Knowl. Based Syst. | 8 |
| 2023 | Distributional and spatial-temporal robust representation learning for transportation activity recognition
Jing Liu 0050, Yang Liu 0246, Xiaoguang Zhu |
Pattern Recognit. | 1 |
| 2023 | OSIN: Object-Centric Scene Inference Network for Unsupervised Video Anomaly DetectionabstractVideo Anomaly Detection (VAD) is an essential yet challenging task in the signal processing community, which aims to understand the spatial and temporal contextual interactions between objects and surrounding scenes to detect unexpected events in surveillance videos. However, existing unsupervised methods either use a single network to learn global prototype patterns without making a unique distinction between foreground objects and background scenes or try to strip objects from frames, ignoring that the essence of anomalies lies in unusual object-scene interactions. To this end, this letter proposes an Object-centric Scene Inference Network (OSIN) that uses a well-designed three-stream structure to learn both global scene normality and local object-specific normal patterns as well as explore the object-scene interactions using scene memory networks. Experimental results on three benchmark datasets demonstrate the effectiveness of the proposed OSIN model, which achieves frame-level AUCs of 91.7%, 79.6%, and 98.3% on the CUHK Avenue, ShanghaiTech, and UCSD Ped2 datasets, respectively. Yang Liu 0246, Zhengliang Guo, Jing Liu 0050, Chengfang Li |
IEEE Signal Process. Lett. | 3 |
| 2023 | Two-Stage Alignments Framework for Unsupervised Domain Adaptation on Time Series DataabstractUnsupervised Domain Adaptation (UDA) aims to free models from labeled information of target domain by minimizing the discrepancy of distributions between different domains. Most existing methods are designed to learn domain-invariant features either by domain discrimination or by matching lower-order moments. However, these methods are not robust due to the limited representation of statistical characteristics for non-Gaussian distributions and thus fail in domain matching. In addition, they often focus on matching distributions while not considering class decision boundaries between domains. To address these issues, we propose a novel Two-Stage Alignments Framework (TSAF) for UAD, which not only performs arbitrary-order moment matching to approximately characterize complex non-Gaussian distributions, but also utilizes domain-specific decision boundaries to align the probabilistic outputs of classifiers. Moreover, the reconstruction-based task is introduced to enhance the representation of the inherent characteristics for specific distribution. Extensive experiments on three real-world time series datasets demonstrate that: 1) our model evidently outperforms many state-of-the-art domain adaptation methods in cross-domain classification tasks; 2) TSAF can learn domain-invariant features efficiently. Xiaowei Xiang, Yang Liu 0246, Gaoyun Fang, Jing Liu 0050, Mengyang Zhao 0002 |
IEEE Signal Process. Lett. | 4 |
| 2022 | Learning Task-Specific Representation for Video Anomaly Detection with Spatial-Temporal AttentionabstractThe automatic detection of abnormal events in surveillance videos with weak supervision has been formulated as a multiple instance learning task, which aims to localize the clips containing abnormal events temporally with the video-level labels. However, most existing methods rely on the features extracted by the pre-trained action recognition models, which are not discriminative enough for video anomaly detection. In this work, we propose a spatial-temporal attention mechanism to learn inter- and intra-correlations of video clips, and the boosted features are encouraged to be task-specific via the mutual cosine embedding loss. Experimental results on standard benchmarks demonstrate the effectiveness of the spatial-temporal attention, and our method achieves superior performance to the state-of-the-art methods. Yang Liu 0246, Jing Liu 0050, Xiaoguang Zhu, Donglai Wei 0002 |
ICASSP | 2 |
| 2022 | Look, Listen and Pay More Attention: Fusing Multi-Modal Information for Video Violence DetectionabstractViolence detection is an essential and challenging problem in the computer vision community. Most existing works focus on single modal data analysis, which is not effective when multi-modality is available. Therefore, we propose a two-stage multi-modal information fusion method for violence detection: 1) the first stage adopts multiple instance learning strategies to refine video-level hard labels into clip-level soft labels, and 2) the next stage uses multi-modal information fused attention module to achieve fusion, and supervised learning is carried out using the soft labels generated at the first stage. Extensive empirical evidence on the XD-Violence dataset shows that our method outperforms the state-of-the-art methods. Donglai Wei 0002, Chen-Geng Liu, Yang Liu 0246, Jing Liu 0050, Xiao-Guang Zhu, Xinhua Zeng |
ICASSP | 4 |
| 2022 | Learning Appearance-Motion Normality for Video Anomaly DetectionabstractVideo anomaly detection is a challenging task in the Computer vision community. Most single task-based methods do not consider the independence of unique spatial and temporal patterns, while two-stream structures lack the exploration of the correlations. In this paper, we propose spatial-temporal memories augmented two-stream auto-encoder framework, which learns the appearance normality and motion normal-ity independently and explores the correlations via adversar-ial learning. Specifically, we first design two proxy tasks to train the two-stream structure to extract appearance and motion features in isolation. Then, the prototypical features are recorded in the corresponding spatial and temporal memory pools. Finally, the encoding-decoding network performs ad-versariallearning with the discriminator to explore the corre-lations between spatial and temporal patterns. Experimental results show that our framework outperforms the state-of-the-art methods, achieving AUCs of 98.1% and 89.8% on UCSD Ped2 and CUHK Avenue datasets. Yang Liu 0246, Jing Liu 0050, Mengyang Zhao 0002, Dingkang Yang, Xiaoguang Zhu |
ICME | 2 |
| 2022 | MAR2MIX: A Novel Model for Dynamic Problem in Multi-agent Reinforcement Learning
Gaoyun Fang, Yang Liu 0246, Jing Liu 0050 |
ICONIP (4) | 3 |
| 2022 | Exploiting Spatial-temporal Correlations for Video Anomaly DetectionabstractVideo anomaly detection (VAD) remains a challenging task in the pattern recognition community due to the ambiguity and diversity of abnormal events. Existing deep learning-based VAD methods usually leverage proxy tasks to learn the normal patterns and discriminate the instances that deviate from such patterns as abnormal. However, most of them do not take full advantage of spatial-temporal correlations among video frames, which is critical for understanding normal patterns. In this paper, we address unsupervised VAD by learning the evolution regularity of appearance and motion in the long and short-term and exploit the spatial-temporal correlations among consecutive frames in normal videos more adequately. Specifically, we proposed to utilize the spatiotemporal long short-term memory (ST-LSTM) to extract and memorize spatial appearances and temporal variations in a unified memory cell. In addition, inspired by the generative adversarial network, we introduce a discriminator to perform adversarial learning with the ST-LSTM to enhance the learning capability. Experimental results on standard benchmarks demonstrate the effectiveness of spatial-temporal correlations for unsupervised VAD. Our method achieves competitive performance compared to the state-of-the-art methods with AUCs of 96.7%, 87.8%, and 73.1% on the UCSD Ped2, CUHK Avenue, and ShanghaiTech, respectively. Mengyang Zhao 0002, Yang Liu 0246, Jing Liu 0050, Xinhua Zeng |
ICPR | 3 |
| 2022 | Abnormal Event Detection with Self-guiding Multi-instance Ranking FrameworkabstractThe detection of abnormal events in surveillance videos with weak supervision is a challenging task, which tries to temporally find abnormal frames using readily accessible video-level labels. In this paper, we propose a self-guiding multi-instance ranking (SMR) framework, which has explored task-specific deep representations and considered the temporal correlations between video clips. Specifically, we apply a clustering algorithm to fine-tune the features extracted by the pre-trained 3D-convolutional-based models. Besides, the clustering module can generate clip-level labels for abnormal videos, and the pseudo-labels are in part used to supervise the training of the multi-instance regression. While implementing the regression module, we compare the effectiveness of various recurrent neural networks, and the results demonstrate the necessity of temporal correlations for weakly supervised video anomaly detection tasks. Experimental results on two standard benchmarks reveal that the SMR framework is comparable to the state-of-the-art approaches, with frame-level AUCs of 81.7% and 92.4% on the UCF-crime and UCSD Ped2 datasets respectively. Additionally, ablation studies and visualization results prove the effectiveness of the component, and our framework can accurately locate abnormal events. Yang Liu 0246, Jing Liu 0050 |
IJCNN | 2 |
| 2022 | Attention-Based Auto-Encoder Framework for Abnormal Driving DetectionabstractWith the popularity of smartphones, abnormal driving detection via smartphone sensors has been proposed in recent years. However, existing methods are insufficient in exploring feature extraction, so the practical value is limited due to the low accuracy. To address this problem, we propose an attention-based auto-encoder framework for abnormal driving detection that combines the advantages of bi-directional long short-term memory and self-attention. Specifically, these two modules are embedded in the auto-encoder for modeling latent vector and exploring the internal correlations of spatial-temporal features, respectively, so as to improve the capability of reconstructing driving time series using small and representative features. We conduct experiments on the real-world datasets, and the results show that the proposed framework achieves significant performance with recall and F1-score of 96.2% and 95.0%, superior to the other baselines. Jing Liu 0050, Yang Liu 0246, Donglai Wei 0002, Xinhua Zeng |
ISCAS | 1 |
| 2022 | Multi-level Attention Fusion for Multimodal Driving Maneuver RecognitionabstractSensor-based driving maneuver recognition (DMR) is a fundamental and challenging task in ubiquitous computing, which uses multimodal signals from embedded sensors such as accelerometers and gyroscopes to recognize driving maneuvers. However, the spatial-temporal features from neural networks are often treated equally, which may limit the performance of the model in predicting maneuvers. In this paper, we propose a novel hybrid neural network model based on multi-level attention fusion for multimodal DMR. The proposed model utilizes convolutional neural networks and gated recurrent unit to extract temporal-spatial features from multimodal sensing signals and propose the multi-level attention fusion to explore the significant patterns over local and global periods. In addition, We design three different levels of fusion (early, late, and full fusion) to explore the effects of different attention fusions on the model. Extensive experiments on the real-world dataset show that the proposed model achieves superior performance to the baseline methods, and multi-level attention fusion brings 6.17% gain to the F1-score. Jing Liu 0050, Yang Liu 0246, Chengwen Tian, Mengyang Zhao 0002, Xinhua Zeng |
ISCAS | 1 |
| 2022 | MSAF: Multimodal Supervise-Attention Enhanced Fusion for Video Anomaly DetectionabstractThe complementarity of multimodal signal is essential for video anomaly detection. However, existing methods either lack exploration to multimodal data or ignore the implicit alignment of multimodal features. In our work, we address this problem using a novel fusion method and propose a Multimodal Supervise-Attention enhanced Fusion (MSAF) framework under weak supervision. Our framework can be divided into two parts: 1) the multimodal labels refinement part refines video-level ground truth into pseudo clip-level labels for subsequent training, 2) the multimodal supervise-attention fusion network enhances features via implicitly aligning different information, then fusing them effectively to predict anomaly scores with the help of refined labels. We validate our framework on four challenging datasets: ShanghaiTech, UCF-Crime, LAD, and XD-Violence. Extensive experiments on the benchmarks demonstrate the effectiveness of our framework, which achieves comparable results on several benchmarks and outperforms current state-of-the-art methods on the XD-Violence audiovisual multimodal dataset. Donglai Wei 0002, Yang Liu 0246, Xiaoguang Zhu, Jing Liu 0050, Xinhua Zeng |
IEEE Signal Process. Lett. | 4 |
| 2021 | Stack Multiple Shallow Autoencoders into a Strong One: A New Reconstruction-Based Method to Detect Anomaly
Hanqi Wang, Xing Hu 0006, Yang Liu 0246, Jing Liu 0050, Linhua Jiang |
ICONIP (1) | 6 |
| 2020 | Learning Progressive Joint Propagation for Human Motion Prediction
Yujun Cai, Lin Huang 0004, Yiwei Wang 0001, Tat-Jen Cham, Jianfei Cai 0001, Junsong Yuan 0001, Jun Liu 0036, Xu Yang 0021, Yiheng Zhu 0003, Xiaohui Shen, Ding Liu 0001, Jing Liu 0050, Nadia Magnenat-Thalmann |
ECCV (7) | 12 |