Nannan Li 0001

dblp:121/0837-1 · DBLP profile ↗
← Back
41ranked-venue papers
7as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 6 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 DEPTH: Disentangled embeddings and priors via two-stage heterogeneous-fusion
Nannan Li 0001, Kan Huang
Expert Syst. Appl.3
2026 Hierarchical sketch encoding and text-guided 3D modeling: a novel approach for single free-hand sketch reconstruction
Qinchuan Lei, Nannan Li 0001, Jiaming Zhong, Mengbao Fan
Vis. Comput.2
2025 3D-Telepathy: Reconstructing 3D Objects from EEG Signals
Yuxiang Ge, Jionghao Cheng, Ruiquan Ge, Zhaojie Fang, Gangyong Jia, Nannan Li 0001, Ahmed El-Azab, Changmiao Wang
IJCNN7
2025 Fabric-DETR: An Efficient Transformer Network for Multi-Scale Fabric Defect Detection in Complex Environments
abstract
Fabric defect detection plays a pivotal role in achieving intelligent quality control within textile manufacturing. However, intricate background textures, the high visual similarity between defects and the background, and the low proportion of small defects in high-resolution images impede detection accuracy. To address these challenges, we introduce Fabric-DETR, an efficient model for multi-scale fabric defect detection in complex environments. This network integrates three key modules: the Bottle Neck Conv2X Block for enhanced backbone feature extraction, Dynamic Attention-based Intra-scale Feature Interaction to improve attention to small targets, and the Zoom Diffuse Pyramid Network for efficient multi-scale feature fusion. Experiments on a six-class fabric defect dataset demonstrate that Fabric-DETR outperforms existing state-of-the-art methods, achieving a 94.8% mAP, representing improvements of 2.8%, 2.8%, 1.9%, 6.1%, 3.1%, and 2.5% over RT-DETR, YOLOv5-m, YOLOv8-m, YOLOv10-m, YOLOv11-m, and YOLOv12-m, respectively.
Fuqin Deng, Qingshan Xia, Lanhui Fu, Yingzhu Wu, Nannan Li 0001, Ningbo Yi, Guangming You
INDIN5
2025 CSRP: Modeling class spatial relation with prototype network for novel class discovery
Nannan Li 0001, Jiuqing Dong, Huiwen Guo, Wenmin Wang 0001, Chuanchuan You
Appl. Intell.2
2025 Causality thinking for large-scale long-tailed video action recognition
Zhengjin Zhang, Nannan Li 0001, Wenmin Wang 0001, Huiwen Guo, Sudan Huang
Eng. Appl. Artif. Intell.2
2025 GNN-based primitive recombination for compositional zero-shot learning
Fuqin Deng, Caiyun Tang, Lanhui Fu, Jiaming Zhong, Hongming Wang, Nannan Li 0001
Image Vis. Comput.7
2025 Multimodal Sensitive Adaptive Transformer for 3D medical image segmentation
abstract
Three-dimensional medical imaging segmentation presents a significant challenge within the field, with the segmentation of multiple organs and lesions in MRI images being particularly demanding. This paper introduces an innovative approach utilizing the Multimodal Sensitive Adaptive Attention (MSAA). We refer to this new structure as the Multimodal Sensitive Adaptive Transformer Network (MSAT), which incorporates downsampling and Multimodal Sensitive Adaptive Attention into the encoding phase and integrate skip connections from different layers, outputs from Multimodal Sensitive Adaptive Attention, and upsampled feature outputs into the decoding phase. The MSAT consists of two primary components. The initial component is designed to extract a richer set of high-dimensional features through an advanced network architecture. This includes integration of different layers skip connections, outputs from the MSAA, and the results of the preceding upsampling layer. The second component features a Multimodal Sensitive Adaptive Attention block, which integrates two types of attention mechanisms: Local Sensitive Adaptive Attention (LSAA) and Spatial Sensitive Adaptive Attention (SSAA). These attention mechanisms work synergistically to blend high and low-dimensional features effectively, thereby enriching the contextual information captured by the model. Our experiments, conducted across several datasets including Synapse, BTCV, ACDC, and the BraTS 2021 dataset, demonstrate that the MSAT outperforms other existing methodologies. The MSAT shows superior segmentation capabilities for 3D multi-organ, cardiac, and brain tumor segmentation tasks.
Zhibing Wang, Wenmin Wang 0001, Nannan Li 0001, Yifan Zhang 0028, Haomei Jia, Shenyong Zhang
Image Vis. Comput.3
2025 Causalseg: investigating causality modeling for semi-supervised video object segmentation
Zhengjin Zhang, Nannan Li 0001, Wenmin Wang 0001, Huiwen Guo
Multim. Syst.2
2025 Local Texture Pattern Estimation for Image Detail Super-Resolution
abstract
In the image super-resolution (SR) field, recovering missing high-frequency textures has always been an important goal. However, deep SR networks based on pixel-level constraints tend to focus on stable edge details and cannot effectively restore random high-frequency textures. It was not until the emergence of the generative adversarial network (GAN) that GAN-based SR models achieved realistic texture restoration and quickly became the mainstream method for texture SR. However, GAN-based SR models still have some drawbacks, such as relying on a large number of parameters and generating fake textures that are inconsistent with ground truth. Inspired by traditional texture analysis research, this paper proposes a novel SR network based on local texture pattern estimation (LTPE), which can restore fine high-frequency texture details without GAN. A differentiable local texture operator is first designed to extract local texture structures, and a texture enhancement branch is used to predict the high-resolution local texture distribution based on the LTPE. Then, the predicted high-resolution texture structure map can be used as a reference for the texture fusion SR branch to obtain high-quality texture reconstruction. Finally, $L_{1}$L1 loss and Gram loss are simultaneously used to optimize the network. Experimental results demonstrate that the proposed method can effectively recover high-frequency texture without using GAN structures. In addition, the restored high-frequency details are constrained by local texture distribution, thereby reducing significant errors in texture generation.
Yang Zhao 0002, Yuan Chen 0012, Nannan Li 0001, Wei Jia 0001, Ronggang Wang
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 Frequency-Domain Convolutional Network With Historical Data Fusion Module for Regional Streamflow Prediction
abstract
Accurate runoff prediction is essential for effective water resource management, particularly in addressing flood control and monitoring drought conditions. However, the diverse nature of land types and varying climate conditions often complicate this task, requiring frequent adaptations to prediction models for local applications. Existing methods primarily focus on modeling for individual regions, while regional runoff prediction models cannot often learn long-term patterns, limiting their regional adaptability. To overcome this challenge, we present the temporal fusion runoff network (TFRN), a new framework designed to enhance long short-term memory (LSTM) models by enabling them to incorporate distant historical information. This innovation offers a promising framework for regional runoff prediction by enhancing model performance and minimizing computational demands. In this study, the proposed TFRN utilizes convolutional networks to extract and integrate both long-term and short-term trends from input sequences, and by merging the strengths of LSTM and Transformer architectures, TFRN achieves a thorough integration of historical data. Specifically, our method employs convolutional networks across both time and frequency domains to capture multi-scale features. Within the Transformer component, we introduce an adaptive fusion module to improve the integration of historical information. We validated the effectiveness of our model using two extensive hydrological datasets for a 7-day runoff prediction task. The results underscore the superiority of our approach, demonstrating its advantages over several leading methods. The source code is available at https://github.com/redtea-code/TFRN.
Yuanhao Chen, Haoqi Yu, Jingrong Dai, Nannan Li 0001, Changmiao Wang, Ahmed El-Azab
IEEE Trans. Geosci. Remote. Sens.6
2024 CCLNet: Causal and Contrastive Learning Framework for Enhanced Pulmonary Embolism Detection
abstract
The fusion of multimodal medical data is crucial for helping doctors make accurate treatment decisions. For example, combining Computed Tomography Pulmonary Angiography (CTPA) with Electronic Health Records (EHR) can significantly improve the accuracy of Pulmonary Embolism (PE) detection, thereby increasing patient survival rates. Although multimodal learning has advantages in PE diagnosis, the heterogeneity of multimodal data poses a significant challenge to accurate diagnosis. The natural semantic and structural differences between data modalities make it difficult to effectively integrate their information. In addition, within a single modality, the existence of redundant and irrelevant information introduces unnecessary variability, making the data more complex, and making stable diagnosis challenging. To address these issues, we propose a new framework called CCLNet, which includes a contrastive learning component for addressing inter-modality heterogeneity and a causal learning component for handling intra-modality heterogeneity. Specifically, we achieve precise alignment between visual and tabular modalities by using global-level information to soften labels during contrastive learning. In addition, by using causal intervention methods to eliminate the influence of heterogeneous factors within the modality, we can accurately reveal the causal relationship between features and targets, thereby improving the accuracy and stability of the model. Experimental results demonstrate that our method performs excellently, achieving the best results. Our code is available at https://github.com/LeavingStarW/CLPE.
Ruiquan Ge, Jianxun Yu, Fei-wei Qin, Nannan Li 0001, Wenwen Min, Ahmed El-Azab, Changmiao Wang
BIBM6
2024 Make an Image Move: Few-Shot Based Video Generation Guided by CLIP
Yonglong Huang, Nannan Li 0001, Fuqin Deng, Ruiquan Ge, Changmiao Wang
ICPR (6)3
2024 Integrating pseudo labeling with contrastive clustering for transformer-based semi-supervised action recognition
Nannan Li 0001, Kan Huang, Qingtian Wu, Yang Zhao 0002
Appl. Intell.1
2024 Multimodal parallel attention network for medical image segmentation
Zhibing Wang, Wenmin Wang 0001, Nannan Li 0001, Shenyong Zhang
Image Vis. Comput.3
2024 Exploiting Memory-Based Cross-Image Contexts for Salient Object Detection in Optical Remote Sensing Images
abstract
Current state-of-the-art methods for salient object detection in optical remote sensing images (RSI-SOD) primarily relies on individual image context to detect salient objects. However, the potential of cross-image contexts remains largely unexplored in existing works, which can provide valuable auxiliary and complementary information for discriminating object representations in RSIs. In this paper, we investigate the utilization of cross-image contextual information for RSI-SOD. We propose a novel memory-based context propagation network (MCP-Net) to harness dataset-level contextual information. MCP-Net incorporates a cross-image dual memory module (CDM) to store dataset-level information and utilize it to generate contextual information for the current image. CDM effectively captures intra-scene variations by leveraging both foreground and background memory banks, resulting in improved object representations. Additionally, we enhance the representations by leveraging scale-aware context information within individual images. To preserve RSI details before memory modules, we introduce a shared attention-guided fusion module (SAF) to align the adjacent network-level features. Extensive evaluation results demonstrate the superior performance of our proposed method compared to state-of-the-art methods on three public benchmarks. These results affirm that the inclusion of cross-image contexts can significantly benefit salient object detection in remote sensing images.
Kan Huang, Nannan Li 0001, Jiarong Huang, Chunwei Tian
IEEE Trans. Geosci. Remote. Sens.2
2024 Driver Drowsiness Detection Based on Joint Human Face and Facial Landmark Localization With Cheap Operations
abstract
Real-time detection of driver drowsiness is critical to reduce the risk of road accidents and fatalities. Current facial landmark-based methods usually use a two-stage paradigm, where faces and facial landmarks are localized separately. Additionally, most methods can be hindered by challenging conditions, such as night driving or eyes closed. To address these challenges, we present a refined YOLO network named YOLOFaceMark that can simultaneously detect faces and their facial landmarks. Furthermore, we introduce a drowsiness detection model based on facial landmarks. This model utilizes extracted eye and mouth information to identify drowsy states. We optimize the original YOLO components through structural re-parameterization, channel shuffling, and the design of a dual-branch detection head with an implicit module. These enhancements are designed to improve the accuracy while maintaining computational efficiency. We validate the real-time performance and accuracy of YOLOFaceMark on public datasets, including 300W and COFW. Additionally, we conduct further validation to demonstrate our ability to achieve effective and robust drowsiness detection solely based on the facial landmarks detected by YOLOFaceMark.
Qingtian Wu, Nannan Li 0001, Liming Zhang 0002, F. Richard Yu
IEEE Trans. Intell. Transp. Syst.2
2023 Motion Context guided Edge-preserving network for video salient object detection
Kan Huang, Chunwei Tian, Zhijing Xu, Nannan Li 0001, Jerry Chun-Wei Lin
Expert Syst. Appl.4
2022 A Unified Weight Initialization Paradigm for Tensorial Convolutional Neural Networks
abstract
Tensorial Convolutional Neural Networks (TCNNs) have attracted much research attention for their power in reducing model parameters or enhancing the generalization ability. However, exploration of TCNNs is hindered even from weight initialization methods. To be specific, general initialization methods, such as Xavier or Kaiming initialization, usually fail to generate appropriate weights for TCNNs. Meanwhile, although there are ad-hoc approaches for specific architectures (e.g., Tensor Ring Nets), they are not applicable to TCNNs with other tensor decomposition methods (e.g., CP or Tucker decomposition). To address this problem, we propose a universal weight initialization paradigm, which generalizes Xavier and Kaiming methods and can be widely applicable to arbitrary TCNNs. Specifically, we first present the Reproducing Transformation to convert the backward process in TCNNs to an equivalent convolution process. Then, based on the convolution operators in the forward and backward processes, we build a unified paradigm to control the variance of features and gradients in TCNNs. Thus, we can derive fan-in and fan-out initialization for various TCNNs. We demonstrate that our paradigm can stabilize the training of TCNNs, leading to faster convergence and better results.
Yu Pan 0005, Zeyong Su, Ao Liu 0008, Jingquan Wang, Nannan Li 0001, Zenglin Xu
ICML5
2022 Weakly-supervised anomaly detection in video surveillance via graph convolutional label noise cleaning
Nannan Li 0001, Jia-Xing Zhong, Xiujun Shu, Huiwen Guo
Neurocomputing1
2020 Spatial-Temporal Context-Aware Online Action Detection and Prediction
abstract
Spatial-temporal action detection in videos is a challenging problem that has attracted considerable attention in recent years. Most current approaches address action detection as an object detection problem, which utilizes successful object detection frameworks such as Faster R-CNN to operate action detection at every single frame first, and then generates action tubes by linking bounding boxes across the whole video in an offline fashion. However, unlike object detection in static images, temporal context information is vital for action detection in videos. Therefore, we propose an online action detection model that leverages the spatial-temporal context information existing in videos to perform action inference and localization. More specifically, we try to depict the spatial-temporal context pattern of actions via an encoder-decoder model that is based on a convolutional recurrent neural network. The model accepts a video snippet as input and encodes the dynamic information inside the snippet in the forward pass. During the backward pass, the decoder resolves the information for action detection with the current appearance or motion cue at each time stamp. In addition, we devise an incremental action-tube construction algorithm that enables our model to accomplish action prediction ahead of time and performs action detection in an online fashion. To evaluate the performance of our method, we conduct experiments on three popular public datasets UCF-101, UCF-Sports, and J-HMDB-21. The experimental results demonstrate that our method can achieve competitive or superior performance when compared to the state-of-the-art methods. To encourage further research, we release our project on “https://github.com.hjjpku.OATD.”
Jingjia Huang, Nannan Li 0001, Thomas H. Li, Shan Liu 0001, Ge Li 0002
IEEE Trans. Circuits Syst. Video Technol.2
2019 Graph Convolutional Label Noise Cleaner: Train a Plug-And-Play Action Classifier for Anomaly Detection
abstract
Video anomaly detection under weak labels is formulated as a typical multiple-instance learning problem in previous works. In this paper, we provide a new perspective, i.e., a supervised learning task under noisy labels. In such a viewpoint, as long as cleaning away label noise, we can directly apply fully supervised action classifiers to weakly supervised anomaly detection, and take maximum advantage of these well-developed classifiers. For this purpose, we devise a graph convolutional network to correct noisy labels. Based upon feature similarity and temporal consistency, our network propagates supervisory signals from high-confidence snippets to low-confidence ones. In this manner, the network is capable of providing cleaned supervision for action classifiers. During the test phase, we only need to obtain snippet-wise predictions from the action classifier without any extra post-processing. Extensive experiments on 3 datasets at different scales with 2 types of action classifiers demonstrate the efficacy of our method. Remarkably, we obtain the frame-level AUC score of 82.12% on UCF-Crime.
Jia-Xing Zhong, Nannan Li 0001, Weijie Kong, Shan Liu 0001, Thomas H. Li, Ge Li 0002
CVPR2
2019 BLP - Boundary Likelihood Pinpointing Networks for Accurate Temporal Action Localization
abstract
Despite tremendous progress achieved in temporal action detection, state-of-the-art methods still suffer from the sharp performance deterioration when localizing the starting and ending temporal action boundaries. Although most methods apply boundary regression paradigm to tackle this problem, we argue that the direct regression lacks detailed enough information to yield accurate temporal boundaries. In this paper, we propose a novel Boundary Likelihood Pinpointing (BLP) network to alleviate this deficiency of boundary regression and improve the localization accuracy. Given a loosely localized search interval that contains an action instance, BLP casts the problem of localizing temporal boundaries as that of assigning probabilities on each equally divided unit of this interval. These generated probabilities provide useful information regarding the boundary location of the action inside this search interval. Based on these probabilities, we introduce a boundary pinpointing paradigm to pinpoint the accurate boundaries under a simple probabilistic framework. Compared with other C3D feature based detectors, extensive experiments demonstrate that BLP significantly improves the localization performance of recent state-of-the-art detectors, and achieves competitive detection mAP on both THUMOS' 14 and ActivityNet datasets, particularly when the evaluation tIoU is high.
Weijie Kong, Nannan Li 0001, Shan Liu 0001, Thomas H. Li, Ge Li 0002
ICASSP2
2019 AttPool: Towards Hierarchical Feature Representation in Graph Convolutional Networks via Attention Mechanism
abstract
Graph convolutional networks (GCNs) are potentially short of the ability to learn hierarchical representation for graph embedding, which holds them back in the graph classification task. Here, we propose AttPool, which is a novel graph pooling module based on attention mechanism, to remedy the problem. It is able to select nodes that are significant for graph representation adaptively, and generate hierarchical features via aggregating the attention-weighted information in nodes. Additionally, we devise a hierarchical prediction architecture to sufficiently leverage the hierarchical representation and facilitate the model learning. The AttPool module together with the entire training structure can be integrated into existing GCNs, and is trained in an end-to-end fashion conveniently. The experimental results on several graph-classification benchmark datasets with various scales demonstrate the effectiveness of our method.
Jingjia Huang, Zhangheng Li, Nannan Li 0001, Shan Liu 0001, Ge Li 0002
ICCV3
2018 SAP: Self-Adaptive Proposal Model for Temporal Action Detection Based on Reinforcement Learning
abstract
Existing action detection algorithms usually generate action proposals through an extensive search over the video at multiple temporal scales, which brings about huge computational overhead and deviates from the human perception procedure. We argue that the process of detecting actions should be naturally one of observation and refinement: observe the current window and refine the span of attended window to cover true action regions. In this paper, we propose a Self-Adaptive Proposal (SAP) model that learns to find actions through continuously adjusting the temporal bounds in a self-adaptive way. The whole process can be deemed as an agent, which is firstly placed at the beginning of the video and traverse the whole video by adopting a sequence of transformations on the current attended region to discover actions according to a learned policy. We utilize reinforcement learning, especially the Deep Q-learning algorithm to learn the agent’s decision policy. In addition, we use temporal pooling operation to extract more effective feature representation for the long temporal window, and design a regression network to adjust the position offsets between predicted results and the ground truth. Experiment results on THUMOS’14 validate the effectiveness of SAP, which can achieve competitive performance with current action detection algorithms via much fewer proposals.
Jingjia Huang, Nannan Li 0001, Tao Zhang 0069, Ge Li 0002, Tiejun Huang 0001, Wen Gao 0001
AAAI2
2018 An Active Action Proposal Method Based on Reinforcement Learning
abstract
Detecting human activities in untrimmed video is a significant yet challenging task. Existing methods usually generate temporal action proposals via searching extensively at multiple preset scales or combining a bunch of short video snippets. However, we argue that the localization of action instances should be a process of observation, refinement and determination: observe the attended temporal window, refine its position and scale, then determine whether a true action region has been accurately found. To this end, we formulate temporal action localization task as a Markov Decision Process, and propose an active temporal action proposal model based on reinforcement learning. Our model learns to localize actions in videos by automatically adjusting the position and span of temporal window via a sequence of transformations. We train an action/non-action binary classifier to determine whether a temporal window contains an action instance. Validation results on THUMOS'14 dataset show that our proposed method achieves competitive performance both in accuracy and efficiency compared with some state-of-the-art methods, while using much less proposals.
Tao Zhang 0069, Nannan Li 0001, Jingjia Huang, Jia-Xing Zhong, Ge Li 0002
ICIP2
2018 Online Action Tube Detection via Resolving the Spatio-temporal Context Pattern
abstract
At present, spatio-temporal action detection in the video is still a challenging problem, considering the complexity of the background, the variety of the action or the change of the viewpoint in the unconstrained environment. Most of current approaches solve the problem via a two-step processing: first detecting actions at each frame; then linking them, which neglects the continuity of the action and operates in an offline and batch processing manner. In this paper, we attempt to build an online action detection model that introduces the spatio-temporal coherence existed among action regions when performing action category inference and position localization. Specifically, we seek to represent the spatio-temporal context pattern via establishing an encoder-decoder model based on the convolutional recurrent network. The model accepts a video snippet as input and encodes the dynamic information of the action in the forward pass. During the backward pass, it resolves such information at each time instant for action detection via fusing the current static or motion cue. Additionally, we propose an incremental action tube generation algorithm, which accomplishes action bounding-boxes association, action label determination and the temporal trimming in a single pass. Our model takes in the appearance, motion or fused signals as input and is tested on two prevailing datasets, UCF-Sports and UCF-101. The experiment results demonstrate the effectiveness of our method which achieves a performance superior or comparable to compared existing approaches.
Jingjia Huang, Nannan Li 0001, Jia-Xing Zhong, Thomas H. Li, Ge Li 0002
ACM Multimedia2
2018 Step-by-step Erasion, One-by-one Collection: A Weakly Supervised Temporal Action Detector
abstract
Weakly supervised temporal action detection is a Herculean task in understanding untrimmed videos, since no supervisory signal except the video-level category label is available on training data. Under the supervision of category labels, weakly supervised detectors are usually built upon classifiers. However, there is an inherent contradiction between classifier and detector; i.e., a classifier in pursuit of high classification performance prefers top-level discriminative video clips that are extremely fragmentary, whereas a detector is obliged to discover the whole action instance without missing any relevant snippet. To reconcile this contradiction, we train a detector by driving a series of classifiers to find new actionness clips progressively, via step-by-step erasion from a complete video. During the test phase, all we need to do is to collect detection results from the one-by-one trained classifiers at various erasing steps. To assist in the collection process, a fully connected conditional random field is established to refine the temporal localization outputs. We evaluate our approach on two prevailing datasets, THUMOS'14 and ActivityNet. The experiments show that our detector advances state-of-the-art weakly supervised temporal action detection results, and even compares with quite a few strongly supervised methods.
Jia-Xing Zhong, Nannan Li 0001, Weijie Kong, Tao Zhang 0069, Thomas H. Li, Ge Li 0002
ACM Multimedia2
2018 Deep Pedestrian Detection Using Contextual Information and Multi-level Features
Weijie Kong, Nannan Li 0001, Thomas H. Li, Ge Li 0002
MMM (1)2
2018 Detecting action tubes via spatial action estimation and temporal path inference
Nannan Li 0001, Jingjia Huang, Thomas H. Li, Huiwen Guo, Ge Li 0002
Neurocomputing1
2017 A Violence Detection Approach Based on Spatio-temporal Hypergraph Transition
Jingjia Huang, Ge Li 0002, Nannan Li 0001, Ronggang Wang, Wenmin Wang 0001
CAIP (2)3
2017 Fast action localization based on spatio-temporal path search
abstract
In this paper, a method is proposed to search for spatio-temporal path for action localization in unconstrained videos. We mainly focus on two requirements, i.e., accurate human extraction and speeding generation of action proposal. The approach first generates human proposals at the frame level, then scores them based on two complementary parts, i.e., posteriori probability evaluated via a fine-tuned Faster-RCNN and template-matching similarity based on the spatiotemporal continuity. Finally, the generation of action proposal is formulated as a Max-Path discovery problem, coupled with dynamic programming to find an optimal path with maximum score. Experiments on UCF-Sports are performed to verify that the proposed method can achieve fast high-quality action proposal and link the missed-detection proposals in successive frames together to form a complete action.
Qingtian Wu, Huiwen Guo, Xinyu Wu 0001, Yimin Zhou 0001, Nannan Li 0001
ICIP5
2017 Deep Metric Learning with False Positive Probability - Trade Off Hard Levels in a Weighted Way
Jia-Xing Zhong, Ge Li 0002, Nannan Li 0001
ICONIP (3)3
2016 Searching Action Proposals via Spatial Actionness Estimation and Temporal Path Inference and Tracking
Nannan Li 0001, Dan Xu 0006, Zhenqiang Ying, Zhihao Li 0002, Ge Li 0002
ACCV (2)1
2016 Tube ConvNets: Better exploiting motion for action recognition
abstract
Motion information is a key factor for action recognition and has been eagerly pursued for decades. How to effectively learn motion features in Convolutional Networks (ConvNets) remains an open issue. Prevalent ConvNets often take several full frames of video as input at a time, which can be a heavy burden for network training. In this paper, we introduce a novel framework called Tube ConvNets, by substituting action tubes for full frames to reduce this burden. Tube ConvNets focus on the regions of interest (ROI) where key motions occur, and thus eliminate the distraction of irrelevant objects. Each action tube is a fraction of spatiotemporal volumes, generated by the techniques of object detection and clustering algorithm. We demonstrate the effectiveness of Tube ConvNets for action classification on UCF-101 dataset, and illustrate its potential to support fine-grained localization on UCF-Sports dataset. Source code is available at https://github.com/wangjinzhuo/tubecnn.
Zhihao Li 0002, Wenmin Wang 0001, Nannan Li 0001, Jinzhuo Wang
ICIP3
2016 Quaternion discrete cosine transformation signature analysis in crowd scenes for abnormal event detection
Huiwen Guo, Xinyu Wu 0001, Shibo Cai, Nannan Li 0001, Jun Cheng 0002, Yen-Lun Chen
Neurocomputing4
2015 Spatio-temporal context analysis within video volumes for anomalous-event detection and localization
Nannan Li 0001, Xinyu Wu 0001, Dan Xu 0006, Huiwen Guo, Wei Feng 0009
Neurocomputing1
2015 Anomaly Detection in Video Surveillance via Gaussian Process
abstract
In this paper, we propose a new approach for anomaly detection in video surveillance. This approach is based on a nonparametric Bayesian regression model built upon Gaussian process priors. It establishes a set of basic vectors describing motion patterns from low-level features via online clustering, and then constructs a Gaussian process regression model to approximate the distribution of motion patterns in kernel space. We analyze different anomaly measure criterions derived from Gaussian process regression model and compare their performances. To reduce false detections caused by crowd occlusion, we utilize supplement information from previous frames to assist in anomaly detection for current frame. In addition, we address the problem of hyperparameter tuning and discuss the method of efficient calculation to reduce computation overhead. The approach is verified on published anomaly detection datasets and compared with other existing methods. The experiment results demonstrate that it can detect various anomalies efficiently and accurately.
Nannan Li 0001, Xinyu Wu 0001, Huiwen Guo, Dan Xu 0006, Yongsheng Ou, Yen-Lun Chen
Int. J. Pattern Recognit. Artif. Intell.1
2014 Multi-scale analysis of contextual information within spatio-temporal video volumes for anomaly detection
abstract
In this paper, we present a novel approach for video anomaly detection in crowded scenes. The proposed approach detects anomalies based on the contextual information analysis within spatio-temporal video volume. Around each pixel, spatio-temporal volumes are built and clustered to construct the activity pattern codebook. Then, the composition information of the volumes within a large spatiotemporal window is described via a dictionary learned by sparse representation. Furthermore, multi-scale analysis is employed to adapt the size change of abnormal events. Finally, the sparse reconstruction cost is designed to evaluate the abnormal level of an input motion pattern. We demonstrate the efficiency of the proposed method on the existing public available anomaly-detection datasets and the performance comparasion with three existing methods validates that the proposed method detects anomalies more accurately.
Nannan Li 0001, Huiwen Guo, Dan Xu 0006, Xinyu Wu 0001
ICIP1
2014 Video anomaly detection based on a hierarchical activity discovery within spatio-temporal contexts
Dan Xu 0006, Rui Song 0002, Xinyu Wu 0001, Nannan Li 0001, Wei Feng 0009, Huihuan Qian
Neurocomputing4
2013 Hierarchical activity discovery within spatio-temporal context for video anomaly detection
abstract
In this paper, we present a novel approach for video anomaly detection in crowded and complicated scenes. The proposed approach detects anomalies based on a hierarchical activity pattern discovery framework comprehensively considering both global and local spatio-temporal contexts. The discovery is a coarse-to-fine learning process with unsupervised ways for automatically constructing normal activity patterns at different levels. An unified anomaly energy function is designed based on these discovered activity patterns to identify the abnormal level of an input motion pattern. We demonstrate the efficiency of the proposed method on the UCSD anomaly detection datasets (Ped1 and Ped2) and compare the performance with existing work.
Dan Xu 0006, Xinyu Wu 0001, Dezhen Song, Nannan Li 0001, Yen-Lun Chen
ICIP4