EDBT 2026 Demo / reviewers in the wild / expert
Faliang Chang
dblp:29/4980
· DBLP profile ↗
41ranked-venue papers
2as first author
27since 2021 · last 2025
0000-0003-1276-2267ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 9 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Long-Short Observation-driven Prediction Network for pedestrian crossing intention prediction with momentary observation
Hui Liu 0040, Chunsheng Liu 0001, Faliang Chang, Yansha Lu, Minhang Liu |
Neurocomputing | 3 |
| 2025 | End-Point Drive and Reverse Enhanced Decoding- Based Traffic Participants Trajectory Prediction Under Bird's Eye ViewabstractTrajectory prediction under Bird’s Eye View (BEV) refers to predicting the movement intention of agents based on the historical observation trajectories, which is of great significance for autonomous driving, driving safety and social navigation. The traffic prediction trajectory is multimodality with multiple reasonable trajectories and various prediction time, which often suffer complex agents interaction, cumulative errors and a large number of agents. To overcome these problems, we explore the BEV-based traffic participants trajectory prediction problem and propose the novel End-point Drive and Reverse Enhanced Decoding Network (EDRED-TPNet), based on the mechanisms of end-point driving and reverse enhanced decoding. Firstly, the Encoder based on Dynamic Spatio-temporal Graph and Multimodality Coding Fusion (ST-MC-Encoder) are constructed to effectively represent complex traffic scenario with changeable agents, encode social interactions with historical trajectories, social interaction, future trajectory and multimodality. Secondly, the End-Point Drive Module is proposed to predict the end point before predicting the complete trajectory, thus providing more accurate trajectory prediction; Lastly, to further improve the long-term prediction performance, the Reverse Enhanced Decoder (RE-Decoder) is proposed to fuse forward and reverse hidden state vectors to obtain diverse trajectories that conform to physical and social acceptability rules. We build the first AAV-captured Trajectory Prediction Dataset (UTP-Dataset) for traffic participants trajectory prediction. Experimental results show that the proposed methods can fulfill the multi-target trajectory prediction task in complex traffic scenarios and achieve high performance. Chunsheng Liu 0001, Jincan Xie, Faliang Chang, Shuang Li 0016, Yinhai Wang |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | UAV-Based Vehicle Re-Identification via Counterfactual Attention Learning and Hard-Sensitive Binomial Joint PromotionabstractAutonomous Aerial Vehicles (AAVs) based Vehicle Re-Identification (ReID) brings high flexibility to the ReID system, and also brings challenges of complicated shooting views and special occlusions. In this study, we propose a novel framework for UAV-based vehicle ReID, calledCounterfactual Attention and Joint Promotion based ReID(CAJP-ReID), to deal with fine-grained feature extraction and occlusions. Based on the mechanism of using randomly generated counterfactual attention intervention to train, theCounterfactual Attention Learning Network(CAL-Net) is proposed to learn the fine-grained features of vehicle images, for distinguishing similar vehicles in different ID categories. In order to enhance the diversity of the dataset and the robustness of the network to occluded images, theCounterfactual Attention Enhancement Module(CAE-Module) is proposed based on counterfactual attention mechanism, by the data-enriching mechanism of cropping and erasing. TheHard-sensitive Binomial Joint Promotion Loss(HBJP-Loss) is proposed to comprehensively consider the relative distance and absolute distance of positive and negative samples of vehicle images, and further improves the accuracy of vehicle re-identification. Experiments on two public datasets show that the proposed method achieves State-of-the-Art performance. Chunsheng Liu 0001, Baoqi Xue, Shuang Li 0016, Faliang Chang, Nanjun Li, Yinhai Wang |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2025 | EraW-Net: Enhance-Refine-Align W-Net for Scene-Associated Driver Attention EstimationabstractAssociating driver attention with driving scene across two fields of view is a challenging cross-domain perception problem, which requires comprehensive consideration of cross-view mapping, dynamic driving scene analysis and driver status tracking. Previous methods typically analyze a single view or map attention to the scene through a two-step projection, failing to exploit their implicit connections and establish accurate associations. Moreover, simple fusion modules are inadequate for modeling the complex relationships between the two views, making information integration complicated. To address these issues, we propose EraW-Net, a novel end-to-end framework for scene-associated driver attention estimation by aggregating information from dual views. This method enhances the most discriminative dynamic cues, refines feature representations, and facilitates semantically aligned cross-domain integration through a W-shaped architecture, termed W-Net. Specifically, a Dynamic Adaptive Filter Module (DAF-Module) is proposed to address the challenges of frequently changing driving environments by extracting vital regions. It suppresses the indiscriminately recorded dynamics and highlights crucial ones by innovative joint frequency-spatial analysis, enhancing the model's ability to parse complex dynamics. Additionally, to track driver states during non-fixed facial poses, we propose a Global Context Sharing Module (GCS-Module) to construct refined feature representations by capturing hierarchical features that adapt to various scales of head and eye movements. Finally, W-Net achieves systematic cross-view information integration through its unique two-stage decoding strategy, addressing semantic misalignment in heterogeneous data integration. Experiments demonstrate that the proposed method robustly and accurately estimates scene-associated driver attention on large public datasets. Chunsheng Liu 0001, Faliang Chang, Penghui Hao, Yiming Huang 0005 |
IEEE Trans. Multim. | 3 |
| 2025 | TODO-Net: Temporally Observed Domain Contrastive Network for 3-D Early Action PredictionabstractEarly action prediction aiming to recognize which classes the actions belong to before they are fully conveyed is a very challenging task, owing to the insufficient discrimination information caused by the domain gaps among different temporally observed domains. Most of the existing approaches focus on using fully observed temporal domains to "guide" the partially observed domains while ignoring the discrepancies between the harder low-observed temporal domains and the easier highly observed temporal domains. The recognition models tend to learn the easier samples from the highly observed temporal domains and may lead to significant performance drops on low-observed temporal domains. Therefore, in this article, we propose a novel temporally observed domain contrastive network, namely, TODO-Net, to explicitly mine the discrimination information from the hard actions samples from the low-observed temporal domains by mitigating the domain gaps among various temporally observed domains for 3-D early action prediction. More specifically, the proposed TODO-Net is able to mine the relationship between the low-observed sequences and all the highly observed sequences belonging to the same action category to boost the recognition performance of the hard samples with fewer observed frames. We also introduce a temporal domain conditioned supervised contrastive (TD-conditioned SupCon) learning scheme to empower our TODO-Net with the ability to minimize the gaps between the temporal domains within the same action categories, meanwhile pushing apart the temporal domains belonging to different action classes. We conduct extensive experiments on two public 3-D skeleton-based activity datasets, and the results show the efficacy of the proposed TODO-Net. Faliang Chang, Chunsheng Liu 0001, Bin Wang 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Hierarchical Diffusion Policy: Manipulation Trajectory Generation via Contact GuidanceabstractDecision-making in robotics using denoising diffusion processes has increasingly become a hot research topic, but end-to-end policies perform poorly in tasks with rich contact and have limited interactivity. This article proposes Hierarchical Diffusion Policy (HDP), a new robot manipulation policy of using contact points to guide the generation of robot trajectories. The policy is divided into two layers: the high-level policy predicts the contact for the robot's next object manipulation based on 3-D information, while the low-level policy predicts the action sequence toward the high-level contact based on the latent variables of observation and contact. We represent both-level policies as conditional denoising diffusion processes, and combine behavioral cloning and Q-learning to optimize the low-level policy for accurately guiding actions towards contact. We benchmark Hierarchical Diffusion Policy across six different tasks and find that it significantly outperforms the existing state-of-the-art imitation learning method Diffusion Policy with an average improvement of 20.8% . We find that contact guidance yields significant improvements, including superior performance, greater interpretability, and stronger interactivity, especially on contact-rich tasks. To further unlock the potential of HDP, this article proposes a set of key technical contributions including one-shot gradient optimization, trajectory augmentation, and prompt guidance, which improve the policy's optimization efficiency, spatial awareness, and interactivity respectively. Finally, real-world experiments verify that HDP can handle both rigid and deformable objects. Chunsheng Liu 0001, Faliang Chang, Yichen Xu 0006 |
IEEE Trans. Robotics | 3 |
| 2024 | Human-robot interaction-oriented video understanding of human actions
Bin Wang 0004, Faliang Chang, Chunsheng Liu 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | An efficient motion visual learning method for video action recognition
Bin Wang 0004, Faliang Chang, Chunsheng Liu 0001, Ruiyi Ma |
Expert Syst. Appl. | 2 |
| 2024 | ACF-net: appearance-guided content filter network for video captioning
Dongmei Liu 0007, Chunsheng Liu 0001, Faliang Chang, Bin Wang 0004 |
Multim. Tools Appl. | 4 |
| 2024 | Drone-captured vehicle re-identification via perspective mask segmentation and hard sample learning
Chunsheng Liu 0001, Baoqi Xue, Shuang Li 0016, Faliang Chang |
Multim. Tools Appl. | 4 |
| 2024 | Egocentric Vulnerable Road Users Trajectory Prediction With Incomplete ObservationabstractVulnerable Road Users (VRUs) trajectory prediction aims to analyze the future movements of pedestrians and cyclists for intelligent driving. Most previous methods just focus on VRUs trajectory prediction using idealized complete observations, and rare consider occlusions and tracking losses. Focus on the incomplete observation problem, we propose a novel Observation Store and Query Fusion Network (OSQF-Net), for VRUs trajectory prediction with incomplete observation. Firstly, based on the external memory bank mechanism and complete-incomplete observation joint training strategy, a Memory Bank-based Feature Store and Query Module (MSQ-Module) is proposed to extract complete motion features, from disrupted motion patterns caused by incomplete observation. Subsequently, based on temporal extraction and attention mechanism, a Spatio-Temporal Fusion Module (STF-Module) is proposed to effectively fuse the pseudo-complete motion features and incomplete motion features in both spatial and temporal dimensions. Finally, with these two modules and a CVAE network, the OSQF-Net can generate a latent space with complete motion patterns, which guides future trajectory prediction. Experimental results demonstrate that OSQF-Net achieves superior prediction performance and real-time inference capability for egocentric VRUs trajectory prediction, under both complete and incomplete observation scenarios. Hui Liu 0040, Chunsheng Liu 0001, Faliang Chang, Yansha Lu, Minhang Liu |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Traffic Scenario Understanding and Video Captioning via Guidance Attention Captioning NetworkabstractDescribing a traffic scenario from the driver’s perspective is a challenging process for Advanced Driving Assistance System (ADAS), involving different sub-tasks of detection, tracking, segmentation, etc. Previous methods mainly focus on independent sub-tasks and have difficulties to comprehensively describe the incidents. In this study, this problem is novelly treated as a video captioning task, and a Guidance Attention Captioning Network (GAC-Network) structure is proposed for describing the incidents in a concise single sentence. In GAC-Network, an Attention based Encoder-Decoder Net (AED-Net) is built as the main network; with the temporal spatial attention mechanisms, the AED-Net make it possible to effectively reject the unimportant traffic behaviors and redundant backgrounds. Considering various driving scenarios, the Spatio-Temporal Layer Normalization is used to improve the generalization ability. To generate captions for incidents in driving, the novel Guidance Module is proposed to boost the encoder-decoder model to generate words in a caption, which have better relationship to the past and future words. Because there is no public dataset for captioning of driving scenarios, the Traffic Video Captioning (TVC) dataset is released for the video captioning task in driving scenarios. Experimental results show that the proposed methods can fulfill the captioning task for complex driving scenarios, and achieve higher performance than the methods for comparison, including at least 2.5%, 1.8%, 3.6%, and 13.1% better results on BLEU_1, METEOR, ROUGE_L and CIDEr, respectively. Chunsheng Liu 0001, Faliang Chang, Shuang Li 0016, Penghui Hao, Yansha Lu, Yinhai Wang |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | DSS-Net: Dynamic Self-Supervised Network for Video Anomaly DetectionabstractVideo Anomaly detection, aiming to detect the abnormal behaviors in surveillance videos, is a challenging task since the anomalous events are diversified and complicated in different situations. And this makes it difficult to use one single static network architecture to extract useful information from diverse abnormal patterns. Therefore, in this article, we propose a novel Dynamic Self-Supervised Network (DSS-Net) to explore both spatial and temporal anomalous information. In our DSS-Net, we design a dynamic network to adaptively select suitable network architecture to extract latent features from different anomalous patterns and normal patterns. Specifically, we generate spatial and temporalpseudo-abnormaldata as the input of the dynamic network to conduct self-supervised learning. And we have a specific design on Hybrid Anomaly Dynamic Convolution (HAD-Conv) to extract features for diversified anomalous events adaptively. We utilize both normal and pseudo-abnormal data to encourage the dynamic network to mine the discriminative information. Furthermore, we design a feature separation loss to maximize the difference between the anomalous and normal videos. We evaluate our proposed method on four public anomaly detection datasets and achieve competitive results compared with the state-of-the-art approaches. Peihao Wu, Faliang Chang, Chunsheng Liu 0001, Bin Wang 0004 |
IEEE Trans. Multim. | 3 |
| 2023 | Human-related anomalous event detection via memory-augmented Wasserstein generative adversarial network with gradient penalty
Nanjun Li, Faliang Chang, Chunsheng Liu 0001 |
Pattern Recognit. | 2 |
| 2023 | Magi-Net: Meta Negative Network for Early Activity PredictionabstractEarly activity prediction/recognition aims to recognize action categories before they are fully conveyed. Compared to full-length action sequences, partial video sequences only provide insufficient discrimination information, which makes predicting the class labels for some similar activities challenging, especially when only very few frames can be observed. To address this challenge, in this paper, we propose a novel meta negative network, namely, Magi-Net, that utilizes a contrastive learning scheme to alleviate the insufficiency of discriminative information. In our Magi-Net model, the positive samples are generated by augmenting an input anchor conditioned on all observation ratios, while the negative samples are selected from a trainable negative look-up memory (LUM) table, which stores the training samples and the corresponding misleading categories. Furthermore, a meta negative sample optimization strategy (MetaSOS) is proposed to boost the training of Magi-Net by encouraging the model to learn from the most informative negative samples via a meta learning scheme. Extensive experiments are conducted on several public skeleton-based activity datasets, and the results show the efficacy of the proposed Magi-Net model. Faliang Chang, Junhao Zhang 0001, Rui Yan 0010, Chunsheng Liu 0001, Bin Wang 0004, Zheng Shou 0001 |
IEEE Trans. Image Process. | 2 |
| 2023 | JHPFA-Net: Joint Head Pose and Facial Action Network for Driver Yawning Detection Across Arbitrary Poses in VideosabstractYawning detection is a key means in driver fatigue detection, which suffers difficulties including head poses, facial expressions, illumination variations, occlusions, etc. Yet, most previous methods mainly focus on frontal faces, and are deficient to deal with different facial actions under arbitrary poses in the actual driving environment. In this study, we propose a novel Joint Head Pose and Facial Action Network (JHPFA-Net) for driver yawning detection across arbitrary poses in videos, with three main parts including a Geometric-based Key-frame Selection Module (GK-Module), a Face Frontalization with Warp Attention Module (FF-Module) and a dual-channel classifier for Head Pose & Facial Action Fusion Module (HF-Module). Firstly, the GK-Module is proposed to extract geometric vectors and to construct a two-stage judgment mechanism, with the purpose of dealing with frame redundancy and improving the efficiency of JHPFA-Net structure. Secondly, distinguished with existing methods, the FF-Module is proposed to synthesize photo-realistic frontal faces, which can be used for capturing the facial actions under arbitrary poses. Finally, the HF-Module is proposed to fuse head pose attributes and facial modalities together, for the purpose of achieving pose-invariant detection and improving accuracy. Extensive experiments show that the proposed JHPFA-Net achieves state-of-the-art results comparing with some representative methods on the public YawDD benchmark, and it performs well in real-time application. Yansha Lu, Chunsheng Liu 0001, Faliang Chang, Hui Liu 0040, Hengqiang Huan |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | GA-Net: A Guidance Aware Network for Skeleton-Based Early Activity RecognitionabstractEarly activity prediction, which aims to recognize class labels before actions are fully performed, is a very challenging task since partially observed action sequences contain insufficient class-discrimination information, and thus, many partial action sequences belonging to different categories may look very similar. Therefore, in this paper, we propose a novelguidance aware network (GA-Net)to boost the ability to distinguish different activities in diversified partially observed action sequences via metric learning. To mitigate the similarity problem of action segments at very early stages, the proposedguided metric learning module (GMLM)is able to encourage the feature extractor to mine class-discriminative information given partially observed sequences. Specifically, the GMLM is able to minimize the intraclass distance with a full-length guided direction approach and maximize the difference between interclass categories with different observation ratios. To enhance the similarities between the partial- and full-length sequences in the same action categories, we further introduce adistribution alignment module (DAM)that employs full-length guidance to pull the partially observed features closer to the global features. We evaluate our proposed method on three public human activity datasets and achieve competitive results compared with the state-of-the-art approaches. Faliang Chang, Chunsheng Liu 0001, Guangxin Li, Bin Wang 0004 |
IEEE Trans. Multim. | 2 |
| 2023 | AE-Net:Adjoint Enhancement Network for Efficient Action Recognition in Video UnderstandingabstractAction recognition in video understanding is a challenging task, largely because of the complexity and difficulty in temporal modeling, making it suffer from motion information loss and misalignment of temporal attention in spatial dimensions. To overcome these difficulties, we propose a novel temporal modeling method calledAdjoint Enhancement Network(AE-Net), which can fully explore clues of motion and time in the long-range structure. The AE-Net mainly consists of two new modules: theInitial Adjoint Enhancement Module(IAE-Module), which deals with shallow features; and theGlobal Adjoint Enhancement Module(GAE-Module), which deals with global features. With a novel mechanism of parallel spatio-temporal convolution and difference fusion, the IAE-Module is to enhance the degree of motion transformation in shallow network features, exciting the potential of motion flow and avoiding motion information loss. The GAE-Module is proposed to improve the local temporal representation in long-range structures by feeding the enhanced feature differences into a spatial cascade module with residuals to resolve the misalignment of temporal attention in the spatial dimension.The experimental results show that our AE-Net can achieve state-of-the-art results in Something-Something V1, UCF-101 and HMDB-51 datasets. Bin Wang 0004, Chunsheng Liu 0001, Faliang Chang, Nanjun Li |
IEEE Trans. Multim. | 3 |
| 2022 | Human-related anomalous event detection via spatial-temporal graph convolutional autoencoder with embedded long short-term memory network
Nanjun Li, Faliang Chang, Chunsheng Liu 0001 |
Neurocomputing | 2 |
| 2022 | Adaptive Short-Temporal Induced Aware Fusion Network for Predicting Attention Regions Like a DriverabstractDriver attention prediction can solve the problem of ‘Where should the driver pay attention?’, Most previous methods are designed to predict regional attention with redundant regions. Furthermore, popular spatial-temporal feature extraction networks such as ConvLSTM and 3D-CNN are difficult to achieve real-time. To overcome these difficulties, we propose an Adaptive Short-temporal Induced Aware Fusion Network (ASIAF-Net) for region-level and object-level driver attention prediction.1In ASIAF-Net, we design anAttention Related Spatial Feature Encoder(AF-Encoder) and anInduced Aware Fusion Network(IAF-Net) as the main network; with anAssociation Analysis Cell(AAC), the AF-Encoder makes it possible to effectively capture the relationship information of different objects. Considering most vital visual cues from moving objects, we propose aSelf-adaptive Short-temporal Feature Extraction Module(SSFE-Module) to obtain inter-frame motion features. In IAF-Net, aMulti-scale Driver Attention Region Prediction Branchis designed to predict the regional attention, and anObject Saliency Estimation Branchis proposed to fuse the perception results and the regional attention map to estimate the object-level attention. Experiments show that the proposed ASIAF-Net can predict driver’s attention on regions and objects more robustly and precisely than state-of-the-art methods on three datasets, and that it achieves real-time on our ADAS platform. Chunsheng Liu 0001, Faliang Chang, Shuang Li 0016, Hui Liu 0040 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Temporal Shift and Spatial Attention-Based Two-Stream Network for Traffic Risk AssessmentabstractOn-board vision based traffic risk assessment is a challenging task for intelligent driving systems, which has some special challenges including spatial-temporal feature extraction, different judgements of risks, real-time requirement, lacking data, etc. To overcome these difficulties, we propose a novelTemporal Shift and Spatial Attention based Two-stream Network(TSSAT-Net) for on-board vision based traffic risk assessment. Firstly, we build new judgement measures that integrate actual driving experience and scenario complexity, and release an on-board vision based traffic risk assessment dataset. Secondly, a novelweighted Temporal Shift Module(weighted-TSM) based two-stream network is proposed; unlike previous methods that rely on complex and time-consuming 3D CNN or LSTM calculation, the proposed two-stream network can effectively extract spatial-temporal features using a weighted temporal shift mechanism with just 2D CNN calculation requirement. Thirdly, a spatial and channel attention mechanism is proposed to make the TSSAT-Net more focus on the features closely related to traffic risks, avoiding redundant information in complex traffic scenarios. Experiments based on the released comprehensive dataset show that our method achieves the state-of-the-art classification accuracy in real-time. Chunsheng Liu 0001, Zijian Li 0013, Faliang Chang, Shuang Li 0016, Jincan Xie |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Posture Calibration Based Cross-View & Hard-Sensitive Metric Learning for UAV-Based Vehicle Re-IdentificationabstractMachine vision based vehicle re-identification (ReID) plays an important role in some Intelligent Transportation Systems (ITS). Yet, most previous methods mainly focus on fixed surveillance cameras instead of Unmanned Aerial Vehicle (UAV). With high flexibility, the UAV-based vehicle ReID problem has some special challenges including complicated shooting angles, low discrimination of top-down features, and large variance in vehicle scales, etc. To overcome these challenges, we propose a novel structure for UAV-based vehicle ReID without license plates. Firstly, a triple-head segmentation net is proposed for segmenting UAV-captured vehicles under different heights and directions. Secondly, a posture calibration model is designed to uniform the vehicle postures based on the segmentation results, with the purpose of reducing the influence of different postures. Thirdly, the novel Cross-View & Hard-Sensitive Metric Learning (CHSML) method is proposed to train a ReID network with cross-view training constraint and hard sensitive principles; the mechanism of CHSML takes the cross-view samples of same ID as a training unit to learn the potential visual relationship in cross-view and builds a hard sensitive weight matrix to make learning more focus on hard samples, which improves the low ReID accuracy brought by cross-view or hard samples. Moreover, to facilitate the research of UAV-based vehicle ReID, a large-scale UAV-based vehicle ReID dataset called VeRi-UAV is released with 17 516 vehicles of 453 IDs. The experiments show that the proposed structure gains a better performance compared with the representative methods in the UAV-based vehicle ReID task. Chunsheng Liu 0001, Ye Song, Faliang Chang, Shuang Li 0016, Ruimin Ke, Yinhai Wang |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2021 | Level set method with Retinex-corrected saliency embedded for image segmentationabstractAbstract It can be a very challenging task when using level set method segmenting natural images with high intensity inhomogeneity and complex background scenes. A new synthesis level set method for robust image segmentation based on the combination of Retinex‐corrected saliency region information and edge information is proposed in this work. First, the Retinex theory is introduced to correct the saliency information extraction. Second, the Retinex‐corrected saliency information is embedded into the level set method due to its advantageous quality which makes a foreground object stand out relative to the backgrounds. Combined with the edge information, the boundary of segmentation will be more precise and smooth. Experiments indicate that the proposed segmentation algorithm is efficient, fast, reliable, and robust. Dongmei Liu 0007, Faliang Chang, Huaxiang Zhang 0001, Li Liu 0031 |
IET Image Process. | 2 |
| 2021 | Adaptive multi-level feature fusion and attention-based network for arbitrary-oriented object detection in remote sensing imagery
Luchang Chen, Chunsheng Liu 0001, Faliang Chang, Shuang Li 0016, Zhaoying Nie |
Neurocomputing | 3 |
| 2021 | Intermediate fused network with multiple timescales for anomaly detection
Faliang Chang, Huadong Mi |
Neurocomputing | 2 |
| 2021 | Bi-Directional Dense Traffic Counting Based on Spatio-Temporal Counting Feature and Counting-LSTM NetworkabstractMachine vision based vehicle counting and traffic flow estimation are challenging problems especially for dense traffic scenarios. Previousline of interest(LOI) counting methods rarely focus on dense scenarios and their performance largely relies on the accuracy of tracking. Avoiding the use of complex tracking methods, an LOI counting framework is proposed to address the bi-directional LOI counting problem in dense scenarios. There are three main contributions. Firstly, instead of treating the LOI vehicle counting problem as a combination of detecting and tracking of individual vehicles, the bi-directional traffic flow is taken as a whole and a novelspatio-temporal counting feature(STCF) is proposed for extracting bi-directional traffic flow features in dense traffic scenarios. Secondly, without relying on a multi-target tracking process for tracking and counting each vehicle, a counting network is proposed, called thecounting Long Short-Term Memory(cLSTM) network, to do analysis of the bi-directional STCF features and vehicle counting in successive video frames. Lastly, an estimation model is designed for estimating traffic flow parameters including speed, volume and density. Experiments performed on the UA-DETRAC dataset and the captured videos show that the proposed vehicle counting method outperforms the tested representative LOI counting methods in both accuracy and speed, and that the proposed framework can efficiently estimate traffic flow parameters including speed, volume and density in real time. Shuang Li 0016, Faliang Chang, Chunsheng Liu 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Spatial-Temporal Cascade Autoencoder for Video Anomaly Detection in Crowded ScenesabstractTime-efficient anomaly detection and localization in video surveillance still remains challenging due to the complexity of “anomaly”. In this paper, we propose a cuboid-patch-based method characterized by a cascade of classifiers called a spatial-temporal cascade autoencoder (ST-CaAE), which makes full use of both spatial and temporal cues from video data. The ST-CaAE has two main stages, defined by two proposed neural networks: a spatial-temporal adversarial autoencoder (ST-AAE) and a spatial-temporal convolutional autoencoder (ST-CAE). First, the ST-AAE is used to preliminarily identify anomalous video cuboids and exclude normal cuboids. The key idea underlying ST-AAE is to obtain a Gaussian model to fit the distribution of the regular data. Then in the second stage, the ST-CAE classifies the specific abnormal patches in each anomalous cuboid with reconstruction error based strategy that takes advantage of the CAE and skip connection. A two-stream framework is utilized to fuse the appearance and motion cues to achieve more complete detection results, taking the gradient and optical flow cuboids as inputs for each stream. The proposed ST-CaAE is evaluated using three public datasets. The experimental results verify that our framework outperforms other state-of-the-art works. Nanjun Li, Faliang Chang, Chunsheng Liu 0001 |
IEEE Trans. Multim. | 2 |
| 2019 | Attention to Head Locations for Crowd Counting
Youmei Zhang, Chunluan Zhou, Faliang Chang, Alex Chichung Kot, Wei Zhang 0021 |
ICIG (2) | 3 |
| 2019 | Video anomaly detection and localization via multivariate gaussian fully convolution adversarial autoencoder
Nanjun Li, Faliang Chang |
Neurocomputing | 2 |
| 2019 | Multi-resolution attention convolutional neural network for crowd countingabstractEstimating crowd counts remains a challenging task due to the problems of scale variations, non-uniform distribution and complex backgrounds. In this paper, we propose a multi-resolution attention convolutional neural network (MRA-CNN) to address this challenging task. Except for the counting task, we exploit an additional density-level classification task during training and combine features learned for the two tasks, thus forming multi-scale, multi-contextual features to cope with the scale variation and non-uniform distribution. Besides, we utilize a multi-resolution attention (MRA) model to generate score maps, where head locations are with higher scores to guide the network to focus on head regions and suppress non-head regions regardless of the complex backgrounds. During the generation of score maps, atrous convolution layers are used to expand the receptive field with fewer parameters, thus getting higher-level features and providing the MRA model more comprehensive information. Experiments on ShanghaiTech, WorldExpo’10 and UCF datasets demonstrate the effectiveness of our method. Youmei Zhang, Chunluan Zhou, Faliang Chang, Alex Chichung Kot |
Neurocomputing | 3 |
| 2019 | A scale adaptive network for crowd counting
Youmei Zhang, Chunluan Zhou, Faliang Chang, Alex Chichung Kot |
Neurocomputing | 3 |
| 2019 | Multi-target tracking with hierarchical data association using main-parts and spatial-temporal feature models
Faliang Chang, Chunsheng Liu 0001 |
Multim. Tools Appl. | 2 |
| 2019 | Hybrid Cascade Structure for License Plate Detection in Large Visual Surveillance ScenesabstractThough license plate detection has been successfully applied in some commercial products, the detection of small and vague license plates in real applications is still an open problem. In this paper, we propose a novel hybrid cascade structure for fast detecting small and vague license plates in large and complex visual surveillance scenes. For rapid license plate candidate extraction, we propose two cascade detectors, including the Cascaded Color Space Transformation of Pixel detector and the Cascaded Contrast-Color Haar-like detector; these two cascade detectors can do coarse-to-fine detection in the front and in the middle of the hybrid cascade. In the end of the hybrid cascade, we propose a cascaded convolutional network structure (Cascaded ConvNet), including two detection-ConvNets and a calibration-ConvNet, which is designed to do fine detection. Through experiments with different evaluation data sets with many small and vague plates, we show that the proposed framework is able to rapidly detect license plates with different resolutions and different sizes in large and complex visual surveillance scenes. Chunsheng Liu 0001, Faliang Chang |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2018 | Auxiliary learning for crowd counting via count-net
Youmei Zhang, Faliang Chang, Mengdi Wang 0007, Fulei Zhang |
Neurocomputing | 2 |
| 2017 | Expression recognition method based on evidence theory and local texture
Wencheng Wang 0002, Faliang Chang, Yunlong Liu 0002, Xiaojin Wu |
Multim. Tools Appl. | 2 |
| 2016 | Fast Traffic Sign Recognition via High-Contrast Region Extraction and Extended Sparse RepresentationabstractIn this paper, we propose a high-performance traffic sign recognition (TSR) framework to rapidly detect and recognize multiclass traffic signs in high-resolution images. This framework includes three parts: a novel region-of-interest (ROI) extraction method called the high-contrast region extraction (HCRE), the split-flow cascade tree detector (SFC-tree detector), and a rapid occlusion-robust traffic sign classification method based on the extended sparse representation classification (ESRC). Unlike the color-thresholding or extreme region extraction methods used by previous ROI methods, the ROI extraction method of the HCRE is designed to extract ROI with high local contrast, which can keep a good balance of the detection rate and the extraction rate. The SFC-tree detector can detect a large number of different types of traffic signs in high-resolution images quickly. The traffic sign classification method based on the ESRC is designed to classify traffic signs with partial occlusion. Instead of solving the sparse representation problem using an overcomplete dictionary, the classification method based on the ESRC utilizes a content dictionary and an occlusion dictionary to sparsely represent traffic signs, which can largely reduce the dictionary size in the occlusion-robust dictionaries and achieve high accuracy. The experiments demonstrate the advantage of the proposed approach, and our TSR framework can rapidly detect and recognize multiclass traffic signs with high accuracy. Chunsheng Liu 0001, Faliang Chang, Zhenxue Chen, Dongmei Liu 0007 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2014 | Illumination Processing in Face RecognitionabstractChanges in light intensity and angle present a major challenge to the creation of reliable face recognition systems. The existence of bright regions and dark regions has been shown to have a serious negative impact on the performance of face recognition systems. This paper proposes a solution to this problem based on self-quotient image (SQI) processing method. In this method, bright and dark areas are processed separately without changing the essential characteristics of the image of the face. The dark and light areas are processed separately by SQI. Experimental results indicate that this Single-Light-Region and Single-Dark-Region SQI method removes the adverse effect of multi-bright and multi-dark areas better than competing methods. Zhenxue Chen, Chengyun Liu, Faliang Chang, Xuzhen Han, Kaifang Wang |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2014 | Rapid Multiclass Traffic Sign Detection in High-Resolution ImagesabstractThis paper describes a traffic sign detection (TSD) framework that is capable of rapidly detecting multiclass traffic signs in high-resolution images while achieving a high detection rate. There are three key contributions. The first is the introduction of two features called multiblock normalization local binary pattern (MN-LBP) and tilted MN-LBP (TMN-LBP), which are able to express multiclass traffic signs effectively. The second is a tree structure called split-flow cascade, which utilizes common features of multiclass traffic signs to construct a coarse-to-fine TSD detector. The third contribution is the Common-Finder AdaBoost (CF.AdaBoost) algorithm, which is designed to find common features of different training sets to develop an efficient Split-Flow Cascade tree (SFC-tree) for multiclass TSD. Through experiments with an evaluation data set of high-resolution images, we show that the proposed framework is able to detect multiclass traffic signs with high detection accuracy in real time and that it outperforms the state-of-the-art approaches at detecting a large number of different types of traffic signs rapidly without using any color information. Chunsheng Liu 0001, Faliang Chang, Zhenxue Chen |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2013 | Chinese License Plate Recognition Based on Human Vision Attention MechanismabstractLicense plate recognition (LPR) is one of the most important elements affecting intelligent transportation systems. A number of LPR techniques have been proposed. Humans are good target recognition systems. In other words, humans easily recognize common objects. In this paper, the researchers present a novel method of recognizing Chinese license plates. The method is based on the Human Vision Attention Mechanism (HVAM) and uses Chinese license plates as the targets. The research consists of three stages. The first stage involved finding and identifying license plates in videos of moving vehicles. The second stage separated each license plate into the seven characters. In the third stage, the character recognizer extracted some salient features of Chinese characters and used a multi-stage classifier to recognize each character on the license plate. In the experiment locating license plates, 1176 images taken from various scenes and conditions were employed. The method failed to identify the license plates in only 27 of the images; resulting in a license plate location rate of success of 97.7%. In the experiment for identifying license characters, 1149 images were used, from which license plates had been successfully located. The method failed to identify the characters in 45 of these images giving a success rate of 96.1%. Combining the above two rates, the overall rate of success for our LPR is 93.9%. Zhenxue Chen, Faliang Chang, Chunsheng Liu 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2007 | A Region-Based Skin Color Detection Algorithm
Faliang Chang |
PAKDD | 1 |
| 2005 | Target Tracking Under Occlusion by Combining Integral-Intensity-Matching with Multi-block-voting
Faliang Chang, Yizheng Qiao |
ICIC (1) | 1 |