VLDB 2026 Research / reviewers in the wild / expert
Lingbo Liu
dblp:20/5299
· DBLP profile ↗
48ranked-venue papers
13as first author
37since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 28 · 6 first-author · 20 since 2021Artificial intelligence and machine learning · 24 · 6 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The computational reproducibility of articles published under the Open Data + FAIR policy of IJGISabstractAcademic journals increasingly require data and code sharing to improve computational reproducibility, but the effectiveness of these policies remains unclear. This study evaluates the changes in reproducibility after the Open Data + FAIR policy was implemented by the International Journal of Geographic Information Science in August 2019 through a systematic audit of 351 articles published between 2020–2024, with a comparison made to 32 articles from the pre-policy period. Using a five-star reproducibility framework based on FAIR principles, we assessed data and code availability, metadata quality, and adherence to open science standards. Results show significant improvement in material availability post-policy, with most articles now including data and code compared to minimal sharing pre-policy. However, computational reproducibility likely remains limited, with most articles achieving only 1–2 star ratings due to inadequate metadata, unclear workflow documentation, and missing details about computational environments. While compliance with basic sharing requirements increased after policy introduction, it is unclear if it has facilitated the comprehensive documentation necessary for independent reproduction. These findings suggest that journal policies focused solely on availability may be insufficient for achieving computational reproducibility in GIScience research, highlighting the need for enhanced standards addressing metadata, workflow documentation, and computational environments. Peter Kedron, Lingbo Liu |
Int. J. Geogr. Inf. Sci. | 3 |
| 2026 | Reconciling 2SFCA and i2SFCA via distance decay parameterizationabstractUnderstanding spatial accessibility and facility crowdedness is central to public service planning, yet existing methods often treat these two metrics separately. The Two-Step Floating Catchment Area (2SFCA) method measures accessibility from the demand side, while the inverted 2SFCA (i2SFCA) assesses crowdedness from the supply side. Without proper integration, these two measures may diverge, raising concerns of their validity. This study introduces a distance decay parameterization framework to reconcile 2SFCA and i2SFCA by optimizing a unified distance decay function through cross-entropy minimization. It demonstrates that aligning demand-side and supply-side flows effectively enforces a behavioral equilibrium between accessibility and crowdedness. A case study using inpatient hospital flow data in Florida shows that the ‘reconciled 2SFCA (r2SFCA) model’ achieves strong alignment between estimated and observed service flows while maintaining simplicity in its formulation. These findings validate the self-organizing nature of human service-seeking behaviors and support a unified, entropy-based calibration strategy for accessibility modeling. Lingbo Liu, Fahui Wang |
Int. J. Geogr. Inf. Sci. | 1 |
| 2025 | Prohibited Items Segmentation via Occlusion-aware Bilayer ModelingabstractInstance segmentation of prohibited items in security X-ray images is a critical yet challenging task. This is mainly caused by the significant appearance gap between prohibited items in X-ray images and natural objects, as well as the severe overlapping among objects in X-ray images. To address these issues, we propose an occlusion-aware instance segmentation pipeline designed to identify prohibited items in X-ray images. Specifically, to bridge the representation gap, we integrate the Segment Anything Model (SAM) into our pipeline, taking advantage of its rich priors and zero-shot generalization capabilities. To address the overlap between prohibited items, we design an occlusion-aware bilayer mask decoder module that explicitly models the occlusion relationships. To supervise occlusion estimation, we manually annotated occlusion areas of prohibited items in two large-scale X-ray image segmentation datasets, PIDray and PIXray. We then reorganized these additional annotations together with the original information as two occlusion-annotated datasets, PIDray-A and PIXray-A. Extensive experimental results on these occlusion-annotated datasets demonstrate the effectiveness of our proposed method. The datasets and codes are available at: https://github.com/Ryh1218/Occ. Yunhan Ren, Ruihuang Li, Lingbo Liu, Chang Wen Chen |
ICME | 3 |
| 2025 | MDPM: Modulating domain-specific prompt memory for multi-domain traffic flow prediction with transformers
Zhuang Zhuang, Lingbo Liu, Kan Guo, Xingtong Yu, Heng Qi, Yanming Shen |
Knowl. Based Syst. | 2 |
| 2025 | CMAAN: Cross-Modal Aggregation Attention Network for Next POI RecommendationabstractNext point-of-interest (POI) recommendation is to explore the historical check-in sequence information in location-based social networks (LBSNs) to recommend the next location that he/she might be interested in. However, most previous methods used only limited information of unimodal data (i.e., check-in sequences), while some recent methods have attempted to explore multimodal data (e.g., textual content) but lacked sufficient interactions between geographic behavior patterns and content behavior patterns. In this work, we argue that users usually consider geographical trajectories and textual content interdependently to determine the next location to visit. To this end, we propose a novel cross-modal aggregation attention network (CMAAN), which interactively learns multiview representations from POI sequence and content sequence for predicting the next POI. Our approach models inter-modal interaction correlations, intra-modal sequence correlations, and intra-modal semantic correlations simultaneously to fully discover contextual potential relations along the trajectories. Specifically, the intra-modal semantic correlations are able to capture the variable location functionalities under different contextual relationships of cross-modal interaction information. Moreover, we apply the aggregation attention to adaptively aggregate multiview representations which represent the comprehensive hidden state of the next POI. Extensive experiments on two large-scale datasets clearly demonstrate that our CMAAN achieves state-of-the-art performance. Zhuang Zhuang, Lingbo Liu, Heng Qi, Yanming Shen |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2025 | FAME: Fusion of Alignment and Multiview Enhancement for Remote Sensing Image-Text RetrievalabstractContemporary advancements in Earth observation technologies have generated substantial data resources for remote sensing image retrieval applications. However, existing models exhibit limitations in extracting local features from image-text pairs and establishing effective connections between local and global information. Furthermore, these models lack fine-grained alignment capabilities between different modalities. To address these challenges, we propose FAME, which integrates four key components: a Progressive Masking (ProgMask) module, an OmniView Fusion (OVF) module, a Global Feature Attention with Multi-view Feature Alignment (GFA-MFA) module, and Regional Clustering Alignment Loss (RCA-Loss). The ProgMask module employs a progressive training strategy with alternating masking of image and text modalities to enhance cross-modal learning robustness. The OVF module employs omni-view attention mechanisms with adaptive weighting to dynamically focus on different local semantic regions of remote sensing images and corresponding textual descriptions. The GFA-MFA module enhances global feature capture through selective filtering while ensuring precise multi-view feature alignment. These modules work synergistically to achieve fine-grained alignment between multi-scale local and global features across modalities. Additionally, RCA-Loss optimizes intra-class feature clustering while minimizing inter-class confusion through class center distance optimization and enhanced feature discrimination. Experimental validation on three benchmark datasets (RSITMD, RSICD, UCM-Captions) demonstrates that FAME achieves superior performance compared to existing state-of-the-art methods in remote sensing image-text retrieval tasks. Yuanhao Su, Daoye Zhu, Zhande Dong, Qifeng Lin, Lingbo Liu, Shuming Bao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | SQLNet: Scale-Modulated Query and Localization Network for Few-Shot Class-Agnostic CountingabstractThe class-agnostic counting (CAC) task has recently been proposed to solve the problem of counting all objects of an arbitrary class with several exemplars given in the input image. To address this challenging task, existing leading methods all resort to density map regression, which renders them impractical for downstream tasks that require object locations and restricts their ability to well explore the scale information of exemplars for supervision. Meanwhile, they generally model the interaction between the input image and the exemplars in an exemplar-by-exemplar way, which is inefficient and may not fully synthesize information from all exemplars. To address these limitations, we propose a novel localization-based CAC approach, termed Scale-modulated Query and Localization Network (SQLNet). It fully explores the scales of exemplars in both the query and localization stages and achieves effective counting by accurately locating each object and predicting its approximate size. Specifically, during the query stage, rich discriminative representations of the target class are acquired by the Hierarchical Exemplars Collaborative Enhancement (HECE) module from the few exemplars through multi-scale exemplar cooperation with equifrequent size prompt embedding. These representations are then fed into the Exemplars-Unified Query Correlation (EUQC) module to interact with the query features in a unified manner and produce the correlated query tensor. In the localization stage, the Scale-aware Multi-head Localization (SAML) module utilizes the query tensor to predict the confidence, location, and size of each potential object. Moreover, a scale-aware localization loss is introduced, which exploits flexible location associations and exemplar scales for supervision to optimize the model performance. Extensive experiments demonstrate that SQLNet outperforms state-of-the-art methods on popular CAC benchmarks, achieving excellent performance not only in counting accuracy but also in localization and bounding box generation. Hefeng Wu, Yandong Chen 0002, Lingbo Liu, Tianshui Chen, Keze Wang, Liang Lin 0004 |
IEEE Trans. Image Process. | 3 |
| 2025 | Continuous Value Assignment: A Doubly Robust Data Augmentation for Off-Policy LearningabstractDeep reinforcement learning (RL) has witnessed remarkable success in a wide range of control tasks. To overcome RL's notorious sample inefficiency, prior studies have explored data augmentation techniques leveraging collected transition data. However, these methods face challenges in synthesizing transitions adhering to the authentic environment dynamics, especially when the transition is high-dimensional and includes many redundant/irrelevant features to the task. In this article, we introduce continuous value assignment (CVA), an innovative optimization-level data augmentation approach that directly synthesizes novel training data in the state-action value space, effectively bypassing the need for explicit transition modeling. The key intuition of our method is that the transition plays an intermediate role in calculating the state-action value during optimization, and therefore directly augmenting the state-action value is more causally related to the optimization process. Specifically, our CVA combines parameterized value prediction and nonparametric value interpolation from neighboring states, resulting in doubly robust target values w.r.t. novel states and actions. Extensive experiments demonstrate CVA's substantial improvements in sample efficiency across complex continuous control tasks, surpassing several advanced baselines. Junfan Lin, Zhongzhan Huang, Keze Wang, Lingbo Liu, Liang Lin 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Road Network-Guided Fine-Grained Urban Traffic Flow InferenceabstractAccurate inference of fine-grained traffic flow from coarse-grained one is an emerging yet crucial problem, which can help greatly reduce the number of the required traffic monitoring sensors for cost savings. In this work, we note that traffic flow has a high correlation with road network, which was either completely ignored or simply treated as an external factor in previous works. To facilitate this problem, we propose a novel road-aware traffic flow magnifier (RATFM) that explicitly exploits the prior knowledge of road networks to fully learn the road-aware spatial distribution of fine-grained traffic flow. Specifically, a multidirectional 1-D convolutional layer is first introduced to extract the semantic feature of the road network. Subsequently, we incorporate the road network feature and coarse-grained flow feature to regularize the short-range spatial distribution modeling of road-relative traffic flow. Furthermore, we take the road network feature as a query to capture the long-range spatial distribution of traffic flow with a transformer architecture. Benefiting from the road-aware inference mechanism, our method can generate high-quality fine-grained traffic flow maps. Extensive experiments on three real-world datasets show that the proposed RATFM outperforms state-of-the-art models under various scenarios. Our code and datasets are released at https://github.com/luimoli/RATFM. Lingbo Liu, Guanbin Li, Junfan Lin, Liang Lin 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Energy-guided test-time adaptation for data shifts in multi-modal perception
Yun Pei 0001, Lingbo Liu, Runqing Jiang, Ye Zhang 0037, Pengpeng Yu, Liang Lin 0004, Yulan Guo |
Vis. Comput. | 2 |
| 2024 | TAU: Trajectory Data Augmentation with Uncertainty for Next POI RecommendationabstractNext Point-of-Interest (POI) recommendation has been proven effective at utilizing sparse, intricate spatial-temporal trajectory data to recommend subsequent POIs to users. While existing methods commonly alleviate the problem of data sparsity by integrating spatial-temporal context information, POI category features, and social relationships, they largely overlook the fact that the trajectory sequences collected in the datasets are often incomplete. This oversight limits the model’s potential to fully leverage historical context. In light of this background, we propose Trajectory Data Augmentation with Uncertainty (TAU) for Next POI Recommendation. TAU is a general graph-based trajectory data augmentation method designed to complete user mobility patterns by marrying uncertainty estimation into the next POI recommendation task. More precisely, TAU taps into the global transition pattern graph to identify sets of intermediate nodes located between every pair of locations, effectively leveraging edge weights as transition probabilities. During trajectory sequence construction, TAU selectively prompts intermediate nodes, chosen based on their likelihood of occurrence as pseudo-labels, to establish comprehensive trajectory sequences. Furthermore, to gauge the certainty and impact of pseudo-labels on the target location, we introduce a novel confidence-aware calibration strategy using evidence deep learning (EDL) for improved performance and reliability. The experimental results clearly indicate that our TAU method achieves consistent performance improvements over existing techniques across two real-world datasets, verifying its effectiveness as the state-of-the-art approach to the task. Zhuang Zhuang, Tianxin Wei, Lingbo Liu, Heng Qi, Yanming Shen |
AAAI | 3 |
| 2024 | Domain-Agnostic Crowd Counting via Uncertainty-Guided Style Diversity AugmentationabstractDomain shift significantly hinders crowd counting performance in unseen domains. Domain adaptation methods tackle this issue using target domain images but falter when acquiring these images is difficult. Moreover, they demand additional training time for fine-tuning. To address this issue, we propose an Uncertainty-Guided Style Diversity Augmentation (UGSDA) method, enabling the models to be trained solely on the source domain and directly generalized to various target domains. It is achieved by generating sufficiently diverse and realistic samples during the training process. Specifically, our UGSDA method incorporates three tailor-designed components: the Global Styling Elements Extraction (GSEE) module, the Local Uncertainty Perturbations (LUP) module, and the Density Distribution Consistency (DDC) loss. The GSEE extracts global style elements from the feature space of the whole source domain. The LUP aims to obtain uncertainty perturbations from the batch-level input to form style distributions beyond the source domain, which used to generate diversified stylized samples together with global style elements. To regulate the extent of perturbations, the DDC loss imposes constraints between the source samples and the stylized samples, ensuring the stylized samples maintain a higher degree of realism and reliability. Comprehensive experiments validate the superiority of our approach, demonstrating its strong generalization capabilities across various datasets and models. Code is available at https://github.com/gcding/UGSDA-pytorch. Guanchen Ding, Lingbo Liu, Zhenzhong Chen 0001, Chang Wen Chen |
ACM Multimedia | 2 |
| 2024 | Heterogeneous Semantic Transfer for Multi-label Recognition with Partial Labels
Tianshui Chen, Tao Pu 0002, Lingbo Liu, Yukai Shi, Zhijing Yang, Liang Lin 0004 |
Int. J. Comput. Vis. | 3 |
| 2023 | Spatio-Temporal Graph Neural Point Process for Traffic Congestion Event PredictionabstractTraffic congestion event prediction is an important yet challenging task in intelligent transportation systems. Many existing works about traffic prediction integrate various temporal encoders and graph convolution networks (GCNs), called spatio-temporal graph-based neural networks, which focus on predicting dense variables such as flow, speed and demand in time snapshots, but they can hardly forecast the traffic congestion events that are sparsely distributed on the continuous time axis. In recent years, neural point process (NPP) has emerged as an appropriate framework for event prediction in continuous time scenarios. However, most conventional works about NPP cannot model the complex spatio-temporal dependencies and congestion evolution patterns. To address these limitations, we propose a spatio-temporal graph neural point process framework, named STGNPP for traffic congestion event prediction. Specifically, we first design the spatio-temporal graph learning module to fully capture the long-range spatio-temporal dependencies from the historical traffic state data along with the road network. The extracted spatio-temporal hidden representation and congestion event information are then fed into a continuous gated recurrent unit to model the congestion evolution patterns. In particular, to fully exploit the periodic information, we also improve the intensity function calculation of the point process with a periodic gated mechanism. Finally, our model simultaneously predicts the occurrence time and duration of the next congestion. Extensive experiments on two real-world datasets demonstrate that our method achieves superior performance in comparison to existing state-of-the-art approaches. Guangyin Jin, Lingbo Liu, Fuxian Li, Jincai Huang 0001 |
AAAI | 2 |
| 2023 | Being Comes from Not-Being: Open-Vocabulary Text-to-Motion Generation with Wordless TrainingabstractText-to-motion generation is an emerging and challenging problem, which aims to synthesize motion with the same semantics as the input text. However, due to the lack of diverse labeled training data, most approaches either limit to specific types of text annotations or require online optimizations to cater to the texts during inference at the cost of efficiency and stability. In this paper, we investigate offline open-vocabulary text-to-motion generation in a zero-shot learning manner that neither requires paired training data nor extra online optimization to adapt for unseen texts. Inspired by the prompt learning in NLP, we pretrain a motion generator that learns to reconstruct the full motion from the masked motion. During inference, instead of changing the motion generator, our method reformulates the input text into a masked motion as the prompt for the motion generator to “reconstruct” the motion. In constructing the prompt, the unmasked poses of the prompt are synthesized by a text-to-pose generator. To supervise the optimization of the text-to-pose generator, we propose the first text-pose alignment model for measuring the alignment between texts and 3D poses. And to prevent the pose generator from over-fitting to limited training texts, we further propose a novel wordless training mechanism that optimizes the text-to-pose generator without any training texts. The comprehensive experimental results show that our method obtains a significant improvement against the baseline methods. The code is available at https://github.com/junfanlin/oohmg. Junfan Lin, Jianlong Chang, Lingbo Liu, Guanbin Li, Liang Lin 0004, Qi Tian 0001, Chang Wen Chen |
CVPR | 3 |
| 2023 | STEERER: Resolving Scale Variations for Counting and Localization via Selective Inheritance LearningabstractScale variation is a deep-rooted problem in object counting, which has not been effectively addressed by existing scale-aware algorithms. An important factor is that they typically involve cooperative learning across multi-resolutions, which could be suboptimal for learning the most discriminative features from each scale. In this paper, we propose a novel method termed STEERER (SelecTivE inhERitance lEaRning) that addresses the issue of scale variations in object counting. STEERER selects the most suitable scale for patch objects to boost feature extraction and only inherits discriminative features from lower to higher resolution progressively. The main insights of STEERER are a dedicated Feature Selection and Inheritance Adaptor (FSIA), which selectively forwards scale-customized features at each scale, and a Masked Selection and Inheritance Loss (MSIL) that helps to achieve high-quality density maps across all scales. Our experimental results on nine datasets with counting and localization tasks demonstrate the unprecedented scale generalization ability of STEERER. Code is available at https://github.com/taohan10200/STEERER. Tao Han 0002, Lei Bai 0001, Lingbo Liu, Wanli Ouyang |
ICCV | 3 |
| 2023 | DenseLight: Efficient Control for Large-scale Traffic Signals with Dense FeedbackabstractTraffic Signal Control (TSC) aims to reduce the average travel time of vehicles in a road network, which in turn enhances fuel utilization efficiency, air quality, and road safety, benefiting society as a whole. Due to the complexity of long-horizon control and coordination, most prior TSC methods leverage deep reinforcement learning (RL) to search for a control policy and have witnessed great success. However, TSC still faces two significant challenges. 1) The travel time of a vehicle is delayed feedback on the effectiveness of TSC policy at each traffic intersection since it is obtained after the vehicle has left the road network. Although several heuristic reward functions have been proposed as substitutes for travel time, they are usually biased and not leading the policy to improve in the correct direction. 2) The traffic condition of each intersection is influenced by the non-local intersections since vehicles traverse multiple intersections over time. Therefore, the TSC agent is required to leverage both the local observation and the non-local traffic conditions to predict the long-horizontal traffic conditions of each intersection comprehensively. To address these challenges, we propose DenseLight, a novel RL-based TSC method that employs an unbiased reward function to provide dense feedback on policy effectiveness and a non-local enhanced TSC agent to better predict future traffic conditions for more precise traffic control. Extensive experiments and ablation studies demonstrate that DenseLight can consistently outperform advanced baselines on various road networks with diverse traffic flows. The code is available at https://github.com/junfanlin/DenseLight. Junfan Lin, Yuying Zhu 0007, Lingbo Liu, Yang Liu 0267, Guanbin Li, Liang Lin 0004 |
IJCAI | 3 |
| 2023 | Long-term Wind Power Forecasting with Hierarchical Spatial-Temporal TransformerabstractWind power is attracting increasing attention around the world due to its renewable, pollution-free, and other advantages. However, safely and stably integrating the high permeability intermittent power energy into electric power systems remains challenging. Accurate wind power forecasting (WPF) can effectively reduce power fluctuations in power system operations. Existing methods are mainly designed for short-term predictions and lack effective spatial-temporal feature augmentation. In this work, we propose a novel end-to-end wind power forecasting model named Hierarchical Spatial-Temporal Transformer Network (HSTTN) to address the long-term WPF problems. Specifically, we construct an hourglass-shaped encoder-decoder framework with skip-connections to jointly model representations aggregated in hierarchical temporal scales, which benefits long-term forecasting. Based on this framework, we capture the inter-scale long-range temporal dependencies and global spatial correlations with two parallel Transformer skeletons and strengthen the intra-scale connections with downsampling and upsampling operations. Moreover, the complementary information from spatial and temporal features is fused and propagated in each other via Contextual Fusion Blocks (CFBs) to promote the prediction further. Extensive experimental results on two large-scale real-world datasets demonstrate the superior performance of our HSTTN over existing solutions. Lingbo Liu, Xinyu Xiong, Guanbin Li, Liang Lin 0004 |
IJCAI | 2 |
| 2023 | CLIP-Count: Towards Text-Guided Zero-Shot Object CountingabstractRecent advances in visual-language models have shown remarkable zero-shot text-image matching ability that is transferable to downstream tasks such as object detection and segmentation. Adapting these models for object counting, however, remains a formidable challenge. In this study, we first investigate transferring vision-language models (VLMs) for class-agnostic object counting. Specifically, we propose CLIP-Count, the first end-to-end pipeline that estimates density maps for open-vocabulary objects with text guidance in a zero-shot manner. To align the text embedding with dense visual features, we introduce a patch-text contrastive loss that guides the model to learn informative patch-level visual representations for dense prediction. Moreover, we design a hierarchical patch-text interaction module to propagate semantic information across different resolution levels of visual features. Benefiting from the full exploitation of the rich image-text alignment knowledge of pretrained VLMs, our method effectively generates high-quality density maps for objects-of-interest. Extensive experiments on FSC-147, CARPK, and ShanghaiTech crowd counting datasets demonstrate state-of-the-art accuracy and generalizability of the proposed method. Code is available: https://github.com/songrise/CLIP-Count. https://github.com/songrise/CLIP-Count. Ruixiang Jiang, Lingbo Liu, Chang Wen Chen |
ACM Multimedia | 2 |
| 2023 | Urban regional function guided traffic flow prediction
Lingbo Liu, Yang Liu 0084, Guanbin Li, Liang Lin 0004 |
Inf. Sci. | 2 |
| 2023 | Online Metro Origin-Destination Prediction via Heterogeneous Information AggregationabstractMetro origin-destination prediction is a crucial yet challenging time-series analysis task in intelligent transportation systems, which aims to accurately forecast two specific types of cross-station ridership, i.e., Origin-Destination (OD) one and Destination-Origin (DO) one. However, complete OD matrices of previous time intervals can not be obtained immediately in online metro systems, and conventional methods only used limited information to forecast the future OD and DO ridership separately. In this work, we proposed a novel neural network module termed Heterogeneous Information Aggregation Machine (HIAM), which fully exploits heterogeneous information of historical data (e.g., incomplete OD matrices, unfinished order vectors, and DO matrices) to jointly learn the evolutionary patterns of OD and DO ridership. Specifically, an OD modeling branch estimates the potential destinations of unfinished orders explicitly to complement the information of incomplete OD matrices, while a DO modeling branch takes DO matrices as input to capture the spatial-temporal distribution of DO ridership. Moreover, a Dual Information Transformer is introduced to propagate the mutual information among OD features and DO features for modeling the OD-DO causality and correlation. Based on the proposed HIAM, we develop a unified Seq2Seq network to forecast the future OD and DO ridership simultaneously. Extensive experiments conducted on two large-scale benchmarks demonstrate the effectiveness of our method for online metro origin-destination prediction. Our code is resealed at https://github.com/HCPLab-SYSU/HIAM. Lingbo Liu, Yuying Zhu 0007, Guanbin Li, Lei Bai 0001, Liang Lin 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Hybrid-Order Representation Learning for Electricity Theft DetectionabstractElectricity theft is the primary cause of electrical losses in power systems, which severely harms the economic benefits of electricity providers and threatens the safety of the power supply. However, due to the inherent complex correlation and periodicity of electricity consumption and the low efficiency of large-scale data processing, detecting anomalies in electricity consumption data accurately and efficiently remains challenging. Existing methods usually focus on first-order information and ignore the second-order representation learning that can efficiently model global temporal dependency and facilitate discriminative representation learning of electricity consumption data. In this article, we propose a novel electricity theft detection framework named hybrid-order representation learning network (HORLN). Specifically, the sequential electricity consumption data is transformed into the matrix format containing weekly consumption records. Then, an inter-and-intra week convolution block is designed to capture multiscale features in a local-to-global manner. Meanwhile, a self-dependency modeling module is proposed to learn the second-order representations from self-correlation matrices, which are finally combined with the first-order representations to predict the anomaly scores of electricity consumers. Extensive experiments on a real-world benchmark demonstrate the advantages of our HORLN over state-of-the-art methods. Yuying Zhu 0007, Lingbo Liu, Yang Liu 0084, Guanbin Li, Mingzhi Mao, Liang Lin 0004 |
IEEE Trans. Ind. Informatics | 3 |
| 2023 | Aerial Images Meet Crowdsourced Trajectories: A New Approach to Robust Road ExtractionabstractLand remote-sensing analysis is a crucial research in earth science. In this work, we focus on a challenging task of land analysis, i.e., automatic extraction of traffic roads from remote-sensing data, which has widespread applications in urban development and expansion estimation. Nevertheless, conventional methods either only utilized the limited information of aerial images, or simply fused multimodal information (e.g., vehicle trajectories), thus cannot well recognize unconstrained roads. To facilitate this problem, we introduce a novel neural network framework termed cross-modal message propagation network (CMMPNet), which fully benefits the complementary different modal data (i.e., aerial images and crowdsourced trajectories). Specifically, CMMPNet is composed of two deep autoencoders for modality-specific representation learning and a tailor-designed dual enhancement module for cross-modal representation refinement. In particular, the complementary information of each modality is comprehensively extracted and dynamically propagated to enhance the representation of another modality. Extensive experiments on three real-world benchmarks demonstrate the effectiveness of our CMMPNet for robust road extraction benefiting from blending different modal data, either using image and trajectory data or image and light detection and ranging (LiDAR) data. From the experimental results, we observe that the proposed approach outperforms current state-of-the-art methods by large margins. Our source code is resealed on the project page http://lingboliu.com/multimodal_road_extraction.html. Lingbo Liu, Zewei Yang, Guanbin Li, Tianshui Chen, Liang Lin 0004 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Probing Visual-Audio Representation for Video Highlight Detection via Hard-Pairs Guided Contrastive Learning
Lingbo Liu, Shinan Liu, Shuai Yi |
BMVC | 4 |
| 2022 | Pyramid Region-based Slot Attention Network for Temporal Action Proposal Generation
Lingbo Liu, Rui Feng 0001 |
BMVC | 5 |
| 2022 | Scale-Prior Deformable Convolution for Exemplar-Guided Class-Agnostic Counting
Wei Lin 0018, Xinzhu Ma, Junyu Gao 0001, Lingbo Liu, Shinan Liu, Shuai Yi, Antoni B. Chan |
BMVC | 5 |
| 2022 | Multimodal Crowd Counting with Mutual Attention TransformersabstractCrowd counting is a fundamental yet challenging task that aims to automatically estimate the number of people in crowded scenes. Nowadays, with the rapid development of thermal and depth sensors, thermal images and depth maps become more accessible, which are proven to be beneficial information in boosting the performance of crowd counting. Consequently, we propose a Mutual Attention Transformer (MAT) module to fully leverage the complementary information of different modalities. Specifically, our MAT employs a cross-modal mutual attention mechanism to utilize the features of one modality to enhance the features of the other. Moreover, to improve performance by learning better visual representation and further exploiting modality-wise comple-mentarity, we design a self-supervised pre-training method based on cross-modal image reconstruction. Extensive experiments on two standard benchmarks (i.e., RGBT-CC and ShanghaiTechRGBD) show that the proposed method is effective and universal for multimodal crowd counting, outper-forming previous state-of-the-art methods. Zhengtao Wu, Lingbo Liu, Mingzhi Mao, Liang Lin 0004, Guanbin Li |
ICME | 2 |
| 2022 | Unconstrained face sketch synthesis via perception-adaptive network and a new benchmark
Lin Nie, Lingbo Liu, Zhengtao Wu, Wenxiong Kang |
Neurocomputing | 2 |
| 2022 | Cross-Domain Facial Expression Recognition: A Unified Evaluation Benchmark and Adversarial Graph LearningabstractFacial expression recognition (FER) has received significant attention in the past decade with witnessed progress, but data inconsistencies among different FER datasets greatly hinder the generalization ability of the models learned on one dataset to another. Recently, a series of cross-domain FER algorithms (CD-FERs) have been extensively developed to address this issue. Although each declares to achieve superior performance, comprehensive and fair comparisons are lacking due to inconsistent choices of the source/target datasets and feature extractors. In this work, we first propose to construct a unified CD-FER evaluation benchmark, in which we re-implement the well-performing CD-FER and recently published general domain adaptation algorithms and ensure that all these algorithms adopt the same source/target datasets and feature extractors for fair CD-FER evaluations. Based on the analysis, we find that most of the current state-of-the-art algorithms use adversarial learning mechanisms that aim to learn holistic domain-invariant features to mitigate domain shifts. However, these algorithms ignore local features, which are more transferable across different datasets and carry more detailed content for fine-grained adaptation. Therefore, we develop a novel adversarial graph representation adaptation (AGRA) framework that integrates graph representation propagation with adversarial learning to realize effective cross-domain holistic-local feature co-adaptation. Specifically, our framework first builds two graphs to correlate holistic and local regions within each domain and across different domains, respectively. Then, it extracts holistic-local features from the input image and uses learnable per-class statistical distributions to initialize the corresponding graph nodes. Finally, two stacked graph convolution networks (GCNs) are adopted to propagate holistic-local features within each domain to explore their interaction and across different domains for holistic-local feature co-adaptation. In this way, the AGRA framework can adaptively learn fine-grained domain-invariant features and thus facilitate cross-domain expression recognition. We conduct extensive and fair comparisons on the unified evaluation benchmark and show that the proposed AGRA framework outperforms previous state-of-the-art methods. Tianshui Chen, Tao Pu 0002, Hefeng Wu, Yuan Xie 0004, Lingbo Liu, Liang Lin 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Video Crowd Localization With Multifocus Gaussian Neighborhood Attention and a Large-Scale BenchmarkabstractVideo crowd localization is a crucial yet challenging task, which aims to estimate exact locations of human heads in the given crowded videos. To model spatial-temporal dependencies of human mobility, we propose a multi-focus Gaussian neighborhood attention (GNA), which can effectively exploit long-range correspondences while maintaining the spatial topological structure of the input videos. In particular, our GNA can also capture the scale variation of human heads well using the equipped multi-focus mechanism. Based on the multi-focus GNA, we develop a unified neural network called GNANet to accurately locate head centers in video clips by fully aggregating spatial-temporal information via a scene modeling module and a context cross-attention module. Moreover, to facilitate future researches in this field, we introduce a large-scale crowd video benchmark named VSCrowd (https://github.com/HopLee6/VSCrowd), which consists of 60K+ frames captured in various surveillance scenes and 2M+ head annotations. Finally, we conduct extensive experiments on three datasets including our VSCrowd, and the experiment results show that the proposed method is capable to achieve state-of-the-art performance for both video crowd localization and counting. Haopeng Li 0001, Lingbo Liu, Shinan Liu, Junyu Gao 0001, Bin Zhao 0001, Rui Zhang 0003 |
IEEE Trans. Image Process. | 2 |
| 2022 | TCGL: Temporal Contrastive Graph for Self-Supervised Video Representation LearningabstractVideo self-supervised learning is a challenging task, which requires significant expressive power from the model to leverage rich spatial-temporal knowledge and generate effective supervisory signals from large amounts of unlabeled videos. However, existing methods fail to increase the temporal diversity of unlabeled videos and ignore elaborately modeling multi-scale temporal dependencies in an explicit way. To overcome these limitations, we take advantage of the multi-scale temporal dependencies within videos and propose a novel video self-supervised learning framework named Temporal Contrastive Graph Learning (TCGL), which jointly models the inter-snippet and intra-snippet temporal dependencies for temporal representation learning with a hybrid graph contrastive learning strategy. Specifically, a Spatial-Temporal Knowledge Discovering (STKD) module is first introduced to extract motion-enhanced spatial-temporal representations from videos based on the frequency domain analysis of discrete cosine transform. To explicitly model multi-scale temporal dependencies of unlabeled videos, our TCGL integrates the prior knowledge about the frame and snippet orders into graph structures, i.e., the intra-/inter-snippet Temporal Contrastive Graphs (TCG). Then, specific contrastive learning modules are designed to maximize the agreement between nodes in different graph views. To generate supervisory signals for unlabeled videos, we introduce an Adaptive Snippet Order Prediction (ASOP) module which leverages the relational knowledge among video snippets to learn the global context representation and recalibrate the channel-wise features adaptively. Experimental results demonstrate the superiority of our TCGL over the state-of-the-art methods on large-scale action recognition and video retrieval benchmarks. The code is publicly available at https://github.com/YangLiu9208/TCGL. Yang Liu 0084, Keze Wang, Lingbo Liu, Haoyuan Lan, Liang Lin 0004 |
IEEE Trans. Image Process. | 3 |
| 2022 | Physical-Virtual Collaboration Modeling for Intra- and Inter-Station Metro Ridership PredictionabstractDue to the widespread applications in real-world scenarios, metro ridership prediction is a crucial but challenging task in intelligent transportation systems. However, conventional methods either ignore the topological information of metro systems or directly learn on physical topology, and cannot fully explore the patterns of ridership evolution. To address this problem, we model a metro system as graphs with various topologies and propose a unified Physical-Virtual Collaboration Graph Network (PVCGN), which can effectively learn the complex ridership patterns from the tailor-designed graphs. Specifically, a physical graph is directly built based on the realistic topology of the studied metro system, while a similarity graph and a correlation graph are built with virtual topologies under the guidance of the inter-station passenger flow similarity and correlation. These complementary graphs are incorporated into a Graph Convolution Gated Recurrent Unit (GC-GRU) for spatial-temporal representation learning. Further, a Fully-Connected Gated Recurrent Unit (FC-GRU) is also applied to capture the global evolution tendency. Finally, we develop a Seq2Seq model with GC-GRU and FC-GRU to forecast the future metro ridership sequentially. Extensive experiments on two large-scale benchmarks (e.g., Shanghai Metro and Hangzhou Metro) well demonstrate the superiority of our PVCGN for station-level metro ridership prediction. Moreover, we apply the proposed PVCGN to address the online origin-destination (OD) ridership prediction and the experiment results show the universality of our method. Our code and benchmarks are available athttps://github.com/HCPLab-SYSU/PVCGN. Lingbo Liu, Hefeng Wu, Jiajie Zhen, Guanbin Li, Liang Lin 0004 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2021 | Cross-Modal Collaborative Representation Learning and a Large-Scale RGBT Benchmark for Crowd CountingabstractCrowd counting is a fundamental yet challenging task, which desires rich information to generate pixel-wise crowd density maps. However, most previous methods only used the limited information of RGB images and cannot well discover potential pedestrians in unconstrained scenarios. In this work, we find that incorporating optical and thermal information can greatly help to recognize pedestrians. To promote future researches in this field, we introduce a large-scale RGBT Crowd Counting (RGBT-CC) benchmark, which contains 2,030 pairs of RGB-thermal images with 138,389 annotated people. Furthermore, to facilitate the multimodal crowd counting, we propose a cross-modal collaborative representation learning framework, which consists of multiple modality-specific branches, a modality-shared branch, and an Information Aggregation-Distribution Module (IADM) to capture the complementary information of different modalities fully. Specifically, our IADM incorporates two collaborative information transfers to dynamically enhance the modality-shared and modality-specific representations with a dual information propagation mechanism. Extensive experiments conducted on the RGBT-CC benchmark demonstrate the effectiveness of our framework for RGBT crowd counting. Moreover, the proposed approach is universal for multimodal crowd counting and is also capable to achieve superior performance on the ShanghaiTechRGBD [22] dataset. Finally, our source code and benchmark have been released at http://lingboliu.com/RGBT_Crowd_Counting.html. Lingbo Liu, Hefeng Wu, Guanbin Li, Chenglong Li 0002, Liang Lin 0004 |
CVPR | 1 |
| 2021 | Representative Local Feature Mining for Few-Shot LearningabstractFew-shot learning aims to recognize unseen images of new classes with only a few training examples. While great progress has been made with deep learning technology, most metric-based works rely on the measurement based on global feature representation of images, which is sensitive to background factors due to the scarcity of training data. Given this, we propose a novel method that chooses representative local features to facilitate few-shot learning. Specifically, we propose a "task-specific guided" strategy to mine local features that are task-specific and discriminative. For each task, we first mine representative local features for labeled images by a loss guided mechanism. Then these local features are used to guide a classifier to mine representative local features for unlabeled images. In this way, task-specific representative local features can be selected for better classification. We empirically show our method can effectively alleviate the negative effect introduced by background factors. Extensive experiments on two few-shot benchmarks show the effectiveness of the proposed method. Kun Yan 0008, Lingbo Liu, Ping Wang 0003 |
ICASSP | 2 |
| 2021 | GroupFormer: Group Activity Recognition with Clustered Spatial-Temporal TransformerabstractGroup activity recognition is a crucial yet challenging problem, whose core lies in fully exploring spatial-temporal interactions among individuals and generating reasonable group representations. However, previous methods either model spatial and temporal information separately, or directly aggregate individual features to form group features. To address these issues, we propose a novel group activity recognition network termed GroupFormer. It captures spatial-temporal contextual information jointly to augment the individual and group representations effectively with a clustered spatial-temporal transformer. Specifically, our GroupFormer has three appealing advantages: (1) A tailor-modified Transformer, Clustered Spatial-Temporal Transformer, is proposed to enhance the individual representation and group representation. (2) It models the spatial and temporal dependencies integrally and utilizes decoders to build the bridge between the spatial and temporal information. (3) A clustered attention mechanism is utilized to dynamically divide individuals into multiple clusters for better learning activity-aware semantic representations. Moreover, experimental results show that the proposed framework outperforms state-of-the-art methods on the Volleyball dataset and Collective Activity dataset. Code is available at https://github.com/xueyee/GroupFormer Qianggang Cao, Lingbo Liu, Shinan Liu, Shuai Yi |
ICCV | 3 |
| 2021 | TRUFM: a Transformer-Guided Framework for Fine-Grained Urban Flow Inference
Xinchi Zhou, Dongzhan Zhou, Lingbo Liu |
ICONIP (4) | 3 |
| 2021 | Dynamic Spatial-Temporal Representation Learning for Traffic Flow PredictionabstractAs a crucial component in intelligent transportation systems, traffic flow prediction has recently attracted widespread research interest in the field of artificial intelligence (AI) with the increasing availability of massive traffic mobility data. Its key challenge lies in how to integrate diverse factors (such as temporal rules and spatial dependencies) to infer the evolution trend of traffic flow. To address this problem, we propose a unified neural network called Attentive Traffic Flow Machine (ATFM), which can effectively learn the spatial-temporal feature representations of traffic flow with an attention mechanism. In particular, our ATFM is composed of two progressive Convolutional Long Short-Term Memory (ConvLSTM[1]) units connected with a convolutional layer. Specifically, the first ConvLSTM unit takes normal traffic flow features as input and generates a hidden state at each time-step, which is further fed into the connected convolutional layer for spatial attention map inference. The second ConvLSTM unit aims at learning the dynamic spatial-temporal representations from the attentionally weighted traffic flow features. Further, we develop two deep learning frameworks based on ATFM to predict citywide short-term/long-term traffic flow by adaptively incorporating the sequential and periodic data as well as other external influences. Extensive experiments on two standard benchmarks well demonstrate the superiority of the proposed method for traffic flow prediction. Moreover, to verify the generalization of our method, we also apply the customized framework to forecast the passenger pickup/dropoff demands in traffic prediction and show its superior performance. Our code and data are available athttps://github.com/liulingbo918/ATFM. Lingbo Liu, Jiajie Zhen, Guanbin Li, Geng Zhan, Zhaocheng He, Bowen Du 0001, Liang Lin 0004 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2020 | Efficient Crowd Counting via Structured Knowledge TransferabstractCrowd counting is an application-oriented task and its inference efficiency is crucial for real-world applications. However, most previous works relied on heavy backbone networks and required prohibitive run-time consumption, which would seriously restrict their deployment scopes and cause poor scalability. To liberate these crowd counting models, we propose a novel Structured Knowledge Transfer (SKT) framework, which fully exploits the structured knowledge of a well-trained teacher network to generate a lightweight but still highly effective student network. Specifically, it is integrated with two complementary transfer modules, including an Intra-Layer Pattern Transfer which sequentially distills the knowledge embedded in layer-wise features of the teacher network to guide feature learning of the student network and an Inter-Layer Relation Transfer which densely distills the cross-layer correlation knowledge of the teacher to regularize the student's feature evolution. Consequently, our student network can derive the layer-wise and cross-layer knowledge from the teacher network to learn compact yet effective features. Extensive evaluations on three benchmarks well demonstrate the effectiveness of our SKT for extensive crowd counting models. In particular, only using around $6%$ of the parameters and computation cost of original models, our distilled VGG-based models obtain at least 6.5× speed-up on an Nvidia 1080 GPU and even achieve state-of-the-art performance. Our code and models are available at https://github.com/HCPLab-SYSU/SKT. Lingbo Liu, Hefeng Wu, Tianshui Chen, Guanbin Li, Liang Lin 0004 |
ACM Multimedia | 1 |
| 2020 | Crowd counting via scale-communicative aggregation networks
Lixian Yuan, Zhilin Qiu, Lingbo Liu, Hefeng Wu, Tianshui Chen, Liang Lin 0004 |
Neurocomputing | 3 |
| 2019 | Crowd Counting With Deep Structured Scale Integration NetworkabstractAutomatic estimation of the number of people in unconstrained crowded scenes is a challenging task and one major difficulty stems from the huge scale variation of people. In this paper, we propose a novel Deep Structured Scale Integration Network (DSSINet) for crowd counting, which addresses the scale variation of people by using structured feature representation learning and hierarchically structured loss function optimization. Unlike conventional methods which directly fuse multiple features with weighted average or concatenation, we first introduce a Structured Feature Enhancement Module based on conditional random fields (CRFs) to refine multiscale features mutually with a message passing mechanism. Specifically, each scale-specific feature is considered as a continuous random variable and passes complementary information to refine the features at other scales. Second, we utilize a Dilated Multiscale Structural Similarity loss to enforce our DSSINet to learn the local correlation of people's scales within regions of various size, thus yielding high-quality density maps. Extensive experiments on four challenging benchmarks well demonstrate the effectiveness of our method. In particular, our DSSINet achieves improvements of 9.5% error reduction on Shanghaitech dataset and 24.9% on UCF-QNRF dataset against the state-of-the-art methods. Lingbo Liu, Zhilin Qiu, Guanbin Li, Shufan Liu, Wanli Ouyang, Liang Lin 0004 |
ICCV | 1 |
| 2019 | Taxi Origin-Destination Demand Prediction with Contextualized Spatial-Temporal NetworkabstractTaxi demand prediction has recently attracted increasing research interest due to its huge potential application in large-scale intelligent transportation systems. However, most of the previous methods only considered the taxi demand prediction in origin regions, while ignoring the modeling of the specific situation of the destination passengers. In this paper, we present a more challenging task, called taxi origin-destination demand prediction, which aims at predicting the taxi demand between all origin-destination (OD) pairs in a future time interval. Its main challenges lie in how to effectively capture the diverse contextual information to learn the demand patterns. We address this problem with a novel Contextualized Spatial-Temporal Network (CSTN), which can effectively capture various context of taxi demand into a unified framework. Specifically, the proposed network consists of three components for the modeling of local spatial context (LSC), temporal evolution context (TEC) and global correlation context (GCC) respectively. Extensive experiments and evaluations on a large-scale dataset well demonstrate the significant superiority of our CSTN over other compared methods of taxi origin-destination demand prediction. Zhilin Qiu, Lingbo Liu, Guanbin Li, Qing Wang 0018, Nong Xiao 0001, Liang Lin 0004 |
ICME | 2 |
| 2019 | Crowd Counting via Multi-view Scale Aggregation NetworksabstractCrowd counting, aiming at estimating the total number of people in unconstrained crowded scenes, has increasingly received attention. But it is greatly challenged by the huge variation in people scale. In this paper, we propose a novel Multi-View Scale Aggregation Network (MVSAN), which handle the scale variation from feature, input and criterion view comprehensively. Firstly, we design a simple but effective Multi-Scale Feature Encoder, which exploits dilated convolution layers with various dilation rates to improve the representation ability and scale diversity of features. Secondly, we feed multiple scales of input images into networks to generate high-quality density maps in a coarse-to-fine manner. Finally, we propose a Multi-Scale Structural Similarity loss to force our networks to learn the local correlation of density maps. Extensive experiments on two standard benchmarks show that the proposed method can generate high-quality crowd density map and accurate count estimation, outperforming the state-of-the-art methods with a large margin. Zhilin Qiu, Lingbo Liu, Guanbin Li, Qing Wang 0018, Nong Xiao 0001, Liang Lin 0004 |
ICME | 2 |
| 2019 | Contextualized Spatial-Temporal Network for Taxi Origin-Destination Demand PredictionabstractTaxi demand prediction has recently attracted increasing research interest due to its huge potential application in large-scale intelligent transportation systems. However, most of the previous methods only considered the taxi demand prediction in origin regions, but neglected the modeling of the specific situation of the destination passengers. We believe it is suboptimal to preallocate the taxi into each region-based solely on the taxi origin demand. In this paper, we present a challenging and worth-exploring task, called taxi origin-destination demand prediction, which aims at predicting the taxi demand between all-region pairs in a future time interval. Its main challenges come from how to effectively capture the diverse contextual information to learn the demand patterns. We address this problem with a novel contextualized spatial-temporal network (CSTN), which consists of three components for the modeling of local spatial context (LSC), temporal evolution context (TEC), and global correlation context (GCC), respectively. First, an LSC module utilizes two convolution neural networks to learn the local spatial dependencies of taxi, demand respectively, from the origin view and the destination view. Second, a TEC module incorporates the local spatial features of taxi demand and the meteorological information to a Convolutional Long Short-term Memory Network (ConvLSTM) for the analysis of taxi demand evolution. Finally, a GCC module is applied to model the correlation between all regions by computing a global correlation feature as a weighted sum of all regional features, with the weights being calculated as the similarity between the corresponding region pairs. The extensive experiments and evaluations on a large-scale dataset well demonstrate the superiority of our CSTN over other compared methods for the taxi origin-destination demand prediction. Lingbo Liu, Zhilin Qiu, Guanbin Li, Qing Wang 0018, Wanli Ouyang, Liang Lin 0004 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2019 | Facial Landmark Machines: A Backbone-Branches Architecture With Progressive Representation LearningabstractFacial landmark localization plays a critical role in face recognition and analysis. In this paper, we propose a novel cascaded backbone-branches fully convolutional neural network (BB-FCN) for rapidly and accurately localizing facial landmarks in unconstrained and cluttered settings. Our proposed BB-FCN generates facial landmark response maps directly from raw images without any preprocessing. BB-FCN follows a coarse-to-fine cascaded pipeline, which consists of a backbone network to roughly detect the locations of all facial landmarks and one branch network for each type of detected landmark to further refine its location. Furthermore, to facilitate the facial landmark localization under unconstrained settings, we propose a large-scale benchmark named SYSU16K, which contains 16 000 faces with large variations in pose, expression, illumination, and resolution. Extensive experimental evaluations demonstrate that our proposed BB-FCN can significantly outperform the state of the art under both constrained (i.e., within detected facial regions only) and unconstrained settings. We further confirm that high-quality facial landmarks localized with our proposed network can also improve the precision and recall of face detection. Lingbo Liu, Guanbin Li, Yuan Xie 0004, Yizhou Yu, Qing Wang 0018, Liang Lin 0004 |
IEEE Trans. Multim. | 1 |
| 2018 | Crowd Counting using Deep Recurrent Spatial-Aware NetworkabstractCrowd counting from unconstrained scene images is a crucial task in many real-world applications like urban surveillance and management, but it is greatly challenged by the camera’s perspective that causes huge appearance variations in people’s scales and rotations. Conventional methods address such challenges by resorting to fixed multi-scale architectures that are often unable to cover the largely varied scales while ignoring the rotation variations. In this paper, we propose a unified neural network framework, named Deep Recurrent Spatial-Aware Network, which adaptively addresses the two issues in a learnable spatial transform module with a region-wise refinement process. Specifically, our framework incorporates a Recurrent Spatial-Aware Refinement (RSAR) module iteratively conducting two components: i) a Spatial Transformer Network that dynamically locates an attentional region from the crowd density map and transforms it to the suitable scale and rotation for optimal crowd estimation; ii) a Local Refinement Network that refines the density map of the attended region with residual learning. Extensive experiments on four challenging benchmarks show the effectiveness of our approach. Specifically, comparing with the existing best-performing methods, we achieve an improvement of 12\% on the largest dataset WorldExpo’10 and 22.8\% on the most challenging dataset UCF\_CC\_50 Lingbo Liu, Hongjun Wang 0005, Guanbin Li, Wanli Ouyang, Liang Lin 0004 |
IJCAI | 1 |
| 2018 | Attentive Crowd Flow MachinesabstractTraffic flow prediction is crucial for urban traffic management and public safety. Its key challenges lie in how to adaptively integrate the various factors that affect the flow changes. In this paper, we propose a unified neural network module to address this problem, called Attentive Crowd Flow Machine~(ACFM), which is able to infer the evolution of the crowd flow by learning dynamic representations of temporally-varying data with an attention mechanism. Specifically, the ACFM is composed of two progressive ConvLSTM units connected with a convolutional layer for spatial weight prediction. The first LSTM takes the sequential flow density representation as input and generates a hidden state at each time-step for attention map inference, while the second LSTM aims at learning the effective spatial-temporal feature expression from attentionally weighted crowd flow features. Based on the ACFM, we further build a deep architecture with the application to citywide crowd flow prediction, which naturally incorporates the sequential and periodic data as well as other external influences. Extensive experiments on two standard benchmarks (i.e., crowd flow in Beijing and New York City) show that the proposed method achieves significant improvements over the state-of-the-art methods. Lingbo Liu, Ruimao Zhang, Jiefeng Peng, Guanbin Li, Bowen Du 0001, Liang Lin 0004 |
ACM Multimedia | 1 |
| 2016 | DISC: Deep Image Saliency Computing via Progressive Representation LearningabstractSalient object detection increasingly receives attention as an important component or step in several pattern recognition and image processing tasks. Although a variety of powerful saliency models have been intensively proposed, they usually involve heavy feature (or model) engineering based on priors (or assumptions) about the properties of objects and backgrounds. Inspired by the effectiveness of recently developed feature learning, we provide a novel deep image saliency computing (DISC) framework for fine-grained image saliency computing. In particular, we model the image saliency from both the coarse-and fine-level observations, and utilize the deep convolutional neural network (CNN) to learn the saliency representation in a progressive manner. In particular, our saliency model is built upon two stacked CNNs. The first CNN generates a coarse-level saliency map by taking the overall image as the input, roughly identifying saliency regions in the global context. Furthermore, we integrate superpixel-based local context information in the first CNN to refine the coarse-level saliency map. Guided by the coarse saliency map, the second CNN focuses on the local context to produce fine-grained and accurate saliency map while preserving object details. For a testing image, the two CNNs collaboratively conduct the saliency computing in one shot. Our DISC framework is capable of uniformly highlighting the objects of interest from complex background while preserving well object details. Extensive experiments on several standard benchmarks suggest that DISC outperforms other state-of-the-art methods and it also generalizes well across data sets without additional training. The executable version of DISC is available online: http://vision.sysu.edu.cn/projects/DISC. Tianshui Chen, Liang Lin 0004, Lingbo Liu, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2011 | Multi-task GLOH feature selection for human age estimationabstractIn this paper, we propose a novel age estimation method based on gradient location and orientation histogram (GLOH) descriptor and multi-task learning (MTL). The GLOH, one of the state-of-the-art local descriptor, is used to capture the age- related local and spatial information of face image. As the extracted GLOH features are often redundant, MTL is designed to select the most informative GLOH bins for age estimation problem, while the corresponding weights are determined by ridge regression. This approach largely reduces the dimensions of feature, which can not only improve performance but also decrease the computational burden. Experiments on the public available FG-NET database show that the proposed method can achieve comparable performance over previous approaches while using much fewer features. Yixiong Liang, Lingbo Liu, Yao Xiang, Beiji Zou 0001 |
ICIP | 2 |