Xiaobo Lu

dblp:93/8545 · DBLP profile ↗
← Back
133ranked-venue papers
2as first author
83since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 64 · 33 since 2021Artificial intelligence and machine learning · 59 · 43 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 8 since 2021Systems, architecture and hardware · 2 · 2 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 CompTrack: Information Bottleneck-Guided Low-Rank Dynamic Token Compression for Point Cloud Tracking
abstract
3D single object tracking (SOT) in LiDAR point clouds is a critical task in computer vision and autonomous driving. Despite great success having been achieved, the inherent sparsity of point clouds introduces a dual-redundancy challenge that limits existing trackers: (1) vast spatial redundancy from background noise impairs accuracy, and (2) informational redundancy within the foreground hinders efficiency. To tackle these issues, we propose CompTrack, a novel end-to-end framework that systematically eliminates both forms of redundancy in point clouds. First, CompTrack incorporates a Spatial Foreground Predictor (SFP) module to filter out irrelevant background noise based on information entropy, addressing spatial redundancy. Subsequently, its core is an Information Bottleneck-guided Dynamic Token Compression (IB-DTC) module that eliminates the informational redundancy within the foreground. Theoretically grounded in low-rank approximation, this module leverages an online SVD analysis to adaptively compress the redundant foreground into a compact and highly informative set of proxy tokens. Extensive experiments on KITTI, nuScenes and Waymo datasets demonstrate that CompTrack achieves top-performing tracking performance with superior efficiency, running at a real-time 90 FPS on a single RTX 3090 GPU.
Sifan Zhou, Yichao Cao, Jiahao Nie 0001, Yuqian Fu, Xiaobo Lu, Shuo Wang 0030
AAAI6
2026 Class label enhanced Wasserstein distance for classification of remote sensing smoke-related scenes
Shikun Chen, Xin Lu 0007, Xiaobo Lu
Eng. Appl. Artif. Intell.3
2026 Toward Free-Form Local Feature Matching
abstract
Existing feature matching methods are strongly coupled to their pre-defined position priors. For instance, sparse matchers are coupled to keypoints, and semi-dense matchers are coupled to grids. The coupled position prior dictates the distribution of matching points and imposes inherent limitations on the matcher. Consequently, sparse matchers suffer from a reliance on keypoint repeatability, while semi-dense matchers lack texture-based precision. Our preliminary work RCM leverages the keypoint prior in the source image and the grid prior in the target image, ensuring texture-based precision with keypoints while eliminating reliance on repeatability. However, RCM still relies heavily on keypoints in the source image, inheriting limitations such as sparsity and poor distribution in challenging scenes. To address these challenges, we introduce RCM+, which presents a novel free-form matching paradigm. By combining a position-agnostic encoder with a parameter-free decoder, we decouple the matcher from any position prior. As a result, the free-form matcher can match arbitrary input positions in a zero-shot manner, including detected keypoints, lines, edges, grids of any resolution, user-specified points, and more. This paradigm offers exceptional flexibility, allowing users to select position priors based on scene properties without retraining. Thus, RCM+ can leverage the advantages of various position priors without over-relying on any single prior, avoiding limitations in specific scenarios. To better match multiple position priors, we propose the Balancer, which reconciles all input position priors to achieve a more favorable point distribution for downstream tasks. Additionally, we enhance the view switcher and conflict-free matching layer introduced in RCM, further improving matching quality. Comprehensive experiments demonstrate the excellent performance, efficiency, and flexibility of RCM+, underscoring its promising potential for applications.
Xiaoyong Lu, Songlin Du, Yaping Yan, Xiaobo Lu, Takeshi Ikenaga
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 Refining the granularity of smoke representation: SAM-powered density-aware progressive smoke segmentation framework
Yichao Cao, Xuanpeng Li, Xiaolin Meng, Xiaobo Lu
Pattern Recognit.5
2026 Incremental mixture of experts: Continual learning for object detection in forestry scenarios
Ximeng Cheng, Qiaonan Zhu, Shukun Jia, Yichao Cao, Xiaobo Lu
Pattern Recognit.5
2026 Tracking by detection and query: An efficient end-to-end framework for multi-object tracking
Shukun Jia, Yichao Cao, Xin Lu 0007, Xiaobo Lu
Pattern Recognit.6
2026 Skeleton-prompt: A cross-dataset transfer learning approach for skeleton action recognition
Xiaobo Lu, Jun Liu 0036
Pattern Recognit.2
2026 Single-domain generalization for fastener detection via sample reconstruction and class-wise domain contrast
Shixiang Su, Songlin Du, Xiaobo Lu
Pattern Recognit.4
2026 SceneGlue: Scene-Aware Transformer for Feature Matching Without Scene-Level Annotation
abstract
Local feature matching plays a critical role in understanding the correspondence between cross-view images. However, traditional methods are constrained by the inherent local nature of feature descriptors, limiting their ability to capture non-local scene information that is essential for accurate cross-view correspondence. In this paper, we introduce SceneGlue, a scene-aware feature matching framework designed to overcome these limitations. SceneGlue leverages a hybridizable matching paradigm that integrates implicit parallel attention and explicit cross-view visibility estimation. The parallel attention mechanism simultaneously exchanges information among local descriptors within and across images, enhancing the scene’s global context. To further enrich the scene awareness, we propose the Visibility Transformer, which explicitly categorizes features into visible and invisible regions, providing an understanding of cross-view scene visibility. By combining explicit and implicit scene-level awareness, SceneGlue effectively compensates for the local descriptor constraints. Notably, SceneGlue is trained using only local feature matches, without requiring scene-level groundtruth annotations. This scene-aware approach not only improves accuracy and robustness but also enhances interpretability compared to traditional methods. Extensive experiments on applications such as homography estimation, pose estimation, image matching, and visual localization validate SceneGlues superior performance. The source code is available at https://github.com/songlindu/ SceneGlue.
Songlin Du, Xiaoyong Lu, Yaping Yan, Guobao Xiao, Xiaobo Lu, Takeshi Ikenaga
IEEE Trans. Circuits Syst. Video Technol.5
2026 FastPillars: A Deployment-Friendly Pillar-Based 3D Detector
abstract
The deployment of 3D detectors strikes one of the major challenges in real-world self-driving scenarios. Existing BEV-based (i.e., Bird Eye View) detectors favor sparse convolutions (known as SPConv) to speed up training and inference, which puts a hard barrier for deployment, especially for on-device applications. In this paper, in order to tackle the challenge of efficient 3D object detection from an industry perspective, we devise a deployment-friendly pillar-based 3D detector, termed FastPillars. Specifically, aiming to compensate the geometric information loss of pillar encoding. First, we design a novel lightweight Max-and-Attention Pillar Encoding (MAPE) module specially for enhancing small objects. Second, we propose a simple yet effective backbone design for pillar-based 3D detection, enhancing pillar representations. We construct FastPillars based on these designs, achieving high performance and low latency without SPConv. Extensive experiments on two large-scale datasets demonstrate the effectiveness and efficiency of FastPillars for on-device 3D detection regarding both performance and speed. Specifically, FastPillars delivers real-time state-of-the-art accuracy on Waymo Open Dataset with 1.8 × speed up and 3.8 mAPH/L2 improvement over CenterPoint (SPConv-based). Code will be opened soon in: https://github.com/StiphyJay/FastPillars.
Sifan Zhou, Xinyu Zhang 0015, Xiangxiang Chu, Bo Zhang 0046, Xiaobo Lu
IEEE Trans. Circuits Syst. Video Technol.6
2026 Revisiting Semantic Correspondence: When Feature Aggregation Hurts Structural Integrity
abstract
Semantic correspondence seeks to establish matches between different instances of the same category. A common paradigm for this task leverages high-quality features from stable diffusion (SD) and DINOv2. However, we identify a widely overlooked yet critical issue: common feature aggregation disrupts the structural integrity of SD features, degrading semantic matching performance. We revisit and analyze this phenomenon and propose structure-aware aggregation (SAA) for SD features as a direct replacement for common feature aggregation methods. SAA uses filtering to decompose SD features into fine texture details and coarse contour structures. It aggregates only the texture components while preserving the contours. This divide-and-conquer mechanism enables SAA to significantly enhance the performance of state-of-the-art semantic correspondence models without increasing trainable parameters or computational overhead. Extensive qualitative and quantitative experiments confirm our analysis and validate the effectiveness of SAA. Moreover, SAA generalizes well to geometric, cross-species, and cross-family semantic correspondence tasks. Code is available at https://github.com/wzhlearning/SAA.
Zenghui Wang 0009, Songlin Du, Xiaobo Lu, Guobao Xiao
IEEE Trans. Image Process.4
2025 FocusTrack: One-Stage Focus-and-Suppress Framework for 3D Point Cloud Object Tracking
Sifan Zhou, Jiahao Nie 0001, Yichao Cao, Xiaobo Lu
ACM Multimedia5
2025 A depth-aware geometric fusion based view synthesis method for sparse RGB-D input
Xiaobo Lu, Sifan Zhou, Patrick Chiang 0001
Expert Syst. Appl.2
2025 Predicted label distribution guided optimal transport for classification of remote sensing smoke-like scenes with noisy labels
Shikun Chen, Xiaobo Lu
Neurocomputing4
2025 HifiDiff: High-fidelity diffusion model for face hallucination from tiny non-frontal faces
Yuguang Shi, Xiaobo Lu
Neurocomputing4
2025 PINR: A physics-integrated neural representation for dynamic fluid scenes
Sifan Zhou, Xiaobo Lu, Jian Qian
Neurocomputing3
2025 A landmarks-assisted diffusion model with heatmap-guided denoising loss for high-fidelity and controllable facial image generation
Shixiang Su, Lei Zhang 0130, Xiaobo Lu
Image Vis. Comput.6
2025 Depth-free view synthesis from diffusion models for monocular 3D detector in autonomous driving
Yuguang Shi, Sifan Zhou, Xiaobo Lu
Multim. Syst.4
2025 DuPt: Rehearsal-based continual learning with dual prompts
Shengqin Jiang, Daolong Zhang, Fengna Cheng, Xiaobo Lu, Qingshan Liu 0001
Neural Networks4
2025 SmokeAgent: Multimodal agent for fine-grained smoke event analysis in large-scale wild environments
Yichao Cao, Xuanpeng Li, Xiaobo Lu
Pattern Recognit.4
2025 MFECNet: Multi-level feature enhancement and correspondence network for few-shot anomaly detection of high-speed trains
Wei Liu 0175, Xiaobo Lu, Zhidan Ran
Pattern Recognit.2
2025 Camera-aware graph multi-domain adaptive learning for unsupervised person re-identification
Zhidan Ran, Xiaobo Lu
Pattern Recognit.2
2025 Rethinking iterative stereo matching from a diffusion bridge model perspective
Yuguang Shi, Sifan Zhou, Xiaobo Lu
Pattern Recognit.4
2025 UPT-Flow: Multi-scale transformer-guided normalizing flow for low-light image enhancement
Lintao Xu, Changhui Hu 0001, Xiaoyuan Jing, Ziyun Cai, Xiaobo Lu
Pattern Recognit.6
2025 Progressive feature injection and global-local discrimination-based image generation network for improving rail surface defect inspection
Dezhou Wang, Shixiang Su, Xiaobo Lu
Pattern Recognit. Lett.3
2025 Modulated deformable convolution based on graph convolution network for rail surface crack detection
Shuzhen Tong, Xiaobo Lu
Signal Process. Image Commun.5
2025 Gator: Accelerating Graph Attention Networks by Jointly Optimizing Attention and Graph Processing
abstract
Graph attention networks (GATs) have advanced performance in various application domains by introducing the attention mechanism into the graph neural networks (GNNs). The inefficiency of running GATs on CPUs or GPUs necessitates specialized hardware designs. Unfortunately, previous specialized architecture designs have focused on either the GNN architecture or the attention mechanism, resulting in limited performance and leaving ample room for improvement. This article presents Gator , a joint optimization approach with software–hardware co-designs for GAT inference. On the software level, Gator leverages degree-weighted graph partitioning and parameter-adaptive feature selection techniques to preprocess the input graph data, mining subgraph-level parallelism and mitigating the computation bottleneck of the dedicated dataflow. On the hardware level, Gator designs a unified processing engine to support various kernels by extracting a common computation pattern and a dimension-aware microarchitecture for efficient partial sum reduction. Extensive experiments show that our approach can achieve 11.5× more efficiency compared to NVIDIA RTX 4090 and provide a speedup of 3× to 9.4×, along with a 2.6× to 4.7× reduction in memory traffic, when compared to six state-of-the-art methods, with minimal accuracy loss.
Xiaobo Lu, Jianbin Fang, Lin Peng 0001, Chun Huang 0006, Zixiao Yu
ACM Trans. Archit. Code Optim.1
2025 Variational Feature Imitation Conditioned on Visual Descriptions for Few-Shot Fine-Grained Recognition
abstract
In few-shot fine-grained recognition (FS-FGR) tasks, the main challenge is to distinguish novel categories with high intra-class variations and low inter-class differences given scarce training data. Existing studies explore discriminative features through a compact network to avoid overfitting, while they achieve marginal performance gain owing to the limited representation capability. Motivated by the significant progress of the vision foundation model, we introduce it to describe visual attributes and boost the performance of the compact feature extractor. A few-shot fine-grained recognition method with Variational Feature Imitation Conditioned on Visual Descriptions, VFI-CVD for short, has been proposed in this paper. It simultaneously exploits the pre-trained knowledge from a vision foundation model and the expert knowledge mined by a feature extractor. Specifically, the intra-class variations shared across object categories are encoded into a common distribution thus we can augment features by sampling latent variables. To enhance the learning of intra-class variations, a condition exchange strategy (CES) is put forward to interact the knowledge between samples through feature cross-imitation. In the inference stage, the learned knowledge is further integrated through the joint prediction of visual descriptions and cross-imitated features. Comprehensive experimental results on four fine-grained benchmark datasets show that the proposed VFI-CVD achieves state-of-the-art performance, e.g., 90.37% under the 5-way 1-shot setting on CUB-200-2011. It surpasses existing methods by a large margin, especially in the challenging 30-way recognition tasks and cross-domain evaluation. The source code is publicly available:https://github.com/Lx-zjwf/VFI-CVD.
Xin Lu 0007, Yixuan Pan, Yichao Cao, Xin Zhou 0030, Xiaobo Lu
IEEE Trans. Circuits Syst. Video Technol.5
2025 Context-Aided Semantic-Aware Self-Alignment for Video-Based Person Re-Identification
abstract
Video-based person re-identification (Re-ID) aims at associating the video sequences of the identical person across multiple cameras. The ubiquitous appearance misalignment poses a major obstacle for video person Re-ID. Existing alignment-based methods generally rely on off-the-shelf semantic parsing models to locate visible human parts, which ignore identifiable personal belongings and cannot handle various interferences (e.g., pedestrian detection errors and occlusions) in video clips. In this work, we propose a novel framework termed Context-Aided Semantic-Aware Self-Alignment (CSSA) for video-based person Re-ID. First, we propose to jointly learn pixel-level part-aligned representations and semantic-aligned global-level representations in an end-to-end manner. Unlike most existing approaches that depend on prior information in terms of pose for part estimation, CSSA can locate different body parts and achieve the pixel-level semantic alignment without extra human topology semantics. Second, a Context-Aided Region Enhancement (CARE) module is proposed to efficiently highlight macro-visual patterns associated with the target pedestrian and suppress noise caused by factors like background clutters and occlusions. Third, we propose a Semantic-Aware Global Feature Alignment (SGFA) method for generating pair-wise semantic-aligned global representations, which play an essential role in both the training and inference phases. Extensive experimental results on multiple challenging benchmarks indicate the superiority and effectiveness of the proposed CSSA.
Zhidan Ran, Zhiyao Xiao, Xiaobo Lu, Wei Liu 0175
IEEE Trans. Circuits Syst. Video Technol.3
2025 Tex2Sem: Learning From Textures to Semantics for Robust Semantic Correspondence
abstract
Recent advances in semantic correspondence have witnessed growing interest in vision foundation models, particularly stable diffusion (SD) and self-distillation with no labels (DINO). However, existing methods underutilize the matching potential of SD and DINOv2 features and show similar background interference patterns. They lack texture-to-semantic learning and intra- and inter-image feature interaction. This study proposes Tex2Sem, a framework learning from textures to semantics, to address the two problems. For the first problem, we propose a texture-to-semantic learning paradigm that achieves texture-semantic trade-offs on features and correlation maps, including progressive fusion and correlation map computation. The SD and DINOv2 features are aggregated from textures to semantics to produce multi-stage progressive fusion features. The resulting multi-stage progressive fusion correlation maps improve semantic correspondence significantly. For the second problem, MamFormer, a hybrid architecture of Mamba-2 and Transformer, is proposed to improve intra- and inter-image feature aggregation and interaction. It enhances foreground focus and background suppression. Given the high computational cost of processing all-stage progressive fusion features, the terminal-stage aggregation and interaction mechanism (TAIM) is proposed to enhance feature learning efficiency. Experiments demonstrate that Tex2Sem achieves state-of-the-art performance on SPair-71k, AP-10K, and PF-PASCAL. Furthermore, Tex2Sem shows remarkable generalization capabilities in cross-species, cross-family, and cross-dataset matching and demonstrates the potential for applications in video swap and human pose estimation. Code is available at https://github.com/wzhlearning/Tex2Sem.
Zenghui Wang 0009, Songlin Du, Yaping Yan, Guobao Xiao, Xiaobo Lu
IEEE Trans. Circuits Syst. Video Technol.5
2025 Input-Regulated Remote Sensing Counting With Region Understanding
abstract
Remote sensing counting aims to automatically estimate the number of objects of interest from high-resolution aerial or satellite imagery, providing critical decision-making support in areas such as urban planning, traffic monitoring, and disaster response. While most existing methods leverage pre-trained models to enhance feature generalization, their performance is often hindered by the severe scarcity of annotated remote sensing data. This limits their generalizability in complex scenarios. To address these challenges, we propose a novel remote sensing counting network that effectively captures informative signals from relatively limited annotated data. Specifically, we first introduce a graph-driven input regulator that constructs a graph structure by modeling relationships among input features, effectively capturing intrinsic contextual dependencies. This structure allows the regulator to assign adaptive pixel-level weights to network inputs, prioritizing relevant signals while mitigating the risk of overfitting to a fixed data distribution. Second, we design a dynamic region-aware module that leverages fuzzy logic to adaptively identify and enhance highly discriminative local regions. In this way, it improves the robustness of the feature representations. Extensive experiments demonstrate the effectiveness of the proposed method compared with several state-of-the-art methods.
Shengqin Jiang, Haojian Long, Fengna Cheng, Yuankai Qi, Xiaobo Lu, Qingshan Liu 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 Momentum Contrastive Teacher for Semi-Supervised Skeleton Action Recognition
abstract
In the field of semi-supervised skeleton action recognition, existing work primarily follows the paradigm of self-supervised training followed by supervised fine-tuning. However, self-supervised learning focuses on exploring data representation rather than label classification. Inspired by Mean Teacher, we explore a novel pseudo-label-based model called SkeleMoCLR. Specifically, we use MoCo v2 as the foundation and extend it into a teacher-student network through a momentum encoder. The generation of high-confidence pseudo-labels requires a well-pretrained model as a prerequisite. In cases where large-scale skeleton data is lacking, we propose leveraging contrastive learning to transfer discriminative action features from large vision-text models to the skeleton encoder. Following the contrastive pre-training, the key encoder branch from MoCo v2 serves as the teacher to generate pseudo-labels for training the query encoder branch. Furthermore, we introduce pseudo-labels into the memory queues, sampling negative samples from different pseudo-label classes to maximize the representation differentiation between different categories. We jointly optimize the classification loss for both labeled and pseudo-labeled data and the contrastive loss for unlabeled data to update model parameters, fully harnessing the potential of pseudo-label semi-supervised learning and self-supervised learning. Extensive experiments conducted on the NTU-60, NTU-120, PKU-MMD, and NW-UCLA datasets demonstrate that our SkeleMoCLR outperforms existing competitive methods in the semi-supervised skeleton action recognition task.
Xiaobo Lu, Jun Liu 0036
IEEE Trans. Image Process.2
2024 Anomaly-Aware Semantic Self-Alignment Framework for Video-Based Person Re-Identification
abstract
Video-based person re-identification (Re-ID) aims at matching the video snippets of the same person across multiple cameras. The ubiquitous appearance misalignment is a critical challenge in video person re-identification. Existing alignment-based methods rely on off-the-shelf human parsing models and cannot handle anomalous appearance information (e.g., obstacles and pedestrian interference) in video sequences. In this paper, we propose Anomaly-Aware Semantic Self-Alignment (ASSA), a novel video-based person Re-ID framework that seeks out body parts without prior human topology information and learns part-based feature representations against anomalous information. The proposed ASSA performs part classifier training and part-aligned representation learning alternately. For the classifier training, we design a Salient Region Extraction module to segment the entire foreground from the background in each input frame. Furthermore, a novel Anomaly-Aware Refinement module is proposed to suppress the influence of anomalous interference. Extensive experiments on three prevalent benchmarks demonstrate the effectiveness and superiority of the proposed framework.
Zhidan Ran, Xiaobo Lu
ICASSP2
2024 LiDAR-PTQ: Post-Training Quantization for Point Cloud 3D Object Detection
abstract
Due to highly constrained computing power and memory, deploying 3D lidar-based detectors on edge devices equipped in autonomous vehicles and robots poses a crucial challenge. Being a convenient and straightforward model compression approach, Post-Training Quantization (PTQ) has been widely adopted in 2D vision tasks. However, applying it directly to 3D lidar-based tasks inevitably leads to performance degradation. As a remedy, we propose an effective PTQ method called LiDAR-PTQ, which is particularly curated for 3D lidar detection (both SPConv-based and SPConv-free). Our LiDAR-PTQ features three main components, (1) a sparsity-based calibration method to determine the initialization of quantization parameters, (2) an adaptive rounding-to-nearest operation to minimize the layerwise reconstruction error, (3) a Task-guided Global Positive Loss (TGPL) to reduce the disparity between the final predictions before and after quantization. Extensive experiments demonstrate that our LiDAR-PTQ can achieve state-of-the-art quantization performance when applied to CenterPoint (both Pillar-based and Voxel-based). To our knowledge, for the very first time in lidar-based 3D detection tasks, the PTQ INT8 model's accuracy is almost the same as the FP32 model while enjoying 3X inference speedup. Moreover, our LiDAR-PTQ is cost-effective being 6X faster than the quantization-aware training method. The code will be released.
Sifan Zhou, Liang Li 0003, Xinyu Zhang 0015, Bo Zhang 0046, Shipeng Bai, Xiaobo Lu, Xiangxiang Chu
ICLR8
2024 Single-Domain Generalization Combining Geometric Context Toward Instance Segmentation of Track Components
abstract
A stable track components segmentation model should have consistent performance across a broad spectrum of railroad conditions, particularly in unfamiliar locations. Despite this, satisfying this requirement proves challenging when working with a limited track dataset, as there is a substantial domain shift between the given dataset and unobserved distributions. The goal of this paper is to improve the generalization ability of track component segmentation in situations where single-domain training data is available. Toward this end, a novel track component instance segmentation method combining the geometric context is proposed. First, we design an initial mask prediction head (IMPH) that utilizes predicted box output from the object detector to generate initial masks by merging the geometric priors of track components. Meanwhile, a multiscale feature fusion structure is introduced to encourage IMPH to better capture the geometric context. Then, a final mask refinement head (FMRH) is introduced to get higher quality masks with a geometry-gated aggregation strategy. On the basis of experiments on track datasets, it was determined that the heads can be incorporated with several types of detection frameworks and have demonstrated consistent generalization enhancements across multiple object detectors. Furthermore, our method substantially enhances segmentation performance on various unobserved domains.
Shixiang Su, Songlin Du, Dezhou Wang, Shuzhen Tong, Xiaobo Lu
IJCNN5
2024 Distilling object detectors with efficient logit mimicking and mask-guided feature imitation
Xin Lu 0007, Yichao Cao, Shikun Chen, Xin Zhou 0030, Xiaobo Lu
Expert Syst. Appl.6
2024 CLDE-Net: crowd localization and density estimation based on CNN and transformer network
Yaocong Hu, Huicheng Yang, Bingyou Liu, Guoyang Wan, Jinwen Hong, Xiaobo Lu
Multim. Syst.9
2024 ILSR-Diff: joint face illumination normalization and super-resolution via diffusion models
Minghao Mu, Yaocong Hu, Xiaobo Lu
Multim. Syst.5
2024 ESDAR-net: towards high-accuracy and real-time driver action recognition for embedded systems
Yaocong Hu, Zhen Shuai, Huicheng Yang, Guoyang Wan, MingQi Lu, Xiaobo Lu
Multim. Tools Appl.8
2024 AMFF-net: adaptive multi-modal feature fusion network for image classification
Xiaobo Lu
Multim. Tools Appl.2
2024 An objectness-aware network for wildlife detection
Xin Lu 0007, Xiaobo Lu
Multim. Tools Appl.3
2024 Towards better small object detection in UAV scenes: Aggregating more object-oriented information
Chenyue Yang, Yichao Cao, Xiaobo Lu
Pattern Recognit. Lett.3
2024 Mentor: A Memory-Efficient Sparse-dense Matrix Multiplication Accelerator Based on Column-Wise Product
abstract
Sparse-dense matrix multiplication (SpMM) is the performance bottleneck of many high-performance and deep-learning applications, making it attractive to design specialized SpMM hardware accelerators. Unfortunately, existing hardware solutions do not take full advantage of data reuse opportunities of the input and output matrices or suffer from irregular memory access patterns. Their strategies increase the off-chip memory traffic and bandwidth pressure, leaving much room for improvement. We present Mentor , a new approach to designing SpMM accelerators. Our key insight is that column-wise dataflow, while rarely exploited in prior works, can address these issues in SpMM computations. Mentor is a software-hardware co-design approach for leveraging column-wise dataflow to improve data reuse and regular memory accesses of SpMM. On the software level, Mentor incorporates a novel streaming construction scheme to preprocess the input matrix for enabling a streaming access pattern. On the hardware level, it employs a fully pipelined design to unlock the potential of column-wise dataflow further. The design of Mentor is underpinned by a carefully designed analytical model to find the tradeoff between performance and hardware resources. We have implemented an FPGA prototype of Mentor . Experimental results show that Mentor achieves speedup by geomean 2.05× (up to 3.98×), reduces the memory traffic by geomean 2.92× (up to 4.93×), and improves bandwidth utilization by geomean 1.38× (up to 2.89×), compared with the state-of-the-art hardware solutions.
Xiaobo Lu, Jianbin Fang, Lin Peng 0001, Chun Huang 0006, Zidong Du, Yongwei Zhao 0001, Zheng Wang 0079
ACM Trans. Archit. Code Optim.1
2024 Cross-Modal Contrastive Pre-Training for Few-Shot Skeleton Action Recognition
abstract
This paper proposes a novel approach for few-shot skeleton action recognition that comprises of two stages: cross-modal pre-training of a skeleton encoder, followed by fine-tuning of a cosine classifier on the support set. The pre-training and fine-tuning approach has been demonstrated to be more effective for handling few-shot tasks compared to utilizing more intricate meta-learning methods. However, its success relies on the availability of a large-scale training dataset, which yet is difficult to obtain. To address this challenge, we introduce a cross-modal pre-training framework based on Bootstrap Your Own Latent (BYOL), which considers skeleton sequences and their corresponding videos as augmented views of the same action in different modalities. By utilizing a simple regression loss, the framework is able to transfer robust and high-quality vision-language representations to the skeleton encoder. This allows the skeleton encoder to gain a comprehensive understanding of action sequences and benefit from the prior knowledge obtained from a vision-language pre-trained model. The representation transfer enhances the feature extraction capability of the skeleton encoder, compensating for the lack of large-scale skeleton datasets. Extensive experiments on the NTU RGB+D, NTU RGB+D 120, PKU-MMD, NW-UCLA, and MSR Action Pairs datasets demonstrate that our proposed approach achieves state-of-the-art performances for few-shot skeleton action recognition.
Siyuan Yang 0001, Xiaobo Lu, Jun Liu 0036
IEEE Trans. Circuits Syst. Video Technol.3
2024 Multiscale Aligned Spatial-Temporal Interaction for Video-Based Person Re-Identification
abstract
Video-based person re-identification (Re-ID) aims at retrieving the video clips of the same person across multiple cameras. Since video clips are captured at various spatial resolutions (scales), learning multi-scale person appearance features while constructing the cross-scale information interaction is pivotal for video-based person Re-ID. In this paper, we propose an efficient framework, Multi-Scale Aligned Spatial-Temporal Interaction (MS-STI), which not only exchanges the spatial-temporal information within a scale, but also mines implicit related complementary knowledge across scales. MS-STI presents a hierarchical multi-branch architecture that designs the branches with fewer convolutional layers for lower spatial resolution inputs. In this way, the framework enables inter-scale feature size matching for exchanging information across multiple scale-specific branches. We share the parameters of branched sub-networks to optimize the efficiency of person feature extraction. Furthermore, we propose two modules, Spatial Interaction (SI) and Multi-Scale Temporal Interaction (MSTI), which can realize spatial-temporal interaction across multiple branches. SI performs point-wise spatial information transfer within a frame. While MSTI focuses on inter-frame and inter-scale information interaction. Extensive experiments on three challenging benchmarks demonstrate the effectiveness and superiority of the proposed MS-STI.
Zhidan Ran, Wei Liu 0175, Xiaobo Lu
IEEE Trans. Circuits Syst. Video Technol.4
2024 F2CENet: Single-Image Object Counting Based on Block Co-Saliency Density Map Estimation
abstract
This paper presents a novel single-image object counting method based on block co-saliency density map estimation, called free-to-count everything network (F2CENet). Image block co-saliency attention is introduced to promote density estimation adaptation, allowing to input any image with arbitrary size for accurate counting using the learned model without requiring manually labeled few shots. The proposed network also outperforms existing crowd counting methods based on geometry-adaptive kernels in complex scenes. A novel module generates multilevel & scale block correlation maps to guide the co-saliency density map estimation. Co-saliency attention maps are then fused for accurately locating block-wise salient objects under guidance of the initial cues. Hence, accurate density maps are generated via comprehensive learning of internal relations in block co-salient features and progressive optimization of local details with saliency-oriented scene understanding. Results from extensive experiments on existing density map estimation datasets with arbitrary challenges verify the effectiveness of the proposed F2CENet and show that it outperforms various state-of-the-art few-shot and crowd counting methods. Mean Absolute Error (MAE) and Root Mean Square Error (RMSE) are used as evaluation metrics to measure the accuracy which are commonly used metrics for counting task. The average predicted MAE and RMSE are 10.88% and 8.44% less compared with the state-of-the-art evaluated on dataset contains sufficiently large and diverse categories used for few-shot and crowd counting.
Xuehui Wu, Huanliang Xu, Henry Leung 0001, Xiaobo Lu
IEEE Trans. Circuits Syst. Video Technol.4
2023 Coarse2Fine: Local Consistency Aware Re-prediction for Weakly Supervised Object Localization
abstract
Weakly supervised object localization aims to localize objects of interest by using only image-level labels. Existing methods generally segment activation map by threshold to obtain mask and generate bounding box. However, the activation map is locally inconsistent, i.e., similar neighboring pixels of the same object are not equally activated, which leads to the blurred boundary issue: the localization result is sensitive to the threshold, and the mask obtained directly from the activation map loses the fine contours of the object, making it difficult to obtain a tight bounding box. In this paper, we introduce the Local Consistency Aware Re-prediction (LCAR) framework, which aims to recover the complete fine object mask from locally inconsistent activation map and hence obtain a tight bounding box. To this end, we propose the self-guided re-prediction module (SGRM), which employs a novel superpixel aggregation network to replace the post-processing of threshold segmentation. In order to derive more reliable pseudo label from the activation map to supervise the SGRM, we further design an affinity refinement module (ARM) that utilizes the original image feature to better align the activation map with the image appearance, and design a self-distillation CAM (SD-CAM) to alleviate the locator dependence on saliency. Experiments demonstrate that our LCAR outperforms the state-of-the-art on both the CUB-200-2011 and ILSVRC datasets, achieving 95.89% and 70.72% of GT-Know localization accuracy, respectively.
Yixuan Pan, Yichao Cao, Chongjin Chen, Xiaobo Lu
AAAI5
2023 Re-mine, Learn and Reason: Exploring the Cross-modal Semantic Correlations for Language-guided HOI detection
abstract
Human-Object Interaction (HOI) detection is a challenging computer vision task that requires visual models to address the complex interactive relationship between humans and objects and predicttriplets. Despite the challenges posed by the numerous interaction combinations, they also offer opportunities for multi-modal learning of visual texts. In this paper, we present a systematic and unified framework (RmLR) that enhances HOI detection by incorporating structured text knowledge. Firstly, we qualitatively and quantitatively analyze the loss of interaction information in the two-stage HOI detector and propose a re-mining strategy to generate more comprehensive visual representation. Secondly, we design more fine-grained sentence- and word-level alignment and knowledge transfer strategies to effectively address the many-to-many matching problem between multiple interactions and multiple texts. These strategies alleviate the matching confusion problem that arises when multiple interactions occur simultaneously, thereby improving the effectiveness of the alignment process. Finally, HOI reasoning by visual features augmented with textual knowledge substantially improves the understanding of interactions. Experimental results illustrate the effectiveness of our approach, where state-of-the-art performance is achieved on public benchmarks.
Yichao Cao, Qingfei Tang, Xiu Su, Shan You, Xiaobo Lu, Chang Xu 0002
ICCV6
2023 DCPB: Deformable Convolution based on the Poincaré Ball for Top-view Fisheye Cameras
abstract
The accuracy of the visual tasks for top-view fisheye cameras is limited by the Euclidean geometry for pose-distorted objects in images. In this paper, we demonstrate the analogy between the fisheye model and the Poincaré ball and that learning the shape of convolution kernels in the Poincaré Ball can alleviate the spatial distortion problem. In particular, we propose the Deformable Convolution based on the Poincaré Ball, named DCPB, which conducts the Graph Convolutional Network (GCN) in the Poincaré ball and calculates the geodesic distances to Poincaré hyperplanes as the offsets and modulation scalars of the modulated deformable convolution. Besides, we explore an appropriate network structure in the baseline with the DCPB. The DCPB markedly improves the neural network’s performance. Experimental results on the public dataset THEODORE show that DCPB obtains a higher accuracy, and its efficiency demonstrates the potential for using temporal information in fisheye videos.
Zhidan Ran, Xiaobo Lu
ICCV3
2023 Attributes Grouping and Mining Hashing for Fine-Grained Image Retrieval
abstract
In recent years, hashing methods have been popular in the large-scale media search for low storage and strong representation capabilities. To describe objects with similar overall appearance but subtle differences, more and more studies focus on hashing-based fine-grained image retrieval. Existing hashing networks usually generate both local and global features through attention guidance on the same deep activation tensor, which limits the diversity of feature representations. To handle this limitation, we substitute convolutional descriptors for attention-guided features and propose an Attributes Grouping and Mining Hashing (AGMH), which groups and embeds the category-specific visual attributes in multiple descriptors to generate a comprehensive feature representation for efficient fine-grained image retrieval. Specifically, an Attention Dispersion Loss (ADL) is designed to force the descriptors to attend to various local regions and capture diverse subtle details. Moreover, we propose a Stepwise Interactive External Attention (SIEA) to mine critical attributes in each descriptor and construct correlations between fine-grained attributes and objects. The attention mechanism is dedicated to learning discrete attributes, which will not cost additional computations in hash codes generation. Finally, the compact binary codes are learned by preserving pairwise similarities. Experimental results demonstrate that AGMH consistently yields the best performance against state-of-the-art methods on fine-grained benchmark datasets.
Xin Lu 0007, Shikun Chen, Yichao Cao, Xin Zhou 0030, Xiaobo Lu
ACM Multimedia5
2023 Detecting Any Human-Object Interaction Relationship: Universal HOI Detector with Spatial Prompt Learning on Foundation Models
abstract
Human-object interaction (HOI) detection aims to comprehend the intricate relationships between humans and objects, predicting <human, action, object> triplets, and serving as the foundation for numerous computer vision tasks. The complexity and diversity of human-object interactions in the real world, however, pose significant challenges for both annotation and recognition, particularly in recognizing interactions within an open world context. This study explores the universal interaction recognition in an open-world setting through the use of Vision-Language (VL) foundation models and large language models (LLMs). The proposed method is dubbed as UniHOI. We conduct a deep analysis of the three hierarchical features inherent in visual HOI detectors and propose a method for high-level relation extraction aimed at VL foundation models, which we call HO prompt-based learning. Our design includes an HO Prompt-guided Decoder (HOPD), facilitates the association of high-level relation representations in the foundation model with various HO pairs within the image. Furthermore, we utilize a LLM (i.e. GPT) for interaction interpretation, generating a richer linguistic understanding for complex HOIs. For open-category interaction recognition, our method supports either of two input types: interaction phrase or interpretive sentence. Our efficient architecture design and learning methods effectively unleash the potential of the VL foundation models and LLMs, allowing UniHOI to surpass all existing methods with a substantial margin, under both supervised and zero-shot settings. The code and pre-trained weights will be made publicly available.
Yichao Cao, Qingfei Tang, Xiu Su, Shan You, Xiaobo Lu, Chang Xu 0002
NeurIPS6
2023 Keypoint-enhanced adaptive weighting model with effective frequency channel attention for driver action recognition
MingQi Lu, Xiaobo Lu
Eng. Appl. Artif. Intell.2
2023 IMDet: Injecting more supervision to CenterNet-like object detection
Shukun Jia, Yichao Cao, Xiaobo Lu
Expert Syst. Appl.4
2023 Uncertainty meets fixed-time control in neural networks
Shengqin Jiang, Yu Liu 0029, Shuiming Cai, Xiaobo Lu
Neurocomputing5
2023 HD-YOLO: Using radius-aware loss function for head detection in top-view fisheye images
Xiaobo Lu
J. Vis. Commun. Image Represent.3
2023 Bridging the MiniBatch and Adversarial Optimal Transport for Cross-Scene Classification of Remote Sensing Smoke-Related Scenes
abstract
Smoke in remote sensing (RS) images is regarded as the indicator of fire disasters and it is vital to distinguish smoke from other RS scenes. Convolutional neural network (CNN) based classification networks have viewed great success but the cross-entropy loss may suffer from domain gap when training data and testing data do not follow independent and identical distributions. Served as an objective function, the optimal transport (OT) distance is able to tackle the domain gap by aligning two distributions in domain adaptation tasks. Generally speaking, the OT-based algorithms can be divided into two categories: methods that compute OT distance on minibatches, and methods which rely on the adversarial training. In this letter, we propose a hybrid OT algorithm, termed Joint MiniBatch and Adversarial Optimal Transport (JMBAOT), to eliminate the domain gap for cross-scene classification of RS smoke-related scenes. JMBAOT takes advantages of both the minibatch and adversarial OT to better distinguish different classes on the feature space and improve the performance of domain adaptation. In a large amount of experiments, JMBAOT shows its effectiveness and achieves the state-of-the-art (SOTA) performance.
Shikun Chen, Xin Lu 0007, Xiaobo Lu
IEEE Geosci. Remote. Sens. Lett.4
2023 RFS-Net: Railway Track Fastener Segmentation Network With Shape Guidance
abstract
The fastener is one of the main components of a rail track system. In recent years, deep learning methods such as image segmentation have greatly boosted the fastener state detection process. However, there is still a need to improve the segmentation accuracy and speed, especially for the fasteners in complex environments. To handle this problem, a fast and accurate fastener semantic segmentation network named RFS-Net is proposed based on shape guidance, which can offer a better speed/accuracy trade-off performance via a very shallow architecture. Specifically, in the encoder, a two-stream structure (i.e., regular stream and shape stream) that processes the fastener and shape image in parallel is introduced. The shape image is created based on the geometric structure of the fastener, and it is served as input to the shape stream to guide the segmentation of the fastener. The decoder integrates deep features from the two-stream encoder and then recovers the shape information by the shape attention blocks with skipping connections. We provide two versions of RFS-Net: RFS-Net_S (1.0M, 1014FPS) and RFS-Net_L (12.01M, 453FPS) on the NVIDIA RTX 3060. Experimental results demonstrate the effectiveness of our method by achieving a promising trade-off between accuracy and inference speed. In particular, our method is faster and more accurate on a challenging dataset, from fast modes: 1014 FPS for RFS-Net_S versus 724 FPS for Segmenter, to high-quality segmentation: better performance than STDC with nearly one percent (92.36% versus 91.48% Mean IoU score).
Shixiang Su, Songlin Du, Xiaobo Lu
IEEE Trans. Circuits Syst. Video Technol.4
2023 Rotational Convolution: Rethinking Convolution for Downside Fisheye Images
abstract
It has long been recognized that the standard convolution is not rotation equivariant and thus not appropriate for downside fisheye images which are rotationally symmetric. This paper introduces Rotational Convolution, a novel convolution that rotates the convolution kernel by characteristics of downside fisheye images. With the four rotation states of the convolution kernel, Rotational Convolution can be implemented on discrete signals. Rotational Convolution improves the performance of different networks in semantic segmentation and object detection markedly, harming the inference speed slightly. Finally, we demonstrate our methods' numerical accuracy, computational efficiency, and effectiveness on the public segmentation dataset THEODORE and our self-built detection dataset SEU-fisheye. Our code is available at: https://github.com/wx19941204/Rotational-Convolution-for-downside-fisheye-images.
Shixiang Su, Xiaobo Lu
IEEE Trans. Image Process.4
2023 Joint Image-to-Image Translation for Traffic Monitoring Driver Face Image Enhancement
abstract
The real traffic monitoring driver face (TMDF) images are with complex multiple degradations, which decline face recognition accuracy in real intelligent transportation systems (ITS). This paper is the first to propose joint image-to-image (I2I) translation to enhance TMDF images of ITS. First, as TMDF images are without corresponding clear ones, identity preserving is critical for TMDF images under unpaired I2I translation. This paper proposes a fast diagonal symmetry pattern (FDSP) to preserve identity structure under unpaired I2I translation. Second, FDSP is introduced into CycleGAN to form FDSP-CG, which aims to learn the degradation mapping (i.e., FDSP-CG-d) from the clarity domain to the degradation domain. FDSP-CG-d can generate massive degradation/clarity image pairs for paired I2I translation training. Third, this paper proposes the dual residual block (DRB) to strengthen Pix2pix for rich face detail features learning (i.e., DRB-P2P), which learns the enhancement mapping from the degradation image to its clear version under paired I2I translation. Finally, the experiments on TMDF (i.e., the brevity name of the face database collected from real ITS) and Chinese famous face (CFF) databases, as well as CelebA and MegaFace databases, indicate that the proposed method can efficiently enhance TMDF images whose degradation variations are learned by FDSP-CG.
Changhui Hu 0001, Lin-Tao Xu, Xiaoyuan Jing, Xiaobo Lu, Wankou Yang, Pan Liu 0013
IEEE Trans. Intell. Transp. Syst.5
2023 HSV-3S and 2D-GDA for High-Saturation Low-Light Image Enhancement in Night Traffic Monitoring
abstract
This paper proposes HSV (hue, saturation, value) with three sectors (HSV-3S) and two-dimensional gradient descent algorithm (2D-GDA) for high-saturation low-light image enhancement in night traffic monitoring (NTM). The saturation of HSV-3S is defined as the ratio of the projection vector length and twice length of the sector start vector, which results in that the saturation of HSV-3S is smaller than that of HSV, and a saturation weakening model is proposed to further decrease the saturation of HSV-3S. The hue of HSV-3S is defined as the cosine value of the included angle between the projection vector and the sector start vector in each of three sectors. HSV-3S is more concise and faster than HSV. Then, 2D-GDA extends the gradient descent algorithm to 2D image domain. 2D-GDA employs the iteration matrix with variable step values (i.e., the step values of the dark regions are less than those of the bright regions), which can improve the pixel distribution of the 2D-GDA enhanced image. Finally, the HSV-3S+2D- GDA based RGB image can be obtained by performing 2D-GDA on the value of HSV-3S with transforming the processed HSV-3S to RGB. The experimental results on NTM (i.e., the brevity name of the database collected from real ITS), LOL, ExDark and SICE databases, indicate that HSV-3S+2D-GDA is fast and efficient for high-saturation low-light image enhancement.
Changhui Hu 0001, Lin-Tao Xu, Yanyong Guo, Xiaoyuan Jing, Xiaobo Lu, Pan Liu 0013
IEEE Trans. Intell. Transp. Syst.5
2023 MFANet: Multifaceted Feature Aggregation Network for Oil Stains Detection of High-Speed Trains
abstract
Oil is of great significance in the key components like the hydraulic damper of high-speed trains and its leaks suggest that some key components may malfunction, bringing potential danger to the safety of the train. It is indispensable to diagnose oil leakage timely and faults can be discovered by detecting oil stains to avoid possible accidents caused by the breakdown. But it is challenging to discover the oil stains due to their irregularity of shape and diversity of size in complex environments. To deal with these intractable problems, we propose the Multifaceted Feature Aggregation Network (MFANet) which is composed of Multifaceted Refinement Feature (MRF) module, Cross Layer Feature Attention (CLFA) module, and Cross Layer Feature Enhancement (CLFE) module. The MRF solves the limitation of the single convolution method’s inability to catch changeable oil stains better and captures more expressive features. The CLFA and the CLFE obtain richer information and compensate for the shortcomings of insufficient information on single-layer features. Furthermore, we propose a loss function to guide the learning of our network in the case of the unbalanced dataset, which surpasses the useful Binary Cross-entropy loss. We conduct experiments based on four widely used backbone networks and achieve cutting-edge performance with fewer FLOPs and Parameters compared with many advanced methods. In addition, we perform our MFANet on several challenging saliency detection datasets, the rail surface defects detection dataset, and the COCO-Stuff segmentation dataset, which outperforms some progressive approaches and demonstrates powerful generalization capability and adaptation.
Wei Liu 0175, Xiaobo Lu, Zhidan Ran
IEEE Trans. Intell. Transp. Syst.2
2022 CNN-Transformer Hybrid Architecture for Early Fire Detection
Chenyue Yang, Yixuan Pan, Yichao Cao, Xiaobo Lu
ICANN (4)4
2022 Searching for Better Spatio-temporal Alignment in Few-Shot Action Recognition
abstract
Spatio-Temporal feature matching and alignment are essential for few-shot action recognition as they determine the coherence and effectiveness of the temporal patterns. Nevertheless, this process could be not reliable, especially when dealing with complex video scenarios. In this paper, we propose to improve the performance of matching and alignment from the end-to-end design of models. Our solution comes at two-folds. First, we encourage to enhance the extracted Spatio-Temporal representations from few-shot videos in the perspective of architectures. With this aim, we propose a specialized transformer search method for videos, thus the spatial and temporal attention can be well-organized and optimized for stronger feature representations. Second, we also design an efficient non-parametric spatio-temporal prototype alignment strategy to better handle the high variability of motion. In particular, a query-specific class prototype will be generated for each query sample and category, which can better match query sequences against all support sequences. By doing so, our method SST enjoys significant superiority over the benchmark UCF101 and HMDB51 datasets. For example, with no pretraining, our method achieves 17.1\% Top-1 accuracy improvement than the baseline TRX on UCF101 5-way 1-shot setting but with only 3x fewer FLOPs.
Yichao Cao, Xiu Su, Qingfei Tang, Shan You, Xiaobo Lu, Chang Xu 0002
NeurIPS5
2022 A pose-aware dynamic weighting model using feature integration for driver action recognition
MingQi Lu, Yaocong Hu, Xiaobo Lu
Eng. Appl. Artif. Intell.3
2022 RMDC: Rotation-mask deformable convolution for object detection in top-view fisheye cameras
Xiaobo Lu
Neurocomputing3
2022 STCNet: spatiotemporal cross network for industrial smoke detection
Yichao Cao, Qingfei Tang, Xiaobo Lu
Multim. Tools Appl.3
2022 QuasiVSD: efficient dual-frame smoke detection
Yichao Cao, Qingfei Tang, Shaosheng Xu, Xiaobo Lu
Neural Comput. Appl.5
2022 Recursive Copy and Paste GAN: Face Hallucination From Shaded Thumbnails
abstract
Existing face hallucination methods based on convolutional neural networks (CNNs) have achieved impressive performance on low-resolution (LR) faces in a normal illumination condition. However, their performance degrades dramatically when LR faces are captured in non-uniform illumination conditions. This paper proposes a Recursive Copy and Paste Generative Adversarial Network (Re-CPGAN) to recover authentic high-resolution (HR) face images while compensating for non-uniform illumination. To this end, we develop two key components in our Re-CPGAN: internal and recursive external Copy and Paste networks (CPnets). Our internal CPnet exploits facial self-similarity information residing in the input image to enhance facial details; while our recursive external CPnet leverages an external guided face for illumination compensation. Specifically, our recursive external CPnet stacks multiple external Copy and Paste (EX-CP) units in a compact model to learn normal illumination and enhance facial details recursively. By doing so, our method offsets illumination and upsamples facial details progressively in a coarse-to-fine fashion, thus alleviating the ambiguity of correspondences between LR inputs and external guided inputs. Furthermore, a new illumination compensation loss is developed to capture illumination from the external guided face image effectively. Extensive experiments demonstrate that our method achieves authentic HR face images in a uniform illumination condition with a 16× magnification factor and outperforms state-of-the-art methods qualitatively and quantitatively.
Yang Zhang 0067, Ivor W. Tsang, Yawei Luo, Changhui Hu 0001, Xiaobo Lu, Xin Yu 0002
IEEE Trans. Pattern Anal. Mach. Intell.5
2022 Camera domain adaptation based on cross-patch transformers for person re-identification
Zhidan Ran, Xiaobo Lu
Pattern Recognit. Lett.2
2022 Pose-guided model for driving behavior recognition using keypoint action learning
MingQi Lu, Yaocong Hu, Xiaobo Lu
Signal Process. Image Commun.3
2022 EFFNet: Enhanced Feature Foreground Network for Video Smoke Source Prediction and Detection
abstract
Smoke detection in video is a challenging task because of the irregular shape of smoke, its complex motion state, which is affected by temperature, wind and other external factors, and background disturbances. Pixel-based foreground modeling method is a crucial step in many smoke detection systems and can be applied to efficiently focus on a certain object or a specific region to detect movements or anomalies. In video analysis, it is a natural idea to move the focus from the pixel-level foreground to the feature-level foreground. In this paper, the feature foreground is generated by the middle layer of a convolutional neural network (CNN) to guide the temporal modeling process for smoke objects. A novel temporal module called the Feature Foreground Module (FFM) is proposed to boost learning of a smoke temporal representation. Consider the problem of smoke analysis in video, we present a novel unifying approach, named an enhanced feature foreground network (EFFNet), that performs both smoke source prediction and detection. Efficient branch networks are designed in EFFNet, to predict the source mask and bounding boxes of smoke plumes in video. To the best of our knowledge, this is the first paper to study the source of smoke using deep learning methods. Finally, experiments on a realistic smoke dataset and a public dataset show that EFFNet method performs much better than do previous state-of-the-art methods.
Yichao Cao, Qingfei Tang, Xuehui Wu, Xiaobo Lu
IEEE Trans. Circuits Syst. Video Technol.4
2022 Combining the Convolution and Transformer for Classification of Smoke-Like Scenes in Remote Sensing Images
abstract
Remote sensing (RS) images are used in a wide range of tasks. In the fire detection field, smoke in RS images is considered as an indicator of wildfires. However, smoke-like scenes, e.g., cloud, in RS images increase the difficulty of smoke recognition. Convolutional neural networks (CNNs) have greatly promoted the development of image processing. CNNs are good at capturing local features; however, their ability to capture global features is relatively weak. Recently, the transformer deep learning model has shown strong potential in vision tasks. The transformer model utilizes self-attention modules to extract global features but may lose local details. Recognition of smoke in RS images depends strongly on the combination of both local and global features. Thus, this article proposes the transformer enhanced convolutional network (TECN) to classify RS smoke-like scenes. The proposed hybrid TECN model exploits the advantages of the CNN and transformer techniques at the same time. In TECN, the feature merge and intelligent aggregation modules are used to promote conversion and aggregation between CNN feature maps and transformer patch embeddings. Experiments are conducted on the USTC_SmokeRS dataset, which is developed for the classification of RS smoke-like scenes. The experimental results demonstrate that the proposed TECN achieves a competitive accuracy of 98.39% on this dataset.
Shikun Chen, Yichao Cao, Xiaobo Lu
IEEE Trans. Geosci. Remote. Sens.4
2022 Pro-UIGAN: Progressive Face Hallucination From Occluded Thumbnails
abstract
In this paper, we study the task of hallucinating an authentic high-resolution (HR) face from an occluded thumbnail. We propose a multi-stage Progressive Upsampling and Inpainting Generative Adversarial Network, dubbed Pro-UIGAN, which exploits facial geometry priors to replenish and upsample ( 8× ) the occluded and tiny faces ( 16×16 pixels). Pro-UIGAN iteratively (1) estimates facial geometry priors for low-resolution (LR) faces and (2) acquires non-occluded HR face images under the guidance of the estimated priors. Our multi-stage hallucination network upsamples and inpaints occluded LR faces via a coarse-to-fine fashion, significantly reducing undesirable artifacts and blurriness. Specifically, we design a novel cross-modal attention module for facial priors estimation, in which an input face and its landmark features are formulated as queries and keys, respectively. Such a design encourages joint feature learning across the input facial and landmark features, and deep feature correspondences will be discovered by attention. Thus, facial appearance features and facial geometry priors are learned in a mutually beneficial manner. Extensive experiments show that our Pro-UIGAN attains visually pleasing completed HR faces, thus facilitating downstream tasks, i.e., face alignment, face parsing, face recognition as well as expression classification.
Yang Zhang 0067, Xin Yu 0002, Xiaobo Lu, Ping Liu 0004
IEEE Trans. Image Process.3
2022 Geometric Constraint and Image Inpainting-Based Railway Track Fastener Sample Generation for Improving Defect Inspection
abstract
Defective fastener images detection is an essential task in the vision-based railway track safety inspection. Although existing methods have achieved some level of success, the detection accuracy in this field suffers from the defective fasteners being far less common than normal fasteners. One way to tackle this problem is to expand the defect sample. However, current state-of-the-art defective fastener generation methods mainly rely on generative adversarial networks or simply augment the defect data through traditional image processing. These methods may not be ideal as it is difficult to produce images with high quality and rich diversity at the same time. This paper proposes a new method for fastener sample generation that actively divides the sample generation into two independent parts: defective foregrounds generation and complete backgrounds generation. The key to this method is to generate foregrounds and backgrounds based on geometric constraint and image inpainting, respectively. Specifically, we adopt a skeleton mapping algorithm to directionally control the generated types of defective foregrounds. Meanwhile, an image inpainting network is employed to expand the background. The experiments show that this enables us to generate better-quality and richer-diversity images by combining deep learning and image processing advantages. To the best of our knowledge, our method is the first to achieve state-of-the-art performance, i.e., the classification accuracy reaches 97.97%, without using real defective fastener images during the defect classification network training process.
Shixiang Su, Songlin Du, Xiaobo Lu
IEEE Trans. Intell. Transp. Syst.3
2021 PTeacher: a Computer-Aided Personalized Pronunciation Training System with Exaggerated Audio-Visual Corrective Feedback
abstract
Second language (L2) English learners often find it difficult to improve their pronunciations due to the lack of expressive and personalized corrective feedback. In this paper, we present Pronunciation Teacher (PTeacher), a Computer-Aided Pronunciation Training (CAPT) system that provides personalized exaggerated audio-visual corrective feedback for mispronunciations. Though the effectiveness of exaggerated feedback has been demonstrated, it is still unclear how to define the appropriate degrees of exaggeration when interacting with individual learners. To fill in this gap, we interview 100 L2 English learners and 22 professional native teachers to understand their needs and experiences. Three critical metrics are proposed for both learners and teachers to identify the best exaggeration levels in both audio and visual modalities. Additionally, we incorporate the personalized dynamic feedback mechanism given the English proficiency of learners. Based on the obtained insights, a comprehensive interactive pronunciation training course is designed to help L2 learners rectify mispronunciations in a more perceptible, understandable, and discriminative manner. Extensive user studies demonstrate that our system significantly promotes the learners’ learning efficiency.
Yaohua Bu, Hang Zhou 0009, Jia Jia 0001, Shengqi Chen 0001, Dachuan Shi, Haozhe Wu, Kun Li 0003, Zhiyong Wu 0001, Yuanchun Shi, Xiaobo Lu, Ziwei Liu 0002
CHI14
2021 Global2Salient: Self-adaptive feature aggregation for remote sensing smoke detection
Shikun Chen, Yichao Cao, Xiaoqiang Feng, Xiaobo Lu
Neurocomputing4
2021 FIN-GAN: Face illumination normalization via retinex-based self-supervised learning and conditional generative adversarial network
Yaocong Hu, MingQi Lu, Xiaobo Lu
Neurocomputing4
2021 Light fixed-time control for cluster synchronization of complex networks
Shengqin Jiang, Yuankai Qi, Shuiming Cai, Xiaobo Lu
Neurocomputing4
2021 Video-based driver action recognition via hybrid spatial-temporal deep learning framework
Yaocong Hu, MingQi Lu, Xiaobo Lu
Multim. Syst.4
2021 Patchwise dictionary learning for video forest fire smoke detection in wavelet domain
Xuehui Wu, Yichao Cao, Xiaobo Lu, Henry Leung 0001
Neural Comput. Appl.3
2021 Face illumination recovery for the deep learning feature under severe illumination variations
Changhui Hu 0001, Jian Yu 0007, Fei Wu 0004, Yang Zhang 0067, Xiaoyuan Jing, Xiaobo Lu, Pan Liu 0013
Pattern Recognit.6
2021 D3D: Dual 3-D Convolutional Network for Real-Time Action Recognition
abstract
Three-dimensional convolutional neural networks (3D CNNs) have been explored to learn spatio-temporal information for video-based human action recognition. Expensive computational cost and memory demand resulted from standard 3D CNNs, however, hinder their application in practical scenarios. In this article, we address the aforementioned limitations by proposing a novel dual 3-D convolutional network (D3DNet) with two complementary lightweight branches. A coarse branch maintains large temporal receptive field by a fast temporal downsampling strategy and simulates the expensive 3-D convolutions using a combination of more efficient spatial convolutions and temporal convolutions. Meanwhile, a fine branch progressively downsamples the video in the temporal domain and adopts 3-D convolutional units with reduced channel capacities to capture multiresolution spatio-temporal information. Instead of learning these two branches independently, a shallow spatiotemporal downsampling module is shared for these two branches for efficient low-level feature learning. Besides, lateral connections are learned to effectively fuse the information from the two branches at multiple stages. The proposed network makes good balance between inference speed and action recognition performance. Based on RGB information only, it achieves competing performance on five popular video-based action recognition datasets, with inference speed of 3200 FPS on a single NVIDIA GTX 2080Ti card.
Shengqin Jiang, Yuankai Qi, Haokui Zhang, Zongwen Bai, Xiaobo Lu, Peng Wang 0023
IEEE Trans. Ind. Informatics5
2021 Face Hallucination With Finishing Touches
abstract
Obtaining a high-quality frontal face image from a low-resolution (LR) non-frontal face image is primarily important for many facial analysis applications. However, mainstreams either focus on super-resolving near-frontal LR faces or frontalizing non-frontal high-resolution (HR) faces. It is desirable to perform both tasks seamlessly for daily-life unconstrained face images. In this paper, we present a novel Vivid Face Hallucination Generative Adversarial Network (VividGAN) for simultaneously super-resolving and frontalizing tiny non-frontal face images. VividGAN consists of coarse-level and fine-level Face Hallucination Networks (FHnet) and two discriminators, i.e., Coarse-D and Fine-D. The coarse-level FHnet generates a frontal coarse HR face and then the fine-level FHnet makes use of the facial component appearance prior, i.e., fine-grained facial components, to attain a frontal HR face image with authentic details. In the fine-level FHnet, we also design a facial component-aware module that adopts the facial geometry guidance as clues to accurately align and merge the frontal coarse HR face and prior information. Meanwhile, two-level discriminators are designed to capture both the global outline of a face image as well as detailed facial characteristics. The Coarse-D enforces the coarsely hallucinated faces to be upright and complete while the Fine-D focuses on the fine hallucinated ones for sharper details. Extensive experiments demonstrate that our VividGAN achieves photo-realistic frontal HR faces, reaching superior performance in downstream tasks, i.e., face recognition and expression classification, compared with other state-of-the-art methods.
Yang Zhang 0067, Ivor W. Tsang, Jun Li 0010, Ping Liu 0004, Xiaobo Lu, Xin Yu 0002
IEEE Trans. Image Process.5
2020 Copy and Paste GAN: Face Hallucination From Shaded Thumbnails
abstract
Existing face hallucination methods based on convolutional neural networks (CNN) have achieved impressive performance on low-resolution (LR) faces in a normal illumination condition. However, their performance degrades dramatically when LR faces are captured in low or non-uniform illumination conditions. This paper proposes a Copy and Paste Generative Adversarial Network (CPGAN) to recover authentic high-resolution (HR) face images while compensating for low and non-uniform illumination. To this end, we develop two key components in our CPGAN: internal and external Copy and Paste nets (CPnets). Specifically, our internal CPnet exploits facial information residing in the input image to enhance facial details; while our external CPnet leverages an external HR face for illumination compensation. A new illumination compensation loss is thus developed to capture illumination from the external guided face image effectively. Furthermore, our method offsets illumination and upsamples facial details alternatively in a coarse-to-fine fashion, thus alleviating the correspondence ambiguity between LR inputs and external HR inputs. Extensive experiments demonstrate that our method manifests authentic HR face images in a uniform illumination condition and outperforms state-of-the-art methods qualitatively and quantitatively.
Yang Zhang 0067, Ivor W. Tsang, Yawei Luo, Changhui Hu 0001, Xiaobo Lu, Xin Yu 0002
CVPR5
2020 DLEP: A Deep Learning Model for Earthquake Prediction
abstract
Earthquakes are one of the most costly natural disasters facing human beings, which happens without an explicit warning, therefore earthquake prediction becomes a very important and challenging task for humanity. Although many existing methods attempt to address this task, most of them use either seismic indicators (explicit features) designed by geologists, or feature vectors (implicit features) extracted by deep learning methods, to characterize an earthquake for earthquake prediction. The problem of combining these two kind of features to improve final earthquake prediction performance remains pretty much open. To this end, we propose a deep learning model named DLEP to effectively fuse the explicit and implicit features for accurate earthquake prediction. In DLEP, we adopt eight precursory pattern-based indicators as the explicit features, and use a convolutional neural network (CNN) to extract implicit features. Then, an attention-based strategy is suggested to fuse these two kinds of features well. In addition, a dynamic loss function is designed to deal with the category imbalance of seismic data. Finally, experimental results on eight datasets from different regions demonstrate the effectiveness of the proposed DLEP for earthquake prediction comparing to several state-of-the-art baselines.
Rui Li 0093, Xiaobo Lu, Shuowei Li, Haipeng Yang, Jianfeng Qiu, Lei Zhang 0060
IJCNN2
2020 Visual-speech Synthesis of Exaggerated Corrective Feedback
abstract
To provide more discriminative feedback for the second language (L2) learners to better identify their mispronunciation, we propose a method for exaggerated visual-speech feedback in computer-assisted pronunciation training (CAPT). The speech exaggeration is realized by an emphatic speech generation neural network based on Tacotron, while the visual exaggeration is accomplished by ADC Viseme Blending, namely increasing Amplitude of movement, extending the phone's Duration and enhancing the color Contrast. User studies show that exaggerated feedback outperforms non-exaggerated version on helping learners with pronunciation identification and pronunciation improvement.
Yaohua Bu, Shengqi Chen 0001, Jia Jia 0001, Kun Li 0003, Xiaobo Lu
ACM Multimedia7
2020 Diagonal Symmetric Pattern Based Illumination Invariant Measure for Severe Illumination Variations
Changhui Hu 0001, Mengjun Ye, Yang Zhang 0067, Xiaobo Lu
PRCV (2)4
2020 Driver action recognition using deformable and dilated faster R-CNN with optimized region proposals
MingQi Lu, Yaocong Hu, Xiaobo Lu
Appl. Intell.3
2020 A three-stage framework for smoky vehicle detection in traffic surveillance videos
Huanjie Tao, Xiaobo Lu
Inf. Sci.4
2020 A motion and lightness saliency approach for forest smoke segmentation and detection
Xuehui Wu, Xiaobo Lu, Henry Leung 0001
Multim. Tools Appl.2
2020 A universal sample-based background subtraction method for traffic surveillance videos
Xiaobo Lu
Multim. Tools Appl.4
2020 SCRM: self-correlated representation model for visual tracking
Shengqin Jiang, Xiaobo Lu, Fengna Cheng
Soft Comput.2
2020 Feature refinement for image-based driver action recognition via multi-scale attention convolutional neural network
Yaocong Hu, MingQi Lu, Xiaobo Lu
Signal Process. Image Commun.3
2020 Driver Drowsiness Recognition via 3D Conditional GAN and Two-Level Attention Bi-LSTM
abstract
Driver drowsiness has currently been a severe issue threatening road safety, hence it is vital to develop an effective drowsiness recognition algorithm to avoid traffic accidents. However, recognizing drowsiness is still very challenging, due to the large intra-class variations in facial expression, head pose and illumination condition. In this paper, a new deep learning framework based on the hybrid of 3D conditional generative adversarial network and two-level attention bidirectional long short-term memory network (3DcGAN-TLABiLSTM) has been proposed for robust driver drowsiness recognition. Aiming at extracting short-term spatial-temporal features with abundant drowsiness-related information, we design a 3D encoder-decoder generator with the condition of auxiliary information to generate high-quality fake image sequences and devise a 3D discriminator to learn drowsiness-related representation from spatial-temporal domain. In addition, for long-term spatial-temporal fusion, we investigate the use of two-level attention mechanism to guide the bidirectional long short-term memory learn the saliency of short-term memory information and long-term temporal information. For experiment, we evaluate our 3DcGAN-TLABiLSTM framework on a public NTHU-DDD dataset. Experimental results show that the proposed approach achieves higher precision of drowsiness recognition compared to the state-of-the-art.
Yaocong Hu, MingQi Lu, Xiaobo Lu
IEEE Trans. Circuits Syst. Video Technol.4
2020 Mask-Aware Networks for Crowd Counting
abstract
Crowd counting problem aims to count the number of objects within an image or a frame in the videos and is usually solved by estimating the density map generated from the object location annotations. The values in the density map, by nature, take two possible states: zero indicating no object around, a non-zero value indicating the existence of objects and the value denoting the local object density. In contrast to traditional methods which do not differentiate the density prediction of these two states, we propose to use a dedicated network branch to predict the object/non-object mask and then combine its prediction with the input image to produce the density map. Our rationale is that the mask prediction could be better modeled as a binary segmentation problem and the difficulty of estimating the density could be reduced if the mask is known. A key to the proposed scheme is the strategy of incorporating the mask prediction into the density map estimator. To this end, we study five possible solutions, and via analysis and experimental validation we identify the most effective one. Through extensive experiments on three public datasets, we demonstrate the superior performance of the proposed approach over the baselines and show that our network could achieve the state-of-the-art performance.
Shengqin Jiang, Xiaobo Lu, Yinjie Lei, Lingqiao Liu
IEEE Trans. Circuits Syst. Video Technol.2
2020 Smoke Vehicle Detection Based on Spatiotemporal Bag-Of-Features and Professional Convolutional Neural Network
abstract
Existing smoke vehicle detection methods are vulnerable to false alarms. To solve this issue, this paper presents two automatic smoke vehicle detection methods based on spatiotemporal bag-of-features (S-BoF) and professional convolutional neural network (P-CNN). In the first method, we propose the S-BoF model to characterize the key regions detected by the visual background extractor (ViBe) algorithm. The S-BoF model contains three groups of features, including color moments on three orthogonal planes (CM-TOP), completed robust local binary pattern on three orthogonal planes (CRLBP-TOP), and histogram of oriented gradient on three orthogonal planes (HOG-TOP). The extracted features are fed to the support vector machine (SVM) and classify the key regions to smoke regions or non-smoke regions to further detect smoke vehicles. In the second method, we propose the P-CNN model to extract more robust and complementary spatiotemporal features by designing three professional models to analyze different kinds of features in the key region sequence on three orthogonal planes. The three professional models, including color CNN (CCNN), texture CNN (TCNN), and gradient CNN (GCNN), are based on three independent CNN128 models with different inputs. The experimental results show that the proposed methods achieve higher detection rates and lower false alarm rates than existing smoke detection methods.
Huanjie Tao, Xiaobo Lu
IEEE Trans. Circuits Syst. Video Technol.2
2020 Toward Driver Face Recognition in the Intelligent Traffic Monitoring Systems
abstract
This paper models the driver face recognition problem under the intelligent traffic monitoring systems as severe illumination variation face recognition with single sample problem. Firstly, in the point of view of numerical value sign, the current illumination invariant unit is derived from the subtraction of two pixels in the face local region, which may be positive or negative, we propose a generalized illumination robust (GIR) model based on positive and negative illumination invariant units to tackle severe illumination variations. Then, the GIR model can be used to generate several GIR images based on the local edge-region or the local block-region, which results in the edge-region based GIR (EGIR) image or the block-region based GIR (BGIR) image. For single GIR image based classification, the GIR image utilizes the saturation function and the nearest neighbor classifier, which can develop EGIR-face and BGIR-face. For multi GIR images based classification, the GIR images employ the extended sparse representation classification (ESRC) as the classifier that can form the EGIR image based classification (GIRC) and the BGIR image based classification (BGIRC). Further, the GIR model is integrated with the pre-trained deep learning (PDL) model to construct the GIR-PDL model. Finally, the performances of the proposed methods are verified on the Extended Yale B, CMU PIE, AR, self-built Driver and VGGFace2 face databases. The experimental results indicate that the proposed methods are efficient to tackle severe illumination variations.
Changhui Hu 0001, Yang Zhang 0067, Fei Wu 0004, Xiaobo Lu, Pan Liu 0013, Xiaoyuan Jing
IEEE Trans. Intell. Transp. Syst.4
2019 Video smoke separation and detection via sparse representation
Xuehui Wu, Xiaobo Lu, Henry Leung 0001
Neurocomputing2
2019 Dilated Light-Head R-CNN using tri-center loss for driving behavior recognition
MingQi Lu, Yaocong Hu, Xiaobo Lu
Image Vis. Comput.3
2019 Smoke vehicle detection based on robust codebook model and robust volume local binary count patterns
Huanjie Tao, Xiaobo Lu
Image Vis. Comput.2
2019 IL-GAN: Illumination-invariant representation learning for single sample face recognition
Yang Zhang 0067, Changhui Hu 0001, Xiaobo Lu
J. Vis. Commun. Image Represent.3
2019 Learning spatial-temporal representation for smoke vehicle detection
Yichao Cao, Xiaobo Lu
Multim. Tools Appl.2
2019 General logarithm difference model for severe illumination variation face recognition
Changhui Hu 0001, Xiaobo Lu, Fei Wu 0004, Songsong Wu, Xiaoyuan Jing
Multim. Tools Appl.2
2019 Adaptive pixel-block based background subtraction using low-rank and block-sparse matrix decomposition
Xuehui Wu, Xiaobo Lu
Multim. Tools Appl.2
2019 Face detection and alignment method for driver on highroad based on improved multi-task cascaded convolutional networks
Yang Zhang 0067, Peihua Lv, Xiaobo Lu, Jun Li 0010
Multim. Tools Appl.3
2019 Driving behaviour recognition from still images by using multi-stream fusion CNN
Yaocong Hu, MingQi Lu, Xiaobo Lu
Mach. Vis. Appl.3
2019 Fast Single-Image Super-Resolution via Deep Network With Component Learning
abstract
Driven by the spectacular success of deep learning, several advanced models based on neural networks have recently been proposed for single-image super-resolution, incrementally revealing their superiority over their alternatives. In this paper, we pursue this latest line of research and present an improved network structure by taking advantage of the proposed component learning. The core idea and difference of this learning strategy are to use the residual extracted from the input to predict its counterpart in the corresponding output. To this end, a global decomposition procedure is designed on the basis of convolutional sparse coding and performed on the input for extracting the low-resolution (LR) residual component from it. Owing to the properties of this decomposition, the represented residual component still stays in the LR space so that the subsequent part is capable of operating it economically in terms of computational complexity. Thorough experimental results demonstrate the merit and effectiveness of the proposed component learning strategy, and our trained model outperforms many state-of-the-art methods in terms of both speed and reconstruction quality.
Xiaobo Lu
IEEE Trans. Circuits Syst. Video Technol.3
2019 Single Sample Face Recognition Under Varying Illumination via QRCP Decomposition
abstract
In this paper, we present a novel high-frequency facial feature and a high-frequency based sparse representation classification to tackle single sample face recognition (SSFR) under varying illumination. Firstly, we propose the assumption that QRCP bases can represent intrinsic face surface features with different frequencies, and their corresponding energy coefficients describe illumination intensities. Based on this assumption, we take QRCP bases with corresponding weighting coefficients (i.e. the major components of energy coefficients) to develop the high-frequency facial feature of the face image, which is named as QRCP-face. The normalized QRCP-face (NQRCPface) is constructed to further constraint illumination effects by normalizing the weighting coefficients of QRCP-face. Moreover, we propose the adaptive QRCP-face (AQRCP-face) that assigns a special parameter to NQRCP-face via the illumination level estimated by the weighting coefficients. Secondly, we consider that the differences of pixel images cannot model the intraclass variations of generic faces with illumination variations, and the specific identification information of the generic face is redundant for the current SSFR with generic learning. To tackle above two issues, we develop a general high-frequency based sparse representation (GHSP) model. Two practical approaches separated high-frequency based sparse representation (SHSP) and unified high-frequency based sparse representation (UHSP) are developed. Finally, the performances of the proposed methods are verified on the Extended Yale B, CMU PIE, AR, LFW and our self-built Driver face databases. The experimental results indicate that the proposed methods outperform previous approaches for SSFR under varying illumination.
Changhui Hu 0001, Xiaobo Lu, Pan Liu 0013, Xiaoyuan Jing, Dong Yue 0001
IEEE Trans. Image Process.2
2018 Spatial-Temporal Fusion Convolutional Neural Network for Simulated Driving Behavior Recognition
abstract
Abnormal driving behaviour is one of the leading cause of terrible traffic accidents endangering human life. Therefore, study on driving behaviour surveillance has become essential to traffic security and public management. In this paper, we conduct this promising research and employ a two stream CNN framework for video-based driving behaviour recognition, in which spatial stream CNN captures appearance information from still frames, whilst temporal stream CNN captures motion information with pre-computed optical flow displacement between a few adjacent video frames. We investigate different spatial-temporal fusion strategies to combine the intra frame static clues and inter frame dynamic clues for final behaviour recognition. So as to validate the effectiveness of the designed spatial-temporal deep learning based model, we create a simulated driving behaviour dataset, containing 1237 videos with 6 different driving behavior for recognition. Experiment result shows that our proposed method obtains noticeable performance improvements compared to the existing methods.
Yaocong Hu, MingQi Lu, Xiaobo Lu
ICARCV3
2018 IcooBook: When the Picture Book for Children Encounters Aesthetics of Interaction
abstract
In this work, we propose a novel PCA (Perception & Cognition & Affection) model from the prospective of aesthetics in interaction. Based on PCA, we establish a new electronic interactive picture book for children, named IcooBook. At the first level of perception, the proposed IcooBook provides interfaces of multi-sensory interaction; at the second level of cognition, IcooBook builds immersive interactive scenes; at the third level of affection, IcooBook creates high-level interaction modes based on automatic emotion recognition. The research on user study had proved the effectiveness of IcooBook in helping children being focusing on reading, getting better understanding about the context, and further encouraging children to appreciate the beauty of deep affective interaction.
Yaohua Bu, Jia Jia 0001, Xiang Li 0105, Suping Zhou, Xiaobo Lu
ACM Multimedia5
2018 An AR system for artistic creativity education
abstract
Creativity and innovation training is the core of the art education. Modern technology provides more effective tools to help students obtain artistic creativity. In this paper, we propose to employ augmented reality technology to assist artistic creativity education. We first analyze the inefficiency of traditional artistic creation training. We then introduce our AR-based smartphone app with technical detail and explain how it can improve accelerate artistic creativity training. We finally show 3 examples created by our AR app to demonstrate the effectiveness of our proposed method.
Jiajia Tan, Boyang Gao, Xiaobo Lu
VRST3
2018 A novel fuzzy linear discriminant analysis for face recognition
abstract
In practical application, the performances of face recognition are always affected by variations of expression, illumination and so on. To address this problem, an interval type-2 fuzzy linear discriminant analysis (IT2FLDA) method is proposed. In this paper, we first propose the supervised interva l type-2 fuzzy C-Means (IT2FCM) algorithm. Moreover, the supervised IT2FCM is incorporated into linear discriminant analysis (LDA). In this method, the membership degree matrix of training samples belonging to each class and means of each class are firstly calculated by the supervised IT2FCM algorithm. They are then applied to the definition of fuzzy within-class scatter matrix and fuzzy between-class scatter matrix, respectively. In doing so, means of each class that are estimated by the supervised IT2FCM can converge to a more desirable location than ones obtained by class sample average and fuzzy k-nearest neighbor (FKNN) method. Furthermore, the IT2FLDA is able to minimize the effects of uncertainties, find the optimal projective directions and make the feature subspace discriminating and robust, which inherits the benefits of the supervised IT2FCM and LDA. The experiment results show that the IT2FLDA improves the recognition rate and reduces sensitivity to variations when compared to results from the previous techniques.
Yijun Du, Xiaobo Lu, Changhui Hu 0001
Intell. Data Anal.2
2018 Correction of micro-CT image geometric artefacts based on marker
abstract
Small geometric misalignments of micro computed tomography (CT) system will cause geometric artefacts in the reconstructed image. A new correction method of geometrical artefacts based on marker and non‐linear optimisation model is proposed. In this method, the simple balls marker and the measured objects are scanned simultaneously, and the geometric parameters of the micro‐CT system are precisely estimated by solving the non‐linear optimisation model which is based on the scanning data. The geometric artefacts caused by geometric parameters are corrected and the authors can reconstruct the image correctly. In addition to estimating geometric parameters for the traditional scanning mode, the proposed method can also be used for the limited angle CT scanning and the half detector CT scanning. Simulated experiments and real experiments verify that the correction method effectively decrease the geometric artefacts of micro‐CT images.
Huanjie Tao, Xiaobo Lu
IET Image Process.2
2018 Learning spatial-temporal features for video copy detection by the combination of CNN and RNN
Yaocong Hu, Xiaobo Lu
J. Vis. Commun. Image Represent.2
2018 Real-time video fire smoke detection by utilizing spatial-temporal ConvNet features
Yaocong Hu, Xiaobo Lu
Multim. Tools Appl.2
2018 Smoky vehicle detection based on multi-feature fusion and ensemble neural networks
Huanjie Tao, Xiaobo Lu
Multim. Tools Appl.2
2018 Bidirectionally aligned sparse representation for single image super-resolution
Shengqin Jiang, Xiaobo Lu
Multim. Tools Appl.4
2018 Face recognition under varying illumination based on singular value decomposition and retina modeling
Yang Zhang 0067, Changhui Hu 0001, Xiaobo Lu
Multim. Tools Appl.3
2018 WeSamBE: A Weight-Sample-Based Method for Background Subtraction
abstract
Background subtraction techniques are often treated as fundamental and significant ways to analyze and understand video content. In this paper, we propose a weight-sample-based method for foreground detection. This method allows us to use a few samples with variable weights to achieve effective change detection. To rapidly adapt to changing scenarios, a minimum-weight update policy is first proposed to replace the most inefficient sample instead of the oldest sample or a random sample. In addition, a reward-and-penalty weighting strategy is put forward to reinforce active samples and punish others. In this way, the weights of relatively effective samples are increased and the false updating of effective samples with smaller weights is reduced. Moreover, some other strategies, such as spatial-diffusion policy and random time subsampling, are also incorporated to ensure the flexibility of the proposed method. Finally, in our experiments, an adaptive feedback technique is incorporated into our algorithm to adapt to more challenging videos, and the final results indicate that our method is superior to the state-of-the-art approaches on the challenging CDnet data set.
Shengqin Jiang, Xiaobo Lu
IEEE Trans. Circuits Syst. Video Technol.2
2018 Optimal Type-2 Fuzzy System For Arterial Traffic Signal Control
abstract
Arterial traffic is the artery of urban transport and loads huge traffic pressure. In order to alleviate its traffic pressure effectively, a coordinated arterial traffic type-2 fuzzy logic control (FLC) method is proposed. First, arterial traffic flow model and evaluation index model are set up, in which the turning vehicles and lane length are given full consideration. The traditional queue spillover phenomenon in the traffic models can be prevented here. Second, aiming at the coordination and dynamic uncertainty problem in arterial traffic, a coordinated arterial traffic type-2 fuzzy coordination control method is put forward. It consists of two-layer type-2 fuzzy controller, the basic control layer and the coordination layer. The former allocates green time according to the traffic situation of each intersection, while the latter adjusts each intersection's green time on basis of the vehicles between the intersection and the downstream intersections for the purpose of enlarging green wave band. Finally, in order to configure the high-dimensional complex parameters of the coordinated two-layer type-2 FLC effectively, the parameters of membership function and the rules of the two controllers are optimized alternately by gravitational search algorithm. The simulation results verify the effectiveness of the proposed method from several aspects.
Yunrui Bi, Xiaobo Lu, Zhe Sun 0010, Dipti Srinivasan, Zhixin Sun
IEEE Trans. Intell. Transp. Syst.2
2017 An adaptive threshold deep learning method for fire and smoke detection
abstract
This paper proposes a novel method for fire and smoke detection using video images. The ViBe method is used to extract a background from the whole video and to update the exact motion areas using frame-by-frame differences. Dynamic and static features extraction are combined to recognize the fire and smoke areas. For static features, we use deep learning to detect most of fire and smoke areas based on a Caffemodel. Another static feature is the degree of irregularity of fire and smoke. An adaptive weighted direction algorithm is further introduced to this paper. To further reduce the false alarm rate and locate the original fire position, every frame image of video is divided into 16×16 grids and the times of smoke and fire occurrences of each part is recorded. All clues are combined to reach a final detection result. Experimental results show that the proposed method in this paper can efficiently detect fire and smoke and reduce the loss and false detection rates.
Xuehui Wu, Xiaobo Lu, Henry Leung 0001
SMC2
2017 Adaptive finite-time control for overlapping cluster synchronization in coupled complex networks
Shengqin Jiang, Xiaobo Lu, Shuiming Cai
Neurocomputing2
2017 Multiscale self-similarity and sparse representation based single image super-resolution
Shengqin Jiang, Xiaobo Lu
Neurocomputing4
2017 Illumination robust single sample face recognition based on ESRC
Changhui Hu 0001, Xiaobo Lu, Mengjun Ye, Yijun Du
Multim. Tools Appl.2
2017 Singular value decomposition and local near neighbors for face recognition under varying illumination
Changhui Hu 0001, Xiaobo Lu, Mengjun Ye
Pattern Recognit.2
2016 An interval type-2 T-S fuzzy classification system based on PSO and SVM for gender recognition
Yijun Du, Xiaobo Lu
Multim. Tools Appl.2
2015 A new face recognition method based on image decomposition for single sample per person problem
Changhui Hu 0001, Mengjun Ye, Saiping Ji, Xiaobo Lu
Neurocomputing5
2015 Image super-resolution employing a spatial adaptive prior model
Xiaobo Lu, Shumin Fei
Neurocomputing2
2015 An adaptive approximation image reconstruction method for single sample problem in face recognition using FLDA
Changhui Hu 0001, Mengjun Ye, Xiaobo Lu
Multim. Tools Appl.4
2014 Type-2 fuzzy multi-intersection traffic signal control with differential evolution optimization
Yunrui Bi, Dipti Srinivasan, Xiaobo Lu, Zhe Sun 0010
Expert Syst. Appl.3
2014 A local structure adaptive super-resolution reconstruction method based on BTV regularization
Xiaobo Lu
Multim. Tools Appl.2
2013 Non-linear fourth-order telegraph-diffusion equation for noise removal
abstract
Fourth‐order partial differential equations (PDEs) for noise removal are able to provide a good trade‐off between noise removal and edge preservation, and can avoid blocky effects often caused by second‐order PDE. In this study, the authors propose a fourth‐order telegraph‐diffusion equation (TDE) for noise removal. In the authors method, a domain‐based fourth‐order PDE is proposed, which takes advantage of statistic characteristics of isolated speckles in the Laplace domain to segment the image domain into two domains: speckle domain and non‐speckle domain. Then, depending on the domain type, they adopt different conductance coefficients in the proposed fourth‐order PDE. The proposed method inherits the advantage of fourth‐order PDE which is able to avoid the blocky effects widely seen in images processed by second‐order PDE. Furthermore, a TDE processing scheme is derived from previously proposed domain‐based fourth‐order PDE by adding second time derivative, which results in better edge preservation, whereas yielding better improvement in signal‐to‐noise ratio and low noise sensitivity. Experimental results show the effectiveness of the proposed method.
Xiaobo Lu, Xianghua Tan
IET Image Process.2
2012 A Generalized DAMRF Image Modeling for Superresolution of License Plates
abstract
In this paper, we propose a novel superresolution (SR) reconstruction algorithm to handle license plate texts in real traffic videos. To make license plate numbers more legible, a generalized discontinuity-adaptive Markov random field (DAMRF) model is proposed based on the recently reported bilateral filtering, which not only preserves edges but is robust to noise as well. Moreover, instead of looking for a fixed value for the regularization parameter, a method for automatically estimating it is applied to the proposed model based on the input images. Information needed to determine the regularization parameter is updated at each iteration step, which is based on the available reconstructed image. Finally, we use the graduated nonconvexity optimization procedure to minimize the cost function. Results on synthetic and real traffic sequences are presented, which show the effectiveness of the proposed method and demonstrate its superiority to the conventional DAMRF SR method.
Xiaobo Lu
IEEE Trans. Intell. Transp. Syst.2