Tianrui Liu 0001

dblp:226/4299-1 · also Tian-Rui Liu 0001 · DBLP profile ↗
← Back
20ranked-venue papers
4as first author
18since 2021 · last 2026
0000-0001-7926-3310ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 12 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CLUENet: Cluster Attention Makes Neural Networks Have Eyes
abstract
Despite the success of convolution- and attention-based models in vision tasks, their rigid receptive fields and complex architectures limit their ability to model irregular spatial patterns and hinder interpretability, thereby posing challenges for tasks requiring high model transparency. Clustering paradigms offer promising interpretability and flexible semantic modeling, but suffer from limited accuracy, low efficiency, and gradient vanishing during training. To address these issues, we propose the CLUster attEntion Network (CLUENet), a transparent deep architecture for visual semantic understanding. Specifically, we introduce three key innovations, including (i) a Global and Soft Feature Aggregation with a Temperature-Scaled Cosine Attention for capturing long-range dependencies and a Gated Fusion Mechanism for enhanced local modeling, (ii) Hard and Shared Feature Dispatching, and (iii) an Improved Cluster Pooling Block. These enhancements significantly improve both classification performance and visual interpretability. Experiments on CIFAR-100 and Mini-ImageNet demonstrate that CLUENet outperforms existing clustering methods and mainstream visual models, offering a compelling balance of accuracy, efficiency, and transparency.
Xiangshuai Song, Tianrui Liu 0001, Chang Tang
AAAI3
2026 Self-Regressive Prototype Refinement: Stepping from Local to Global Prototypes in Few-Shot Image Classification
abstract
Metric-based methods, such as ProtoNet, excel in few-shot image classification by encouraging similarity to class prototypes. However, prototypes built from limited samples often capture only partial class information, limiting performance. Recent distribution estimation-based methods attempt to enhance performance by leveraging similar base class distributions. Yet, these approaches struggle when the distributions of base and novel classes differ significantly. Empirical analysis reveals that conceptually related categories share a local-global semantic invariance even under large distribution gaps. Based on this insight, a Self-Regressive Prototype Refinement (SRPR) is proposed to address the issue of incomplete prototype representations in few-shot learning. SRPR estimates an optimization direction for local embeddings, progressively refining them toward more global representations by exploiting local-global semantic invariance in base class data. The conservative use of coarse-grained local-global semantic structures, rather than relying on similar distributions, enhances SRPR’s applicability. With minimal computational overhead per refinement step, SRPR significantly improves classification performance and achieves state-of-the-art results across multiple few-shot benchmarks, particularly in the challenging 1-shot setting. Code is available at: https://github.com/giraffe2021/SRPR .
Qing Liao 0001, Tianrui Liu 0001, Lei Luo 0002, Xinwang Liu 0002, En Zhu
ACM Trans. Multim. Comput. Commun. Appl.3
2025 EASEMVC: Efficient Dual Selection Mechanism for Deep Multi-View Clustering
abstract
Multi-view clustering (MVC) has emerged as a leading paradigm in unsupervised learning, gaining significant attention. Central to this framework is the concept of view-pair contrastive learning, which aims to maximize mutual information between pairs of views, thereby facilitating consistent latent representations. Nevertheless, two critical challenges remain: i) Identifying the most suitable pairs of views for contrastive learning becomes difficult when more than two views are available, especially in the absence of prior knowledge; ii)Including all available views in contrastive learning can degrade performance due to the presence of low-quality views. To address these issues, we propose a novel mechanism, EASEMVC(Efficient DuAl Selection MEchanism for Deep Multi-View Clustering). EASEMVC begins by constructing a view graph using the Optimal Transport (OT) distance between bipartite graphs of individual views. A view selection module is then designed to perform efficient view-level selection based on the topological relationships within the view graph. Additionally, a cross-view sample graph is built at the sample level, where the topological relationships among samples are used to generate reliable learning weights. Leveraging the selected view pairs and sample weights, contrastive learning is employed to obtain consistent representations across views. Extensive experiments across six benchmark datasets demonstrate that EASEMVC outperforms current state-of-the-art methods.
Baili Xiao, Zhibin Dong, Ke Liang 0006, Suyuan Liu, Siwei Wang 0001, Tianrui Liu 0001, Xingchen Hu 0001, En Zhu, Xinwang Liu 0002
CVPR6
2025 Generalized Deep Multi-View Clustering Via Causal Learning With Partially Aligned Cross-View Correspondence
abstract
Multi-view clustering (MVC) aims to explore the common clustering structure across multiple views. Many existing MVC methods heavily rely on the assumption of view consistency, where alignments for corresponding samples across different views are ordered in advance. However, real-world scenarios often present a challenge as only partial data is consistently aligned across different views, restricting the overall clustering performance. In this work, we consider the model performance decreasing phenomenon caused by data order shift (i.e., from fully to partially aligned) as a generalized multi-view clustering problem. To tackle this problem, we design a causal multi-view clustering network, termed CauMVC. We adopt a causal modeling approach to understand multi-view clustering procedure. To be specific, we formulate the partially aligned data as an intervention and multi-view clustering with partially aligned data as an post-intervention inference. However, obtaining invariant features directly can be challenging. Thus, we design a Variational Auto-Encoder for causal learning by incorporating an encoder from existing information to estimate the invariant features. Moreover, a decoder is designed to perform the post-intervention inference. Lastly, we design a contrastive regularizer to capture sample correlations. To the best of our knowledge, this paper is the first work to deal generalized multi-view clustering via causal learning. Empirical experiments on both fully and partially aligned data illustrate the strong generalization and effectiveness of CauMVC.
Xihong Yang, Siwei Wang 0001, Jiaqi Jin, Fangdi Wang, Tianrui Liu 0001, Yueming Jin, Xinwang Liu 0002, En Zhu, Kunlun He
ICCV5
2025 DFDUN: Deep Infrared and Visible Image Fusion with Diffusion Prior Unfolding Network
abstract
Infrared and Visible Image Fusion (IVF) intends to aggregate information from infrared and visible modalities, generating comprehensive images. While deep-unfolding-based methods and diffusion-based methods show satisfactory performances, the former suffers from insufficiently deterministic priors, and latter is limited by the inaccurate prior generation or opaque working mechanisms. In this paper, we propose Deep infrared and visible image Fusion with Diffusion prior Unfolding Network (DFDUN), aiming for effective fusion with the generative diffusion prior in a transparent mechanism. DFDUN starts with a model-based fusion optimization formulation, which is unfolded into a Denoising Diffusion Module (DDM) for generating informative diffusion prior, and a Data Consistent Module (DCM) that transparently and effectively aggregates complementary information from diffusion prior and source modalities. Moreover, DFDUN employs a hypernetwork for adaptive dictionary parameter generation in DCM, enhancing fusion flexibility. Experimental results indicate that DFDUN outperforms existing methods, providing superior IVF performance with efficient and transparent fusion. Code is available at https://github.com/XiongMaoyi2001/DFDUN.
Maoyi Xiong, Tianrui Liu 0001, Xueqiong Li, Yuhua Tang
ICME4
2025 VSumMamba: Mamba Empowered Efficient Video Summarization with Multi-Scale Spatial-Temporal Modeling
abstract
The exponential growth of video content necessitates efficient summarization techniques that balance local redundancy reduction and global dependency modeling. In this work, we introduce VSumMamba, an innovative video summarization approach that leverages Selective State Space Models to address the quadratic complexity limitations of Transformer based approaches meanwhile surpassing CNNs' restricted long-range modeling capabilities. The proposed framework comprises three core components: 1) a Multi-Scale Aggregator, 2) a Cascaded Temporal Modeling Module with bi-directional Mamba blocks for temporal representation enhancement, and 3) a Parallel Spatial Modeling Module employing spatial Mamba blocks, operating in concert to effectively refine spatiotemporal video representations. Through three specialized multi-scale spatial-temporal modeling schemes, VSumMamba demonstrate the ability to balance computational efficiency and summarization performance. Comprehensive evaluations on benchmarks datasets demonstrate VSumMamba's superior performance, achieving 67.5% and 56.0% F1-scores on TVSum and SumMe respectively, while maintaining lower computational cost compared to existing state-of-the-art methods.
Yamiao Ding, Tianrui Liu 0001, Zhizhou Lu, Junjie Huang 0001, Xinwang Liu 0002, Meng Wang 0001
ACM Multimedia2
2025 A Lightweight Deep Exclusion Unfolding Network for Single Image Reflection Removal
abstract
Single Image Reflection Removal (SIRR) is a canonical blind source separation problem and refers to the issue of separating a reflection-contaminated image into a transmission and a reflection image. The core challenge lies in minimizing the commonalities among different sources. Existing deep learning approaches either neglect the significance of feature interactions or rely on heuristically designed architectures. In this paper, we propose a novel Deep Exclusion unfolding Network (DExNet), a lightweight, interpretable, and effective network architecture for SIRR. DExNet is principally constructed by unfolding and parameterizing a simple iterative Sparse and Auxiliary Feature Update (i-SAFU) algorithm, which is specifically designed to solve a new model-based SIRR optimization formulation incorporating a general exclusion prior. This general exclusion prior enables the unfolded SAFU module to inherently identify and penalize commonalities between the transmission and reflection features, ensuring more accurate separation. The principled design of DExNet not only enhances its interpretability but also significantly improves its performance. Comprehensive experiments on four benchmark datasets demonstrate that DExNet achieves state-of-the-art visual and quantitative results while utilizing only approximately 8% of the parameters required by leading methods.
Junjie Huang 0001, Tianrui Liu 0001, Xinwang Liu 0002, Meng Wang 0001, Pier Luigi Dragotti
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Recalling Unknowns Without Losing Precision: An Effective Solution to Large Model-Guided Open World Object Detection
abstract
Open World Object Detection (OWOD) aims to adapt object detection to an open-world environment, so as to detect unknown objects and learn knowledge incrementally. Existing OWOD methods typically leverage training sets with a relatively small number of known objects. Due to the absence of generic object knowledge, they fail to comprehensively perceive objects beyond the scope of training sets. Recent advancements in large vision models (LVMs), trained on extensive large-scale data, offer a promising opportunity to harness rich generic knowledge for the fundamental advancement of OWOD. Motivated by Segment Anything Model (SAM), a prominent LVM lauded for its exceptional ability to segment generic objects, we first demonstrate the possibility to employ SAM for OWOD and establish the very first SAM-Guided OWOD baseline solution. Subsequently, we identify and address two fundamental challenges in SAM-Guided OWOD and propose a pioneering SAM-Guided Robust Open-world Detector (SGROD) method, which can significantly improve the recall of unknown objects without losing the precision on known objects. Specifically, the two challenges in SAM-Guided OWOD include: 1) Noisy labels caused by the class-agnostic nature of SAM; 2) Precision degradation on known objects when more unknown objects are recalled. For the first problem, we propose a dynamic label assignment (DLA) method that adaptively selects confident labels from SAM during training, evidently reducing the noise impact. For the second problem, we introduce cross-layer learning (CLL) and SAM-based negative sampling (SNS), which enable SGROD to avoid precision loss by learning robust decision boundaries of objectness and classification. Experiments on public datasets show that SGROD not only improves the recall of unknown objects by a large margin (~20%), but also preserves highly-competitive precision on known objects. The program codes are available at https://github.com/harrylin-hyl/SGROD.
Yu-Lin He, Wei Chen 0009, Siqi Wang 0001, Tianrui Liu 0001, Meng Wang 0001
IEEE Trans. Image Process.4
2025 Sampling Enhanced Contrastive Multi-View Remote Sensing Data Clustering With Long-Short Range Information Mining
abstract
Multi-view clustering (MVC) for remote sensing data has demonstrated significant potential in Earth observation, given its ability to aggregate multi-source information without relying on labels. Despite achieving compelling results through the combination of deep encoders and contrastive learning, existing algorithms still face two limitations: inadequate exploration of diverse spatial relationships and inability to guide the selection of sample pairs leads to blind sampling, both of which lead to suboptimal clustering performance. To tackle these challenges, we propose a sampling enhanced contrastive multi-view clustering method for remote sensing data, namely SEC-LSRM. The proposed method incorporates long- and short-range information mining to enhance clustering performance. By aggregating shortrange information extracted through autoencoders and longrange information obtained via graph autoencoders, our method improves the sampling quality of positive and negative sample pairs. To render the extracted features more compact, a multiview correlation reduction strategy is devised to filter out irrelevant information. With the extracted comprehensive features, an adaptive sampling strategy is designed to obtain high-quality positive and negative samples. Subsequently, we select positive and negative sample pairs based on these affinity matrices with idempotence and block diagonal constraints. Moreover, we integrate the optimization of these sample pairs and contrastive learning within the same framework to achieve iterative updates of both. Experiments conducted on multiple multi-view remote sensing datasets illustrate that our proposed SEC-LSRM method achieves excellent and reliable clustering performance.
Renxiang Guan, Tianrui Liu 0001, Wenxuan Tu, Chang Tang, Wenhan Luo, Xinwang Liu 0002
IEEE Trans. Knowl. Data Eng.2
2025 Category Alignment Mechanism for Few-Shot Image Classification
abstract
While humans can excel at image classification tasks by comparing a few images, existing metric-based few-shot classification methods are still not well adapted to novel tasks. Performance declines rapidly when encountering new patterns, as feature embeddings cannot effectively encode discriminative properties. Moreover, existing matching methods inadequately utilize support set samples, focusing only on comparing query samples to category prototypes without exploiting contrastive relationships across categories for discriminative features. In this work, we propose a method where query samples select their most category-representative features for matching, making feature embeddings adaptable and category-related. We introduce a category alignment mechanism (CAM) to align query image features with different categories. CAM ensures features chosen for matching are distinct and strongly correlated to intra- and inter-contrastive relationships within categories, making extracted features highly related to their respective categories. CAM is parameter-free, requires no extra training to adapt to new tasks, and adjusts features for matching when task categories change. We also implement a cross-validation-based feature selection technique for support samples, generating more discriminative category prototypes. We implement two versions of inductive and transductive inference and conduct extensive experiments on six datasets to demonstrate the effectiveness of our algorithm. The results indicate that our method consistently yields performance improvements on benchmark tasks and surpasses the current state-of-the-art methods.
Lei Luo 0002, Tianrui Liu 0001, Qing Liao 0001, Xinwang Liu 0002, En Zhu
IEEE Trans. Neural Networks Learn. Syst.3
2024 DURRNET: Deep Unfolded Single Image Reflection Removal Network with Joint Prior
abstract
Single image reflection removal (SIRR) problem can be interpreted as a canonical blind source separation problem and is highly ill-posed. A parameter effective, fast learning and interpretable reflection removal algorithm is essential for many vision analysis applications. In this paper, we propose a novel model-inspired and learning-based SIRR method called Deep Unfolded Reflection Removal Network (DURRNet). It combines the merits of both model-based and learning-based paradigms, leading to a more interpretable and effective deep architecture. To achieve this, we first propose a model-based optimization approach and then obtain DURRNet by unfolding an iterative step into a Unfolded Separation Block (USB) based on proximal gradient descent. Key features of DURR-Net include the use of Invertible Neural Networks to impose the transform-based exclusion prior on the basis of natural image prior, as well as a coarse-to-fine architecture to fine-grain the reflection removal process. Extensive experiments on public datasets demonstrate that DURRNet achieves state-of-the-art results not only visually, quantitatively, but also effectively.
Junjie Huang 0001, Tianrui Liu 0001, Jingyuan Xia, Meng Wang 0001, Pier Luigi Dragotti
ICASSP2
2024 View Gap Matters: Cross-view Topology and Information Decoupling for Multi-view Clustering
abstract
Multi-view clustering, a pivotal technology in multimedia research, aims to leverage complementary information from diverse perspectives to enhance clustering performance. The current multi-view clustering methods normally enforce the reduction of distances between any pair of views, overlooking the heterogeneity between views, thereby sacrificing the diverse and valuable insights inherent in multi-view data. In this paper, we propose a Tree-Based View-Gap Maintaining Multi-View Clustering (TGM-MVC) method. Our approach introduces a novel conceptualization of multiple views as a graph structure. In this structure, each view corresponds to a node, with the view gap, calculated by the cosine distance between views, acting as the edge. Through graph pruning, we derive the minimum spanning tree of the views, reflecting the neighbouring relationships among them. Specifically, we applied a share-specific learning framework, and generate view trees for both view-shared and view-specific information. Concerning shared information, we only narrow the distance between adjacent views, while for specific information, we maintain the view gap between neighboring views. Theoretical analysis highlights the risks of eliminating the view gap, and comprehensive experiments validate the efficacy of our proposed TGM-MVC method.
Fangdi Wang, Jiaqi Jin, Zhibin Dong, Xihong Yang, Xinwang Liu 0002, Xinzhong Zhu, Siwei Wang 0001, Tianrui Liu 0001, En Zhu
ACM Multimedia9
2024 Alleviate Anchor-Shift: Explore Blind Spots with Cross-View Reconstruction for Incomplete Multi-View Clustering
abstract
Incomplete multi-view clustering aims to learn complete correlations among samples by leveraging complementary information across multiple views for clustering. Anchor-based methods further establish sample-level similarities for representative anchor generation, effectively addressing scalability issues in large-scale scenarios. Despite efficiency improvements, existing methods overlook the misguidance in anchors learning induced by partial missing samples, i.e., the absence of samples results in shift of learned anchors, further leading to sub-optimal clustering performance. To conquer the challenges, our solution involves a cross-view reconstruction strategy that not only alleviate the anchor shift problem through a carefully designed cross-view learning process, but also reconstructs missing samples in a way that transcends the limitations imposed by convex combinations. By employing affine combinations, our method explores areas beyond the convex hull defined by anchors, thereby illuminating blind spots in the reconstruction of missing samples. Experimental results on four benchmark datasets and three large-scale datasets validate the effectiveness of our proposed method.
Suyuan Liu, Siwei Wang 0001, Ke Liang 0006, Junpu Zhang, Zhibin Dong, Tianrui Liu 0001, En Zhu, Xinwang Liu 0002, Kunlun He
NeurIPS6
2024 DeMPAA: Deployable Multi-Mini-Patch Adversarial Attack for Remote Sensing Image Classification
abstract
Deep Neural Networks (DNNs) have demonstrated excellent performance in image classification, yet remain vulnerable to adversarial attacks. Generating deployable adversarial patches represents a promising approach to safeguard critical facilities against DNN-based classifiers used for Remote Sensing Images (RSI). While existing adversarial patch attack methods are designed for natural images, they typically generate a single and large patch which is impractically oversize for RSI applications. In this paper, we propose a Deployable Multi-Mini-Patch Adversarial Attack (DeMPAA) method for RSI classification task, which deploys multiple small adversarial patches on key locations considering both the feasibility and the effectiveness. The proposed DeMPAA method formulates the problem as a constrained optimization problem that jointly optimizes patch locations and adversarial patches. The proposed DeMPAA method takes a searching and optimization strategy to tackle it. The DeMPAA framework consists of a Feasible and Effective Map Generation (FEMG) module and a Patch Generation (PG) module. The FEMG module generates a location map to guide the adversarial patch location sampling by excluding the infeasible locations and considering the location effectiveness. In the PG module, a Probability guided Random Sampling based patch location selection (PRSamp) method is used to search better locations, then we optimize the adversarial patches using gradient descent with respect to an adversarial classification loss and an imperceptibility loss. Extensive experimental results conducted on Aerial Image Dataset show that the proposed DeMPAA method achieves 94.80% attacking success rate against ResNet50 using 16 small patches, which significantly outperforms other adversarial patch methods.
Junjie Huang 0001, Tianrui Liu 0001, Wenhan Luo, Meng Wang 0001
IEEE Trans. Geosci. Remote. Sens.3
2022 Video Summarization Through Reinforcement Learning With a 3D Spatio-Temporal U-Net
abstract
Intelligent video summarization algorithms allow to quickly convey the most relevant information in videos through the identification of the most essential and explanatory content while removing redundant video frames. In this paper, we introduce the 3DST-UNet-RL framework for video summarization. A 3D spatio-temporal U-Net is used to efficiently encode spatio-temporal information of the input videos for downstream reinforcement learning (RL). An RL agent learns from spatio-temporal latent scores and predicts actions for keeping or rejecting a video frame in a video summary. We investigate if real/inflated 3D spatio-temporal CNN features are better suited to learn representations from videos than commonly used 2D image features. Our framework can operate in both, a fully unsupervised mode and a supervised training mode. We analyse the impact of prescribed summary lengths and show experimental evidence for the effectiveness of 3DST-UNet-RL on two commonly used general video summarization benchmarks. We also applied our method on a medical video summarization task. The proposed video summarization method has the potential to save storage costs of ultrasound screening videos as well as to increase efficiency when browsing patient video data during retrospective analysis or audit without loosing essential information.
Tianrui Liu 0001, Qingjie Meng, Junjie Huang 0001, Athanasios Vlontzos, Daniel Rueckert, Bernhard Kainz
IEEE Trans. Image Process.1
2022 MulViMotion: Shape-Aware 3D Myocardial Motion Tracking From Multi-View Cardiac MRI
abstract
Recovering the 3D motion of the heart from cine cardiac magnetic resonance (CMR) imaging enables the assessment of regional myocardial function and is important for understanding and analyzing cardiovascular disease. However, 3D cardiac motion estimation is challenging because the acquired cine CMR images are usually 2D slices which limit the accurate estimation of through-plane motion. To address this problem, we propose a novel multi-view motion estimation network (MulViMotion), which integrates 2D cine CMR images acquired in short-axis and long-axis planes to learn a consistent 3D motion field of the heart. In the proposed method, a hybrid 2D/3D network is built to generate dense 3D motion fields by learning fused representations from multi-view images. To ensure that the motion estimation is consistent in 3D, a shape regularization module is introduced during training, where shape information from multi-view images is exploited to provide weak supervision to 3D motion estimation. We extensively evaluate the proposed method on 2D cine CMR images from 580 subjects of the UK Biobank study for 3D motion tracking of the left ventricular myocardium. Experimental results show that the proposed method quantitatively and qualitatively outperforms competing methods.
Qingjie Meng, Chen Qin, Wenjia Bai, Tianrui Liu 0001, Antonio M. Simoes Monteiro de Marvao, Declan P. O'Regan, Daniel Rueckert
IEEE Trans. Medical Imaging4
2021 Detecting Hypo-plastic Left Heart Syndrome in Fetal Ultrasound via Disease-Specific Atlas Maps
Samuel Budd, Matthew Sinclair, Thomas G. Day, Athanasios Vlontzos, Jeremy Tan, Tianrui Liu 0001, Jacqueline Matthew, Emily Skelton, John M. Simpson, Reza Razavi, Ben Glocker, Daniel Rueckert, Emma C. Robinson, Bernhard Kainz
MICCAI (7)6
2021 Coupled Network for Robust Pedestrian Detection With Gated Multi-Layer Feature Extraction and Deformable Occlusion Handling
abstract
Pedestrian detection methods have been significantly improved with the development of deep convolutional neural networks. Nevertheless, detecting ismall-scaled pedestrians and occluded pedestrians remains a challenging problem. In this paper, we propose a pedestrian detection method with a couple-network to simultaneously address these two issues. One of the sub-networks, the gated multi-layer feature extraction sub-network, aims to adaptively generate discriminative features for pedestrian candidates in order to robustly detect pedestrians with large variations on scale. The second sub-network targets on handling the occlusion problem of pedestrian detection by using deformable regional region of interest (RoI)-pooling. We investigate two different gate units for the gated sub-network, namely, the channel-wise gate unit and the spatio-wise gate unit, which can enhance the representation ability of the regional convolutional features among the channel dimensions or across the spatial domain, repetitively. Ablation studies have validated the effectiveness of both the proposed gated multi-layer feature extraction sub-network and the deformable occlusion handling sub-network. With the coupled framework, our proposed pedestrian detector achieves promising results on both two pedestrian datasets, especially on detecting small or occluded pedestrians. On the CityPersons dataset, the proposed detector achieves the lowest missing rates (i.e. 40.78% and 34.60%) on detecting small and occluded pedestrians, surpassing the second best comparison method by 6.0% and 5.87%, respectively.
Tianrui Liu 0001, Wenhan Luo, Lin Ma 0002, Junjie Huang 0001, Tania Stathaki, Tianhong Dai
IEEE Trans. Image Process.1
2020 Gated Multi-Layer Convolutional Feature Extraction Network for Robust Pedestrian Detection
abstract
Pedestrian detection methods have been significantly improved with the development of deep convolutional neural networks. Nevertheless, it remains a challenging problem how to robustly detect pedestrians of varied sizes and with occlusions. In this paper, we propose a gated multi-layer convolutional feature extraction method which can adaptively generate discriminative features for candidate pedestrian regions. The proposed gated feature extraction framework consists of squeeze units, gate units and concatenation layers which perform feature dimension squeezing, feature manipulation and features combination from multiple CNN layers, respectively. We proposed two different gate models that can manipulate the regional feature maps in a channel-wise selection manner and a spatial-wise selection manner, respectively. Experiments on the challenging CityPersons dataset demonstrate the effectiveness of the proposed method, especially on detecting small-size and occluded pedestrians.
Tianrui Liu 0001, Junjie Huang 0001, Tianhong Dai, Guangyu Ren, Tania Stathaki
ICASSP1
2018 SAM-RCNN: Scale-Aware Multi-Resolution Multi-Channel Pedestrian Detection
Tianrui Liu 0001, Mohamed Elmikaty, Tania Stathaki
BMVC1