VLDB 2026 Research / reviewers in the wild / expert
Meiqing Wu
dblp:119/4173
· DBLP profile ↗
30ranked-venue papers
3as first author
17since 2021 · last 2026
0000-0002-4978-2969ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 since 2021Artificial intelligence and machine learning · 8 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 4 since 2021Security and privacy · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing Stereo Matching Domain Generalization With Adversarial Domain AlignmentabstractRecently, state-of-the-art stereo-matching networks trained on large-scale synthetic data have shown remarkable performance. However, their capacity to extrapolate effectively to unseen real-world data,i.e.different domains, remains a challenge. The major difficulty resides in the unforeseeable domain gap when generalizing from synthetic data to real-world data. In this paper, we introduceADASM, an approach using adversarial domain alignment, designed to enhance the robustness and generalization of stereo-matching networks. It mainly consists of two modules: an end-to-end robustness optimizer and a domain-invariant feature learner. First, we adapt adversarial training into the stereo-matching task to reduce models' sensitivity to the perturbation in real-world samples. By introducing worst cases into the training space, we take unseen data into account and achieve robust disparity estimation for the end-to-end model. Then, via simulating the real-world noise with gradient-based perturbation, we construct a fictitious domain, which is taken as a referential distribution of the real-world noisy data, for further domain alignment. Specifically, we propose to utilize Maximum Mean Discrepancy to realize domain regularization between the original domain and the fictitious one. Finally, we fuse all aforementioned objectives and propose a unified, simple but effective loss function that can be adapted toallstereo-matching networks. The extensive experiments show that our method achieves a superior disparity estimation performance on various real-world benchmarks, including KITTI, Middlebury, and DrivingStereo. More importantly,ADASMobtains competitive or even better performance than the fine-tuning strategy, revealing its fine-tuning-free character. Shiqian Zhao, Meiqing Wu, Kangjie Chen, Yi Xie 0011, Tianlin Li, Siew-Kei Lam, Guowen Xu, Anran Li 0001 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2025 | Taylor Series-Inspired Local Structure Fitting Network for Few-shot Point Cloud Semantic SegmentationabstractFew-shot point cloud semantic segmentation aims to accurately segment "unseen" new categories in point cloud scenes using limited labeled data. However, pretraining-based methods not only introduce excessive time overhead but also overlook the local structure representation among irregular point clouds. To address these issues, we propose a pretraining-free local structure fitting network for few-shot point cloud semantic segmentation, named TaylorSeg. Specifically, inspired by Taylor series, we treat the local structure representation of irregular point clouds as a polynomial fitting problem and propose a novel local structure fitting convolution, called TaylorConv. This convolution learns the low-order basic information and high-order refined information of point clouds from explicit encoding of local geometric structures. Then, using TaylorConv as the basic component, we construct two variant of TaylorSeg: a non-parametric TaylorSeg-NN and a parametric TaylorSeg-PN. The former can achieve performance comparable to existing parametric models without pretraining. For the latter, we equip it with an Adaptive Push-Pull (APP) module to mitigate the feature distribution differences between the query set and the support set. Extensive experiments validate the effectiveness of the proposed method. Notably, under the 2-way 1-shot setting, TaylorSeg-PN achieves improvements of +2.28% and +4.37% mIoU on the S3DIS and ScanNet datasets respectively, compared to the previous state-of-the-art methods. Changshuo Wang 0001, Shuting He, Meiqing Wu, Siew-Kei Lam, Prayag Tiwari |
AAAI | 4 |
| 2025 | VFM-Depth: Leveraging Vision Foundation Model for Self-Supervised Monocular Depth EstimationabstractSelf-supervised monocular depth estimation has exploited semantics to reduce depth ambiguities in texture-less regions and object boundaries. However, existing methods struggle to obtain universal semantics across scenes for effective depth estimation. This paper proposes VFM-Depth, a novel self-supervised teacher-student framework, that effectively leverages the vision foundation model as semantic regularization to significantly improve the accuracy of monocular depth estimation. Firstly, we propose a novel Geometric-Semantic Aggregation Encoding, integrating universal semantic constraints from the foundation model to reduce ambiguities in the teacher model. Specifically, semantic features from the foundation model and geometric features from the depth model are first encoded and then fused through cross-modal aggregation. Secondly, we introduce a novel Multi-Alignment for Depth Distillation to distill semantic constraints from the teacher, further leveraging knowledge from the foundation model. We obtain a lightweight yet effective student model through an innovative approach that combines distance category alignment with complementary feature and depth imitation. Extensive experiments on KITTI, Cityscapes, and Make3D datasets demonstrate that VFM-Depth (both teacher and student) outperforms state-of-the-art self-supervised methods by a large margin. Shangshu Yu, Meiqing Wu, Siew-Kei Lam |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Looking Clearer With Text: A Hierarchical Context Blending Network for Occluded Person Re-IdentificationabstractExisting occluded person re-identification (re-ID) methods mainly learn limited visual information for occluded pedestrians from images. However, textual information, which can describe various human appearance attributes, is rarely fully utilized in the task. To address this issue, we propose a Text-guided Hierarchical Context Blending Network ( THCB-Net) for occluded person re-ID. Specifically, at the data level, informative multi-modal inputs are first generated to make full use of the auxiliary role of textual information and make image data have a strong inductive bias for occluded environments. At the feature expression level, we design a novel Hierarchical Context Blending (HCB) module that can adaptively integrate shallow appearance features obtained by CNNs and multi-scale semantic features from visual transformer encoder. At the model optimization level, a Multi-modal Feature Interaction (MFI) module is proposed to learn the multi-modal information of pedestrians from texts and images, then guide the visual transformer encoder and HCB module to further learn discriminative identity information for occluded pedestrians through Image-Multimodal Contrastive (IMC) learning. Extensive experiments on standard occluded person re-ID benchmarks demonstrate that the proposed THCB-Net outperforms state-of-the-art methods. Changshuo Wang 0001, Xingyu Gao 0001, Meiqing Wu, Siew-Kei Lam, Shuting He, Prayag Tiwari |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | EDS-Depth: Enhancing Self-Supervised Monocular Depth Estimation in Dynamic ScenesabstractSelf-supervised monocular depth estimation usually assumes that training samples contain only static objects, which leads to poor performance in real-world environments. The presence of dynamic objects incurs camera motion estimation errors, motion blur, and occlusions, which induce significant challenges for network training. To address these issues, we introduce EDS-Depth, a self-supervised learning framework, that improves monocular depth estimation in dynamic scenes. Firstly, we propose a novel TCE (Temporal Continuity Enhancement) strategy to reduce camera motion estimation errors and motion blur caused by dynamic objects. Video frames are interpolated to generate more continuous frames in order to smooth dynamic changes and enrich motion details. Secondly, we design a novel IPDM (Iterative Pseudo Depth Masking) module to address inaccurate object motion and occlusions in dynamic scenes. The module integrates multiple optical flows from different frames for triangulation, generating optimal depth as pseudo-supervision labels in dynamic regions. Extensive experiments on Cityscapes and KITTI datasets demonstrate the effectiveness of EDS-Depth, which surpasses state-of-the-art self-supervised monocular depth estimation methods, particularly in dynamic scenes. Shangshu Yu, Meiqing Wu, Siew-Kei Lam, Changshuo Wang 0001, Ruiping Wang 0005 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | GPSFormer: A Global Perception and Local Structure Fitting-Based Transformer for Point Cloud Understanding
Changshuo Wang 0001, Meiqing Wu, Siew-Kei Lam, Xin Ning 0001, Shangshu Yu, Ruiping Wang 0005, Weijun Li 0002, Thambipillai Srikanthan |
ECCV (8) | 2 |
| 2024 | CurricularVPR: Curricular Contrastive Loss for Visual Place RecognitionabstractVisual Place Recognition (VPR) techniques commonly utilize Contrastive Losses (CL) to train models that generate compact and discriminative global descriptors for images. These models often result in poor performance due to one of the following reasons during training: 1) loss functions that focus primarily on easier samples, 2) reliance on time-consuming hard sample mining methods to identify informative supervisory samples, which hinders effective learning from large-scale datasets. To enhance both learning efficiency and effectiveness, we propose a Curricular Contrastive Loss (CCL) and use graded similarity labels as a measure of sample difficulty. Inspired by human learning that begin with easier concepts and progressively tackle more challenging ones, our CCL dynamically emphasizes easier samples during the initial training stages to achieve rapid convergence. The learning gradually focuses on harder samples in later training stages to bolster robustness of the models under challenging conditions. Our proposed method has been extensively evaluated on popular datasets, and the results demonstrate its superior performance compared to the CL and Generalized CL functions. Dongshuo Zhang, Nanhua Chen, Meiqing Wu, Siew-Kei Lam |
IROS | 3 |
| 2024 | Hierarchical Object-Aware Dual-Level Contrastive Learning for Domain Generalized Stereo MatchingabstractStereo matching algorithms that leverage end-to-end convolutional neural networks have recently demonstrated notable advancements in performance. However, a common issue is their susceptibility to domain shifts, hindering their ability in generalizing to diverse, unseen realistic domains. We argue that existing stereo matching networks overlook the importance of extracting semantically and structurally meaningful features. To address this gap, we propose an effective hierarchical object-aware dual-level contrastive learning (HODC) framework for domain generalized stereo matching. Our framework guides the model in extracting features that support semantically and structurally driven matching by segmenting objects at different scales and enhances correspondence between intra- and inter-scale regions from the left feature map to the right using dual-level contrastive loss. HODC can be integrated with existing stereo matching models in the training stage, requiring no modifications to the architecture. Remarkably, using only synthetic datasets for training, HODC achieves state-of-the-art generalization performance with various existing stereo matching network architectures, across multiple realistic datasets. Yikun Miao, Meiqing Wu, Siew-Kei Lam, Thambipillai Srikanthan |
NeurIPS | 2 |
| 2023 | Training-Free Attentive-Patch Selection for Visual Place RecognitionabstractVisual Place Recognition (VPR) utilizing patch descriptors from Convolutional Neural Networks (CNNs) has shown impressive performance in recent years. Existing works either perform exhaustive matching of all patch descriptors, or employ complex networks to select good candidate patches for further geometric verification. In this work, we develop a novel two-step training-free patch selection method that is fast, while being robust to large occlusions and extreme viewpoint variations. In the first step, a self-attention mechanism is used to select sparse and evenly distributed discriminative patches in the query image. Next, a novel spatial-matching method is used to rapidly select corresponding patches with high similar appearances between the query and each reference image. The proposed method is inspired by how humans perform place recognition by first identifying prominent regions in the query image, and then relying on back-and-forth visual inspection of the query and reference image to attentively identify similar regions while ignoring dissimilar ones. Extensive experiment results show that our proposed method outperforms state-of-the-art (SOTA) methods in both place recognition precision and runtime, on various challenging conditions. Dongshuo Zhang, Meiqing Wu, Siew-Kei Lam |
IROS | 2 |
| 2022 | Decoupled self-supervised label augmentation for fully-supervised image classification
Wanshun Gao, Meiqing Wu, Siew-Kei Lam, Qihui Xia, Jianhua Zou |
Knowl. Based Syst. | 2 |
| 2022 | Fast Semantic-Aware Motion State Detection for Visual SLAM in Dynamic EnvironmentabstractExisting visual SLAM (vSLAM) systems fail to perform well in dynamic environments as they cannot effectively ignore moving objects during pose estimation and mapping. We propose a lightweight approach to improve the robustness of existing feature based RGB-D and stereo vSLAM by accurately removing dynamic outliers in the scene that contribute to failures in pose estimation and mapping. First, a novel motion state detection algorithm using the depth and feature flow information is presented to identify regions in the scene with high moving probability. This information is then fused with semantic cues via a probability framework to enable accurate and robust moving object extraction to retain the useful features for pose estimation and mapping. To reduce the computational complexity of extracting semantic information in every frame, we propose to extract semantics only on keyframes with significant changes in image content. Semantic propagation is used to compensate for the changes in the intermediate frames (i.e., non-keyframes). This is achieved by computing the dense transformation map using the available feature flow vectors. The proposed techniques can be integrated into existing vSLAM systems to increase their robustness in dynamic environments without incurring much computation cost. Our work highlights the importance of distinguishing between motion states of potential moving objects for vSLAM in highly dynamic environments. We provide extensive experimental results on four well-known RGB-D and stereo datasets to show that the proposed technique outperforms existing vSLAM methods in indoor and outdoor environments under various dynamic scenarios including crowded scenes. We also perform our experiments on a low-cost embedded platform, i.e., Jetson TX1, to demonstrate the computational efficiency of our method. Gaurav Singh 0009, Meiqing Wu, Do Van Minh, Siew-Kei Lam |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | A Unified Multi-Task Learning Architecture for Fast and Accurate Pedestrian DetectionabstractWe present a unified multi-task learning architecture for fast and accurate pedestrian detection. Different from existing methods which often focus on either a new loss function or architecture, we propose an improved multi-task convolutional neural network learning architecture to effectively and efficiently interfuse the task of pedestrian detection and semantic segmentation. To achieve this, we integrate a lightweight semantic segmentation branch to Faster R-CNN detection framework that enables end-to-end hard parameter sharing in order to boost the detection performance and maintain computational efficiency as follows. Firstly, a Semantic Segmentation to Feature Module (SS2FM) refines the convolutional features in RPN stage by integrating the features generated from the semantic segmentation branch. Secondly, a Semantic Segmentation to Confidence Module (SS2CM) refines the classification confidence in RPN stage by fusing it with the semantic segmentation confidence. We also introduce an effective anchor matching point transform to alleviate the problem of feature misalignment for heavily occluded pedestrians. The proposed unified multi-task learning architecture lends itself well to more robust pedestrian detection in diverse scenarios with negligible computation overhead. In addition, the proposed architecture can achieve high detection performance with low resolution input images, which significantly reduces the computational complexity. Experiment results on CityPersons and Caltech datasets show that our method is the fastest among all state-of-the-art pedestrian detection methods while exhibiting competitive detection performance. Chengju Zhou, Meiqing Wu, Siew-Kei Lam |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2022 | Enhanced Multi-Task Learning Architecture for Detecting Pedestrian at Far DistanceabstractExisting pedestrian detection methods suffer from performance degradation in the presence of small-scale pedestrians who are positioned at far distance from the camera. We present a pedestrian detection framework that is not only robust to small- and large-scale pedestrians, but is also significantly faster than state-of-the-art methods. The proposed framework incorporates semantic segmentation to confidence modules for RPN (Region Proposal Network) head and R-FCN (Region-based Fully Convolutional Networks) head, and a cascaded R-FCN head. The semantic segmentation confidence is extracted and utilized as auxiliary classification prior knowledge for RPN proposal selection and R-FCN head prediction. Finally, the cascaded R-FCN head progressively refine the pedestrian prediction accuracy with negligible computation overhead. The proposed framework is also capable of maintaining high detection performance on down-sampled input images, which leads to further reduction in overall computational complexity. Experiment results on CityPersons and MOT17Det datasets show that the proposed framework achieves competitive detection performance with about$3\times $speedup over state-of-the-art methods. Chengju Zhou, Meiqing Wu, Siew-Kei Lam |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2021 | Edge Accelerator for Lifelong Deep Learning using Streaming Linear Discriminant AnalysisabstractLifelong deep learning models are expected to continuously adapt and acquire new knowledge in dynamic environments. This capability is essential for numerous vision tasks in robotics and drones, and the models must be deployed on the edge to achieve real-time performance. We propose a FPGA accelerator of a streaming classifier for lifelong deep learning, which is based on streaming linear discriminant analysis (SLDA). When combined with a frozen Convolutional Neural Network (CNN) model, the proposed system is capable of class incremental lifelong learning for object classification. Duvindu Piyasena, Siew-Kei Lam, Meiqing Wu |
FCCM | 3 |
| 2021 | Accelerating Continual Learning on Edge FPGAabstractReal-time edge AI systems operating in dynamic environments must learn quickly from streaming input samples without needing to undergo offline model training. We propose an FPGA accelerator for continual learning based on streaming linear discriminant analysis (SLDA), which is capable of class-incremental object classification. The proposed SLDA accelerator employs application-specific parallelism, efficient data reuse, resource sharing, and approximate computing to achieve high performance and power efficiency. Additionally, we introduce a new variant of SLDA and discuss the accuracy-efficiency trade-offs. The proposed SLDA accelerator is combined with a Convolutional Neural Network (CNN). which is implemented on Xilinx DPU to achieve full continual learning capability at nearly the same latency as inference. Experiments based on popular datasets for continual learning, CoRE50 and CUB200, demonstrate that the proposed SLDA accelerator outperforms the embedded CPU and GPU counterparts, in terms of speed and energy efficiency. Duvindu Piyasena, Siew-Kei Lam, Meiqing Wu |
FPL | 3 |
| 2021 | CAP: Context-Aware Pruning for Semantic SegmentationabstractNetwork pruning for deep convolutional neural networks (CNNs) has recently achieved notable research progress on image-level classification. However, most existing pruning methods are not catered to or evaluated on semantic segmentation networks. In this paper, we advocate the importance of contextual information during channel pruning for semantic segmentation networks by presenting a novel Context-aware Pruning framework. Concretely, we formulate the embedded contextual information by leveraging the layer-wise channels interdependency via the Context-aware Guiding Module (CAGM) and introduce the Context-aware Guided Sparsification (CAGS) to adaptively identify the informative channels on the cumbersome model by inducing channel-wise sparsity on the scaling factors in batch normalization (BN) layers. The resulting pruned models require significantly lesser operations for inference while maintaining comparable performance to (at times outperforming) the original models. We evaluated our framework on widely-used benchmarks and showed its effectiveness on both large and lightweight models. On Cityscapes dataset, our framework reduces the number of parameters by 32%, 47%, 54%, and 63%, on PSPNet101, PSPNet50, ICNet, and SegNet, respectively, while preserving the performance. Wei He 0025, Meiqing Wu, Mingfu Liang, Siew-Kei Lam |
WACV | 2 |
| 2021 | ACSL: Adaptive correlation-driven sparsity learning for deep neural network compression
Wei He 0025, Meiqing Wu, Siew-Kei Lam |
Neural Networks | 2 |
| 2020 | Dynamically Growing Neural Network Architecture for Lifelong Deep Learning on the EdgeabstractConventional deep learning models are trained once and deployed. However, models deployed in agents operating in dynamic environments need to constantly acquire new knowledge, while preventing catastrophic forgetting of previous knowledge. This ability is commonly referred to as lifelong learning. In this paper, we address the performance and resource challenges for realizing lifelong learning on edge devices. We propose a FPGA based architecture for a Self-Organization Neural Network (SONN), that in combination with a Convolutional Neural Network (CNN) can perform class-incremental lifelong learning for object classification. The proposed SONN architecture is capable of performing unsupervised learning on input features from the CNN by dynamically growing neurons and connections. In order to meet the tight constraints of edge computing, we introduce efficient scheduling methods to maximize resource reuse and parallelism, as well as approximate computing strategies. Experiments based on the Core50 dataset for continuous object recognition from video sequences demonstrated that the proposed FPGA architecture significantly outperforms CPU and GPU based implementations. Duvindu Piyasena, Miyuru Thathsara, Sathursan Kanagarajah, Siew-Kei Lam, Meiqing Wu |
FPL | 5 |
| 2020 | Fusing Semantics and Motion State Detection for Robust Visual SLAMabstractAchieving robust pose tracking and mapping in highly dynamic environments is a major challenge faced by existing visual SLAM (vSLAM) systems. In this paper, we increase the robustness of existing vSLAM by accurately removing moving objects from the scene so that they will not contribute to pose estimation and mapping. Specifically, semantic information is fused with motion states of the scene via a probability framework to enable accurate and robust moving object extraction in order to retain the useful features for pose estimation and mapping. Our work highlights the importance of distinguishing between motion states of potential moving objects for vSLAM in highly dynamic environments. The proposed method can be integrated into existing vSLAM systems to increase their robustness in dynamic environments without incurring much computation cost. We provide extensive experimental results on three well-known datasets to show that the proposed technique outperforms existing vSLAM methods in indoor and outdoor environments, under various scenarios such as crowded scenes. Gaurav Singh 0009, Meiqing Wu, Siew-Kei Lam |
WACV | 2 |
| 2020 | Group Cost-Sensitive BoostLR With Vector Form Decorrelated Filters for Pedestrian DetectionabstractPedestrian detection has achieved notable progress in the field of computer vision over the past decade. However, existing top-performing approaches suffer from high computational complexity which prohibits their realization on embedded platforms with low computational capabilities. In this paper, we propose a robust and fast pedestrian detection framework which is based on the Filtered Channel Feature (FCF) approach. The proposed framework exploits vector-form decorrelated filters to extract more discriminative channel features while benefiting from low computational complexity. A novel group cost-sensitive BoostLR (Boosting with Loss Regularization) algorithm is proposed to train the classifier. The proposed training strategy provides more emphasis to the harder samples by exploring the variations of negatives selected from different rounds in hard negative mining processing, and hence is able to boost the overall detection performance. In addition, the proposed method also benefits from the BoostLR framework to achieve better generalization. Experiments on the well-known Caltech, INRIA and CityPersons pedestrian detection datasets show that our proposed approach achieves the best detection performance among all of the state-of-the-art non-deep learning methods and can run one order of magnitude faster than classical FCF methods (e.g. Checkerboards). Chengju Zhou, Meiqing Wu, Siew-Kei Lam |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2019 | Reducing Dynamic Power in Streaming CNN Hardware Accelerators by Exploiting Computational RedundanciesabstractConvolutional neural networks (CNNs) have achieved tremendous successes in various application domains such as computer vision. However, current implementations are characterized by large memory requirements and accesses, which pose an impediment towards their deployment on low cost embedded devices with fast runtime requirements. Recently, FPGA based streaming CNN hardware accelerators have been reported for alleviating these memory bottlenecks. However, these implementations suffer from large number of convolution operations which incur high power consumption. In this paper, we investigate methods to exploit the redundancies in the activation layers in order to reduce the dynamic power. We propose a computationally efficient approximation method to reduce the overall convolution operations with marginal accuracy loss. Experimental results of our FPGA implementation based on image classification datasets show that the proposed method leads to considerable power savings. Duvindu Piyasena, Rukshan Wickramasinghe, Debdeep Paul, Siew-Kei Lam, Meiqing Wu |
FPL | 5 |
| 2019 | Lowering Dynamic Power of a Stream-based CNN Hardware AcceleratorabstractCustom hardware accelerators of Convolutional Neural Networks (CNN) provide a promising solution to meet real-time constraints for a wide range of applications on low-cost embedded devices. In this work, we aim to lower the dynamic power of a stream-based CNN hardware accelerator by reducing the computational redundancies in the CNN layers. In particular, we investigate the redundancies due to the downsampling effect of max pooling layers which are prevalent in state-of-the-art CNNs, and propose an approximation method to reduce the overall computations. The experimental results show that the proposed method leads to lower dynamic power without sacrificing accuracy. Duvindu Piyasena, Rukshan Wickramasinghe, Debdeep Paul, Siew-Kei Lam, Meiqing Wu |
MMSP | 5 |
| 2018 | Stream-Based ORB Feature Extractor with Dynamic Power OptimizationabstractThe Oriented Fast and Rotated BRIEF (ORB) feature extractor, which consists of key-point detection and descriptor computation, is a key module in many computer vision systems. Existing hardware implementations of ORB feature extractor only focus on increasing performance with power optimization as a post consideration. In this paper, we present a stream-based ORB feature extractor that incorporates mechanisms to lower the dynamic power consumption. These mechanisms exploit the fact that the number of detected keypoints is typically small. The proposed solution significantly lowers the switching activity of the key-point detection and descriptor computation stages by early pruning of non-likely key-points and gating the descriptor computation stages. Further power reduction and resource minimization are achieved by employing a threshold-guided bit-width optimization strategy to truncate the redundant bits in the key-point detection stage. Finally, we propose an approximation method to achieve rotation invariance of the descriptors. FPGA implementation targeting the Altera Aria V device shows that the proposed strategies lead to over 25% reduction in dynamic power and lower resource utilization, with only marginal loss in accuracy. Thinh Hung Pham, Siew-Kei Lam, Meiqing Wu, Bhavan A. Jasani |
FPT | 4 |
| 2018 | Threshold-Guided Design and Optimization for Harris Corner Detector ArchitectureabstractHigh-speed corner detection is an essential step in many real-time computer vision applications, e.g., object recognition, motion analysis, and stereo matching. Hardware implementation of corner detection algorithms, such as the Harris corner detector (HCD) has become a viable solution for meeting real-time requirements of the applications. A major challenge lies in the design of power, energy and area efficient architectures that can be deployed in tightly constrained embedded systems while still meeting real-time requirements. In this paper, we proposed a bit-width optimization strategy for designing hardware-efficient HCD that exploits the thresholding step in the algorithm to determine interest points from the corner responses. The proposed strategy relies on the threshold as a guide to truncate the bit-widths of the operators at various stages of the HCD pipeline with only marginal loss of accuracy. Synthesis results based on 65-nm CMOS technology show that the proposed strategy leads to power-delay reduction of 35.2%, and area reduction of 35.4% over the baseline implementation. In addition, through careful retiming, the proposed implementation achieves over 2.2 times increase in maximum frequency while achieving an area reduction of 35.1% and power-delay reduction of 35.7% over the baseline implementation. Finally, we performed repeatability tests to show that the optimized HCD architecture achieves comparable accuracy with the baseline implementation (average decrease of repeatability is less than 0.6%). Bhavan A. Jasani, Siew-Kei Lam, Pramod Kumar Meher, Meiqing Wu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2017 | Group Cost-sensitive Boosting with Multi-scale Decorrelated Filters for Pedestrian Detection
Chengju Zhou, Meiqing Wu, Siew-Kei Lam |
BMVC | 2 |
| 2017 | Lowering dynamic power in stream-based harris corner detection architectureabstractStream-based image processing architectures are extremely attractive as they can achieve high throughput and do not require external memories for storing input video frames. However, a major challenge in designing stream-based architectures lies in lowering the dynamic power consumption since all the processing elements are typically in continuous operation to keep up with the rate of incoming pixel streams. In this work, we show that the dynamic power of the stream-based Harris corner detector (HCD) can be reduced by inhibiting redundant signal activity in the complex calculations of the corner scores. Specifically, we perform simple approximations to detect non-likely corners at the early stages of the pipeline for putting the subsequent pipeline stages into a dormant state. Synthesis results on the Altera Cyclone V FPGA show that the proposed strategy leads to an average dynamic power reduction of over 10% compared to the conventional implementation with similar performance and negligible increase in resources. In addition, we performed repeatability tests to show that the proposed stream-based HCD architecture achieves comparable accuracy with the conventional implementation. Siew-Kei Lam, Rakesh Kumar Bijarniya, Meiqing Wu |
FPT | 3 |
| 2017 | Fast and Accurate Pedestrian Detection using Dual-Stage Group Cost-Sensitive RealBoost with Vector Form FiltersabstractDespite significant research efforts in pedestrian detection over the past decade, there is still a ten-fold performance gap between the state-of-the-art methods and human perception. Deep learning methods can provide good performance but suffers from high computational complexity which prohibits their deployment on affordable systems with limited computational resources. In this paper, we propose a pedestrian detection framework that provides a major fillip to the robustness and run-time efficiency of the recent top performing non-deep learning Filtered Channel Feature (FCF) approach. The proposed framework overcomes the computational bottleneck of existing FCF methods by exploiting vector form filters to efficiently extract more discriminative channel features for pedestrian detection. A novel dual-stage group cost-sensitive RealBoost algorithm is used to explore different costs among different types of misclassification in the boosting process in order to improve detection performance. In addition, we propose two strategies, selective classification and selective scale processing, to further accelerate the detection process at the channel feature level and image pyramid level respectively. Experiments on the Caltech and INRIA datasets show that the proposed method achieves the highest detection performance among all the state-of-the-art non-CNN methods and is about 148X faster than the existing best performing FCF method on the Caltech dataset. Chengju Zhou, Meiqing Wu, Siew-Kei Lam |
ACM Multimedia | 2 |
| 2017 | A Framework for Fast and Robust Visual OdometryabstractKnowledge of the ego-vehicle's motion state is essential for assessing the collision risk in advanced driver assistance systems or autonomous driving. Vision-based methods for estimating the ego-motion of vehicle, i.e., visual odometry, face a number of challenges in uncontrolled realistic urban environments. Existing solutions fail to achieve a good tradeoff between high accuracy and low computational complexity. In this paper, a framework for ego-motion estimation that integrates runtime-efficient strategies with robust techniques at various core stages in visual odometry is proposed. First, a pruning method is employed to reduce the computational complexity of Kanade-Lucas-Tomasi (KLT) feature detection without compromising on the quality of the features. Next, three strategies, i.e., smooth motion constraint, adaptive integration window technique, and automatic tracking failure detection scheme, are introduced into the conventional KLT tracker to facilitate generation of feature correspondences in a robust and runtime efficient way. Finally, an early termination condition for the random sample consensus (RANSAC) algorithm is integrated with the Gauss-Newton optimization scheme to enable rapid convergence of the motion estimation process while achieving robustness. Experimental results based on the KITTI odometry data set show that the proposed technique outperforms the state-of-the-art visual odometry methods by producing more accurate ego-motion estimation in notably lesser amount of time. Meiqing Wu, Siew-Kei Lam, Thambipillai Srikanthan |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2015 | Nonparametric Technique Based High-Speed Road Surface DetectionabstractIt has been well recognized that detecting road surface in a realistic environment is a challenging problem that is also computationally intensive. Existing road surface detection methods attempt to fit the road surface into rigid models (e.g., planar, clothoid, or B-Spline), thereby restricting to road surfaces that match specific models. In addition, the curve-fitting strategies employed in such techniques incur high computational complexity, making them unsuitable for in-vehicle deployments. In this paper, we propose an efficient nonparametric road surface detection algorithm that exploits the depth cue. The proposed method relies on four intrinsic road scene attributes observed under stereo geometry and has been shown to reliably detect both planar and nonplanar road surfaces efficiently. Extensive evaluations are performed on three widely used benchmarks (i.e., enpeda, KITTI, and Daimler), encompassing many complex road scenarios. The experimental results show that the proposed algorithm significantly outperforms the well-known techniques both in terms of detection accuracy and runtime performance. Meiqing Wu, Siew-Kei Lam, Thambipillai Srikanthan |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2012 | Low-complexity pruning for accelerating corner detectionabstractIn this paper, we present a novel and computationally efficient pruning technique to speed up the Shi-Tomasi and Harris corner detectors. The proposed technique quickly prunes non-corners and selects a small corner candidate set by approximating the complex corner measure of Shi-Tomasi and Harris. The actual corner measure is then applied only to the reduced candidate set. Experimental results on the NiOS-II platform show that the proposed technique achieves an average execution time savings of 90% for Shi-Tomasi and 70% for Harris detectors for 500 corners with no loss in accuracy. Meiqing Wu, Nirmala Ramakrishnan, Siew-Kei Lam, Thambipillai Srikanthan |
ISCAS | 1 |