Siew-Kei Lam

dblp:74/1907 · also Siew Kei Lam · DBLP profile ↗
← Back
118ranked-venue papers
8as first author
52since 2021 · last 2026
0000-0002-8346-2635ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 56 · 7 first-author · 17 since 2021Artificial intelligence and machine learning · 20 · 16 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 since 2021Databases, data management, data science and information retrieval · 6 · 4 since 2021Computer networks · 5 · 1 first-author · 1 since 2021Security and privacy · 5 · 5 since 2021Software engineering, systems software and programming languages · 4 · 1 since 2021
YearPublicationVenuePosition
2026 Towards Automatic Incremental Learning: A Self-adaptive Framework for Continual Instruction Tuning
Peiyi Lin, Fukai Zhang, Kai Niu 0007, Siew-Kei Lam
ICIC (23)4
2026 Towards Effective Prompt Stealing Attack against Text-to-Image Diffusion Models
Shiqian Zhao, Chong Wang 0013, Yiming Li 0004, Yihao Huang 0001, Wenjie Qu 0001, Siew-Kei Lam, Yi Xie 0011, Kangjie Chen, Jie Zhang 0073, Tianwei Zhang 0004
NDSS6
2026 Securing automated insulin delivery systems: A review of security threats and protective strategies
Siew-Kei Lam
Comput. Secur.2
2026 Improving adversarial transferability and imperceptibility with loss landscape and diffusion model
Wenbo Zhou 0004, Ee-Chien Chang, Siew-Kei Lam
Pattern Recognit.6
2026 DAGSIS: A DAG-Aware MAGIC-Based Synthesis Framework for In-Memory Computing
abstract
This paper presents a comprehensive synthesis framework, named DAGSIS, for memristor-aided logic (MAGIC)-based in-memory computing system. DAGSIS addresses the limitations of prior works, such as overlooking the benefits of MAGIC’s high fan-in capability and the impact of global properties of netlists on the scheduling of computation sequence (CS). DAGSIS achieves the optimization in two synthesis stages. In the technology-independent optimization stage, DAGSIS encourages the merging of nodes in the network to reduce circuit size, by utilizing equivalent transformation of multiplexer (MUX). In the CS scheduling stage, DAGSIS introduces two schemes for optimizing area overhead and latency, respectively. For area optimization, DAGSIS maximizes the utilization of memristive cells by erasing the expired data as early as possible. For latency optimization, DAGSIS aims to minimize erasing operations, by maximizing the number of erased cells in each epoch of filling the memory. To achieve better CS scheduling, DAGSIS introduces two design rules to guide CS scheduling, which fully considers the global attributes of circuit design, such as critical path and high fan-out nodes. Experiment results show that DAGSIS reduces the circuit size by 6.69% on ISCAS’85 benchmarks compared to ABC tool, an open-source logic synthesis framework. Compared to the state-of-the-art works, DAGSIS achieves a reduction of 40.68% and 12.67% in area overhead and erasing operations respectively, on ISCAS’85 and EPFL benchmarks. The improvements are further translated into the reduction in energy consumption by up to 13.7%.
Lian Yao, Jigang Wu, Peng Liu 0045, Siew-Kei Lam
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2026 A Novel Memristive Combinational Logic for Accelerating N-bit Adders
abstract
Memristors are anticipated to replace CMOS technology due to their low power consumption and high-speed in-memory processing capabilities. However, most existing in-memory technologies implement traditional Boolean functions to achieve complex functions, which lead to increased logical depth and long delays due to the repeated iterations. In this paper, we propose a novel combinational logic, namely the AND-OR gate, which integrates the functionality of AND and OR logic into a single function. The proposed gate is able to implement multiple commonly-used logic functions within a single cycle. To highlight the advantages of the proposed AND-OR gate, twoN-bit adders, based on parallel prefix algorithms, are designed by integrating the AND-OR gate into a memristive crossbar array. Benefiting from the proposed AND-OR gate that supports prefix computation within a single cycle, the latency of the adders is significantly reduced to$O(log(N))$. Compared with the fastest reported adder that uses Majority gate for the implementation, our proposed adder achieves notable performance improvements of$1.2\times $and$2.7\times $in terms of latency and area, respectively. Moreover, the Figure of Merits (FoMs) are employed for fair comparison, in which both latency and area are considered simultaneously. Simulation results demonstrate a remarkable improvement of$40\times $over state-of-the-art circuit (i.e., the carry-select adder).
Lian Yao, Jigang Wu, Peng Liu 0045, Siew-Kei Lam
IEEE Trans. Circuits Syst. I Regul. Pap.4
2026 Enhancing Stereo Matching Domain Generalization With Adversarial Domain Alignment
abstract
Recently, state-of-the-art stereo-matching networks trained on large-scale synthetic data have shown remarkable performance. However, their capacity to extrapolate effectively to unseen real-world data,i.e.different domains, remains a challenge. The major difficulty resides in the unforeseeable domain gap when generalizing from synthetic data to real-world data. In this paper, we introduceADASM, an approach using adversarial domain alignment, designed to enhance the robustness and generalization of stereo-matching networks. It mainly consists of two modules: an end-to-end robustness optimizer and a domain-invariant feature learner. First, we adapt adversarial training into the stereo-matching task to reduce models' sensitivity to the perturbation in real-world samples. By introducing worst cases into the training space, we take unseen data into account and achieve robust disparity estimation for the end-to-end model. Then, via simulating the real-world noise with gradient-based perturbation, we construct a fictitious domain, which is taken as a referential distribution of the real-world noisy data, for further domain alignment. Specifically, we propose to utilize Maximum Mean Discrepancy to realize domain regularization between the original domain and the fictitious one. Finally, we fuse all aforementioned objectives and propose a unified, simple but effective loss function that can be adapted toallstereo-matching networks. The extensive experiments show that our method achieves a superior disparity estimation performance on various real-world benchmarks, including KITTI, Middlebury, and DrivingStereo. More importantly,ADASMobtains competitive or even better performance than the fine-tuning strategy, revealing its fine-tuning-free character.
Shiqian Zhao, Meiqing Wu, Kangjie Chen, Yi Xie 0011, Tianlin Li, Siew-Kei Lam, Guowen Xu, Anran Li 0001
IEEE Trans. Dependable Secur. Comput.6
2026 Maximizing Edge Throughput in Collaborative Multi-Task Inference With Shareable Model Structures
abstract
Recent studies in collaborative edge computing fail to take advantage of shareable structures in multi-task learning (MTL) models and the potential of MTL models sharing at the edge. This leads to resource under-utilization at the edge. Thus, this paper focuses on shareable-structure-aware model deployment and task scheduling in collaborative multi-task inference, so as to fully utilize the low-latency potential of edge computing. Specifically, we formally define the problem with an objective to maximize edge throughput under multiple constraints (e.g., resource constraint, model integrity, etc.), and prove that it is NP-hard. To solve the problem, we first propose an approximation algorithm based on randomized rounding to generate sub-optimal solutions. We then present an adjustment strategy that generates a feasible solution when the approximation algorithm violates any of the given constraints. To evaluate the proposed algorithms, we conduct comprehensive simulations based on state-of-the-art MTL models, Google cluster-usage trace, and four kinds of computing units. Extensive experiments show that the proposed approximation algorithm coupled with the adjustment strategy, outperforms state-of-the-art methods for all cases, in terms of edge throughput.
Yalan Wu, Jigang Wu, Longkun Guo, Siew-Kei Lam
IEEE Trans. Serv. Comput.6
2025 Taylor Series-Inspired Local Structure Fitting Network for Few-shot Point Cloud Semantic Segmentation
abstract
Few-shot point cloud semantic segmentation aims to accurately segment "unseen" new categories in point cloud scenes using limited labeled data. However, pretraining-based methods not only introduce excessive time overhead but also overlook the local structure representation among irregular point clouds. To address these issues, we propose a pretraining-free local structure fitting network for few-shot point cloud semantic segmentation, named TaylorSeg. Specifically, inspired by Taylor series, we treat the local structure representation of irregular point clouds as a polynomial fitting problem and propose a novel local structure fitting convolution, called TaylorConv. This convolution learns the low-order basic information and high-order refined information of point clouds from explicit encoding of local geometric structures. Then, using TaylorConv as the basic component, we construct two variant of TaylorSeg: a non-parametric TaylorSeg-NN and a parametric TaylorSeg-PN. The former can achieve performance comparable to existing parametric models without pretraining. For the latter, we equip it with an Adaptive Push-Pull (APP) module to mitigate the feature distribution differences between the query set and the support set. Extensive experiments validate the effectiveness of the proposed method. Notably, under the 2-way 1-shot setting, TaylorSeg-PN achieves improvements of +2.28% and +4.37% mIoU on the S3DIS and ScanNet datasets respectively, compared to the previous state-of-the-art methods.
Changshuo Wang 0001, Shuting He, Meiqing Wu, Siew-Kei Lam, Prayag Tiwari
AAAI5
2025 Evaluating Differentially Private Generation of Domain-Specific Text
abstract
Generative AI offers transformative potential for high-stakes domains such as healthcare and finance, yet privacy and regulatory barriers hinder the use of real-world data. To address this, differentially private synthetic data generation has emerged as a promising alternative. In this work, we introduce a unified benchmark to systematically evaluate the utility and fidelity of text datasets generated under formal Differential Privacy (DP) guarantees. Our benchmark addresses key challenges in domain-specific benchmarking, including choice of representative data and realistic privacy budgets, accounting for pre-training and a variety of evaluation metrics. We assess state-of-the-art privacy-preserving generation methods across five domain-specific datasets, revealing significant utility and fidelity degradation compared to real data, especially under strict privacy constraints. These findings underscore the limitations of current approaches, outline the need for advanced privacy-preserving data sharing methods and set a precedent regarding their evaluation in realistic scenarios.
Viktor Schlegel, Srinivasan Nandakumar, Iqra Zahid, Yuping Wu 0001, Warren Del-Pinto, Goran Nenadic, Siew-Kei Lam, Jie Zhang 0073, Anil A. Bharath
CIKM8
2025 FPGA Stereo Visual Slam with Efficient Stereo Feature Matching and Key-Frame Generation
Miyuru Thathsara, Damith Anhettigama, Siew-Kei Lam
FPL3
2025 Motion-Feat: Motion Blur-Aware Local Feature Description for Image Matching
abstract
Local feature description is crucial for robotic tasks, yet existing methods struggle with motion blur, a prevalent challenge in high-dynamic and low-light environments. While effective on sharp images, they suffer significant degradation under blur. To address this issue, we propose Motion-Feat, an end-to-end motion blur-aware feature description method. Our approach introduces a Motion Deformable Block (MDB) that adaptively adjusts the receptive field based on pixel-wise motion information at different stages of the network, enhancing multi-scale feature descriptor robustness in blurred conditions. Additionally, we construct synthetic blurred datasets to systematically benchmark feature matching performance across varying blur intensities. Extensive experiments demonstrate that Motion-Feat outperforms state-of-the-art methods on blurred images while maintaining competitive performance on sharp images for relative camera pose estimation and homography estimation tasks. Both code and datasets are available at https://github.com/AndreGao08/Motion-Feat.
Dongshuo Zhang, Qing Gao 0001, Zhijun Xu, Siew-Kei Lam, Jinhu Lü 0001
IROS6
2025 T2V-OptJail: Discrete Prompt Optimization for Text-to-Video Jailbreak Attacks
abstract
In recent years, fueled by the rapid advancement of diffusion models, text-to-video (T2V) generation models have achieved remarkable progress, with notable examples including Pika, Luma, Kling, and Open-Sora. Although these models exhibit impressive generative capabilities, they also expose significant security risks due to their vulnerability to jailbreak attacks, where the models are manipulated to produce unsafe content such as pornography, violence, or discrimination. Existing works such as T2VSafetyBench provide preliminary benchmarks for safety evaluation, but lack systematic methods for thoroughly exploring model vulnerabilities. To address this gap, we are the first to formalize the T2V jailbreak attack as a discrete optimization problem and propose a joint objective-based optimization framework, called \emph{T2V-OptJail}. This framework consists of two key optimization goals: bypassing the built-in safety filtering mechanisms to increase the attack success rate, preserving semantic consistency between the adversarial prompt and the unsafe input prompt, as well as between the generated video and the unsafe input prompt, to enhance content controllability. In addition, we introduce an iterative optimization strategy guided by prompt variants, where multiple semantically equivalent candidates are generated in each round, and their scores are aggregated to robustly guide the search toward optimal adversarial prompts. We conduct large-scale experiments on several T2V models, covering both open-source models (\textit{e.g.}, Open-Sora) and real commercial closed-source models (\textit{e.g.}, Pika, Luma, Kling). The experimental results show that the proposed method improves 11.4\% and 10.0\% over the existing state-of-the-art method (SoTA) in terms of attack success rate assessed by GPT-4, attack success rate assessed by human accessors, respectively, verifying the significant advantages of the method in terms of attack effectiveness and content control. This study reveals the potential abuse risk of the semantic alignment mechanism in the current T2V model and provides a basis for the design of subsequent jailbreak defense methods.
Siyuan Liang 0004, Shiqian Zhao, Rongcheng Tu, Wenbo Zhou 0004, Aishan Liu, Dacheng Tao, Siew-Kei Lam
NeurIPS8
2025 VFM-Depth: Leveraging Vision Foundation Model for Self-Supervised Monocular Depth Estimation
abstract
Self-supervised monocular depth estimation has exploited semantics to reduce depth ambiguities in texture-less regions and object boundaries. However, existing methods struggle to obtain universal semantics across scenes for effective depth estimation. This paper proposes VFM-Depth, a novel self-supervised teacher-student framework, that effectively leverages the vision foundation model as semantic regularization to significantly improve the accuracy of monocular depth estimation. Firstly, we propose a novel Geometric-Semantic Aggregation Encoding, integrating universal semantic constraints from the foundation model to reduce ambiguities in the teacher model. Specifically, semantic features from the foundation model and geometric features from the depth model are first encoded and then fused through cross-modal aggregation. Secondly, we introduce a novel Multi-Alignment for Depth Distillation to distill semantic constraints from the teacher, further leveraging knowledge from the foundation model. We obtain a lightweight yet effective student model through an innovative approach that combines distance category alignment with complementary feature and depth imitation. Extensive experiments on KITTI, Cityscapes, and Make3D datasets demonstrate that VFM-Depth (both teacher and student) outperforms state-of-the-art self-supervised methods by a large margin.
Shangshu Yu, Meiqing Wu, Siew-Kei Lam
IEEE Trans. Circuits Syst. Video Technol.3
2025 Looking Clearer With Text: A Hierarchical Context Blending Network for Occluded Person Re-Identification
abstract
Existing occluded person re-identification (re-ID) methods mainly learn limited visual information for occluded pedestrians from images. However, textual information, which can describe various human appearance attributes, is rarely fully utilized in the task. To address this issue, we propose a Text-guided Hierarchical Context Blending Network ( THCB-Net) for occluded person re-ID. Specifically, at the data level, informative multi-modal inputs are first generated to make full use of the auxiliary role of textual information and make image data have a strong inductive bias for occluded environments. At the feature expression level, we design a novel Hierarchical Context Blending (HCB) module that can adaptively integrate shallow appearance features obtained by CNNs and multi-scale semantic features from visual transformer encoder. At the model optimization level, a Multi-modal Feature Interaction (MFI) module is proposed to learn the multi-modal information of pedestrians from texts and images, then guide the visual transformer encoder and HCB module to further learn discriminative identity information for occluded pedestrians through Image-Multimodal Contrastive (IMC) learning. Extensive experiments on standard occluded person re-ID benchmarks demonstrate that the proposed THCB-Net outperforms state-of-the-art methods.
Changshuo Wang 0001, Xingyu Gao 0001, Meiqing Wu, Siew-Kei Lam, Shuting He, Prayag Tiwari
IEEE Trans. Inf. Forensics Secur.4
2025 EDS-Depth: Enhancing Self-Supervised Monocular Depth Estimation in Dynamic Scenes
abstract
Self-supervised monocular depth estimation usually assumes that training samples contain only static objects, which leads to poor performance in real-world environments. The presence of dynamic objects incurs camera motion estimation errors, motion blur, and occlusions, which induce significant challenges for network training. To address these issues, we introduce EDS-Depth, a self-supervised learning framework, that improves monocular depth estimation in dynamic scenes. Firstly, we propose a novel TCE (Temporal Continuity Enhancement) strategy to reduce camera motion estimation errors and motion blur caused by dynamic objects. Video frames are interpolated to generate more continuous frames in order to smooth dynamic changes and enrich motion details. Secondly, we design a novel IPDM (Iterative Pseudo Depth Masking) module to address inaccurate object motion and occlusions in dynamic scenes. The module integrates multiple optical flows from different frames for triangulation, generating optimal depth as pseudo-supervision labels in dynamic regions. Extensive experiments on Cityscapes and KITTI datasets demonstrate the effectiveness of EDS-Depth, which surpasses state-of-the-art self-supervised monocular depth estimation methods, particularly in dynamic scenes.
Shangshu Yu, Meiqing Wu, Siew-Kei Lam, Changshuo Wang 0001, Ruiping Wang 0005
IEEE Trans. Intell. Transp. Syst.3
2024 GPSFormer: A Global Perception and Local Structure Fitting-Based Transformer for Point Cloud Understanding
Changshuo Wang 0001, Meiqing Wu, Siew-Kei Lam, Xin Ning 0001, Shangshu Yu, Ruiping Wang 0005, Weijun Li 0002, Thambipillai Srikanthan
ECCV (8)3
2024 Graph Mining under Data scarcity
abstract
Multitude of deep learning models have been proposed for node classification in graphs. However, they tend to perform poorly under labeled-data scarcity. Although Few-shot learning for graphs has been introduced to overcome this problem, the existing models are not easily adaptable for generic graph learning frameworks like Graph Neural Networks (GNNs). Our work proposes an Uncertainty Estimator framework that can be applied on top of any generic GNN backbone network (which are typically designed for supervised/semi-supervised node classification) to improve the node classification performance. A neural network is used to model the Uncertainty Estimator as a probability distribution rather than probabilistic discrete scalar values. We train these models under the classic episodic learning paradigm in the n-way, k-shot fashion, in an end-to-end setting.Our work demonstrates that implementation of the uncertainty estimator on a GNN backbone network improves the classification accuracy under Few-shot setting without any meta-learning specific architecture. We conduct experiments on multiple datasets under different Few-shot settings and different GNN-based backbone networks. Our method outperforms the baselines, which demonstrates the efficacy of the Uncertainty Estimator for Few-shot node classification on graphs with a GNN.
Appan Rakaraddi, Siew-Kei Lam, Mahardhika Pratama, Marcus de Carvalho
IJCNN2
2024 CurricularVPR: Curricular Contrastive Loss for Visual Place Recognition
abstract
Visual Place Recognition (VPR) techniques commonly utilize Contrastive Losses (CL) to train models that generate compact and discriminative global descriptors for images. These models often result in poor performance due to one of the following reasons during training: 1) loss functions that focus primarily on easier samples, 2) reliance on time-consuming hard sample mining methods to identify informative supervisory samples, which hinders effective learning from large-scale datasets. To enhance both learning efficiency and effectiveness, we propose a Curricular Contrastive Loss (CCL) and use graded similarity labels as a measure of sample difficulty. Inspired by human learning that begin with easier concepts and progressively tackle more challenging ones, our CCL dynamically emphasizes easier samples during the initial training stages to achieve rapid convergence. The learning gradually focuses on harder samples in later training stages to bolster robustness of the models under challenging conditions. Our proposed method has been extensively evaluated on popular datasets, and the results demonstrate its superior performance compared to the CL and Generalized CL functions.
Dongshuo Zhang, Nanhua Chen, Meiqing Wu, Siew-Kei Lam
IROS4
2024 Streamlining DNN Obfuscation to Defend Against Model Stealing Attacks
abstract
Side-channel-based Deep Neural Network (DNN) model stealing has become a major concern with the advent of learning-based attacks. In respond to this threat, defence mechanisms have been presented to obfuscate the DNN execution, making it difficult to infer the correlation between side-channel information and DNN architecture. However, state-of-the-art (SOTA) DNN obfuscation is time-consuming, requires expert-level changes in existing DNN compilers (e.g., Tensor Virtual Machine (TVM)), and often relies on prior knowledge of the attack models. In this work, we study the impact of various obfuscation levels on the defence effectiveness, and present a streamlined DNN obfuscation process that is extremely fast and is agnostic to any attack models. Our study reveals that by just modifying the scheduling of DNN operations on the GPU, we can achieve comparable defense performance as the SOTA in an attack agnostic manner. We also propose a simple algorithm that determines an effective scheduling configuration for mitigating DNN model stealing at a fraction of a time required by SOTA obfuscation methods. Our method can be easily integrated into existing DNN compilers as a security feature, even by non-experts, to protect their DNN against side-channel attacks.
Siew-Kei Lam, Guiyuan Jiang, Peilan He
ISCAS2
2024 Hardware Accelerator for Feature Matching with Binary Search Tree
abstract
Feature matching is an essential step for autonomous robot to localize itself during navigation. However, it is often difficult to achieve the matching in real-time due to limited on-board computing resources. We present an FPGA implementation for stream-processing based feature matching called Binary Search Tree Matcher (BSTM) that utilizes a balanced binary search tree. To improve the matching precision, we integrated a ratio-test (RT) outlier rejection mechanism in our design. Our proposed FPGA-based BSTM is scalable and resource efficient, and significantly faster than the best-performing hardware implementation in the literature. When compared to the commonly used Linear Exhaustive Search (LES) method running on the FPGA, our proposed design is approximately 12X faster.
Miyuru Thathsara, Siew-Kei Lam, Damith Kawshan, Duvindu Piyasena
ISCAS2
2024 Hierarchical Object-Aware Dual-Level Contrastive Learning for Domain Generalized Stereo Matching
abstract
Stereo matching algorithms that leverage end-to-end convolutional neural networks have recently demonstrated notable advancements in performance. However, a common issue is their susceptibility to domain shifts, hindering their ability in generalizing to diverse, unseen realistic domains. We argue that existing stereo matching networks overlook the importance of extracting semantically and structurally meaningful features. To address this gap, we propose an effective hierarchical object-aware dual-level contrastive learning (HODC) framework for domain generalized stereo matching. Our framework guides the model in extracting features that support semantically and structurally driven matching by segmenting objects at different scales and enhances correspondence between intra- and inter-scale regions from the left feature map to the right using dual-level contrastive loss. HODC can be integrated with existing stereo matching models in the training stage, requiring no modifications to the architecture. Remarkably, using only synthetic datasets for training, HODC achieves state-of-the-art generalization performance with various existing stereo matching network architectures, across multiple realistic datasets.
Yikun Miao, Meiqing Wu, Siew-Kei Lam, Thambipillai Srikanthan
NeurIPS3
2024 Employing feature mixture for active learning of object detection
Licheng Zhang 0004, Siew-Kei Lam, Dingsheng Luo, Xihong Wu
Neurocomputing2
2024 Layer Sequence Extraction of Optimized DNNs Using Side-Channel Information Leaks
abstract
Deep neural network (DNN) intellectual property (IP) models must be kept undisclosed to avoid revealing trade secrets. Recent works have devised machine learning techniques that leverage on side-channel information leakage of the target platform to reverse engineer DNN architectures. However, these works fail to perform successful attacks on DNNs that have undergone performance optimizations (i.e., operator fusion) using DNN compilers, e.g., Apache tensor virtual machine (TVM). We propose a two-phase attack framework to infer the layer sequences of optimized DNNs through side-channel information leakage. In the first phase, we use a recurrent network with multihead attention components to learn the intra and interlayer fusion patterns from GPU traces of TVM-optimized DNNs, in order to accurately predict the operation distribution. The second phase uses a model to learn the run-time temporal correlations between operations and layers, which enables the prediction of layer sequence. An encoding strategy is proposed to overcome the convergence issues faced by existing learning-based methods when inferring the layer sequences of optimized DNNs. Extensive experiments show that our learning-based framework outperforms state-of-the-art DNN model extraction techniques. Our framework is also the first to effectively reverse engineer both convolutional neural networks (CNNs) and recurrent neural networks (RNNs) using side-channel leakage.
Guiyuan Jiang, Xinwang Liu 0002, Peilan He, Siew-Kei Lam
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2024 Share-Aware Joint Model Deployment and Task Offloading for Multi-Task Inference
abstract
In vehicular edge computing, efficient strategies for model deployment and task offloading offer tremendous potential to reduce response time for machine learning inference. However, existing works do not pay much attention to that there are shared structures among different types of inference tasks. This limits the improvement in response time. This paper aims to fill this gap by investigating a share-aware joint model deployment and task offloading problem for multi-task inference in vehicular edge computing. We formulate the problem with an objective to minimize the total response time of all inference requests, under constraints of per task response time, per roadside unit storage capacity, etc. We prove that the formulated problem is NP-hard. To solve the problem, a time period aware algorithm, called TPA, is proposed with guaranteed approximation ratio. In TPA, an iterative approach is designed to solve the problem of maximizing system throughput during a certain time period. Then, the certain time period approximates to the minimum time period of completing all requests. The algorithms are evaluated in the environment comprising two CPUs, two GPUs, state-of-the-art multi-task learning models and the dataset of Google cluster-usage trace. Simulation results derived from this environment show that, the proposed TPA outperforms the state-of-the-art methods for all cases, in terms of the total response time of all requests. For example, TPA can significantly reduce the total response time by at least$73.72\%$for different numbers of RSUs considered, compared with state-of-the-art methods.
Yalan Wu, Jigang Wu, Long Chen 0006, Bosheng Liu, Mianyang Yao, Siew-Kei Lam
IEEE Trans. Intell. Transp. Syst.6
2023 Training-Free Attentive-Patch Selection for Visual Place Recognition
abstract
Visual Place Recognition (VPR) utilizing patch descriptors from Convolutional Neural Networks (CNNs) has shown impressive performance in recent years. Existing works either perform exhaustive matching of all patch descriptors, or employ complex networks to select good candidate patches for further geometric verification. In this work, we develop a novel two-step training-free patch selection method that is fast, while being robust to large occlusions and extreme viewpoint variations. In the first step, a self-attention mechanism is used to select sparse and evenly distributed discriminative patches in the query image. Next, a novel spatial-matching method is used to rapidly select corresponding patches with high similar appearances between the query and each reference image. The proposed method is inspired by how humans perform place recognition by first identifying prominent regions in the query image, and then relying on back-and-forth visual inspection of the query and reference image to attentively identify similar regions while ignoring dissimilar ones. Extensive experiment results show that our proposed method outperforms state-of-the-art (SOTA) methods in both place recognition precision and runtime, on various challenging conditions.
Dongshuo Zhang, Meiqing Wu, Siew-Kei Lam
IROS3
2023 Computer vision framework for crack detection of civil infrastructure - A review
Dihao Ai, Guiyuan Jiang, Siew-Kei Lam, Peilan He, Chengwu Li
Eng. Appl. Artif. Intell.3
2023 CoTree: A Side-Channel Collision Tool to Push the Limits of Conquerable Space
abstract
By introducing collision information into divide-and-conquer distinguishers, the existing collision-optimized side-channel attacks transform the given candidate space into a significantly smaller collision space, thus achieving more efficient key recovery. However, the candidates of the first several subkeys shared by collision chains are still repeatedly detected, which happens very frequently and brings huge computational overhead. To alleviate this, we propose a highly efficient collision-optimized attack named collision tree (CoTree). This collision detection tool exploits tree structure to store the chains created from the same subchain on the same branch, thus significantly reducing the storage requirements. It then benefits from the properties of both tree and collisions and exploits a top-down tree building procedure and traverses each node only once when detecting their collisions with a candidate of the subkey currently under consideration. Finally, unlike the traditional top-down node removal, CoTree launches a bottom-up branch removal procedure to remove the chains unsatisfying the collision conditions from the tree after traversing all the considered candidates of this subkey, thus avoiding the traversal of the branches satisfying the collision condition. These strategies make our CoTree significantly alleviate the repetitive collision detection, and our experiments verify that it significantly outperforms the existing works.
Changhai Ou, Debiao He, Kexin Qiao, Shihui Zheng, Siew-Kei Lam, Fan Zhang 0010
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2023 Two-Level Scheduling Algorithms for Deep Neural Network Inference in Vehicular Networks
abstract
In vehicular networks, task scheduling at the microarchitecture-level and network-level offers tremendous potential to improve the quality of computing services for deep neural network (DNN) inference. However, existing task scheduling works only focus on either one of the two levels, which results in inefficient utilization of computing resources. This paper aims to fill this gap by formulating a two-level scheduling problem for DNN inference tasks in a vehicular network, with an objective of minimizing total weighted sum of response time and energy consumption for all tasks under the following constraints: per task response time, per vehicle energy consumption, per vehicle storage capacity. We first formulate the problem and prove that it is NP-hard. A group transformation based algorithm, called GTA, is proposed. GTA makes scheduling decisions at the network-level using the group transformation based approach, and at the microarchitecture-level using a greedy strategy. In addition, an algorithm, denoted as DRL, is proposed to decrease total weighted sum of response time and energy consumption for all tasks. DRL trains two models with deep reinforcement learning to achieve two-level scheduling. The proposed algorithms are evaluated on a platform consisting of a desktop, Raspberry Pi, Eyeriss, OSM, SUMO, NS-3. Simulation results show that DRL outperforms the state-of-the-art methods for all cases, while the proposed GTA outperforms the state-of-the-art methods for most cases, in terms of total weighted sum of response time and energy consumption. Compared with four baseline algorithms, GTA and DRL reduce the total weighted sum of response time and energy consumption by 41.49% and 62.38%, on average respectively, for different numbers of tasks.
Yalan Wu, Jigang Wu, Mianyang Yao, Bosheng Liu, Long Chen 0006, Siew-Kei Lam
IEEE Trans. Intell. Transp. Syst.6
2023 Blockchain-Based Secure Key Management for Mobile Edge Computing
abstract
Mobile edge computing (MEC) is a promising edge technology to provide high bandwidth and low latency shared services and resources to mobile users. However, the MEC infrastructure raises major security concerns when the shared resources involve sensitive and private data of users. This paper proposes a novel blockchain-based key management scheme for MEC that is essential for ensuring secure group communication among the mobile devices as they dynamically move from one subnetwork to another. In the proposed scheme, when a mobile device joins a subnetwork, it first generates lightweight key pairs for digital signature and communication, and broadcasts its public key to neighbouring peer users in the subnetwork blockchain. The blockchain miner in the subnetwork packs all the public key of mobile devices into a block that will be sent to other users in the subnetwork. This enables the mobile device to communicate with its peers in the subnetwork by encrypting the data with the public key stored in the blockchain. When the mobile device moves to another subnetwork in the tree network, all the mobile devices of the new subnetwork can quickly verify its identity by checking its record in the local or higher hierarchy subnetwork blockchain. Furthermore, when the mobile device leaves the subnetwork, it does not need to do anything and its records will remain in the blockchain which is an append-only database. Theoretical security analysis shows that the proposed scheme can defend against the 51 percent attack and malicious entities in the blockchain network utilizing Proof-of-Work consensus mechanism. Moreover, the backward and forward secrecy is also preserved. Experimental results demonstrate that the proposed scheme outperforms two baselines in terms of computation, communication and storage.
Jiaxing Li 0009, Jigang Wu, Long Chen 0006, Jin Li 0002, Siew-Kei Lam
IEEE Trans. Mob. Comput.5
2022 Reinforced Continual Learning for Graphs
abstract
Graph Neural Networks (GNNs) have become the backbone for a myriad of tasks pertaining to graphs and similar topological data structures. While many works have been established in domains related to node and graph classification/regression tasks, they mostly deal with a single task. Continual learning on graphs is largely unexplored and existing graph continual learning approaches are limited to the task-incremental learning scenarios. This paper proposes a graph continual learning strategy that combines the architecture-based and memory-based approaches. The structural learning strategy is driven by reinforcement learning, where a controller network is trained in such a way to determine an optimal number of nodes to be added/pruned from the base network when new tasks are observed, thus assuring sufficient network capacities. The parameter learning strategy is underpinned by the concept of Dark Experience replay method to cope with the catastrophic forgetting problem. Our approach is numerically validated with several graph continual learning benchmark problems in both task-incremental learning and class-incremental learning settings. Compared to recently published works, our approach demonstrates improved performance in both the settings. The implementation code can be found at https://github.com/codexhammer/gcl.
Appan Rakaraddi, Siew-Kei Lam, Mahardhika Pratama, Marcus de Carvalho
CIKM2
2022 ML-MMAS: Self-learning ant colony optimization for multi-criteria journey planning
Peilan He, Guiyuan Jiang, Siew-Kei Lam
Inf. Sci.3
2022 Decoupled self-supervised label augmentation for fully-supervised image classification
Wanshun Gao, Meiqing Wu, Siew-Kei Lam, Qihui Xia, Jianhua Zou
Knowl. Based Syst.3
2022 Exploring Public Transport Transfer Opportunities for Pareto Search of Multicriteria Journeys
abstract
Multimodal public transport networks (MMPTNs) in modern cities are becoming increasingly complex. This makes finding optimal journey routes challenging due to a large number of transfer options that need to be properly considered. Furthermore, the complexity of the problem is compounded when multiple conflicting travel criteria are considered (e.g., travel time, walking distance, travel fare, etc.). This paper proposes a transfer graph (TG) model to explore the transfer opportunities of the MMPTN to support efficient journey route planning. TG considers all possible transfer opportunities, while employing a representative mechanism to optimize the TG structure that supports efficient route planning algorithms. Based on the proposed TG, we develop two exact algorithms to search the Pareto-optimal solutions for multi-criteria journey planning (MCJP) over the MMPTN. The first algorithm runs faster by eliminating many partial solutions at an early stage, which is more suited for lowering computation time at the expense of marginal degradation in output quality. In contrast, the second algorithm provides a more dependable solution by incorporating accurate journey time prediction that caters to the evolving traffic conditions. We also develop techniques to accelerate the TEDE and TEAE algorithms. Experiments on real-world public transport networks and traffic data demonstrate the effectiveness of our approach for MCJP. Experiment results also reveal interesting insights on the impact of the TOs, number of transfers, and number of travel criteria on MCJP algorithms, which can contribute to better public transportation planning.
Peilan He, Guiyuan Jiang, Siew-Kei Lam, Fangxin Ning
IEEE Trans. Intell. Transp. Syst.3
2022 A Multi-Scale Attributes Attention Model for Transport Mode Identification
abstract
Transport mode identification (TMI), which infers the travel modes of user trajectories, is essential to facilitate an understanding of urban mobility patterns and passengers’ choice behaviors with the goal of improving urban transportation systems. To achieve higher accuracy, existing TMI methods usually rely on mobility features obtained from densely sampled GPS trajectory points (e.g. 1 second per GPS point) or data measurements of additional inertial measurement unit (IMU) sensors (e.g. accelerometer, gyroscope, rotation vector). However, these lead to high energy consumption of the users’ mobile devices. In this paper, we propose a novel deep learning framework, Multi-Scale Attributes Attention (MSAA) model, to extract discriminating trajectory features from GPS data only, without the need to increase its sampling rate. The proposed model first partitions the trajectories into different scales and extract the latent representation of local attributes at each scale. The MSAA model relies on Convolutional Neural Network (CNN) to capture the spatial correlation of different trajectory segments, and utilizes attention mechanism to select the most suitable local attributes on the different trajectory scales that can effectively characterize the various transport modes. Since the learned latent local attributes are significantly different from the global features (e.g. average/min/max travel speeds which are measurable quantities), an ensemble model based on Neural Decision Forest (NDF) is employed to fuse the heterogeneous features consisting of both measurable quantities and non-measurable elements for determining the transport mode. Experiments on real-world datasets demonstrate the competitive performance of the proposed approach compared to several state-of-the-art baselines, with average improvements in accuracy ranging from 0.76% to 6.4%. In addition, the proposed multi-scale local attributes well complement the global features. Our results show that by incorporating the local attributes, the detection performance improved by 2.3% on average compared to using only global features.
Guiyuan Jiang, Siew-Kei Lam, Peilan He, Changhai Ou, Dihao Ai
IEEE Trans. Intell. Transp. Syst.2
2022 Fast Semantic-Aware Motion State Detection for Visual SLAM in Dynamic Environment
abstract
Existing visual SLAM (vSLAM) systems fail to perform well in dynamic environments as they cannot effectively ignore moving objects during pose estimation and mapping. We propose a lightweight approach to improve the robustness of existing feature based RGB-D and stereo vSLAM by accurately removing dynamic outliers in the scene that contribute to failures in pose estimation and mapping. First, a novel motion state detection algorithm using the depth and feature flow information is presented to identify regions in the scene with high moving probability. This information is then fused with semantic cues via a probability framework to enable accurate and robust moving object extraction to retain the useful features for pose estimation and mapping. To reduce the computational complexity of extracting semantic information in every frame, we propose to extract semantics only on keyframes with significant changes in image content. Semantic propagation is used to compensate for the changes in the intermediate frames (i.e., non-keyframes). This is achieved by computing the dense transformation map using the available feature flow vectors. The proposed techniques can be integrated into existing vSLAM systems to increase their robustness in dynamic environments without incurring much computation cost. Our work highlights the importance of distinguishing between motion states of potential moving objects for vSLAM in highly dynamic environments. We provide extensive experimental results on four well-known RGB-D and stereo datasets to show that the proposed technique outperforms existing vSLAM methods in indoor and outdoor environments under various dynamic scenarios including crowded scenes. We also perform our experiments on a low-cost embedded platform, i.e., Jetson TX1, to demonstrate the computational efficiency of our method.
Gaurav Singh 0009, Meiqing Wu, Do Van Minh, Siew-Kei Lam
IEEE Trans. Intell. Transp. Syst.4
2022 Learning Traffic Network Embeddings for Predicting Congestion Propagation
abstract
Traffic congestion has become a global concern due to continuous increase in traffic demand and limited road capacity. The ability to predict traffic congestion propagation, which depicts the spatiotemporal evolution of the congestion scenario, is essential for developing smart traffic management systems and enabling road users to make informed route choices. In this work, we study the behavior of congestion propagation at the road segment level, and leverage this to develop a novel machine learning framework that characterizes and predicts the congestion evolution among different road segments in the traffic network. In particular, our framework can infer the likelihood of congestion propagation between any pair of road segments through single or multiple propagation paths. The proposed framework relies on a network embedding module to learn a representation for each road segment, and a propagation model which calculates the congestion propagation likelihood based on the learned representations. Specifically, an asymmetric embedding of local proximity and global tendency (AE-LPGT) is relied upon for learning low dimension embeddings of the road segments which incorporate various realistic properties of congestion propagations, such as the local proximity property, global propagation tendency, and asymmetric transitivity of congestion propagations. Experimental results with Singapore traffic data show that our method significantly outperforms the state-of-the-art, and the congestion propagation properties in our embeddings have significant impact on the prediction performance.
Guiyuan Jiang, Siew-Kei Lam, Peilan He
IEEE Trans. Intell. Transp. Syst.3
2022 A Unified Multi-Task Learning Architecture for Fast and Accurate Pedestrian Detection
abstract
We present a unified multi-task learning architecture for fast and accurate pedestrian detection. Different from existing methods which often focus on either a new loss function or architecture, we propose an improved multi-task convolutional neural network learning architecture to effectively and efficiently interfuse the task of pedestrian detection and semantic segmentation. To achieve this, we integrate a lightweight semantic segmentation branch to Faster R-CNN detection framework that enables end-to-end hard parameter sharing in order to boost the detection performance and maintain computational efficiency as follows. Firstly, a Semantic Segmentation to Feature Module (SS2FM) refines the convolutional features in RPN stage by integrating the features generated from the semantic segmentation branch. Secondly, a Semantic Segmentation to Confidence Module (SS2CM) refines the classification confidence in RPN stage by fusing it with the semantic segmentation confidence. We also introduce an effective anchor matching point transform to alleviate the problem of feature misalignment for heavily occluded pedestrians. The proposed unified multi-task learning architecture lends itself well to more robust pedestrian detection in diverse scenarios with negligible computation overhead. In addition, the proposed architecture can achieve high detection performance with low resolution input images, which significantly reduces the computational complexity. Experiment results on CityPersons and Caltech datasets show that our method is the fastest among all state-of-the-art pedestrian detection methods while exhibiting competitive detection performance.
Chengju Zhou, Meiqing Wu, Siew-Kei Lam
IEEE Trans. Intell. Transp. Syst.3
2022 Enhanced Multi-Task Learning Architecture for Detecting Pedestrian at Far Distance
abstract
Existing pedestrian detection methods suffer from performance degradation in the presence of small-scale pedestrians who are positioned at far distance from the camera. We present a pedestrian detection framework that is not only robust to small- and large-scale pedestrians, but is also significantly faster than state-of-the-art methods. The proposed framework incorporates semantic segmentation to confidence modules for RPN (Region Proposal Network) head and R-FCN (Region-based Fully Convolutional Networks) head, and a cascaded R-FCN head. The semantic segmentation confidence is extracted and utilized as auxiliary classification prior knowledge for RPN proposal selection and R-FCN head prediction. Finally, the cascaded R-FCN head progressively refine the pedestrian prediction accuracy with negligible computation overhead. The proposed framework is also capable of maintaining high detection performance on down-sampled input images, which leads to further reduction in overall computational complexity. Experiment results on CityPersons and MOT17Det datasets show that the proposed framework achieves competitive detection performance with about$3\times $speedup over state-of-the-art methods.
Chengju Zhou, Meiqing Wu, Siew-Kei Lam
IEEE Trans. Intell. Transp. Syst.3
2021 XDIVINSA: eXtended DIVersifying INStruction Agent to Mitigate Power Side-Channel Leakage
abstract
Side-channel analysis (SCA) attacks pose a major threat to embedded systems due to their ease of accessibility. Realising SCA resilient cryptographic algorithms on embedded systems under tight intrinsic constraints, such as low area cost, limited computational ability, etc., is extremely challenging and often not possible. We propose a seamless and effective approach to realise a generic countermeasure against SCA attacks. XDIVINSA, an extended diversifying instruction agent, is introduced to realise the countermeasure at the microarchitecture level based on the combining concept of diversified instruction set extension (ISE) and hardware diversification. XDIVINSA is developed as a lightweight co-processor that is tightly coupled with a RISC-V processor. The proposed method can be applied to various algorithms without the need for software developers to undertake substantial design efforts hardening their implementations against SCA. XDIVINSA has been implemented on the SASEBO G-III board which hosts a Kintex-7 XC7K160T FPGA device for SCA mitigation evaluation. Experimental results based on non-specific t-statistic tests show that our solution can achieve leakage mitigation on the power side channel of different cryptographic kernels, i.e., Speck, ChaCha20, AES, and RSA with an acceptable performance overhead compared to existing countermeasures.
Thinh Hung Pham, Ben Marshall, Alexander Fell, Siew-Kei Lam, Dan Page
ASAP4
2021 Cache-Aware Dynamic Skewed Tree for Fast Memory Authentication
abstract
Memory integrity trees are widely-used to protect external memories in embedded systems against bus attacks. However, existing methods often result in high performance overheads incurred during memory authentication. To reduce memory accesses during authentication, the tree nodes are cached on-chip. In this paper, we propose a cacheaware technique to dynamically skew the integrity tree based on the application workloads in order to reduce the performance overhead. The tree is initialized using Van-Emde Boas (vEB) organization to take advantage of locality of reference. At run time, the nodes of the integrity tree are dynamically positioned based on their memory access patterns. In particular, frequently accessed nodes are placed closer to the root to reduce the memory access overheads. The proposed technique is compared with existing methods on Multi2Sim using benchmarks from SPEC-CPU2006, SPLASH-2 and PARSEC to demonstrate its performance benefits.
Saru Vig, Siew-Kei Lam, Rohan Juneja
ASP-DAC2
2021 Edge Accelerator for Lifelong Deep Learning using Streaming Linear Discriminant Analysis
abstract
Lifelong deep learning models are expected to continuously adapt and acquire new knowledge in dynamic environments. This capability is essential for numerous vision tasks in robotics and drones, and the models must be deployed on the edge to achieve real-time performance. We propose a FPGA accelerator of a streaming classifier for lifelong deep learning, which is based on streaming linear discriminant analysis (SLDA). When combined with a frozen Convolutional Neural Network (CNN) model, the proposed system is capable of class incremental lifelong learning for object classification.
Duvindu Piyasena, Siew-Kei Lam, Meiqing Wu
FCCM2
2021 Accelerating Continual Learning on Edge FPGA
abstract
Real-time edge AI systems operating in dynamic environments must learn quickly from streaming input samples without needing to undergo offline model training. We propose an FPGA accelerator for continual learning based on streaming linear discriminant analysis (SLDA), which is capable of class-incremental object classification. The proposed SLDA accelerator employs application-specific parallelism, efficient data reuse, resource sharing, and approximate computing to achieve high performance and power efficiency. Additionally, we introduce a new variant of SLDA and discuss the accuracy-efficiency trade-offs. The proposed SLDA accelerator is combined with a Convolutional Neural Network (CNN). which is implemented on Xilinx DPU to achieve full continual learning capability at nearly the same latency as inference. Experiments based on popular datasets for continual learning, CoRE50 and CUB200, demonstrate that the proposed SLDA accelerator outperforms the embedded CPU and GPU counterparts, in terms of speed and energy efficiency.
Duvindu Piyasena, Siew-Kei Lam, Meiqing Wu
FPL2
2021 Self-Growing Spatial Graph Network for Context-Aware Pedestrian Trajectory Prediction
abstract
Pedestrian trajectory prediction is an active research area with recent works undertaken to embed accurate models of pedestrians social interactions and their contextual compliance into dynamic spatial graphs. However, existing works rely on spatial assumptions about the scene and dynamics, which entails a significant challenge to adapt the graph structure in unknown environments for an online system. In addition, there is a lack of assessment approach for the relational modeling impact on prediction performance. To fill this gap, we propose Social Trajectory Recommender-Gated Graph Recurrent Neighborhood Network (STR-GGRNN), which uses data-driven adaptive online neighborhood recommendation based on the contextual scene features and pedestrian visual cues. The neighborhood recommendation is achieved by online Nonnegative Matrix Factorization (NMF) to construct the graph adjacency matrices for predicting the pedestrians’ trajectories. Experiments based on widely-used datasets show that our method outperforms the state-of-the-art. Our best performing model achieves 12 cm ADE and $\sim 15$ cm FDE on ETH-UCY dataset.
Sirin Haddad, Siew-Kei Lam
ICIP2
2021 Predicting Traffic Congestion Evolution: A Deep Meta Learning Approach
abstract
Many efforts are devoted to predicting congestion evolution using propagation patterns that are mined from historical traffic data. However, the prediction quality is limited to the intrinsic properties that are present in the mined patterns. In addition, these mined patterns frequently fail to sufficiently capture many realistic characteristics of true congestion evolution (e.g., asymmetric transitivity, local proximity). In this paper, we propose a representation learning framework to characterize and predict congestion evolution between any pair of road segments (connected via single or multiple paths). Specifically, we build dynamic attributed networks (DAN) to incorporate both dynamic and static impact factors while preserving dynamic topological structures. We propose a Deep Meta Learning Model (DMLM) for learning representations of road segments which support accurate prediction of congestion evolution. DMLM relies on matrix factorization techniques and meta-LSTM modules to exploit temporal correlations at multiple scales, and employ meta-Attention modules to merge heterogeneous features while learning the time-varying impacts of both dynamic and static features. Compared to all state-of-the-art methods, our framework achieves significantly better prediction performance on two congestion evolution behaviors (propagation and decay) when evaluated using real-world dataset.
Guiyuan Jiang, Siew-Kei Lam, Peilan He
IJCAI3
2021 CAP: Context-Aware Pruning for Semantic Segmentation
abstract
Network pruning for deep convolutional neural networks (CNNs) has recently achieved notable research progress on image-level classification. However, most existing pruning methods are not catered to or evaluated on semantic segmentation networks. In this paper, we advocate the importance of contextual information during channel pruning for semantic segmentation networks by presenting a novel Context-aware Pruning framework. Concretely, we formulate the embedded contextual information by leveraging the layer-wise channels interdependency via the Context-aware Guiding Module (CAGM) and introduce the Context-aware Guided Sparsification (CAGS) to adaptively identify the informative channels on the cumbersome model by inducing channel-wise sparsity on the scaling factors in batch normalization (BN) layers. The resulting pruned models require significantly lesser operations for inference while maintaining comparable performance to (at times outperforming) the original models. We evaluated our framework on widely-used benchmarks and showed its effectiveness on both large and lightweight models. On Cityscapes dataset, our framework reduces the number of parameters by 32%, 47%, 54%, and 63%, on PSPNet101, PSPNet50, ICNet, and SegNet, respectively, while preserving the performance.
Wei He 0025, Meiqing Wu, Mingfu Liang, Siew-Kei Lam
WACV4
2021 Passenger-centric vehicle routing for first-mile transportation considering request uncertainty
Fangxin Ning, Guiyuan Jiang, Siew-Kei Lam, Changhai Ou, Peilan He
Inf. Sci.3
2021 ACSL: Adaptive correlation-driven sparsity learning for deep neural network compression
Wei He 0025, Meiqing Wu, Siew-Kei Lam
Neural Networks3
2021 The Science of Guessing in Collision-Optimized Divide-and-Conquer Attacks
abstract
Recovering keys ranked in very deep candidate space efficiently is a very important but challenging issue in side-channel attacks (SCAs). State-of-the-art collision-optimized divide-and-conquer attacks (CODCAs) extract collision information from a collision attack to optimize the key recovery of a divide-and-conquer attack, and transform the very huge guessing space to a much smaller collision space. However, the inefficient collision detection makes them time consuming. The very limited collisions exploited and large performance difference between the collision attack and the divide-and-conquer attack in CODCAs also prevent their application in much larger spaces. In this article, we propose a Minkowski distance enhanced collision attack (MDCA) with performance closer to template attack (TA) compared to traditional correlation-enhanced collision attack (CECA), thus making the optimization more practical and meaningful. Next, we build a more advanced CODCA named full-collision chain (FCC) from TA and MDCA to exploit all collisions. Moreover, to minimize the thresholds while guaranteeing a high success probability of key recovery, we propose a fault-tolerant scheme to optimize FCC. The full key is divided into several big “blocks,” on which a fault-tolerant vector (FTV) is exploited to flexibly adjust its chain space. Finally, guessing theory is exploited to optimize thresholds determination and search order of subkeys. Experimental results show that FCC notably outperforms the existing CODCAs.
Changhai Ou, Siew-Kei Lam, Guiyuan Jiang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 SNR-Centric Power Trace Extractors for Side-Channel Attacks
abstract
Existing power trace extractors consider the case where the number of power traces available to the attacker is sufficient to guarantee successful attacks, and the goal of power trace extraction is to extract a small part of traces with high signal-to-noise ratio (SNR) to reduce the complexity of attacks rather than to increase the success rates. Although strict theoretical proofs are given, the existing power trace extractors are too simple and leakage characteristics of Points-of-Interest (POIs) have not been thoroughly analyzed. They only maximize the variance of the data-dependent power consumption component and ignore the noise component, which results in very limited SNR that hampers the performance of extractors. In this article, we provide a rigorous theoretical analysis of SNR of power traces, and propose a simple yet efficient SNR-centric extractor, named shortest distance first (SDF), to extract power traces with the smallest estimated noise by taking advantage of known plaintexts. In addition, to maximize the variance of the exploitable component while minimizing the noise, we refer to the SNR estimation model and propose another novel extractor named maximizing estimated SNR first (MESF). Finally, we further propose an advanced extractor called mean-optimized MESF (MMESF) that exploits the mean power consumption of each plaintext byte value to more accurately and reasonably estimate the data-dependent power consumption of the corresponding samples. Experiments on both simulated power traces and measurements from an ATmega328p micro-controller demonstrate the superiority of our new extractors.
Changhai Ou, Siew-Kei Lam, Degang Sun, Xinping Zhou, Kexin Qiao, Qu Wang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2021 Information Entropy-Based Leakage Profiling
abstract
An accurate leakage model is critical to side-channel attacks and evaluations. Leakage certification plays an important role to address the following question: “how good is my leakage model?” Moreover, most of the current leakage model profiling only exploit the information from lower orders of moments. They still need to tolerate assumption error and estimation error from unknown leakage models. There are many probability density functions (PDFs) satisfying given moment constraints. As such, finding an unbiased, objective, and reasonable model still remains an unresolved problem. In this article, we address a more fundamental question: “which model can approach the leakage infinitely and is the optimal in theory?” In particular, we extract information from higher order moments and propose maximum entropy distribution (MED) to estimate the leakage model as MED is an unbiased, objective, and theoretically the most reasonable PDF conditioned upon the available information. MED is a moment-based statistical PDF model in side-channel attacks. It can theoretically use information on arbitrary higher order moments to infinitely approximate the leakage distribution, and well compensates the theory vacancy of model profiling and evaluation. Experimental results demonstrate the superiority of our proposed method for approximating the leakage model using MED estimation.
Changhai Ou, Xinping Zhou, Siew-Kei Lam, Chengju Zhou, Fangxin Ning
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2021 Multiple-Differential Mechanism for Collision-Optimized Divide-and-Conquer Attacks
abstract
Several combined attacks have shown promising results in recovering cryptographic keys by introducing collision information into divide-and-conquer attacks to transform a part of the best key candidates within given thresholds into a much smaller collision space. However, these Collision-Optimized Divide-and-Conquer Attacks (CODCAs) uniformly demarcate the thresholds for all sub-keys, which is unreasonable. Moreover, the inadequate exploitation of collision information and backward fault tolerance mechanisms of CODCAs also lead to low attack efficiency. Finally, existing CODCAs mainly focus on improving collision detection algorithms but lack theoretical basis. We exploit Correlation-Enhanced Collision Attack (CECA) to optimize Template Attack (TA). To overcome the above-mentioned problems, we first introduce guessing theory into TA to enable the quick estimation of success probability and the corresponding complexity of key recovery. Next, a novel Multiple-Differential mechanism for CODCAs (MD-CODCA) is proposed. The first two differential mechanisms construct collision chains satisfying the given number of collisions from several sub-keys with the fewest candidates under a fixed probability provided by guessing theory, then exploit them to vote for the remaining sub-keys. This guarantees that the number of remaining chains is minimal, and makes MD-CODCA suitable for very high thresholds. Our third differential mechanism simply divides the key into several large non-overlapping “blocks” to further exploit intra-block collisions from the remaining candidates and properly ignore the inter-block collisions, thus facilitating the later key enumeration. The experimental results show that MD-CODCA significantly reduces the candidate space and lowers the complexity of collision detection, without considerably reducing the success probability of attacks.
Changhai Ou, Chengju Zhou, Siew-Kei Lam, Guiyuan Jiang
IEEE Trans. Inf. Forensics Secur.3
2020 DISSECT: Dynamic Skew-and-Split Tree for Memory Authentication
abstract
Memory integrity trees are widely-used to protect external memories in embedded systems against replay, splicing and spoofing attacks. However, existing methods often result in high-performance overhead that is proportional to the height of the tree. Reducing the height of the integrity tree by increasing its arity, however, leads to frequent overflowing of the counters that are used for encryption in the tree. We will show that increasing the tree arity of a widely-use integrity tree from 2 to 8 can result in over 200% increase in memory authentication overhead for some benchmark applications, despite the reduction in tree height. In this paper, we propose DISSECT, a memory authentication framework which utilizes a dynamic memory integrity tree that can adapt to the memory access patterns of the application by progressively adjusting the tree height and arity in order to significantly reduce performance overhead. This is achieved by 1) initializing an integrity tree structure with the largest arity possible to meet the security requirements, 2) dynamically skewing the tree such that the more frequently accessed memory locations are positioned closer to the tree root (overcomes the tree height problem), and 3) dynamically splitting the tree at nodes with counters that are about to overflow (overcomes the counter overflow problem). Experimental results undertaken using Multi2Sim on benchmarks from SPEC-CPU2006, SPLASH-2, and PARSEC demonstrate the performance benefits of our proposed memory integrity tree.
Saru Vig, Rohan Juneja, Siew-Kei Lam
DATE3
2020 Dynamically Growing Neural Network Architecture for Lifelong Deep Learning on the Edge
abstract
Conventional deep learning models are trained once and deployed. However, models deployed in agents operating in dynamic environments need to constantly acquire new knowledge, while preventing catastrophic forgetting of previous knowledge. This ability is commonly referred to as lifelong learning. In this paper, we address the performance and resource challenges for realizing lifelong learning on edge devices. We propose a FPGA based architecture for a Self-Organization Neural Network (SONN), that in combination with a Convolutional Neural Network (CNN) can perform class-incremental lifelong learning for object classification. The proposed SONN architecture is capable of performing unsupervised learning on input features from the CNN by dynamically growing neurons and connections. In order to meet the tight constraints of edge computing, we introduce efficient scheduling methods to maximize resource reuse and parallelism, as well as approximate computing strategies. Experiments based on the Core50 dataset for continuous object recognition from video sequences demonstrated that the proposed FPGA architecture significantly outperforms CPU and GPU based implementations.
Duvindu Piyasena, Miyuru Thathsara, Sathursan Kanagarajah, Siew-Kei Lam, Meiqing Wu
FPL4
2020 Evaluating the Merits of Ranking in Structured Network Pruning
abstract
Pruning of channels in trained deep neural networks has been widely used to implement efficient DNNs that can be deployed on embedded/mobile devices. Majority of existing techniques employ criteria-based sorting of the channels to preserve salient channels during pruning as well as to automatically determine the pruned network architecture. However, recent studies on widely used DNNs, such as VGG-16, have shown that selecting and preserving salient channels using pruning criteria is not necessary since the plasticity of the network allows the accuracy to be recovered through fine-tuning. In this work, we further explore the value of the ranking criteria in pruning to show that if channels are removed gradually and iteratively, alternating with fine-tuning on the target dataset, ranking criteria are indeed not necessary to select redundant channels. Experimental results confirm that even a random selection of channels for pruning leads to similar performance (accuracy). In addition, we demonstrate that even a simple pruning technique that uniformly removes channels from all layers in the network, performs similar to existing ranking criteria-based approaches, while leading to lower inference time (GFLOPs). Our extensive evaluations include the context of embedded implementations of DNNs - specifically, on small networks such as SqueezeNet and at aggressive pruning percentages. We leverage these insights, to propose a GFLOPs-aware iterative pruning strategy that does not rely on any ranking criteria and yet can further lead to lower inference time by 15% without sacrificing accuracy.
Nirmala Ramakrishnan, Alok Prakash, Siew-Kei Lam, Thambipillai Srikanthan
ICDCS4
2020 Self-Growing Spatial Graph Networks for Pedestrian Trajectory Prediction
abstract
Intelligent vehicles and social robots need to navigate in crowded environments while avoiding collisions with pedestrians. To achieve this, pedestrian trajectory prediction is essential. However, predicting pedestrians' trajectory in crowded environments is nontrivial as human-to-human interactions among the crowd participants influence their motion. In this work, we propose a novel end-to-end graph-centric gated learning model to estimate the existence of interactions between individuals. Accordingly, the model predicts pedestrians' future locations and velocities. Recent methods based on LSTM networks used thresholding techniques to define neighborhood boundaries and relationships. Other graph-structured methods grow edges in polynomial size. In contrast, our graph-based GRU network model employs an online data-driven criterion that can learn from interactions and grow connections between pedestrian nodes. The proposed model yields outperforming prediction accuracy over state-of-the-art works in two public datasets, i.e. Crowds and SDD.
Sirin Haddad, Siew-Kei Lam
WACV2
2020 Fusing Semantics and Motion State Detection for Robust Visual SLAM
abstract
Achieving robust pose tracking and mapping in highly dynamic environments is a major challenge faced by existing visual SLAM (vSLAM) systems. In this paper, we increase the robustness of existing vSLAM by accurately removing moving objects from the scene so that they will not contribute to pose estimation and mapping. Specifically, semantic information is fused with motion states of the scene via a probability framework to enable accurate and robust moving object extraction in order to retain the useful features for pose estimation and mapping. Our work highlights the importance of distinguishing between motion states of potential moving objects for vSLAM in highly dynamic environments. The proposed method can be integrated into existing vSLAM systems to increase their robustness in dynamic environments without incurring much computation cost. We provide extensive experimental results on three well-known datasets to show that the proposed technique outperforms existing vSLAM methods in indoor and outdoor environments, under various scenarios such as crowded scenes.
Gaurav Singh 0009, Meiqing Wu, Siew-Kei Lam
WACV3
2020 Multiple-Choice Hardware/Software Partitioning for Tree Task-Graph on MPSoC
abstract
Abstract Hardware/software (HW/SW) partitioning, that decides which components of an application are implemented in hardware and which ones in software, is a crucial step in embedded system design. On modern heterogeneous embedded system platform, each component of application can typically have multiple feasible configurations/implementations, trading off quality aspects (e.g. energy consumption, completion time) with usage for various types of resources. This provides new opportunities for further improving the overall system performance, but few works explore the potential opportunity by incorporating the multiple choices of hardware implementation in the partitioning process. This paper proposes three algorithms for multiple-choice HW/SW partitioning of tree-shape task graph on multiple processors system on chip (MPSoC) with the objective of minimizing execution time, while meeting area constraint. Firstly, an efficient heuristic algorithm is proposed to rapidly generate an approximate solution. The obtained solution produced by the first algorithm is then further refined by a customized Tabu search algorithm. We also propose a dynamic programming algorithm to calculate the exact solutions for relatively smaller scale instances. Simulation results show that the proposed heuristic algorithm is able to quickly generate good approximate solutions, and the solutions become very close to the exact solutions after refined by the proposed Tabu search algorithm, in comparison to the exact solutions produced by the dynamic programming algorithm.
Wenjun Shi, Jigang Wu, Guiyuan Jiang, Siew-Kei Lam
Comput. J.4
2020 Learning heterogeneous traffic patterns for travel time prediction of bus journeys
Peilan He, Guiyuan Jiang, Siew-Kei Lam
Inf. Sci.3
2020 A Lightweight Detection Algorithm For Collision-Optimized Divide-and-Conquer Attacks
abstract
By introducing collision information into divide-and-conquer attacks, several existing works transform the original candidate space, which may be too large to enumerate, into a significantly smaller collision space, thus making key recovery possible. However, the inefficient collision detection algorithms and fault tolerance mechanisms make them time-consuming and their success rate low. Moreover, they may still leave very huge chain spaces that makes it difficult for key recovery. In this article, we exploit collision attack to optimize Template Attack (TA), and propose a Lightweight Collision Detection (LCD) algorithm. The proposed method exploits a jump detection mechanism to efficiently reduce the repetitive collision detections on chains with the same prefix sub-chains. We then introduce guessing theory to reorder the collision detection of the sub-keys according to their guessing lengths, and provide us with an evaluation tool. Finally, we design a highly efficient fault tolerance mechanism for our LCD to allow flexible thresholds adjustment, and further optimize sieving mechanism to efficiently extract the best chains with the largest number of collisions. Experimental results fully demonstrate LCD's superiority.
Changhai Ou, Siew-Kei Lam, Chengju Zhou, Guiyuan Jiang, Fan Zhang 0010
IEEE Trans. Computers2
2020 A First Study of Compressive Sensing for Side-Channel Leakage Sampling
abstract
An important prerequisite for side-channel attacks (SCAs) is leakage sampling where the side-channel measurements (i.e., power traces) of the cryptographic device are collected for further analysis. However, as the operating frequency of cryptographic devices continues to increase due to advancing technology, leakage sampling will impose higher requirements on the sampling rate and storage capacity of the sampling equipment. This article undertakes the first study to show that effective leakage sampling can be achieved without relying on sophisticated equipments through compressive sensing (CS). As long as the information is leaked in the low-frequency component, CS can obtain low-dimensional samples by simply projecting the high-dimensional signals onto the observation matrix. The power traces can then be reconstructed in a workstation for further analysis and storage. With this approach, the sampling rate to obtain power traces is no longer limited by the operating frequency of the cryptographic device and the Nyquist sampling theorem. Instead, it depends on the sparsity of the leakage signal. As such, CS can employ a much lower sampling rate and yet obtain equivalent leakage sampling performance, which significantly lowers the requirement of sampling equipments. The feasibility of our approach is verified theoretically and through experiments.
Changhai Ou, Chengju Zhou, Siew-Kei Lam
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2020 Hardware Performance Counter-Based Fine-Grained Malware Detection
abstract
Detection of malicious programs using hardware-based features has gained prominence recently. The tamper-resistant hardware metrics prove to be a better security feature than the high-level software metrics, which can be easily obfuscated. Hardware Performance Counters (HPC), which are inbuilt in most of the recent processors, are often the choice of researchers amongst hardware metrics. However, a lack of determinism in their counts, thereby affecting the malware detection rate, minimizes the advantages of HPCs. To overcome this problem, in our work, we propose a three-step methodology for fine-grained malware detection. In the first step, we extract the HPCs of each system call of an unknown program. Later, we make a dimensionality reduction of the fine-grained data to identify the components that have maximum variance. Finally, we use a machine learning based approach to classify the nature of the unknown program into benign or malicious. Our proposed methodology has obtained a 98.4% detection rate, with a 3.1% false positive. It has improved the detection rate significantly when compared to other recent works in hardware-based anomaly detection.
Sai Praveen Kadiyala, Pranav Jadhav, Siew-Kei Lam, Thambipillai Srikanthan
ACM Trans. Embed. Comput. Syst.3
2020 Peak-Hour Vehicle Routing for First-Mile Transportation: Problem Formulation and Algorithms
abstract
The first-mile transportation provides a transit service using ridesharing-based vehicles, e.g., feeder buses, for passengers to travel from their homes, workplaces, or public institutions to the nearest public transportation depots (rapid-transit metro or appropriated bus stations) which are located beyond comfortable walking distance. This paper studies the vehicle routing problem (VRP) for the first-mile transportation, which aims at finding the optimal travel routes for a vehicle fleet to deliver passengers from their doorstep to the depots, where the passengers can continue their journeys using fixed-route buses or trains. We focus on the Peak-Hour VRP (PHVRP) for a limited vehicle fleet capacity to serve a large volume of travel requests, with the aim of maximizing the number of served passengers. The PHVRP generalizes the VRP with time window by considering multiple alternative depots for each travel request, such that a request is satisfied if the passenger is taken to one of his/her nearest depots. We formally formulate the PHVRP with constraints on vehicle capacity, pickup time windows, and quality of service regarding riding time, where a novel trip-based constraint model is used. We proposed an ant-colony optimization algorithm for the PHVRP, which is initialized with pheromone information that jointly considers the temporal-spatial distance as well as depot similarity among different travel requests. We introduced a novel scheme (called trip-by-trip scheme) to construct the travel routes by repeatedly forming a single trip for the vehicle with earliest end time until no vehicle can accept any more trips. In constructing a single trip, the algorithm intelligently decides whether or not to end the trip instead of taking more passengers. The effectiveness of the proposed methods is evaluated by comparing with optimal solutions on small size instances and with heuristic solutions on large-size instances, using road network in Singapore and synthetic travel requests that are generated based on real bus travel demands.
Guiyuan Jiang, Siew-Kei Lam, Fangxin Ning, Peilan He, Jidong Xie
IEEE Trans. Intell. Transp. Syst.2
2020 Group Cost-Sensitive BoostLR With Vector Form Decorrelated Filters for Pedestrian Detection
abstract
Pedestrian detection has achieved notable progress in the field of computer vision over the past decade. However, existing top-performing approaches suffer from high computational complexity which prohibits their realization on embedded platforms with low computational capabilities. In this paper, we propose a robust and fast pedestrian detection framework which is based on the Filtered Channel Feature (FCF) approach. The proposed framework exploits vector-form decorrelated filters to extract more discriminative channel features while benefiting from low computational complexity. A novel group cost-sensitive BoostLR (Boosting with Loss Regularization) algorithm is proposed to train the classifier. The proposed training strategy provides more emphasis to the harder samples by exploring the variations of negatives selected from different rounds in hard negative mining processing, and hence is able to boost the overall detection performance. In addition, the proposed method also benefits from the BoostLR framework to achieve better generalization. Experiments on the well-known Caltech, INRIA and CityPersons pedestrian detection datasets show that our proposed approach achieves the best detection performance among all of the state-of-the-art non-deep learning methods and can run one order of magnitude faster than classical FCF methods (e.g. Checkerboards).
Chengju Zhou, Meiqing Wu, Siew-Kei Lam
IEEE Trans. Intell. Transp. Syst.3
2020 Designing Energy-Efficient MPSoC with Untrustworthy 3PIP Cores
abstract
The adoption of large-scale MPSoCs and the globalization of the IC design flow give rise to two major concerns: high power density due to continuous technology scaling and security due to the untrustworthiness of the third-party intellectual property (3PIP) cores. However, little work has been undertaken to consider these two critical issues jointly during the design stage. In this paper, we propose a design methodology that minimizes the energy consumption while simultaneously protecting the MPSoC against the effects of hardware trojans. The proposed methodology consists of three main stages: 1) Task scheduling to introduce core diversity in the MPSoC in order to detect the presence of malicious modifications in the cores, or mute their effects at runtime, 2) Vendor assignment to the cores using a novel heuristic that chooses vendor-specific cores with operating speed that minimizes the total energy consumption of the MPSoC, and 3) Explore optimization opportunities for further energy savings by minimizing idle periods on the cores, which are caused by the inter-task data dependencies. Experimental results show that our solutions consume only 1/3 energy of existing solutions without increasing schedule length while satisfying the security constraints.
Guiyuan Jiang, Siew-Kei Lam, Fangxin Ning
IEEE Trans. Parallel Distributed Syst.3
2019 TAD: time side-channel attack defense of obfuscated source code
abstract
Program obfuscation is widely used to protect commercial software against reverse-engineering. However, an adversary can still download, disassemble and analyze binaries of the obfuscated code executed on an embedded System-on-Chip (SoC), and by correlating execution times to input values, extract secret information from the program. In this paper, we show (1) the impact of widely-used obfuscation methods on timing leakage, and (2) that well-known software countermeasures to reduce timing leakage of programs, are not always effective for low-noise environments found in embedded systems. We propose two methods for mitigating timing leakage in obfuscated codes. The first is a compiler driven method, called TAD, which removes conditional branches with distinguishable execution times for an input program. In the second method (TADCI), TAD is combined with dynamic hardware diversity by replacing primitive instructions with Custom Instructions (CIs) that exhibit non-deterministic execution times at runtime. Experimental results on the RISC-V platform show that the information leakage is reduced by 92% and 82% when TADCI is applied to the original and obfuscated source code, respectively.
Alexander Fell, Thinh Hung Pham, Siew-Kei Lam
ASP-DAC3
2019 Reducing Dynamic Power in Streaming CNN Hardware Accelerators by Exploiting Computational Redundancies
abstract
Convolutional neural networks (CNNs) have achieved tremendous successes in various application domains such as computer vision. However, current implementations are characterized by large memory requirements and accesses, which pose an impediment towards their deployment on low cost embedded devices with fast runtime requirements. Recently, FPGA based streaming CNN hardware accelerators have been reported for alleviating these memory bottlenecks. However, these implementations suffer from large number of convolution operations which incur high power consumption. In this paper, we investigate methods to exploit the redundancies in the activation layers in order to reduce the dynamic power. We propose a computationally efficient approximation method to reduce the overall convolution operations with marginal accuracy loss. Experimental results of our FPGA implementation based on image classification datasets show that the proposed method leads to considerable power savings.
Duvindu Piyasena, Rukshan Wickramasinghe, Debdeep Paul, Siew-Kei Lam, Meiqing Wu
FPL4
2019 Lowering Dynamic Power of a Stream-based CNN Hardware Accelerator
abstract
Custom hardware accelerators of Convolutional Neural Networks (CNN) provide a promising solution to meet real-time constraints for a wide range of applications on low-cost embedded devices. In this work, we aim to lower the dynamic power of a stream-based CNN hardware accelerator by reducing the computational redundancies in the CNN layers. In particular, we investigate the redundancies due to the downsampling effect of max pooling layers which are prevalent in state-of-the-art CNNs, and propose an approximation method to reduce the overall computations. The experimental results show that the proposed method leads to lower dynamic power without sacrificing accuracy.
Duvindu Piyasena, Rukshan Wickramasinghe, Debdeep Paul, Siew-Kei Lam, Meiqing Wu
MMSP4
2019 Bus Travel Speed Prediction using Attention Network of Heterogeneous Correlation Features
abstract
Accurate bus travel speed prediction can lead to improved urban mobility by enabling passengers to reliably plan their trips in advance and traffic administrators to manage the bus operations more effectively. However, the increasing complexity of public transportation networks pose a significant challenge to existing prediction methods as the bus operations are affected by numerous factors such as varying traffic conditions, tight bus operation schedules, wide-ranging travel demands, frequent accelerations/decelerations at bus stops, delays at intersections, etc. This paper aims to achieve accurate bus speed prediction by identifying important intrinsic and extrinsic features that impact the bus speed, and their significance in specific situations. We propose to jointly incorporate multiple feature components that provide discriminating information to train the prediction model by exploring the spatial correlation, temporal correlation, as well as contextual information (e.g. road characteristics and weather conditions). In particular, we introduce an attribute-driven attention network model to integrate the feature components, which considers the heterogeneous influence of different feature components on bus speed and dynamically assigns weights to the learned latent features based on specific traffic situations. Extensive experiments using real bus travel data involving 42 bus services show that our proposed method outperforms six well-known methods.
Guiyuan Jiang, Siew-Kei Lam, Shicheng Chen, Peilan He
SDM3
2019 Collaborative Task Offloading with Computation Result Reusing for Mobile Edge Computing
abstract
Abstract The task offloading problem, which aims to balance the energy consumption and latency for Mobile Edge Computing (MEC), is still a challenging problem due to the dynamic changing system environment. To reduce energy while guaranteeing delay constraint for mobile applications, we propose an access control management architecture for 5G heterogeneous network by making full use of Base Station’s storage capability and reusing repetitive computational resource for tasks. For applications that rely on real-time information, we propose two algorithms to offload tasks with consideration of both energy efficiency and computation time constraint. For the first scenario, i.e. the rarely changing system environment, an optimal static algorithm is proposed based on dynamic programming technique to get the exact solution. For the second scenario, i.e. the frequently changing system environment, a two-stage online algorithm is proposed to adaptively obtain the current optimal solution in real time. Simulation results demonstrate that the exact algorithm in the first scenario runs 4 times faster than the enumeration method. In the second scenario, the proposed online algorithm can reduce the energy consumption and computation time violation rate by 16.3% and 25% in comparison with existing methods.
Zikai Zhang 0004, Jigang Wu, Long Chen 0006, Guiyuan Jiang, Siew-Kei Lam
Comput. J.5
2019 Travel-Time Prediction of Bus Journey With Multiple Bus Trips
abstract
Accurate travel-time prediction of public transport is essential for reliable journey planning in urban transportation systems. However, existing studies on bus travel-/arrival-time prediction often focus only on improving the prediction accuracy of a single bus trip. This is inadequate in modern public transportation systems, where a bus journey usually consists of multiple bus trips. In this paper, we investigate the problem of travel-time prediction for bus journeys that takes into account a passenger's riding time on multiple bus trips, and also his/her waiting time at transfer points (interchange stations or bus stops). A novel framework is proposed to separately predict the riding and waiting time of a given journey from multiple datasets (i.e., historical bus trajectories, bus route, and road network), and combining the results to form the final travel-time prediction. We empirically determine the impact factors of bus riding times and develop a long short-term memory model that can accurately predict the riding time of each segment of the bus lines/routes. We also demonstrate that the waiting time at transfer points significantly impacts the total journey travel time, and estimating the waiting time is non-trivial as we cannot assume a fixed distribution waiting time. In order to accurately predict the waiting time, we introduce a novel interval-based historical average method that can efficiently address the correlation and sensitivity issues in waiting time prediction. Experiments on real-world data show that the proposed method notably outperforms six baseline approaches for all the scenarios considered.
Peilan He, Guiyuan Jiang, Siew-Kei Lam, Dehua Tang
IEEE Trans. Intell. Transp. Syst.3
2019 High-Throughput and Area-Optimized Architecture for rBRIEF Feature Extraction
abstract
Feature matching is a fundamental step in many real-time computer vision applications such as simultaneous localization and mapping, motion analysis, and stereo correspondence. The performance of these applications depends on the distinctiveness of the visual feature descriptors used, and the speed at which they can be extracted from video frames. When combined with standard key-point detectors, the rotation-aware binary robust independent elementary features (rBRIEF) descriptor has been shown to outperform its counterparts. In this paper, we present a deep-pipelined stream processing architecture that is capable of extracting rBRIEF features from high-throughput video frames. To achieve high processing rate and low complexity hardware, the proposed architecture incorporates an enhanced moving summation strategy to calculate the key-points' patch moments and employs approximate computations to achieve patch rotation. Multiplier-less circuitry is introduced throughout the architecture to avoid the use of costly multipliers. Implementation on the Altera Aria V device demonstrates that the proposed architecture leads to 53.3% reduction in hardware resources (adaptive logic modules), while achieving 50% higher accuracy (in terms of average Hamming distance) when compared to the state-of-the-art architecture. In addition, the proposed architecture is able to process high-resolution ($1920 \times 1080$ ) images at 60 fps, while consuming only 456.15 mW power.
Thinh Hung Pham, Siew-Kei Lam
IEEE Trans. Very Large Scale Integr. Syst.3
2019 Framework for Fast Memory Authentication Using Dynamically Skewed Integrity Tree
abstract
Integrity trees are widely used in computer systems to prevent replay, splicing, and spoofing attacks on memories. Such mechanisms incur excessive performance and energy overhead. We propose a memory authentication framework that combines architecture-specific optimizations of the integrity tree with mechanisms that enable it to restructure at runtime based on memory access patterns. The integrity tree structure is customized based on the cache configuration in order to minimize the performance and energy overhead through speculative authentication. At runtime, the tree nodes that are accessed more frequently will be dynamically shifted closer to the root such that fewer levels of the tree are accessed during authentication. The framework is simulated with Multi2Sim and compared with other existing mechanisms [i.e., tamper-evident counter (TEC) tree and ASSURE] to demonstrate its performance and energy benefits. Experimental results using benchmarks from SPEC-CPU2006, SPLASH-2, and PARSEC show that the proposed dynamic integrity tree leads to an average reduction in instruction per cycle of 13% and 10% over TEC tree and ASSURE, respectively. The corresponding average reduction in authentication time is 30% and 20%, respectively. We show that the proposed framework facilitates the selection of a processor with a smaller cache size such that the energy consumption is reduced without sacrificing performance.
Saru Vig, Rohan Juneja, Guiyuan Jiang, Siew-Kei Lam, Changhai Ou
IEEE Trans. Very Large Scale Integr. Syst.4
2019 Efficient three-stage auction schemes for cloudlets deployment in wireless access network
Gangqiang Zhou, Jigang Wu, Long Chen 0006, Guiyuan Jiang, Siew-Kei Lam
Wirel. Networks5
2018 HiMap: A hierarchical mapping approach for enhancing lifetime reliability of dark silicon manycore systems
abstract
Technology scaling into the nano-scale CMOS regime has resulted in increased leakage and roadblock on voltage scaling, which has led to several issues like high power density and elevated on-chip temperature. This consequently aggravates device aging, compromising lifetime reliability of the manycore systems. This paper proposes HiMap, a dynamic hierarchical mapping approach to maximize lifetime reliability of manycore systems while satisfying performance, power, and thermal constraints. HiMap is process variation- and aging-aware. It comprises of two levels: (1) it identifies a region of cores suitable for mapping, and (2) it maps threads in the region and intersperses dark cores for thermal mitigation while considering the current health of the cores. Both the levels strive to reduce aging variance across the chip. We evaluated HiMap for 64-core and 256-core systems. Results demonstrate an improved system lifetime reliability by up to 2 years at the end of 3.25 years of use, as compared to the state-of-the-art.
Vijeta Rathore, Vivek Chaturvedi, Amit Kumar Singh 0002, Thambipillai Srikanthan, R. Rohith, Siew-Kei Lam, Muhammad Shafique 0001
DATE6
2018 Dynamic skewed tree for fast memory integrity verification
abstract
Memory authentication techniques often employ an integrity tree as a countermeasure against replay, spoofing and splicing attacks. However, the balanced memory integrity trees used in existing approaches lead to excessive memory access overheads for runtime verification. In this paper, we propose a framework to dynamically construct a customized integrity tree based on the data access patterns to reduce the overhead of runtime verification. The proposed framework can adapt the memory integrity tree structure at runtime such that the nodes that correspond to frequently accessed data are placed closer to the root. We validated the effectiveness of our approach on the Altera NIOS II processor with an external DRAM. Experimental results based on applications from widely used CHStone and SNU Real-Time benchmarks demonstrate that the proposed approach can lead to an average performance gain of 30% compared to the conventional means of using balanced memory integrity trees. In addition, to preserve data confidentiality, we implemented the encryption/decryption operations using custom instructions on the NIOS II processor to notably reduce the overall overhead of memory security.
Saru Vig, Guiyuan Jiang, Siew-Kei Lam
DATE3
2018 CIDPro: Custom Instructions for Dynamic Program Diversification
abstract
Timing side-channel attacks pose a major threat to embedded systems due to their ease of accessibility. We propose CIDPro, a framework that relies on dynamic program diversification to mitigate timing side-channel leakage. The proposed framework integrates the widely used LLVM compiler infrastructure and the increasingly popular RISC-V FPGA soft-processor. The compiler automatically generates custom instructions in the security critical segments of the program, and the instructions execute on the RISC-V custom co-processor to produce diversified timing characteristics on each execution instance. CIDPro has been implemented on the Zynq7000 XC7Z020 FPGA device to study the performance overhead and security tradeoffs. Experimental results show that our solution can achieve 80% and 86% timing side-channel capacity reduction for two benchmarks with an acceptable performance overhead compared to existing solutions. In addition, the proposed method incurs only a negligible hardware area overhead of 1% slices of the entire RISC-V system.
Thinh Hung Pham, Alexander Fell, Arnab Kumar Biswas, Siew-Kei Lam, Nandeesha Veeranna
FPL4
2018 Stream-Based ORB Feature Extractor with Dynamic Power Optimization
abstract
The Oriented Fast and Rotated BRIEF (ORB) feature extractor, which consists of key-point detection and descriptor computation, is a key module in many computer vision systems. Existing hardware implementations of ORB feature extractor only focus on increasing performance with power optimization as a post consideration. In this paper, we present a stream-based ORB feature extractor that incorporates mechanisms to lower the dynamic power consumption. These mechanisms exploit the fact that the number of detected keypoints is typically small. The proposed solution significantly lowers the switching activity of the key-point detection and descriptor computation stages by early pruning of non-likely key-points and gating the descriptor computation stages. Further power reduction and resource minimization are achieved by employing a threshold-guided bit-width optimization strategy to truncate the redundant bits in the key-point detection stage. Finally, we propose an approximation method to achieve rotation invariance of the descriptors. FPGA implementation targeting the Altera Aria V device shows that the proposed strategies lead to over 25% reduction in dynamic power and lower resource utilization, with only marginal loss in accuracy.
Thinh Hung Pham, Siew-Kei Lam, Meiqing Wu, Bhavan A. Jasani
FPT3
2018 Algorithms for Replica Placement and Update in Tree Network
abstract
A critical issue in data replication is to wisely place data replicas which involves identifying the best possible nodes to duplicate data. Facing dynamics of data requests, this paper investigates the problem of replica placement and update in tree networks, where part of nodes have pre-existing replicas. We aim to develop efficient algorithms to accelerate the replica placement and update without causing obvious degradation in solution quality via reusing pre-existing replicas. Firstly, an efficient heuristic algorithm GRP is proposed to quickly place replicas when users change their requests dynamically, under the Closest policy where a client must be served by the closest server. Then, a Tabu search algorithm TSRP is customized to further refine the solution obtained by GRP. Furthermore, we propose a heuristic algorithm MPFSF for the replica placement and update problem, under the Multiple policy where requests of a client are served by multiple servers. Simulation results show that, GRP and TSRP can accelerate existing dynamic programming algorithm by 87.97% while quality degradation is bounded by 2.49%. MPFSF can achieve the best improvement for about 84.6% than existing heuristic algorithm.
Jigang Wu, Long Chen 0006, Guiyuan Jiang, Siew-Kei Lam, Thambipillai Srikanthan
Comput. J.5
2018 Efficient hybrid multicast approach in wireless data center network
Longting Zhu, Jigang Wu, Guiyuan Jiang, Long Chen 0006, Siew-Kei Lam
Future Gener. Comput. Syst.5
2018 Threshold-Guided Design and Optimization for Harris Corner Detector Architecture
abstract
High-speed corner detection is an essential step in many real-time computer vision applications, e.g., object recognition, motion analysis, and stereo matching. Hardware implementation of corner detection algorithms, such as the Harris corner detector (HCD) has become a viable solution for meeting real-time requirements of the applications. A major challenge lies in the design of power, energy and area efficient architectures that can be deployed in tightly constrained embedded systems while still meeting real-time requirements. In this paper, we proposed a bit-width optimization strategy for designing hardware-efficient HCD that exploits the thresholding step in the algorithm to determine interest points from the corner responses. The proposed strategy relies on the threshold as a guide to truncate the bit-widths of the operators at various stages of the HCD pipeline with only marginal loss of accuracy. Synthesis results based on 65-nm CMOS technology show that the proposed strategy leads to power-delay reduction of 35.2%, and area reduction of 35.4% over the baseline implementation. In addition, through careful retiming, the proposed implementation achieves over 2.2 times increase in maximum frequency while achieving an area reduction of 35.1% and power-delay reduction of 35.7% over the baseline implementation. Finally, we performed repeatability tests to show that the optimized HCD architecture achieves comparable accuracy with the baseline implementation (average decrease of repeatability is less than 0.6%).
Bhavan A. Jasani, Siew-Kei Lam, Pramod Kumar Meher, Meiqing Wu
IEEE Trans. Circuits Syst. Video Technol.2
2018 Rapid Memory-Aware Selection of Hardware Accelerators in Programmable SoC Design
abstract
Programmable Systems-on-Chips (SoCs) are expected to incorporate a larger number of application-specific hardware accelerators with tightly integrated memories in order to meet stringent performance-power requirements of embedded systems. As data sharing between the accelerator memories and the processor is inevitable, it is of paramount importance that the selection of application segments for hardware acceleration must be undertaken such that the communication overhead of data transfers do not impede the advantages of the accelerators. In this paper, we propose a novel memory-aware selection algorithm that is based on an iterative approach to rapidly recommend a set of hardware accelerators that will provide high performance gain under varying area constraint. In order to significantly reduce the algorithm runtime while still guaranteeing near-optimal solutions, we propose a heuristic to estimate the penalties incurred when the processor accesses the accelerator memories. In each iteration of the proposed algorithm, a two-pass method is employed where a set of good hardware accelerator candidates is selected using a greedy approach in the first pass, and a “sliding window” approach is used in the second pass to refine the solution. The two-pass method is iteratively performed on a bounded set of candidate hardware accelerators to limit the search space and to avoid local maxima. In order to validate the benefits of the proposed selection algorithm, an exhaustive search algorithm is also developed. Experimental results using the popular CHStone benchmark suite show that the performance achieved by the accelerators recommended by the proposed algorithm closely matches the performance of the exhaustive algorithm, with close to 99% accuracy, while being orders of magnitude faster.
Alok Prakash, Christopher T. Clarke, Siew-Kei Lam, Thambipillai Srikanthan
IEEE Trans. Very Large Scale Integr. Syst.3
2017 Group Cost-sensitive Boosting with Multi-scale Decorrelated Filters for Pedestrian Detection
Chengju Zhou, Meiqing Wu, Siew-Kei Lam
BMVC3
2017 Lowering dynamic power in stream-based harris corner detection architecture
abstract
Stream-based image processing architectures are extremely attractive as they can achieve high throughput and do not require external memories for storing input video frames. However, a major challenge in designing stream-based architectures lies in lowering the dynamic power consumption since all the processing elements are typically in continuous operation to keep up with the rate of incoming pixel streams. In this work, we show that the dynamic power of the stream-based Harris corner detector (HCD) can be reduced by inhibiting redundant signal activity in the complex calculations of the corner scores. Specifically, we perform simple approximations to detect non-likely corners at the early stages of the pipeline for putting the subsequent pipeline stages into a dormant state. Synthesis results on the Altera Cyclone V FPGA show that the proposed strategy leads to an average dynamic power reduction of over 10% compared to the conventional implementation with similar performance and negligible increase in resources. In addition, we performed repeatability tests to show that the proposed stream-based HCD architecture achieves comparable accuracy with the conventional implementation.
Siew-Kei Lam, Rakesh Kumar Bijarniya, Meiqing Wu
FPT1
2017 QoE-Aware Task Offloading for Time Constraint Mobile Applications
abstract
In this paper, we develop an access controller management model which provides new opportunities for further reducing the computation repetition and data transmission redundancy for Mobile Edge Computing (MEC) in 5G network. We propose novel algorithms for solving the offloading problem with consideration of tradeoff between energy consumption and the amount of offloaded data under constraint of overall task computation time. For sequential topology applications, we develop a dynamic programming algorithm to produce optimal solutions. For general topology applications, a critical-path based heuristic algorithm is proposed by repeatedly identifying partial critical path (PCP) from the application task graph and calculating optimal solution for the PCP by performing the proposed dynamic programming algorithm. In addition, the interference of parallel data transmission between tasks (one-to-many, manyto-one and many-to-many) using single channel is taken into consideration. Experimental results demonstrate the effectiveness of our proposed method.
Zikai Zhang 0004, Jigang Wu, Guiyuan Jiang, Long Chen 0006, Siew-Kei Lam
LCN5
2017 Fast and Accurate Pedestrian Detection using Dual-Stage Group Cost-Sensitive RealBoost with Vector Form Filters
abstract
Despite significant research efforts in pedestrian detection over the past decade, there is still a ten-fold performance gap between the state-of-the-art methods and human perception. Deep learning methods can provide good performance but suffers from high computational complexity which prohibits their deployment on affordable systems with limited computational resources. In this paper, we propose a pedestrian detection framework that provides a major fillip to the robustness and run-time efficiency of the recent top performing non-deep learning Filtered Channel Feature (FCF) approach. The proposed framework overcomes the computational bottleneck of existing FCF methods by exploiting vector form filters to efficiently extract more discriminative channel features for pedestrian detection. A novel dual-stage group cost-sensitive RealBoost algorithm is used to explore different costs among different types of misclassification in the boosting process in order to improve detection performance. In addition, we propose two strategies, selective classification and selective scale processing, to further accelerate the detection process at the channel feature level and image pyramid level respectively. Experiments on the Caltech and INRIA datasets show that the proposed method achieves the highest detection performance among all the state-of-the-art non-CNN methods and is about 148X faster than the existing best performing FCF method on the Caltech dataset.
Chengju Zhou, Meiqing Wu, Siew-Kei Lam
ACM Multimedia3
2017 A Framework for Fast and Robust Visual Odometry
abstract
Knowledge of the ego-vehicle's motion state is essential for assessing the collision risk in advanced driver assistance systems or autonomous driving. Vision-based methods for estimating the ego-motion of vehicle, i.e., visual odometry, face a number of challenges in uncontrolled realistic urban environments. Existing solutions fail to achieve a good tradeoff between high accuracy and low computational complexity. In this paper, a framework for ego-motion estimation that integrates runtime-efficient strategies with robust techniques at various core stages in visual odometry is proposed. First, a pruning method is employed to reduce the computational complexity of Kanade-Lucas-Tomasi (KLT) feature detection without compromising on the quality of the features. Next, three strategies, i.e., smooth motion constraint, adaptive integration window technique, and automatic tracking failure detection scheme, are introduced into the conventional KLT tracker to facilitate generation of feature correspondences in a robust and runtime efficient way. Finally, an early termination condition for the random sample consensus (RANSAC) algorithm is integrated with the Gauss-Newton optimization scheme to enable rapid convergence of the motion estimation process while achieving robustness. Experimental results based on the KITTI odometry data set show that the proposed technique outperforms the state-of-the-art visual odometry methods by producing more accurate ego-motion estimation in notably lesser amount of time.
Meiqing Wu, Siew-Kei Lam, Thambipillai Srikanthan
IEEE Trans. Intell. Transp. Syst.2
2017 Joint Charging Tour Planning and Depot Positioning for Wireless Sensor Networks Using Mobile Chargers
abstract
Recent breakthrough in wireless energy transfer technology has enabled wireless sensor networks (WSNs) to operate with zero-downtime through the use of mobile energy chargers (MCs), that periodically replenish the energy supply of the sensor nodes. Due to the limited battery capacity of the MCs, a significant number of MCs and charging depots are required to guarantee perpetual operations in large scale networks. Existing methods for reducing the number of MCs and charging depots treat the charging tour planning and depot positioning problems separately even though they are inter-dependent. This paper is the first to jointly consider charging tour planning and MC depot positioning for large-scale WSNs. The proposed method solves the problem through the following three stages: charging tour planning, candidate depot identification and reduction, and depot deployment and charging tour assignment. The proposed charging scheme also considers the association between the MC charging cycle and the operational lifetime of the sensor nodes, in order to maximize the energy efficiency of the MCs. This overcomes the limitations of existing approaches, wherein MCs with small battery capacity ends up charging sensor nodes more frequently than necessary, while MCs with large battery capacity return to the depots to replenish themselves before they have fully transferred their energy to the sensor nodes. Compared with existing approaches, the proposed method leads to an average reduction in the number of MCs by 64%, and an average increase of 19.7 times on the ratio of total charging time over total traveling time.
Guiyuan Jiang, Siew-Kei Lam, Lijia Tu, Jigang Wu
IEEE/ACM Trans. Netw.2
2016 Bounded iterative thresholding for lumen region detection in endoscopic images
abstract
The development of a fully automated robotic endoscopic steering system has been an active area of research for more than a decade. This paper aims at proposing a hardware-efficient iterative thresholding strategy to locate the lumen region in captured endoscopic images in order to enhance traditional endoscopes with certain degree of autonomy and intelligence. The proposed method is characterized by a definite requirement on the number of iterations of thresholding in order to detect the lumen region. The proposed algorithm has been demonstrated to be robust against varying characteristics using real endoscopic sample images. The reduction in the number of operations required by the proposed method can be up to 71% compared to a previously reported method. FPGA synthesis results of the proposed approach confirm its viability for real-time realization.
Pon Nidhya Elango, Siew-Kei Lam
ICARCV2
2016 Exploiting Configuration Dependencies for Rapid Area-efficient Customization of Soft-core Processors
abstract
The large number of possible configurations in modern soft-core processors make it tedious and time consuming to select the optimal configuration for a given application. In this paper, we propose a framework for rapid area-efficient customization of soft-core processors that exploits the dependencies between the various configuration options to prune the design space. Additionally, the proposed technique relies on rapid and accurate estimation models instead of the time consuming synthesis and execution techniques proposed in the existing work. Experimental results based on hand-coded applications and applications from the popular CHStone benchmark suite show that the proposed framework can rapidly and reliably select the best processor configuration for a given application and save an average of 47.58% area over the processor with all the configuration options enabled while achieving similar performance.
Deshya Wijesundera, Alok Prakash, Siew-Kei Lam, Thambipillai Srikanthan
SCOPES3
2016 Real-time road traffic density estimation using block variance
abstract
The increasing demand for urban mobility calls for a robust real-time traffic monitoring system. In this paper we present a vision-based approach for road traffic density estimation which forms the fundamental building block of traffic monitoring systems. Existing techniques based on vehicle counting and tracking suffer from low accuracy due to sensitivity to illumination changes, occlusions, congestions etc. In addition, existing holistic-based methods cannot be implemented in real-time due to high computational complexity. In this paper we propose a block based holistic approach to estimate traffic density which does not rely on pixel based analysis, therefore significantly reducing the computational cost. The proposed method employs variance as a means for detecting the occupancy of vehicles on pre-defined blocks and incorporates a shadow elimination scheme to prevent false positives. In order to take into account varying illumination conditions, a low-complexity scheme for continuous background update is employed. Empirical evaluations on publicly available datasets demonstrate that the proposed method can achieve real-time performance and has comparable accuracy with existing high complexity holistic methods.
Kratika Garg, Siew-Kei Lam, Thambipillai Srikanthan, Vedika Agarwal
WACV2
2015 Fast Replica Placement and Update Strategies in Tree Networks
abstract
Data replication enhances data availability and thereby improves the system reliability and efficiency while reduces access latency and communication cost. A critical issue in data replication is to wisely place data replicas which involves identifying the best possible nodes to duplicate data according to user requests. In this paper, we address the problem of replica placement and update in tree networks, where some nodes of the network contain pre-existing replicas. It is obvious that reusing a pre-existing replica leads to smaller cost than creating a new replica, thus it is necessary to take full advantage of the pre-existing replicas. The only previous work that consider the same problem tries to find the optimal solution by developing a dynamic programming algorithm which runs in O(N5). However, this approach is not suitable for practical situation where the user requests change frequently. In this paper, we develop efficient algorithms to accelerate the replica placement without causing obvious degradation in solution quality. We first propose an efficient heuristic algorithm (named GreedyRP) for quickly placing replicas when users change their requests. Then a tabu search algorithm (named TSRP) is customized to further refine the solution obtained by algorithm GreedyRP. Experimental results show that, on tree networks with 600 nodes and 150 pre-existing replicas, the proposed algorithms can accelerate the previous work by 87.97% while the quality degradation is bounded by 2.49% in comparison to the optimal solution.
Jigang Wu, Guiyuan Jiang, Siew-Kei Lam, Thambipillai Srikanthan
CCGRID4
2015 Adaptive Window Strategy for High-Speed and Robust KLT Feature Tracker
Nirmala Ramakrishnan, Thambipillai Srikanthan, Siew-Kei Lam, Gauri Ravindra Tulsulkar
PSIVT3
2015 Nonparametric Technique Based High-Speed Road Surface Detection
abstract
It has been well recognized that detecting road surface in a realistic environment is a challenging problem that is also computationally intensive. Existing road surface detection methods attempt to fit the road surface into rigid models (e.g., planar, clothoid, or B-Spline), thereby restricting to road surfaces that match specific models. In addition, the curve-fitting strategies employed in such techniques incur high computational complexity, making them unsuitable for in-vehicle deployments. In this paper, we propose an efficient nonparametric road surface detection algorithm that exploits the depth cue. The proposed method relies on four intrinsic road scene attributes observed under stereo geometry and has been shown to reliably detect both planar and nonplanar road surfaces efficiently. Extensive evaluations are performed on three widely used benchmarks (i.e., enpeda, KITTI, and Daimler), encompassing many complex road scenarios. The experimental results show that the proposed algorithm significantly outperforms the well-known techniques both in terms of detection accuracy and runtime performance.
Meiqing Wu, Siew-Kei Lam, Thambipillai Srikanthan
IEEE Trans. Intell. Transp. Syst.2
2015 Algorithmic aspects of graph reduction for hardware/software partitioning
Guiyuan Jiang, Jigang Wu, Siew-Kei Lam, Thambipillai Srikanthan
J. Supercomput.3
2014 Rapid evaluation of custom instruction selection approaches with FPGA estimation
abstract
The main aim of this article is to demonstrate that a fast and accurate FPGA estimation engine is indispensable in design flows for custom instruction (template) selection. The need for a FPGA estimation engine stems from the difficulty in predicting the FPGA performance measures of selected custom instructions. We will present a FPGA estimation technique that partitions the high-level representation of custom instructions into clusters based on the structural organization of the target FPGA, while taking into account general logic synthesis principles adopted by FPGA tools. In this work, we have evaluated a widely used graph covering algorithm with various heuristics for custom instruction selection. In addition, we present an algorithm called Refined Largest Fit First (RLFF) that relies on a graph covering heuristic to select non-overlapping superset templates, which typically incorporate frequently used basic templates. The initial solution is further refined by considering overlapping templates that were ignored previously to see if their introduction could lead to higher performance. While RLFF provides the most efficient cover compared to the ILP method and other graph covering heuristics, FPGA estimation results reveals that RLFF leads to the worst performance in certain applications. It is therefore a worthy proposition to equip design flows with accurate FPGA estimation in order to rapidly determine the most profitable custom instruction approach for a given application.
Siew-Kei Lam, Thambipillai Srikanthan, Christopher T. Clarke
ACM Trans. Embed. Comput. Syst.1
2014 Parallel reconfiguration algorithms for mesh-connected processor arrays
Jigang Wu, Guiyuan Jiang, Yuze Shen, Siew-Kei Lam, Thambipillai Srikanthan
J. Supercomput.4
2014 Exploiting FPGA-Aware Merging of Custom Instructions for Runtime Reconfiguration
abstract
Runtime reconfiguration is a promising solution for reducing hardware cost in embedded systems, without compromising on performance. We present a framework that aims to increase the performance benefits of reconfigurable processors that support full or partial runtime reconfiguration. The proposed framework achieves this by: (1) providing a means for choosing suitable custom instruction selection heuristics, (2) leveraging FPGA-aware merging of custom instructions to maximize the reconfigurable logic block utilization in each configuration, and (3) incorporating a hierarchical loop partitioning strategy to reduce runtime reconfiguration overhead. We show that the performance gain can be improved by employing suitable custom instruction selection heuristics that, in turn, depend on the reconfigurable resource constraints and the merging factor (extent to which the selected custom instructions can be merged). The hierarchical loop partitioning strategy leads to an average performance gain of over 31% and 46% for full and partial runtime reconfiguration, respectively. Performance gain can be further increased to over 52% and 70% for full and partial runtime reconfiguration, respectively, by exploiting FPGA-aware merging of custom instructions.
Siew-Kei Lam, Christopher T. Clarke, Thambipillai Srikanthan
ACM Trans. Reconfigurable Technol. Syst.1
2013 Modelling communication overhead for accessing local memories in hardware accelerators
abstract
Local memories increase the efficiency of hardware accelerators by enabling fast accesses to frequently used data. In addition, the access latencies of local memories are deterministic which allows for more accurate evaluation of the system performance during design exploration. We have previously proposed local memories with an un-cached memory slave interface that permits program running on the processor to access the locally stored variables in the hardware accelerator. While this has relaxed the memory constraints for porting code sections to hardware accelerators, there is now a need to consider the read/write access penalties of local memories from the processor during design exploration. In order to facilitate the selection of profitable hardware accelerators, we need an accurate performance model that takes into account these read/write access penalties. In this paper, we propose a novel model to estimate the penalty incurred due to memory dependencies between the program running on the processor and the local memories in the FPGA hardware accelerator. This model can be used in an automated design exploration framework for heterogeneous FPGA platforms to select profitable hardware accelerators with local memories.
Alok Prakash, Siew-Kei Lam, Thambipillai Srikanthan, Christopher T. Clarke
ASAP2
2013 Preprocessing technique for accelerating reconfiguration of degradable VLSI arrays
abstract
This paper presents a heuristic approach to accelerate the reconfiguration of two-dimensional degradable VLSI arrays linked by 4-port switches in presence of faulty processing elements (PEs). In particular, we proposed a technique to preprocess the host array by 1) identifying fault-free PEs that cannot form the target array due to their proximity to faulty PEs, and 2) labeling these fault-free PEs as faults. The proposed preprocessing method minimizes the number of PEs that will be considered for reconfiguration, thus accelerating the reconfiguration process. Simulation results show that the runtime of two well-known algorithms are significantly reduced by employing the preprocessing technique. In addition, we demonstrate the scalability of the proposed technique by showing that the runtime reduction rate increases with increasing fault density.
Yuanbo Zhu, Jigang Wu, Siew-Kei Lam, Thambipillai Srikanthan
ISCAS3
2013 Efficient heuristic and tabu search for hardware/software partitioning
Jigang Wu, Siew-Kei Lam, Thambipillai Srikanthan
J. Supercomput.3
2012 Area-time estimation of C-based functions for design space exploration
abstract
Rapid evaluation of design metrics is essential for hardware-software co-design of hybrid systems on FPGAs. However, acquisition of design metrics from high-level programs is costly and/or time-consuming, and this prohibits rapid design space exploration. We will present a rapid area-time estimation technique that is capable of obtaining hardware design metrics of all the functions of the given C-based application in a fraction of the time required by FPGA implementation. We will demonstrate the proposed area-time estimation technique as part of an open source high-level synthesis tool. For the application considered, we show that the proposed method, which takes into account the effects of hardware binding during estimation, leads to a reduction in estimation error of more than 35 and 8 times for Altera Cyclone II and Stratix IV FPGA respectively.
Yan Lin Aung, Siew-Kei Lam, Thambipillai Srikanthan
FPT2
2012 Reconfiguration Algorithms for Degradable VLSI Arrays with Switch Faults
abstract
The problem of reconfiguring two-dimensional VLSI arrays with faults is to find a maximum logical array without faults. The existing algorithms only consider faults associated with processing elements, and all switches and links are assumed to be fault-free. But switch faults may often occur in the network-on-chips with high density. In this paper, two novel approaches are proposed to tackle the reconfiguration problem of degradable VLSI arrays with switch faults. The first approach extends the well-known existing algorithm with simple pre-processing and row bypass scheme. The second one employs a novel row and column rerouting scheme to maximize the size of the logical array. Simulation results show that the proposed two approaches can effectively generate the logical arrays on the given host array with switch faults, and the second algorithm performs more favorably with the increasing number of the switch faults.
Yuanbo Zhu, Jigang Wu, Siew-Kei Lam, Thambipillai Srikanthan
ICPADS3
2012 Exploiting stable features for iris recognition of defocused images
abstract
We present a novel approach for recognizing defocused iris images captured outside the Depth of Field (DOF) of cameras. Unlike existing approaches, we do not rely on special hardware or on computationally expensive image restoration algorithms. Instead, the proposed recognition approach exploits stable bits in the iris code representation which are robust to imaging noise. Experimental results based on over 15,000 images show that when compared to iris recognition of defocused images that relies on the entire code representation, the proposed method achieves an average recognition performance gain of over 2 times. Due to its low computational requirements, the proposed method is well suited for use as part of a multi-biometric system in ubiquitous systems.
Siew-Kei Lam, Thambipillai Srikanthan, Weiqi Yuan
ISCAS2
2012 Low-complexity pruning for accelerating corner detection
abstract
In this paper, we present a novel and computationally efficient pruning technique to speed up the Shi-Tomasi and Harris corner detectors. The proposed technique quickly prunes non-corners and selects a small corner candidate set by approximating the complex corner measure of Shi-Tomasi and Harris. The actual corner measure is then applied only to the reduced candidate set. Experimental results on the NiOS-II platform show that the proposed technique achieves an average execution time savings of 90% for Shi-Tomasi and 70% for Harris detectors for 500 corners with no loss in accuracy.
Meiqing Wu, Nirmala Ramakrishnan, Siew-Kei Lam, Thambipillai Srikanthan
ISCAS3
2012 Iris Recognition of Defocused Images for Mobile phones
abstract
In this paper, we introduce a novel iris recognition approach for mobile phones, which takes into account imaging noise arising from image capture outside the depth of field (DOF) of cameras. Unlike existing approaches that rely on special hardware to extend the DOF or computationally expensive algorithms to restore the defocused images prior to recognition, the proposed method performs recognition on the defocused images based on the stable bits in the iris code representation that are robust to imaging noise. To the best of our knowledge, our work is the first to investigate the characteristics of iris features for varying degree of image defocus when the images are captured outside the DOF of cameras. Based on our findings, we present a method to determine the stable bits of an enrolled image. When compared to iris recognition of defocused images that relies on the entire code representation, the proposed recognition method increases the inter-class variability while reducing the intra-class variability of the samples considered. This leads to smaller intersections between the intra-class and inter-class distance distributions, which results in higher recognition performance. Experimental results based on over 15,000 images show that the proposed method achieves an average recognition performance gain of about two times. It is envisioned that the proposed method can be incorporated as part of a multi-biometric system for mobile phones due to its lightweight computational requirements, which is well suited for power sensitive solutions.
Siew-Kei Lam, Thambipillai Srikanthan, Weiqi Yuan
Int. J. Pattern Recognit. Artif. Intell.2
2011 Architecture-Aware Technique for Mapping Area-Time Efficient Custom Instructions onto FPGAs
abstract
Area-time efficient custom instructions are desirable for maximizing the performance of reconfigurable processors. Existing data path merging techniques based on resource sharing can be deployed to improve area efficiency of custom instructions. However, these techniques lead to large increase in the critical path delay. In this paper, we propose a novel strategy that takes into account the architectural constraints of the FPGA device in order to realize custom instructions with low-area delay product. The proposed strategy is based on partitioning the custom instruction data paths into a set of basic clusters such that they can be combined using a heuristic-based cluster merging process to maximize the utilization of FPGA logic blocks. Unlike the resource sharing method, the proposed cluster merging process does not maximize sharing of common resources and this leads to lesser reliance on multiplexers for implementing custom instructions. Resource sharing is only applied sparingly at the final stage to increase utilization of logic blocks. We show that the proposed technique leads to more than 34 percent, 34 percent, and 42 percent average reduction in area costs for Spartan-3, Virtex-4, and Virtex-5 architectures, respectively, when compared to optimizations achieved through commercial synthesis tool. We have also shown that the proposed technique leads to more than 18 percent, 17 percent, and 13 percent average reduction in area costs for Spartan-3, Virtex-4, and Virtex-5, respectively, when compared to results obtained using one of the most efficient resource sharing-based method reported in the literature. In addition, the proposed technique outperforms the resource sharing-based method in terms of area-delay product, with average reductions of more than 27 percent, 34 percent, and 19 percent for Spartan-3, Virtex-4, and Virtex-5, respectively.
Siew-Kei Lam, Thambipillai Srikanthan, Christopher T. Clarke
IEEE Trans. Computers1
2010 Performance estimation framework for FPGA-based processors
abstract
Modern FPGA devices can implement a variety of processors with numerous configurable options. Rapid performance estimation of FPGA processors plays a vital role in embedded systems design to select a processor that best fits the application requirements. Traditional performance evaluation techniques such as running the software application on the target processor or using cycle accurate instruction set simulator are time-consuming and poses a threat in meeting the stringent time-to-market pressure. In this paper, we propose a framework to rapidly estimate the performance of a wide range of FPGA processors. The proposed method relies on the LLVM compiler infrastructure and its backend code generator to accurately estimate the software performance within seconds. Experimental results show that the proposed framework can reliably estimate the performance of a widely used FPGA processor with an average accuracy of over 90% for a number of benchmark applications.
Yan Lin Aung, Siew-Kei Lam, Thambipillai Srikanthan
FPT2
2010 An efficient edge and corner detector
abstract
This paper describes an efficient method for image edge and corner detection. Edges are detected before extracting corner points so that the background or noise of the image can be effectively removed. We propose an edge detector that utilizes edge features to localize the edge points. The corner points of an image are then identified by evaluating the degree of turning in the edge direction. The proposed edge and corner detection utilizes low complexity computations. Based on the experiments considered, it can be observed that the proposed edge and corner detector achieves more accurate results than several other well known algorithms. In addition, the proposed method performs faster than most of the other well known algorithms.
Siew-Kei Lam, Thambipillai Srikanthan
ICARCV2
2010 Selecting profitable custom instructions for reconfigurable processors
Tao Li 0008, Jigang Wu, Siew-Kei Lam, Thambipillai Srikanthan, Xicheng Lu
J. Syst. Archit.3
2009 Rapid design exploration framework for application-aware customization of soft core processors
abstract
Off-the-shelf soft core processors are becoming increasingly popular in embedded systems design today as they provide for application specific customization, in particular through instruction subsetting. However, choosing the right processor configuration remains a challenge as the search space becomes prohibitively large when the configurable options increase. In this paper we propose a framework to rapidly explore the processor configuration design space for a given application. Unlike existing approaches that require time-consuming synthesis process, the proposed method relies only on a single-pass output of the LLVM compiler infrastructure. Experimental results based on widely used benchmarks show that the proposed framework can reliably predict the actual performance and area trends of various configurable options.
Alok Prakash, Siew-Kei Lam, Amit Kumar Singh 0002, Thambipillai Srikanthan
FPL2
2009 Rapid design of area-efficient custom instructions for reconfigurable embedded processing
Siew-Kei Lam, Thambipillai Srikanthan
J. Syst. Archit.1
2008 A Short Course on Implementing FPGA Based Digital Systems
abstract
The rapid advances in the FPGA technology along with high-levels of system integration have made FPGAs the preferred platform not only for rapid prototyping but also for production of digital embedded systems. This paper presents the experience of a team of instructors in designing and conducting a short course on implementing FPGA-based digital systems for industry professionals. The selection of topics, course organization, the issues involved in designing effective hands-on exercises and the response of the students to the course are discussed.
George Rosario Jagadeesh, Siew-Kei Lam, Thambipillai Srikanthan
ICPADS2
2007 Estimating Area Costs of Custom Instructions for FPGA-based Reconfigurable Processors
abstract
FPGA (field programmable gate array) based reconfigurable processor has been shown to meet the increasingly challenging performance targets and shorter time-to-market pressures. In this paper, we propose a method to rapidly estimate the FPGA area costs of custom instructions without the need for hardware synthesis. The proposed estimation technique relies on a novel approach to partition the custom instruction data-paths into a set of clusters, where each cluster can be realized using an FPGA logic element or a coarse-grained arithmetic unit. Experiments based on 20 custom instructions reveal that the estimation results show an average of only 8% increase in the area costs when compared with the corresponding hardware synthesized results. In addition, we show that the maximum FPGA area utilized by custom instructions of each of the seven applications examined is equivalent to about 1000 Xilinx FPGA logic elements.
Siew-Kei Lam, Thambipillai Srikanthan
ASAP1
2006 Efficient management of custom instructions for run-time reconfigurable instruction set processors
abstract
The instruction set extension capability of RISPs (reconfigurable instruction set processors) provides an attractive means to meet the flexibility, performance, and cost demands of ubiquitous computing devices. Run-time reconfiguration can further increase the cost efficiency and hardware specialization of these processors by dynamically changing the configuration of the reconfigurable logic to the required functionality. In this paper, we propose the use of a heuristic that leads to the selection of large custom instructions for increased performance gain. Result analysis of six applications from the MiBench embedded benchmark suite show that efficient data-path merging can be applied to the custom instructions to reduce the average number of configurations to less than 8 in a run-time RISP. In addition, there is only a small difference in the average number of configurations when compared to a custom instruction selection strategy that results in lower performance
Siew-Kei Lam, Bharathi N. Krishnan, Thambipillai Srikanthan
FPT1
2004 High-throughput image rotation using sign-prediction based redundant cordic algorithm
Suchitra Sathyanarayana, Siew-Kei Lam, Thambipillai Srikanthan
ICIP2
2004 Area-Time Efficient Sign Detection Technique for Binary Signed-Digit Number System
abstract
Computer arithmetic operations based on the BSD (binary signed-digit) number representation system lend themselves well to high-speed computations due to the facilitation of limited carry addition/subtraction. We propose an area-time efficient method for sign detection in a BSD number system based on optimized reverse tree structure. When compared to other popular approaches, such as the most significant carry detection-based CLA (carry look-ahead) and MRC (multilevel reverse carry) implementations, the proposed method is superior to both area and time costs in VLSI. Synthesis results for different word lengths show that the proposed approach to sign detection in the BSD number system continues to maintain its advantage over area and time measures.
Thambipillai Srikanthan, Siew-Kei Lam, Mishra Suman
IEEE Trans. Computers2
2000 Dynamic multicast routing in VLSI
Siew-Kei Lam, Thambipillai Srikanthan
Comput. Commun.1