EDBT 2026 Demo / reviewers in the wild / expert
Thambipillai Srikanthan
dblp:23/1694 · also Srikanthan Thambipillai
· DBLP profile ↗
118ranked-venue papers
2as first author
5since 2021 · last 2024
0000-0003-3664-4345ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 70 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 1 since 2021Artificial intelligence and machine learning · 10 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 10Software engineering, systems software and programming languages · 5Theory of computation · 5Computer networks · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | GPSFormer: A Global Perception and Local Structure Fitting-Based Transformer for Point Cloud Understanding
Changshuo Wang 0001, Meiqing Wu, Siew-Kei Lam, Xin Ning 0001, Shangshu Yu, Ruiping Wang 0005, Weijun Li 0002, Thambipillai Srikanthan |
ECCV (8) | 8 |
| 2024 | Hierarchical Object-Aware Dual-Level Contrastive Learning for Domain Generalized Stereo MatchingabstractStereo matching algorithms that leverage end-to-end convolutional neural networks have recently demonstrated notable advancements in performance. However, a common issue is their susceptibility to domain shifts, hindering their ability in generalizing to diverse, unseen realistic domains. We argue that existing stereo matching networks overlook the importance of extracting semantically and structurally meaningful features. To address this gap, we propose an effective hierarchical object-aware dual-level contrastive learning (HODC) framework for domain generalized stereo matching. Our framework guides the model in extracting features that support semantically and structurally driven matching by segmenting objects at different scales and enhances correspondence between intra- and inter-scale regions from the left feature map to the right using dual-level contrastive loss. HODC can be integrated with existing stereo matching models in the training stage, requiring no modifications to the architecture. Remarkably, using only synthetic datasets for training, HODC achieves state-of-the-art generalization performance with various existing stereo matching network architectures, across multiple realistic datasets. Yikun Miao, Meiqing Wu, Siew-Kei Lam, Thambipillai Srikanthan |
NeurIPS | 5 |
| 2022 | Reconfiguration algorithms for synchronous communication on switch based degradable arrays
Yalan Wu, Jigang Wu, Peng Liu 0045, Yinhe Han 0001, Thambipillai Srikanthan |
Parallel Comput. | 5 |
| 2021 | Longevity Framework: Leveraging Online Integrated Aging-Aware Hierarchical Mapping and VF-Selection for Lifetime Reliability Optimization in Manycore ProcessorsabstractRapid device aging in the nano era threatens system lifetime reliability, posing a major intrinsic threat to system functionality. Traditional techniques to overcome the aging-induced device slowdown, such as guardbanding are static and incur performance, power, and area penalties. In a manycore processor, the system-level design abstraction offers dynamic opportunities through the control of task-to-core mappings and per-core operation frequency towards more balanced core aging profile across the chip, optimizing the system lifetime reliability while meeting the application performance requirements. This article presents Longevity Framework (LF) that leverages online integrated aging-aware hierarchical mapping and voltage frequency (VF)-selection for lifetime reliability optimization in manycore processors. The mapping exploration is hierarchical to achieve scalability. The VF-selection builds on the trade-offs involved between power, performance, and aging as the VF is scaled while leveraging the per-core DVFS capabilities. The methodology takes the chip-wide process variation into account. Extensive experimentation, comparing the proposed approach with two state-of-the-art methods, for 64-core and 256-core systems running applications from PARSEC and SPLASH-2 benchmark suites, show an improvement of up to 3.2 years in the system lifetime reliability and 4× improvement in the average core health. Vijeta Rathore, Vivek Chaturvedi, Amit Kumar Singh 0002, Thambipillai Srikanthan, Muhammad Shafique 0001 |
IEEE Trans. Computers | 4 |
| 2021 | Power-Efficient Mapping of Large Applications on Modern Heterogeneous FPGAsabstractThe increasing size of modern field programmable gate arrays (FPGAs) allows for ever more complex applications to be mapped onto them. However, long design implementation times for large designs can severely affect design productivity. A modular design methodology can improve design productivity in a divide and conqueror fashion but at the expense of degraded performance and power consumption of the resulting implementation. To reduce the dominant power dissipation component in FPGAs, the routing power, methodologies have been proposed that consider data communication between modules during module formation and placement on the FPGA. Selecting proper mapping region on target FPGAs, on the other hand, is becoming a critical process because of the heterogeneous resources and column arrangements in modern FPGAs. Selecting inappropriate FPGA regions for mapping could lead to degraded performance. Hence, we propose a methodology that uses communication-aware module placement, such that modules are mapped by selecting the best shape and region on the FPGA factoring the columnar resource arrangements. Additionally, techniques for module locking and splitting have been proposed for deterministic convergence of the algorithm and for improved module placement. This methodology exhibits nearly 19% routing power reduction with respect to commercial CAD flows without any degradation in achievable performance. Kalindu Herath, Alok Prakash, Suhaib A. Fahmy, Thambipillai Srikanthan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2020 | Evaluating the Merits of Ranking in Structured Network PruningabstractPruning of channels in trained deep neural networks has been widely used to implement efficient DNNs that can be deployed on embedded/mobile devices. Majority of existing techniques employ criteria-based sorting of the channels to preserve salient channels during pruning as well as to automatically determine the pruned network architecture. However, recent studies on widely used DNNs, such as VGG-16, have shown that selecting and preserving salient channels using pruning criteria is not necessary since the plasticity of the network allows the accuracy to be recovered through fine-tuning. In this work, we further explore the value of the ranking criteria in pruning to show that if channels are removed gradually and iteratively, alternating with fine-tuning on the target dataset, ranking criteria are indeed not necessary to select redundant channels. Experimental results confirm that even a random selection of channels for pruning leads to similar performance (accuracy). In addition, we demonstrate that even a simple pruning technique that uniformly removes channels from all layers in the network, performs similar to existing ranking criteria-based approaches, while leading to lower inference time (GFLOPs). Our extensive evaluations include the context of embedded implementations of DNNs - specifically, on small networks such as SqueezeNet and at aggressive pruning percentages. We leverage these insights, to propose a GFLOPs-aware iterative pruning strategy that does not rely on any ranking criteria and yet can further lead to lower inference time by 15% without sacrificing accuracy. Nirmala Ramakrishnan, Alok Prakash, Siew-Kei Lam, Thambipillai Srikanthan |
ICDCS | 5 |
| 2020 | Blockchain-based public auditing for big data in cloud storage
Jiaxing Li 0009, Jigang Wu, Guiyuan Jiang, Thambipillai Srikanthan |
Inf. Process. Manag. | 4 |
| 2020 | LAMBDA: Lightweight Assessment of Malware for emBeddeD ArchitecturesabstractSecurity is a critical aspect in many of the latest embedded and IoT systems. Malware is one of the severe threats of security for such devices. There have been enormous efforts in malware detection and analysis; however, occurrences of newer varieties of malicious codes prove that it is an extremely difficult problem given the nature of these surreptitious codes. In this article, instead of addressing a general solution, we aim at malware detection for platforms that have more than one core for performance enhancement. We investigate the utility of multiple cores from the point of view of security, where one of the cores operate as a watchdog. We define a notion of a new metric called LAMBDA (Lightweight Assessment of Malware for emBeddeD Architectures), denoted by λ, indicating a conceptual boundary between the programs which are allowed to run on a given platform, with the codes that are suspected as malwares. The metric λ is computed using carefully chosen monitors or features, which are tuples of high-level programs representing OS resources, along with low-level hardware performance counters. In comparison to heavy-weight machine learning techniques, we use an online hypothesis testing, in the form of t -test, to classify a given program-under-test. For applications where security is of prime concern, we propose an additional step based on multivariate analysis to classify the unknown programs that are closer to the threshold with a high degree of confidence. We present experimental results focusing on an ARM-based platform which validate that the proposed approach provides a lightweight, accurate assessment of malware codes for embedded platforms. In addition to it, we also present a security analysis to show the difficulty of a mimicry attack attempting to bypass LAMBDA. Sai Praveen Kadiyala, Manaar Alam, Yash Shrivastava, Sikhar Patranabis, Muhamed Fauzi Bin Abbas, Arnab Kumar Biswas, Debdeep Mukhopadhyay, Thambipillai Srikanthan |
ACM Trans. Embed. Comput. Syst. | 8 |
| 2020 | Hardware Performance Counter-Based Fine-Grained Malware DetectionabstractDetection of malicious programs using hardware-based features has gained prominence recently. The tamper-resistant hardware metrics prove to be a better security feature than the high-level software metrics, which can be easily obfuscated. Hardware Performance Counters (HPC), which are inbuilt in most of the recent processors, are often the choice of researchers amongst hardware metrics. However, a lack of determinism in their counts, thereby affecting the malware detection rate, minimizes the advantages of HPCs. To overcome this problem, in our work, we propose a three-step methodology for fine-grained malware detection. In the first step, we extract the HPCs of each system call of an unknown program. Later, we make a dimensionality reduction of the fine-grained data to identify the components that have maximum variance. Finally, we use a machine learning based approach to classify the nature of the unknown program into benign or malicious. Our proposed methodology has obtained a 98.4% detection rate, with a 3.1% false positive. It has improved the detection rate significantly when compared to other recent works in hardware-based anomaly detection. Sai Praveen Kadiyala, Pranav Jadhav, Siew-Kei Lam, Thambipillai Srikanthan |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2020 | Rapid and Robust Background Modeling Technique for Low-Cost Road Traffic Surveillance SystemsabstractFast and accurate detection of vehicles on road traffic scenes captured by traffic surveillance cameras, is essential for large-scale deployment of automated traffic surveillance systems. The state-of-the-art techniques typically employ background modeling for low-complexity foreground detection. However, this is a challenging problem as these methods need to be robust to varying road scene conditions (such as illumination changes, camera jitter, stationary vehicles, and heavy traffic) leading to huge computation cost. In this paper, we propose a highly accurate yet low-complexity foreground (i.e., vehicle) detection technique, which can effectively deal with the varying road scene conditions, and generate accurate pixel-level foreground masks in real-time. We propose a novel robust block-based feature suitable for modeling road background and detecting vehicles as foreground, and employ Bayesian probabilistic modeling on these features. The experimental evaluations on widely used traffic datasets demonstrate that the proposed method can achieve comparable accuracy to the existing state-of-the-art techniques but at a much higher processing frame rate (40x speedup over PAWCS). The real-time performance of the proposed system has also been demonstrated by implementing it on a low-cost embedded platform, Odroid XU-4, that still achieves a frame rate of over 80 frames/s, thereby enabling the real-time detection of foreground objects in road scenes. Kratika Garg, Nirmala Ramakrishnan, Alok Prakash, Thambipillai Srikanthan |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2019 | LifeGuard: A Reinforcement Learning-Based Task Mapping Strategy for Performance-Centric Aging ManagementabstractDevice scaling to subdeca nanometer has pushed device aging as a primary design concern. In manycore systems, inevitable process variation further adds to delay degradation and, coupled with the scalability issues in manycores, makes aging management, while meeting performance demands, a complex problem. LifeGuard is a performance-centric reinforcement learning-based task mapping strategy that leverages the different impact of applications on aging for improving system health. Experimental results, comparing LifeGuard with two state-of-the-art aging optimizing techniques, on a 256-core system, showed that LifeGuard led to improved health for, respectively, 57% and 74% of the cores, and also an enhanced aggregate core frequency. Vijeta Rathore, Vivek Chaturvedi, Amit Kumar Singh 0002, Thambipillai Srikanthan, Muhammad Shafique 0001 |
DAC | 4 |
| 2019 | Towards Scalable Lifetime Reliability Management for Dark Silicon Manycore SystemsabstractAggressive technology scaling enabled very high integration density. Unfortunately, it also led to issues such as process variation, increased power density and consequently rising chip temperature resulting in accelerated device aging and poor lifetime reliability of different components in a manycore system. Moreover, thermal and power limitations let only a fraction of the chip function at full speed; the rest is the dark silicon. Most of the lifetime reliability enhancement solutions for the multi-/manycore systems in the literature are heuristic-based, while some use standard compute-intensive methods to solve the optimization problem making them not scale well with the manycore size. The heuristic-based solutions are formulated to search through the design space of a fine granularity making it huge, limiting their scalability. Also, these approaches do not account for the impact of different applications' execution behavior on the aging of the underlying cores, and their performance requirement distribution across the cores to their advantage. In this paper, we present our resource management strategies towards building scalable lifetime reliability enhancement solutions for dark silicon manycore systems. The first technique, Hierarchical Mapping approach (HiMap), maps a periodic workload employing a block-based hierarchical method that leverages dark cores for thermal mitigation. The second approach, LifeGuard, uses reinforcement learning to learn the applications' aging behavior, and is aware of the performance requirement pattern onto the core frequencies. It maps randomly arriving requests and is scalable to the number of applications and the size of a manycore. Vijeta Rathore, Vivek Chaturvedi, Amit Kumar Singh 0002, Thambipillai Srikanthan, Muhammad Shafique 0001 |
IOLTS | 4 |
| 2019 | Rapid Technique to Eliminate Moving Shadows for Accurate Vehicle DetectionabstractElimination of moving shadows is an essential step to achieve accurate vehicle detection and localization in automated traffic surveillance systems that aim to detect vehicles on road scenes captured by surveillance cameras. However, this is still a challenging problem as existing pixel based methods miss parts of vehicles and region-based methods, while accurate, incur higher computations. In this paper, we propose a highly accurate yet low-complexity block-based moving shadow elimination technique, which can effectively deal with varying shadow conditions. A novel shadow elimination pipeline is proposed that employs computationally lean features to quickly classify distinct vehicles from shadows, and uses a more sophisticated interior edge feature only for classification of difficult scenarios. Extensive evaluations on freely available and self-collected datasets demonstrate that the proposed technique achieves higher accuracy than other state-of-the-art techniques in varying scenarios. Additionally, it also achieves over 20 times speedup on a low-cost embedded platform, Odroid XU-4, over a state-of-the-art technique that achieves comparable accuracy. Experimental results confirm the realtime capability of the proposed approach while achieving robustness to varying shadow scenarios. Kratika Garg, Nirmala Ramakrishnan, Alok Prakash, Thambipillai Srikanthan, Punit Bhatt |
WACV | 4 |
| 2018 | HiMap: A hierarchical mapping approach for enhancing lifetime reliability of dark silicon manycore systemsabstractTechnology scaling into the nano-scale CMOS regime has resulted in increased leakage and roadblock on voltage scaling, which has led to several issues like high power density and elevated on-chip temperature. This consequently aggravates device aging, compromising lifetime reliability of the manycore systems. This paper proposes HiMap, a dynamic hierarchical mapping approach to maximize lifetime reliability of manycore systems while satisfying performance, power, and thermal constraints. HiMap is process variation- and aging-aware. It comprises of two levels: (1) it identifies a region of cores suitable for mapping, and (2) it maps threads in the region and intersperses dark cores for thermal mitigation while considering the current health of the cores. Both the levels strive to reduce aging variance across the chip. We evaluated HiMap for 64-core and 256-core systems. Results demonstrate an improved system lifetime reliability by up to 2 years at the end of 3.25 years of use, as compared to the state-of-the-art. Vijeta Rathore, Vivek Chaturvedi, Amit Kumar Singh 0002, Thambipillai Srikanthan, R. Rohith, Siew-Kei Lam, Muhammad Shafique 0001 |
DATE | 4 |
| 2018 | Wibheda: Framework for Data Dependency-Aware Multi-Constrained Hardware-Software Partitioning in FPGA-Based SoCs for IoT DevicesabstractThe increasing popularity of FPGA-based system-on-chip (SoC) devices for Internet of Things (IoT) applications calls for hardware-software partitioning solutions optimized for performance under stringent area and power constraints. In this work, we propose Wibheda, a heuristic based framework for data dependency-aware multi-constrained hardware-software partitioning at fine-granularity that can be employed to partition designs for FPGA-based SoCs used in IoT. Wibheda, evaluated on 6 applications from the popular CHStone benchmark suite has been shown to find solutions with 98.7% accuracy within several milliseconds compared to several minutes or hours in an existing state-of-the-art work and an exhaustive approach respectively. Deshya Wijesundera, Alok Prakash, Thilina Perera, Kalindu Herath, Thambipillai Srikanthan |
FCCM | 5 |
| 2018 | A Hybrid Methodology for Optimal Fleet Management in an Electric Vehicle Based Flexible Bus ServiceabstractThe ever-increasing traffic congestion and CO2emission caused by rapid urbanization, calls for smarter and energy efficient transit services. Conventional public transit lacks the ability to meet these diversified needs. As a result, intelligent transit systems, influenced by the digital revolution have created a profound impact by enhancing the user-experience of transit services. Consequently, demand responsive transit (DRT) services, which operate with flexible routes and schedules have become a common option among commuters. Thus, in this work, we propose an electric vehicle (EV) based flexible bus service, a variant of DRT, that satisfies passenger demand in a given geographical zone. Next, we present a hybrid methodology to optimally manage the EV fleet minimizing the total vehicle miles travelled (VMT). Experimental results with a real map show that the proposed hybrid method achieves near-optimal results with 120x improvement in computation time. Further, the flexible bus service reduces VMT by over 70% in comparison to single occupancy vehicles, thus reducing both traffic congestion and CO2emissions. Thilina Perera, Alok Prakash, Thambipillai Srikanthan |
ICARCV | 3 |
| 2018 | Algorithms for Replica Placement and Update in Tree NetworkabstractA critical issue in data replication is to wisely place data replicas which involves identifying the best possible nodes to duplicate data. Facing dynamics of data requests, this paper investigates the problem of replica placement and update in tree networks, where part of nodes have pre-existing replicas. We aim to develop efficient algorithms to accelerate the replica placement and update without causing obvious degradation in solution quality via reusing pre-existing replicas. Firstly, an efficient heuristic algorithm GRP is proposed to quickly place replicas when users change their requests dynamically, under the Closest policy where a client must be served by the closest server. Then, a Tabu search algorithm TSRP is customized to further refine the solution obtained by GRP. Furthermore, we propose a heuristic algorithm MPFSF for the replica placement and update problem, under the Multiple policy where requests of a client are served by multiple servers. Simulation results show that, GRP and TSRP can accelerate existing dynamic programming algorithm by 87.97% while quality degradation is bounded by 2.49%. MPFSF can achieve the best improvement for about 84.6% than existing heuristic algorithm. Jigang Wu, Long Chen 0006, Guiyuan Jiang, Siew-Kei Lam, Thambipillai Srikanthan |
Comput. J. | 6 |
| 2018 | Framework for Rapid Performance Estimation of Embedded Soft Core ProcessorsabstractThe large number of embedded soft core processors available today make it tedious and time consuming to select the best processor for a given application. This task is even more challenging due to the numerous configuration options available for a single soft core processor while optimizing for contradicting design requirements such as performance and area. In this article, we propose a generic framework for rapid performance estimation of applications on soft core processors. The proposed technique is scalable to the large number of configuration options available in modern soft core processors by relying on rapid and accurate estimation models instead of time-consuming FPGA synthesis and execution-based techniques. Experimental results on two leading commercial soft core processors executing applications from the widely used CHStone benchmark suite show an average error of less than 6% while running in the order of minutes when compared to hours taken by synthesis-based techniques. Deshya Wijesundera, Alok Prakash, Thambipillai Srikanthan, Achintha Ihalage |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2018 | Rapid Memory-Aware Selection of Hardware Accelerators in Programmable SoC DesignabstractProgrammable Systems-on-Chips (SoCs) are expected to incorporate a larger number of application-specific hardware accelerators with tightly integrated memories in order to meet stringent performance-power requirements of embedded systems. As data sharing between the accelerator memories and the processor is inevitable, it is of paramount importance that the selection of application segments for hardware acceleration must be undertaken such that the communication overhead of data transfers do not impede the advantages of the accelerators. In this paper, we propose a novel memory-aware selection algorithm that is based on an iterative approach to rapidly recommend a set of hardware accelerators that will provide high performance gain under varying area constraint. In order to significantly reduce the algorithm runtime while still guaranteeing near-optimal solutions, we propose a heuristic to estimate the penalties incurred when the processor accesses the accelerator memories. In each iteration of the proposed algorithm, a two-pass method is employed where a set of good hardware accelerator candidates is selected using a greedy approach in the first pass, and a “sliding window” approach is used in the second pass to refine the solution. The two-pass method is iteratively performed on a bounded set of candidate hardware accelerators to limit the search space and to avoid local maxima. In order to validate the benefits of the proposed selection algorithm, an exhaustive search algorithm is also developed. Experimental results using the popular CHStone benchmark suite show that the performance achieved by the accelerators recommended by the proposed algorithm closely matches the performance of the exhaustive algorithm, with close to 99% accuracy, while being orders of magnitude faster. Alok Prakash, Christopher T. Clarke, Siew-Kei Lam, Thambipillai Srikanthan |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2017 | Communication-aware Partitioning for Energy Optimization of Large FPGA DesignsabstractModern FPGAs integrate multi-million logic resources that allow the realization of increasingly large designs. However, state-of-the-art simulated annealing based CAD tools for FPGA suffer from long runtime, poor performance and sub-optimal routing and placement decisions, especially for large applications, leading to less energy efficient designs. In this paper, we present a partitioning methodology that divides large application into smaller subsystems based on the communication frequency between these subsystems. We leverage the existing CAD tools to compile the large design, which is now annotated with their subsystems, to obtain the final bitstream. Experiments show that the proposed strategy can lead to a performance gain of over 60% while still achieving more than 20% reduction in energy consumption. Kalindu Herath, Alok Prakash, Guiyuan Jiang, Thambipillai Srikanthan |
ACM Great Lakes Symposium on VLSI | 4 |
| 2017 | Online Map-Matching of Noisy and Sparse Location Data With Hidden Markov and Route Choice ModelsabstractWith the growing use of crowdsourced location data from smartphones for transportation applications, the task of map-matching raw location sequence data to travel paths in the road network becomes more important. High-frequency sampling of smartphone locations using accurate but power-hungry positioning technologies is not practically feasible as it consumes an undue amount of the smartphone's bandwidth and battery power. Hence, there exists a need to develop robust algorithms for map-matching inaccurate and sparse location data in an accurate and timely manner. This paper addresses the above-mentioned need by presenting a novel map-matching solution that combines the widely used approach based on a hidden Markov model (HMM) with the concept of drivers' route choice. Our algorithm uses an HMM tailored for noisy and sparse data to generate partial map-matched paths in an online manner. We use a route choice model, estimated from real drive data, to reassess each HMM-generated partial path along with a set of feasible alternative paths. We evaluated the proposed algorithm with real world as well as synthetic location data under varying levels of measurement noise and temporal sparsity. The results show that the map-matching accuracy of our algorithm is significantly higher than that of the state of the art, especially at high levels of noise. George Rosario Jagadeesh, Thambipillai Srikanthan |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2017 | A Framework for Fast and Robust Visual OdometryabstractKnowledge of the ego-vehicle's motion state is essential for assessing the collision risk in advanced driver assistance systems or autonomous driving. Vision-based methods for estimating the ego-motion of vehicle, i.e., visual odometry, face a number of challenges in uncontrolled realistic urban environments. Existing solutions fail to achieve a good tradeoff between high accuracy and low computational complexity. In this paper, a framework for ego-motion estimation that integrates runtime-efficient strategies with robust techniques at various core stages in visual odometry is proposed. First, a pruning method is employed to reduce the computational complexity of Kanade-Lucas-Tomasi (KLT) feature detection without compromising on the quality of the features. Next, three strategies, i.e., smooth motion constraint, adaptive integration window technique, and automatic tracking failure detection scheme, are introduced into the conventional KLT tracker to facilitate generation of feature correspondences in a robust and runtime efficient way. Finally, an early termination condition for the random sample consensus (RANSAC) algorithm is integrated with the Gauss-Newton optimization scheme to enable rapid convergence of the motion estimation process while achieving robustness. Experimental results based on the KITTI odometry data set show that the proposed technique outperforms the state-of-the-art visual odometry methods by producing more accurate ego-motion estimation in notably lesser amount of time. Meiqing Wu, Siew-Kei Lam, Thambipillai Srikanthan |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2016 | Rapid design space exploration for soft core processor customization and selectionabstractThe large number of soft core processors available today make it challenging to select the best processor for a given application. This issue is further exacerbated by the numerous configuration options available for soft core processors while customizing for contradictory design requirements such as performance and area. In this paper, we propose a framework for rapid application-specific soft core processor customization as well as selection to achieve best performance within user-specified area constraint. The proposed technique is scalable to the large number of configuration options available in modern soft core processors by relying on rapid and accurate estimation models instead of time consuming FPGA synthesis and execution steps. Experimental results on two leading commercial soft core processors executing applications from the popular CHStone benchmark suite show over 93% accuracy in processor performance estimation, completed in order of seconds, leading to a rapid and reliable soft core processor customization and selection. Deshya Wijesundera, Alok Prakash, Thambipillai Srikanthan |
FPT | 3 |
| 2016 | Performance Constraint-Aware Task Mapping to Optimize Lifetime Reliability of Manycore SystemsabstractNegative bias temperature instability (NBTI) has emerged as a critical challenge to lifetime reliability of computing systems. Traditionally, temperature-aware methodologies are used to mitigate the impact of NBTI on aging and degradation of computing systems. However, in the presence of process variation, which is the norm in manycore processors, temperature-aware techniques are inefficient in improving lifetime reliability and can result in poor performance. In this paper, we propose a novel performance constraint-aware task mapping technique to improve lifetime reliability by mitigating NBTI considering on-chip process variation. Our approach consists of two phases, namely design-time and run-time. During design time, we generate Pareto-optimal mappings. Following which, our run-time technique judiciously intervenes to perform workload migration to save the weakest processing core. We compare our approach with performance-greedy and thermal-aware task mapping techniques. The experiment results demonstrate that our approach outperforms other two techniques and improves lifetime reliability of a manycore system as much as 54% without violating the throughput constraint. Vijeta Rathore, Vivek Chaturvedi, Thambipillai Srikanthan |
ACM Great Lakes Symposium on VLSI | 3 |
| 2016 | Exploiting Configuration Dependencies for Rapid Area-efficient Customization of Soft-core ProcessorsabstractThe large number of possible configurations in modern soft-core processors make it tedious and time consuming to select the optimal configuration for a given application. In this paper, we propose a framework for rapid area-efficient customization of soft-core processors that exploits the dependencies between the various configuration options to prune the design space. Additionally, the proposed technique relies on rapid and accurate estimation models instead of the time consuming synthesis and execution techniques proposed in the existing work. Experimental results based on hand-coded applications and applications from the popular CHStone benchmark suite show that the proposed framework can rapidly and reliably select the best processor configuration for a given application and save an average of 47.58% area over the processor with all the configuration options enabled while achieving similar performance. Deshya Wijesundera, Alok Prakash, Siew-Kei Lam, Thambipillai Srikanthan |
SCOPES | 4 |
| 2016 | Real-time road traffic density estimation using block varianceabstractThe increasing demand for urban mobility calls for a robust real-time traffic monitoring system. In this paper we present a vision-based approach for road traffic density estimation which forms the fundamental building block of traffic monitoring systems. Existing techniques based on vehicle counting and tracking suffer from low accuracy due to sensitivity to illumination changes, occlusions, congestions etc. In addition, existing holistic-based methods cannot be implemented in real-time due to high computational complexity. In this paper we propose a block based holistic approach to estimate traffic density which does not rely on pixel based analysis, therefore significantly reducing the computational cost. The proposed method employs variance as a means for detecting the occupancy of vehicles on pre-defined blocks and incorporates a shadow elimination scheme to prevent false positives. In order to take into account varying illumination conditions, a low-complexity scheme for continuous background update is employed. Empirical evaluations on publicly available datasets demonstrate that the proposed method can achieve real-time performance and has comparable accuracy with existing high complexity holistic methods. Kratika Garg, Siew-Kei Lam, Thambipillai Srikanthan, Vedika Agarwal |
WACV | 3 |
| 2016 | Cost-efficient Acceleration of Hardware Trojan Detection Through Fan-Out Cone Analysis and Weighted Random Pattern TechniqueabstractFabless semiconductor industry and government agencies have raised serious concerns about tampering with inserting hardware Trojans (HTs) in an integrated circuit supply chain in recent years. In this paper, a low hardware overhead acceleration method of the detection of HTs based on the insertion of 2-to-1 MUXs as test points is proposed. In the proposed method, the fact that one logical gate has a significant impact on the transition probability of the logical gates in its logical fan-out cone is utilized to optimize the number of the inserted MUXs. The nets which have smaller transition probability than the user-specified threshold and minimal logical depth from the primary inputs are selected as the candidate nets. As for each candidate net, only its input net with smallest signal probability is required to be inserted the MUXs-based test points. The procedure repeats until the minimal transition probability of the entire circuit is not smaller than the threshold value. In order to further optimize the number of required insertions and reduce the overhead, the weighted random pattern technique is also applied. Experiment results on ISCAS'89 benchmark circuits show that our proposed method can achieve remarkable improvement of transition probability with on average 9.50% power, 2.37% delay, and 10.26% area penalty. Wei Zhang 0012, Thambipillai Srikanthan, Jason Teo Kian Jin, Vivek Chaturvedi, Tao Luo 0014 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2015 | Fast Replica Placement and Update Strategies in Tree NetworksabstractData replication enhances data availability and thereby improves the system reliability and efficiency while reduces access latency and communication cost. A critical issue in data replication is to wisely place data replicas which involves identifying the best possible nodes to duplicate data according to user requests. In this paper, we address the problem of replica placement and update in tree networks, where some nodes of the network contain pre-existing replicas. It is obvious that reusing a pre-existing replica leads to smaller cost than creating a new replica, thus it is necessary to take full advantage of the pre-existing replicas. The only previous work that consider the same problem tries to find the optimal solution by developing a dynamic programming algorithm which runs in O(N5). However, this approach is not suitable for practical situation where the user requests change frequently. In this paper, we develop efficient algorithms to accelerate the replica placement without causing obvious degradation in solution quality. We first propose an efficient heuristic algorithm (named GreedyRP) for quickly placing replicas when users change their requests. Then a tabu search algorithm (named TSRP) is customized to further refine the solution obtained by algorithm GreedyRP. Experimental results show that, on tree networks with 600 nodes and 150 pre-existing replicas, the proposed algorithms can accelerate the previous work by 87.97% while the quality degradation is bounded by 2.49% in comparison to the optimal solution. Jigang Wu, Guiyuan Jiang, Siew-Kei Lam, Thambipillai Srikanthan |
CCGRID | 5 |
| 2015 | Identifying epileptic seizures based on a template-based eyeball detection techniqueabstractMany types of epileptic seizures are accompanied by characteristic eye movement, along with the movement of other parts of the body. Vision-based detection of eye movement in the context of seizures is less explored. The contribution of this paper is three-fold: (1) a compute-efficient algorithm for eyeball detection is proposed and evaluated on a standard database and yields an accuracy of 95%. (2) A new video dataset for absence seizure has been created. (3) The proposed algorithm for eyeball detection is applied on this dataset in order to detect the occurrence of absence seizure. Supriya Sathyanarayana, Ravi Kumar Satzoda, Suchitra Sathyanarayana, Thambipillai Srikanthan |
ICIP | 4 |
| 2015 | Critical-path optimization for efficient hardware realization of lifting and flipping DWTsabstractThe pair of scaling constants of lifting scheme for DWT computation are perfect inverse of each other while those of flipping scheme are not perfect inverse of each other. However, flipping-scheme is preferred over the lifting-scheme for area-delay efficient hardware implementation of 2-D DWT due to its smaller critical-path. In this paper we present critical-path analysis of lifting and flipping schemes and propose a low-complexity data-path for lifting based DWT. We have shown that the critical-path delay (CPD) of lifting DWT is higher by only 2 full-adder (FA) delay than that of flipping DWT. Moreover, due to the saving of multipliers, lifting scheme offers a better area delay efficient structure than the flipping scheme for parallel realization of 2-D DWT. In this paper, we propose efficient realization of both lifting and flipping DWT. Compared with the best of the available flipping-based 2-D DWT structure, the proposed lifting-based and flipping-based structures for block-size 16 involve 74% less and 73% less area-delay-product (ADP), and offer nearly 6.38 and 6.72 times higher throughput, respectively. Compared with similar lifting based existing design, the proposed lifting-based structure involves 15.38% less ADP for the same block-size. Basant K. Mohanty, Pramod Kumar Meher, Thambipillai Srikanthan |
ISCAS | 3 |
| 2015 | Adaptive Window Strategy for High-Speed and Robust KLT Feature Tracker
Nirmala Ramakrishnan, Thambipillai Srikanthan, Siew-Kei Lam, Gauri Ravindra Tulsulkar |
PSIVT | 2 |
| 2015 | Nonparametric Technique Based High-Speed Road Surface DetectionabstractIt has been well recognized that detecting road surface in a realistic environment is a challenging problem that is also computationally intensive. Existing road surface detection methods attempt to fit the road surface into rigid models (e.g., planar, clothoid, or B-Spline), thereby restricting to road surfaces that match specific models. In addition, the curve-fitting strategies employed in such techniques incur high computational complexity, making them unsuitable for in-vehicle deployments. In this paper, we propose an efficient nonparametric road surface detection algorithm that exploits the depth cue. The proposed method relies on four intrinsic road scene attributes observed under stereo geometry and has been shown to reliably detect both planar and nonplanar road surfaces efficiently. Extensive evaluations are performed on three widely used benchmarks (i.e., enpeda, KITTI, and Daimler), encompassing many complex road scenarios. The experimental results show that the proposed algorithm significantly outperforms the well-known techniques both in terms of detection accuracy and runtime performance. Meiqing Wu, Siew-Kei Lam, Thambipillai Srikanthan |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2015 | Algorithmic aspects of graph reduction for hardware/software partitioning
Guiyuan Jiang, Jigang Wu, Siew-Kei Lam, Thambipillai Srikanthan |
J. Supercomput. | 4 |
| 2014 | Area-delay efficient architecture for MP algorithm using reconfigurable inner-product circuitsabstractMatching pursuit (MP) algorithm is popularly used as a low-cost alternative to the orthogonal matching pursuit (OMP) algorithm for the reconstruction of signal from compressively sensed samples. In this paper, we have proposed an efficient scheduling of computation along with a novel reconfigurable inner-product (IP) unit and buffer units to provide regular inflow of input, and storage of intermediate results for efficient implementation of MP algorithm. The proposed reconfigurable IP unit can compute inner-products of different lengths by reusing the arithmetic components with very low reconfiguration overhead. We have customized on-chip buffers to exploit desired level of parallelism with lower latency and higher hardware utilization efficiency. The proposed structure of MP algorithm for the reconstruction of compressively sensed data is found to involve nearly 15% less critical path delay, 7% less area, 10% less reconstruction time, and 17% less area-delay-product (ADP) than the existing structure. Pramod Kumar Meher, Basant K. Mohanty, Thambipillai Srikanthan |
ISCAS | 3 |
| 2014 | Thermal-aware task scheduling for peak temperature minimization under periodic constraint for 3D-MPSoCsabstract3D-MPSoC offer great performance and scalability benefits. However, due to strong vertical thermal correlation and increased power density, thermal challenges in 3D-MPSoC are critical. In this paper, we propose a novel thermal aware task scheduling technique that combine intelligent task mapping with DVFS to minimize the peak temperature of the system. Particularly, our approach leverages on the fundamental thermal characteristics of 3D architecture when mapping tasks to processing cores and employing DVFS at design time followed by a simple thermal optimization step at run time. Our experiments validate the efficiency of our approach in peak temperature minimization up to 14°C compared to other existing methods. Vivek Chaturvedi, Amit Kumar Singh 0002, Wei Zhang 0012, Thambipillai Srikanthan |
RSP | 4 |
| 2014 | Rapid evaluation of custom instruction selection approaches with FPGA estimationabstractThe main aim of this article is to demonstrate that a fast and accurate FPGA estimation engine is indispensable in design flows for custom instruction (template) selection. The need for a FPGA estimation engine stems from the difficulty in predicting the FPGA performance measures of selected custom instructions. We will present a FPGA estimation technique that partitions the high-level representation of custom instructions into clusters based on the structural organization of the target FPGA, while taking into account general logic synthesis principles adopted by FPGA tools. In this work, we have evaluated a widely used graph covering algorithm with various heuristics for custom instruction selection. In addition, we present an algorithm called Refined Largest Fit First (RLFF) that relies on a graph covering heuristic to select non-overlapping superset templates, which typically incorporate frequently used basic templates. The initial solution is further refined by considering overlapping templates that were ignored previously to see if their introduction could lead to higher performance. While RLFF provides the most efficient cover compared to the ILP method and other graph covering heuristics, FPGA estimation results reveals that RLFF leads to the worst performance in certain applications. It is therefore a worthy proposition to equip design flows with accurate FPGA estimation in order to rapidly determine the most profitable custom instruction approach for a given application. Siew-Kei Lam, Thambipillai Srikanthan, Christopher T. Clarke |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2014 | Parallel reconfiguration algorithms for mesh-connected processor arrays
Jigang Wu, Guiyuan Jiang, Yuze Shen, Siew-Kei Lam, Thambipillai Srikanthan |
J. Supercomput. | 6 |
| 2014 | Dataflow Graph Partitioning for Area-Efficient High-Level Synthesis with Systems PerspectiveabstractArea efficiency in datapath synthesis is a widely accepted goal of high-level synthesis. Applications represented by their dataflow graphs are synthesized using resource sharing principles to reduce the area. However, existing resource sharing algorithms focus on absolute area reduction and maximal resource sharing. This kind of a design approach leads to constraints on how often, in terms of number of clock cycles, a new set of input data can be fed to an application. It also leads to very large multiplexers in case of very big dataflow graphs with hundreds of nodes. An adaptive dataflow graph partitioning algorithm is proposed that partitions a graph taking into account a user-defined constraint on how often a new set of input data (generally referred to as data initiation interval) is available. At the same time, a resource sharing algorithm is applied to such partitions in order to reduce area. Multiple design points are generated for a given dataflow graph with different area and time measures to enable a designer to make decisions. We demonstrate our graph partitioning algorithm using synthetically generated large dataflow graphs and on some benchmark applications. Sharad Sinha, Thambipillai Srikanthan |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2014 | Constructing Sub-Arrays with ShortInterconnects from Degradable VLSI ArraysabstractReducing the interconnection length of VLSI arrays leads to less capacitance, power dissipation and dynamic communication cost between the processing elements (PEs). This paper develops efficient algorithms for constructing tightly-coupled subarrays from the mesh-connected VLSI arrays with faulty PEs. For a given size r·s of the target (logical) array, the proposed algorithm searches and reroutes a physical r×s subarray that has the least number of faults, resulting in an approximate target array, which is subsequently extended to the desired target array. Experimental results show that over 65 percent redundant interconnects can be reduced for a 64×64 target array on the 512×512 host array with no more than 1 percent faults. In addition, we propose a recursive divide-and-conquer algorithm for constructing the maximum target array (MTA). The lower bound of the total interconnection length of the MTA has been established. Experimental results show that the proposed algorithm is capable of reducing the long interconnects by over 33 percent for the MTA derived from the 512×512 host array with no more than 1 percent faults. Moreover, the proposed total interconnection length of target array is close to the lower bound for the cases with relatively fewer number of faults. Jigang Wu, Thambipillai Srikanthan, Guiyuan Jiang |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2014 | Exploiting FPGA-Aware Merging of Custom Instructions for Runtime ReconfigurationabstractRuntime reconfiguration is a promising solution for reducing hardware cost in embedded systems, without compromising on performance. We present a framework that aims to increase the performance benefits of reconfigurable processors that support full or partial runtime reconfiguration. The proposed framework achieves this by: (1) providing a means for choosing suitable custom instruction selection heuristics, (2) leveraging FPGA-aware merging of custom instructions to maximize the reconfigurable logic block utilization in each configuration, and (3) incorporating a hierarchical loop partitioning strategy to reduce runtime reconfiguration overhead. We show that the performance gain can be improved by employing suitable custom instruction selection heuristics that, in turn, depend on the reconfigurable resource constraints and the merging factor (extent to which the selected custom instructions can be merged). The hierarchical loop partitioning strategy leads to an average performance gain of over 31% and 46% for full and partial runtime reconfiguration, respectively. Performance gain can be further increased to over 52% and 70% for full and partial runtime reconfiguration, respectively, by exploiting FPGA-aware merging of custom instructions. Siew-Kei Lam, Christopher T. Clarke, Thambipillai Srikanthan |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2013 | Modelling communication overhead for accessing local memories in hardware acceleratorsabstractLocal memories increase the efficiency of hardware accelerators by enabling fast accesses to frequently used data. In addition, the access latencies of local memories are deterministic which allows for more accurate evaluation of the system performance during design exploration. We have previously proposed local memories with an un-cached memory slave interface that permits program running on the processor to access the locally stored variables in the hardware accelerator. While this has relaxed the memory constraints for porting code sections to hardware accelerators, there is now a need to consider the read/write access penalties of local memories from the processor during design exploration. In order to facilitate the selection of profitable hardware accelerators, we need an accurate performance model that takes into account these read/write access penalties. In this paper, we propose a novel model to estimate the penalty incurred due to memory dependencies between the program running on the processor and the local memories in the FPGA hardware accelerator. This model can be used in an automated design exploration framework for heterogeneous FPGA platforms to select profitable hardware accelerators with local memories. Alok Prakash, Siew-Kei Lam, Thambipillai Srikanthan, Christopher T. Clarke |
ASAP | 3 |
| 2013 | Preprocessing technique for accelerating reconfiguration of degradable VLSI arraysabstractThis paper presents a heuristic approach to accelerate the reconfiguration of two-dimensional degradable VLSI arrays linked by 4-port switches in presence of faulty processing elements (PEs). In particular, we proposed a technique to preprocess the host array by 1) identifying fault-free PEs that cannot form the target array due to their proximity to faulty PEs, and 2) labeling these fault-free PEs as faults. The proposed preprocessing method minimizes the number of PEs that will be considered for reconfiguration, thus accelerating the reconfiguration process. Simulation results show that the runtime of two well-known algorithms are significantly reduced by employing the preprocessing technique. In addition, we demonstrate the scalability of the proposed technique by showing that the runtime reduction rate increases with increasing fault density. Yuanbo Zhu, Jigang Wu, Siew-Kei Lam, Thambipillai Srikanthan |
ISCAS | 4 |
| 2013 | Block-Based Search Space Reduction Technique for Face Detection Using Shoulder and Head Curves
Supriya Sathyanarayana, Ravi Kumar Satzoda, Suchitra Sathyanarayana, Thambipillai Srikanthan |
PSIVT | 4 |
| 2013 | CADSE: communication aware design space exploration for efficient run-time MPSoC management
Amit Kumar Singh 0002, Akash Kumar 0001, Jigang Wu, Thambipillai Srikanthan |
Frontiers Comput. Sci. | 4 |
| 2013 | Perceptual Full-Reference Quality Assessment of Stereoscopic Images by Considering Binocular Visual CharacteristicsabstractPerceptual quality assessment is a challenging issue in 3D signal processing research. It is important to study 3D signal directly instead of studying simple extension of the 2D metrics directly to the 3D case as in some previous studies. In this paper, we propose a new perceptual full-reference quality assessment metric of stereoscopic images by considering the binocular visual characteristics. The major technical contribution of this paper is that the binocular perception and combination properties are considered in quality assessment. To be more specific, we first perform left-right consistency checks and compare matching error between the corresponding pixels in binocular disparity calculation, and classify the stereoscopic images into non-corresponding, binocular fusion, and binocular suppression regions. Also, local phase and local amplitude maps are extracted from the original and distorted stereoscopic images as features in quality assessment. Then, each region is evaluated independently by considering its binocular perception property, and all evaluation results are integrated into an overall score. Besides, a binocular just noticeable difference model is used to reflect the visual sensitivity for the binocular fusion and suppression regions. Experimental results show that compared with the relevant existing metrics, the proposed metric can achieve higher consistency with subjective assessment of stereoscopic images. Feng Shao 0001, Weisi Lin, Shanbo Gu, Gangyi Jiang, Thambipillai Srikanthan |
IEEE Trans. Image Process. | 5 |
| 2013 | Efficient heuristic and tabu search for hardware/software partitioning
Jigang Wu, Siew-Kei Lam, Thambipillai Srikanthan |
J. Supercomput. | 4 |
| 2012 | Custom instructions with local memory elements without expensive DMA transfersabstractTraditionally, Instruction set extension (ISE) algorithms have treated memory and control flow as invalid operations during custom instruction identification to ensure deterministic latency of these extended instructions. In order to overcome these constraints some work has been done to incorporate local memory for custom instructions with memory operations. Such architectures have invariably relied on the expensive DMA protocol for data transfer. Cache-coherence management poses another challenge in such systems and requires additional hardware and/or software intervention. We propose a novel custom instruction architecture capable of incorporating certain types of memory and control-flow operations. Unlike existing architectures, the proposed design eliminates the need for expensive Direct Memory Access (DMA) transfers and additional cache management sub-systems, thereby saving significant time and energy. Our method is focused mainly on accelerating code segments with static variables as well as the ones allocated on the stack, which are widely prevalent in embedded applications. Experimental results show that the proposed method achieves a substantial performance gain of upto 47% over base processor implementation. Alok Prakash, Christopher T. Clarke, Thambipillai Srikanthan |
FPL | 3 |
| 2012 | Dataflow graph partitioning for high level synthesisabstractThis paper presents a dataflow graph (DFG) partitioning methodology for effective high level synthesis in the presence of constraints like data initiation interval (II) and area. It also focuses on handling large DFGs for high level synthesis with area reduction as a requirement. An algorithm for dataflow graph partitioning is presented that aims to reduce area utilization as well as ensure that data initiation interval constraint is met. The algorithm works so as to fit a design into the design space between fully pipelined design and fully resource shared design in order to meet the initiation interval constraint and reduce area only as much as required compared to a fully pipelined design where the area is wasted in the presence of II constraint and a fully resource shared design where the extreme reduction in area puts additional unnecessary constraint on data initiation interval. Sharad Sinha, Thambipillai Srikanthan |
FPL | 2 |
| 2012 | Area-time estimation of C-based functions for design space explorationabstractRapid evaluation of design metrics is essential for hardware-software co-design of hybrid systems on FPGAs. However, acquisition of design metrics from high-level programs is costly and/or time-consuming, and this prohibits rapid design space exploration. We will present a rapid area-time estimation technique that is capable of obtaining hardware design metrics of all the functions of the given C-based application in a fraction of the time required by FPGA implementation. We will demonstrate the proposed area-time estimation technique as part of an open source high-level synthesis tool. For the application considered, we show that the proposed method, which takes into account the effects of hardware binding during estimation, leads to a reduction in estimation error of more than 35 and 8 times for Altera Cyclone II and Stratix IV FPGA respectively. Yan Lin Aung, Siew-Kei Lam, Thambipillai Srikanthan |
FPT | 3 |
| 2012 | Reconfiguration Algorithms for Degradable VLSI Arrays with Switch FaultsabstractThe problem of reconfiguring two-dimensional VLSI arrays with faults is to find a maximum logical array without faults. The existing algorithms only consider faults associated with processing elements, and all switches and links are assumed to be fault-free. But switch faults may often occur in the network-on-chips with high density. In this paper, two novel approaches are proposed to tackle the reconfiguration problem of degradable VLSI arrays with switch faults. The first approach extends the well-known existing algorithm with simple pre-processing and row bypass scheme. The second one employs a novel row and column rerouting scheme to maximize the size of the logical array. Simulation results show that the proposed two approaches can effectively generate the logical arrays on the given host array with switch faults, and the second algorithm performs more favorably with the increasing number of the switch faults. Yuanbo Zhu, Jigang Wu, Siew-Kei Lam, Thambipillai Srikanthan |
ICPADS | 4 |
| 2012 | Exploiting stable features for iris recognition of defocused imagesabstractWe present a novel approach for recognizing defocused iris images captured outside the Depth of Field (DOF) of cameras. Unlike existing approaches, we do not rely on special hardware or on computationally expensive image restoration algorithms. Instead, the proposed recognition approach exploits stable bits in the iris code representation which are robust to imaging noise. Experimental results based on over 15,000 images show that when compared to iris recognition of defocused images that relies on the entire code representation, the proposed method achieves an average recognition performance gain of over 2 times. Due to its low computational requirements, the proposed method is well suited for use as part of a multi-biometric system in ubiquitous systems. Siew-Kei Lam, Thambipillai Srikanthan, Weiqi Yuan |
ISCAS | 3 |
| 2012 | Low-complexity pruning for accelerating corner detectionabstractIn this paper, we present a novel and computationally efficient pruning technique to speed up the Shi-Tomasi and Harris corner detectors. The proposed technique quickly prunes non-corners and selects a small corner candidate set by approximating the complex corner measure of Shi-Tomasi and Harris. The actual corner measure is then applied only to the reduced candidate set. Experimental results on the NiOS-II platform show that the proposed technique achieves an average execution time savings of 90% for Shi-Tomasi and 70% for Harris detectors for 500 corners with no loss in accuracy. Meiqing Wu, Nirmala Ramakrishnan, Siew-Kei Lam, Thambipillai Srikanthan |
ISCAS | 4 |
| 2012 | Detection & classification of arrow markings on roads using signed edge signaturesabstractIn this paper, we propose a novel method to robustly identify and classify arrow markings in road images. In the proposed method, simple and unique signatures are first derived for the various arrow types, based on signed edge maps and decomposing the arrows into smaller parts. The signed edge maps are processed using Hough Transform (HT), and the resulting Hough spaces are analyzed systematically, using a set of simple rules. The signatures are rotation-invariant and scale-invariant, thereby making the approach robust to variations in the appearance of the arrow markings. It is shown that the method yields a high detection and classification accuracy, of as high as 97% in the test images considered. Suchitra Sathyanarayana, Ravi Kumar Satzoda, Thambipillai Srikanthan |
Intelligent Vehicles Symposium | 3 |
| 2012 | Robust extraction of lane markings using gradient angle histograms and directional signed edgesabstractIn this paper, we propose novel block-based techniques for robust extraction of lane marking edges in complex scenarios, such as in the presence of shadows, vehicles, other road markings etc. The techniques are based on the properties of lane markings and involve a two-stage processing: (1) generation of customized edge maps using histograms of gradient angles, and (2) directional signed edges in combination with Hough Transform to identify lane markings. It is shown that the proposed techniques show a detection accuracy of as high as 98% on test data collected on real road scenarios, representing the various complex cases. Ravi Kumar Satzoda, Suchitra Sathyanarayana, Thambipillai Srikanthan |
Intelligent Vehicles Symposium | 3 |
| 2012 | Iris Recognition of Defocused Images for Mobile phonesabstractIn this paper, we introduce a novel iris recognition approach for mobile phones, which takes into account imaging noise arising from image capture outside the depth of field (DOF) of cameras. Unlike existing approaches that rely on special hardware to extend the DOF or computationally expensive algorithms to restore the defocused images prior to recognition, the proposed method performs recognition on the defocused images based on the stable bits in the iris code representation that are robust to imaging noise. To the best of our knowledge, our work is the first to investigate the characteristics of iris features for varying degree of image defocus when the images are captured outside the DOF of cameras. Based on our findings, we present a method to determine the stable bits of an enrolled image. When compared to iris recognition of defocused images that relies on the entire code representation, the proposed recognition method increases the inter-class variability while reducing the intra-class variability of the samples considered. This leads to smaller intersections between the intra-class and inter-class distance distributions, which results in higher recognition performance. Experimental results based on over 15,000 images show that the proposed method achieves an average recognition performance gain of about two times. It is envisioned that the proposed method can be incorporated as part of a multi-biometric system for mobile phones due to its lightweight computational requirements, which is well suited for power sensitive solutions. Siew-Kei Lam, Thambipillai Srikanthan, Weiqi Yuan |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2012 | Psychoacoustic Model Compensation for Robust Speaker Verification in Environmental NoiseabstractWe investigate the problem of speaker verification in noisy conditions in this paper. Our work is motivated by the fact that environmental noise severely degrades the performance of speaker verification systems. We present a model compensation scheme based on the psychoacoustic principles that adapts the model parameters in order to reduce the training and verification mismatch. To deal with scenarios where accurate noise estimation is difficult, a modified multiconditioning scheme is proposed. The new algorithm was tested on two speech databases. The first database is the TIMIT database corrupted with white and pink noise and the noise estimation is fairly easy in this case. The second database is the MIT Mobile Device Speaker Verification Corpus (MITMDSVC) containing realistic noisy speech data which makes the noise estimation difficult. The proposed scheme achieves significant performance gain over the baseline system in both cases. Ashish Panda, Thambipillai Srikanthan |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Accelerating throughput-aware runtime mapping for heterogeneous MPSoCsabstractModern embedded systems need to support multiple time-constrained multimedia applications that often employ multiprocessor-systems-on-chip (MPSoCs). Such systems need to be optimized for resource usage and energy consumption. It is well understood that a design-time approach cannot provide timing guarantees for all the applications due to its inability to cater for dynamism in applications. However, a runtime approach consumes large computation requirements at runtime and hence may not lend well to constrained-aware mapping. In this article, we present a hybrid approach for efficient mapping of applications in such systems. For each application to be supported in the system, the approach performs extensive design-space exploration (DSE) at design time to derive multiple design points representing throughput and energy consumption at different resource combinations. One of these points is selected at runtime efficiently, depending upon the desired throughput while optimizing for energy consumption and resource usage. While most of the existing DSE strategies consider a fixed multiprocessor platform architecture, our DSE considers a generic architecture, making DSE results applicable to any target platform. All the compute-intensive analysis is performed during DSE, which leaves for minimum computation at runtime. The approach is capable of handling dynamism in applications by considering their runtime aspects and providing timing guarantees. The presented approach is used to carry out a DSE case study for models of real-life multimedia applications: H.263 decoder, H.263 encoder, MPEG-4 decoder, JPEG decoder, sample rate converter, and MP3 decoder. At runtime, the design points are used to map the applications on a heterogeneous MPSoC. Experimental results reveal that the proposed approach provides faster DSE, better design points, and efficient runtime mapping when compared to other approaches. In particular, we show that DSE is faster by 83% and runtime mapping is accelerated by 93% for some cases. Further, we study the scalability of the approach by considering applications with large numbers of tasks. Amit Kumar Singh 0002, Akash Kumar 0001, Thambipillai Srikanthan |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2011 | A hybrid strategy for mapping multiple throughput-constrained applications on MPSoCsabstractModern embedded systems are based on Multiprocessor-Systems-on-Chip (MPSoCs) to meet the strict timing deadlines of multiple applications. MPSoC resources must be utilized efficiently by mapping the applications in throughput-aware manner in order to meet throughput constraints for each of them. A design-time methodology is applicable only to predefined set of applications with static behavior, which is incapable of handling dynamism in applications. On the other hand, a run-time approach can cater to the dynamism but cannot provide timing guarantees for all the applications due to large computation requirements at run-time. This paper presents a hybrid flow which performs compute intensive analysis at design-time to derive multiple resource-throughput trade-off points and selects one of these at run-time subject to available resources and desired throughput. Experimental results show that the design-time analysis is faster by 39%, provides better trade-off points and the run-time mapping is speeded up by 93% when compared to state-of-the-art techniques. Amit Kumar Singh 0002, Akash Kumar 0001, Thambipillai Srikanthan |
CASES | 3 |
| 2011 | Architecture-Aware Technique for Mapping Area-Time Efficient Custom Instructions onto FPGAsabstractArea-time efficient custom instructions are desirable for maximizing the performance of reconfigurable processors. Existing data path merging techniques based on resource sharing can be deployed to improve area efficiency of custom instructions. However, these techniques lead to large increase in the critical path delay. In this paper, we propose a novel strategy that takes into account the architectural constraints of the FPGA device in order to realize custom instructions with low-area delay product. The proposed strategy is based on partitioning the custom instruction data paths into a set of basic clusters such that they can be combined using a heuristic-based cluster merging process to maximize the utilization of FPGA logic blocks. Unlike the resource sharing method, the proposed cluster merging process does not maximize sharing of common resources and this leads to lesser reliance on multiplexers for implementing custom instructions. Resource sharing is only applied sparingly at the final stage to increase utilization of logic blocks. We show that the proposed technique leads to more than 34 percent, 34 percent, and 42 percent average reduction in area costs for Spartan-3, Virtex-4, and Virtex-5 architectures, respectively, when compared to optimizations achieved through commercial synthesis tool. We have also shown that the proposed technique leads to more than 18 percent, 17 percent, and 13 percent average reduction in area costs for Spartan-3, Virtex-4, and Virtex-5, respectively, when compared to results obtained using one of the most efficient resource sharing-based method reported in the literature. In addition, the proposed technique outperforms the resource sharing-based method in terms of area-delay product, with average reductions of more than 27 percent, 34 percent, and 19 percent for Spartan-3, Virtex-4, and Virtex-5, respectively. Siew-Kei Lam, Thambipillai Srikanthan, Christopher T. Clarke |
IEEE Trans. Computers | 2 |
| 2011 | Accelerating UNISIM-Based Cycle-Level Microarchitectural Simulations on Multicore PlatformsabstractUNISIM has been shown to ease the development of simulators for multi-/many-core systems. However, UNISIM cycle-level simulations of large-scale multiprocessor systems could be very time consuming. In this article, we propose a systematic framework for accelerating UNISIM cycle-level simulations on multicore platforms. The proposed framework relies on exploiting the fine-grained parallelism within the simulated cycles using POSIX threads. A multithreaded simulation engine has been devised from the single-threaded UNISIM SystemC engine to facilitate the exploitation of inherent parallelism. An adaptive technique that manages the overall computation workload by adjusting the number of threads employed at any given time is proposed. In addition, we have introduced a technique to balance the workloads of multithreaded executions. This load balancing involves the distributions of SystemC objects among threads. A graph-partitioning-based technique has been introduced to automate such distributions. Finally, two strategies are proposed for realizing nonautomated and fully automated adaptive multithreaded simulations, respectively. Our investigations show that notable acceleration can be achieved by deploying the proposed framework. In particular, we show that simulations on an 8-core multicore platform can provide for close to 6X speedups when simulating many-core systems with large number of cores. Xiongfei Liao, Thambipillai Srikanthan |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2010 | Performance estimation framework for FPGA-based processorsabstractModern FPGA devices can implement a variety of processors with numerous configurable options. Rapid performance estimation of FPGA processors plays a vital role in embedded systems design to select a processor that best fits the application requirements. Traditional performance evaluation techniques such as running the software application on the target processor or using cycle accurate instruction set simulator are time-consuming and poses a threat in meeting the stringent time-to-market pressure. In this paper, we propose a framework to rapidly estimate the performance of a wide range of FPGA processors. The proposed method relies on the LLVM compiler infrastructure and its backend code generator to accurately estimate the software performance within seconds. Experimental results show that the proposed framework can reliably estimate the performance of a widely used FPGA processor with an average accuracy of over 90% for a number of benchmark applications. Yan Lin Aung, Siew-Kei Lam, Thambipillai Srikanthan |
FPT | 3 |
| 2010 | Mapping real-life applications on run-time reconfigurable NoC-based MPSoC on FPGAabstractMultiprocessor systems-on-chip (MPSoC) are required to fulfill the performance demand of modern real-life embedded applications. These MPSoCs are employing Network-on-Chip (NoC) for reasons of efficiency and scalability. Additionally, these systems need to support run-time reconfiguration of their components to cater to dynamically changing demands of the system. Designing and programming such systems for real-life applications prove to be a major challenge. This paper demonstrates the designing of reconfigurable NoC-based MPSoC and programming it for real-life applications. The NoC is reconfigured at run-time to support different combinations of multiple applications at different times. The platform is verified with a case study executing the parallelized C-codes of a simple producer-consumer and JPEG decoder applications on a NoC-based MPSoC on a Xilinx FPGA. Based on our investigations to map the applications on a 3 × 3 platform, we show that the NoC reconfiguration overhead is kept at a minimum and the platform utilizes 85% of the total available slices of Virtex-5 FPGA. Moreover, we show that the proposed approach is highly scalable when targeting for large number of applications. Amit Kumar Singh 0002, Akash Kumar 0001, Thambipillai Srikanthan, Yajun Ha |
FPT | 3 |
| 2010 | An efficient edge and corner detectorabstractThis paper describes an efficient method for image edge and corner detection. Edges are detected before extracting corner points so that the background or noise of the image can be effectively removed. We propose an edge detector that utilizes edge features to localize the edge points. The corner points of an image are then identified by evaluating the degree of turning in the edge direction. The proposed edge and corner detection utilizes low complexity computations. Based on the experiments considered, it can be observed that the proposed edge and corner detector achieves more accurate results than several other well known algorithms. In addition, the proposed method performs faster than most of the other well known algorithms. Siew-Kei Lam, Thambipillai Srikanthan |
ICARCV | 3 |
| 2010 | Selecting profitable custom instructions for reconfigurable processors
Tao Li 0008, Jigang Wu, Siew-Kei Lam, Thambipillai Srikanthan, Xicheng Lu |
J. Syst. Archit. | 4 |
| 2010 | Communication-aware heuristics for run-time task mapping on NoC-based MPSoC platforms
Amit Kumar Singh 0002, Thambipillai Srikanthan, Akash Kumar 0001, Jigang Wu |
J. Syst. Archit. | 2 |
| 2010 | Algorithmic Aspects of Hardware/Software Partitioning: 1D Search AlgorithmsabstractHardware/software (HW/SW) partitioning is one of the key challenges in HW/SW codesign. This paper presents efficient algorithms for the HW/SW partitioning problem, which has been proved to be NP-hard. We reduce the HW/SW partitioning problem to a variation of knapsack problem that is approximately solved by searching 1D solution space, instead of searching 2D solution space in the latest work cited in this paper, to reduce time complexity. Three heuristic algorithms are proposed to determine suitable partitions to satisfy HW/SW partitioning constraints. We have shown that the time complexity for partitioning a graph with n nodes and m edges is significantly reduced from O(dx· dy· n3) to O(n log n + d · (n + m)), where d and dx· dyare the number of the fragments of the searched 1D solution space and the searched 2D solution space, respectively. The lower bound on the solution quality is also proposed based on the new computing model to show that it is comparable to that reported in the literature. Moreover, empirical results show that the proposed algorithms produce comparable and often better solutions when compared to the latest algorithm while reducing the time complexity significantly. Jigang Wu, Thambipillai Srikanthan |
IEEE Trans. Computers | 2 |
| 2010 | Preprocessing and Partial Rerouting Techniques for Accelerating Reconfiguration of Degradable VLSI ArraysabstractThis paper presents novel techniques to accelerate the reconfiguration of degradable very large scale integration arrays. A preprocessing step is used to derive the upper and lower size bounds of the maximum logical array (MLA) such that only those subarrays that possibly contain the MLA are reconfigured, thereby reducing the reconfiguration time and also obtaining a same-sized logical array. In addition, the partial rerouting approach is generalized so that as many as possible previous routing results can be reused in the current rerouting step. The reconfiguration time is reduced from$ O((1-\rho)\cdot \beta \cdot m \cdot n)$to its lower bound$ O((1-\rho)\cdot m\cdot n)$for$ m\times n$host arrays with small fault density$ \rho $, where$ \beta $is the expected routing length required per logical column. Jigang Wu, Thambipillai Srikanthan, Xiaogang Han |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2009 | Mapping Algorithms for NoC-Based Heterogeneous MPSoC PlatformsabstractMapping of applications onto multiprocessor system-on-chip (MPSoC) can be realized either at design-time or run-time. At any time the number of tasks executing in MPSoC platform can exceed the available resources, requiring efficient run-time mapping techniques to meet the real-time constraints of the applications. This paper presents two run-time mapping heuristics for mapping the tasks of an application in close proximity so as to minimize the communication overhead. In particular, the communication overhead between two adjacent hardware tasks is eliminated by mapping them onto the same reconfigurable processing node. We show that the proposed approach is capable of alleviating network-on-chip (NoC) congestion bottlenecks to optimize the overall performance. Based on our investigations to map the tasks of applications' at run-time onto an 8×8 NoC-based heterogeneous MPSoC, our mapping heuristics are capable of reducing total execution time and average channel load of applications when compared to state-of-the-art runtime mapping heuristics. Amit Kumar Singh 0002, Jigang Wu, Alok Prakash, Thambipillai Srikanthan |
DSD | 4 |
| 2009 | Rapid design exploration framework for application-aware customization of soft core processorsabstractOff-the-shelf soft core processors are becoming increasingly popular in embedded systems design today as they provide for application specific customization, in particular through instruction subsetting. However, choosing the right processor configuration remains a challenge as the search space becomes prohibitively large when the configurable options increase. In this paper we propose a framework to rapidly explore the processor configuration design space for a given application. Unlike existing approaches that require time-consuming synthesis process, the proposed method relies only on a single-pass output of the LLVM compiler infrastructure. Experimental results based on widely used benchmarks show that the proposed framework can reliably predict the actual performance and area trends of various configurable options. Alok Prakash, Siew-Kei Lam, Amit Kumar Singh 0002, Thambipillai Srikanthan |
FPL | 4 |
| 2009 | Gradient angle histograms for efficient linear Hough transformabstractNon-collinear edge pixels are equivalent to noise for the linear Hough transform (LHT). Existing methods that reduce the number of points for Hough voting are based on random and/or probabilistic selection. Such methods select both collinear and noisy pixels, thereby incurring unwanted computational costs. In this paper, we propose a novel gradient angle histogram based technique to generate modified straight line edge map (SLEM), which largely retains the straight line edges and eliminates noisy edge pixels. A block-based SLEM generation is proposed to increase the robustness of straight line extraction and validated on test images. Further, effect of varying block sizes on accuracy of straight line detection is studied and appropriate block settings are derived. The proposed gradient angle histogram based method reduces the number of edge pixels by as much as 85%. Ravi Kumar Satzoda, Suchitra Sathyanarayana, Thambipillai Srikanthan |
ICIP | 3 |
| 2009 | Modeling RTOS Components for Instruction Cache Hit Rate EstimationabstractRTOS components have a direct impact on the instruction cache performance - an aspect that has been reported as a bottleneck in several studies. Therefore, there exists a need for insight into the instruction cache hit rates for various RTOS components at design time for efficient design space exploration. In this paper, we propose a technique to model RTOS components for instruction cache hit rate estimation. Our methods rely on rapid generation of hit rate values for different cache sizes for the RTOS components. We then fit the generated instruction cache hit rates using multivariate regression schemes to account for all parameters influencing the hit rate. The estimation techniques were tested for a large range of cache sizes and numerous parametric as well as non-parametric RTOS components. Comparison with hit rates generated by a cache simulator shows that these models can accurately estimate the hit rates with the mean difference in hit rates ranging from 0.00 to 0.05. The proposed technique also offers significant speed-up as compared to other modeling approaches because it does not require exhaustive cache hierarchy simulation for building the models. Santanu Kumar Dash 0001, Thambipillai Srikanthan |
ISCAS | 2 |
| 2009 | Techniques for Area-time Efficient Image RotationabstractReal-time image rotation is a necessary task in many vision based applications. Most available image rotation architectures are serial in nature, while the existing parallel techniques suffer from computational bottlenecks, limiting their performance. In this paper, we propose techniques that not only speed up the rotation process but also reduce area, yielding an area-time efficient architecture. They exploit properties of symmetry in image coordinates for parallel image rotation. We show that the number of computations, the total computation time and area cost of the proposed architecture are lesser by as much as 80%, 55% and 67% respectively when compared to the most efficient parallel architecture in recent literature. Ravi Kumar Satzoda, Suchitra Sathyanarayana, Thambipillai Srikanthan |
ISCAS | 3 |
| 2009 | Minimizing interconnect length on reconfigurable meshes
Jigang Wu, Thambipillai Srikanthan |
Frontiers Comput. Sci. China | 2 |
| 2009 | Rapid design of area-efficient custom instructions for reconfigurable embedded processing
Siew-Kei Lam, Thambipillai Srikanthan |
J. Syst. Archit. | 2 |
| 2009 | Exploiting Inherent Parallelisms for Accelerating Linear Hough TransformabstractAccelerating Hough transform in hardware has been of interest due its popularity in real-time capable image processing applications. In most existing linear Hough transform architectures, an m xm edge map is serially read for processing, resulting in a total computation time of at least m(2) cycles. In this paper, we propose a novel parallel Hough transform computation method called the Additive Hough transform (AHT), wherein the image is divided using a k x k grid to reduce the total computation time by a factor of k(2). We have also proposed an efficient implementation of the AHT consisting of a look-up table (LUT) and two-operand adder arrays for every angle. Techniques to condense the LUT size have also been proposed to further reduce area utilization by as much as 50%. Our investigations based on employing an 8 x 8 grid shows a 1000 x speedup compared to existing architectures for a range of image sizes. Area-time trade-off analysis has been presented to demonstrate that the area-time product of the proposed AHT-based implementation is at least 43% lower than other implementations reported in the literature. We have also included and characterized a hierarchical addition step in order to generate a global accumulation space equivalent to that of the conventional HT. It is shown that the proposed implementation with the hierarchical addition step remains superior to other methods in terms of both performance and area-time product metrics. Finally, we show that the proposed solution is equally efficient when applied on rectangular images. Suchitra Sathyanarayana, Ravi Kumar Satzoda, Thambipillai Srikanthan |
IEEE Trans. Image Process. | 3 |
| 2008 | Rapid estimation of instruction cache hit rates using loop profilingabstractEstimation of the hit rate curve for an application is the first step in application specific cache tuning. Several techniques have been proposed to meet this objective however most of these have dealt with the data cache with little attention to the instruction cache. In this paper, we propose a novel, lightweight and highly scalable technique for rapid estimation of the instruction cache hit rate curve for a given application. Our technique works at the basic block level and relies on a one-time loop profiling of the weighted control flow graph of the application followed by estimation of the hit rate for different cache sizes. It accounts for the spatial and temporal locality separately and is sensitive to the cache size as well as block size. The proposed technique is highly accurate and when compared with results from an actual cache simulator, the mean error in estimation ranged from 1.11% to 2.46% for the benchmarks tested. Santanu Kumar Dash 0001, Thambipillai Srikanthan |
ASAP | 2 |
| 2008 | A Short Course on Implementing FPGA Based Digital SystemsabstractThe rapid advances in the FPGA technology along with high-levels of system integration have made FPGAs the preferred platform not only for rapid prototyping but also for production of digital embedded systems. This paper presents the experience of a team of instructors in designing and conducting a short course on implementing FPGA-based digital systems for industry professionals. The selection of topics, course organization, the issues involved in designing effective hands-on exercises and the response of the students to the course are discussed. George Rosario Jagadeesh, Siew-Kei Lam, Thambipillai Srikanthan |
ICPADS | 3 |
| 2008 | Finding minimum interconnect sub-arrays in reconfigurable VLSI arraysabstractShorter total interconnect and fewer switches in a VLSI array definitely lead to less capacitance, power dissipation and dynamic communication cost between the processing elements (PEs). This paper presents techniques to find a logical (target) array that has shorter interconnect and fewer switches in a reconfigurable VLSI array with faulty PEs. The proposed algorithm initially searches for a sub-array on the host array, which contains the minimum number of the faults. Then it reroutes the sub-array, rather than the whole host array as was done in previous algorithm, to an approximate target array whose size is less then but close to the size of the target array. Finally, the target array is obtained by simple extension of the approximate target array. Experimental results show that the proposed algorithm can construct a target array with much shorter total interconnect. The improvement over the previous work is up to 68% in terms of the interconnect redundancy for the case of the cluster faults. Jigang Wu, Thambipillai Srikanthan |
ISCAS | 2 |
| 2008 | A temperature-aware virtual submesh allocation scheme for noc-based manycore chipsabstractVarious continuous and non-continuous submesh allocation schemes have been proposed for traditional mesh-connected multiprocessor systems. Discussions of these schemes are limited to issues such as fragmentation, system performance and algorithmic complexity. NoC-based manycore chips with mesh topology are envisaged to enter the domain of parallel computing and cores of them are allocated to applications in the forms of submeshes. However, unlike the processors in traditional systems, these cores may face severe runtime thermal non-uniformities and a promising strategy to avoid these thermal crises is keeping good heat balance throughout chips at runtime. One opportunity for good heat balance is to include thermally favorable cores when submeshes are allocated. This requires a suitable temperature sensitive submesh allocation scheme for manycore chips. Xiongfei Liao, Jigang Wu, Thambipillai Srikanthan |
SPAA | 3 |
| 2008 | Energy-efficient cluster-based scheme for failure management in sensor networksabstractWireless sensor networks are characterised by dense deployment of energy constrained nodes. Owing to the deployment of large number of sensor nodes in uncontrolled hostile environments and unmonitored operation, it is common for the nodes to exhaust its energy and become inactive. The failing nodes create holes in the network topology causing connectivity loss, which may lead to critical information loss. To avoid degradation of performance, it is necessary that the failures are detected well in advance and appropriate measures are taken to sustain the network operation. An energy-efficient cluster-based technique is proposed to detect failures and recover the cluster structure. The proposed technique relies on the cluster members to detect the failures in the cluster and recover the connectivity. The proposed failure detection and recovery technique recovers the cluster structure in less than one-fourth of the time taken by the Gupta algorithm and is also proven to be 70% more energy-efficient than the same. The proposed cluster-based failure detection and recovery scheme proves to be an efficient and quick solution to robust and scalable sensor network for long and sustained operation. Gayathri Venkataraman, Sabu Emmanuel, Thambipillai Srikanthan |
IET Commun. | 3 |
| 2008 | New Model and Algorithm for Hardware/Software Partitioning
Jigang Wu, Thambipillai Srikanthan, Guang-Wei Zou |
J. Comput. Sci. Technol. | 2 |
| 2008 | Parallelizing the Hough Transform ComputationabstractThe Hough transform, a widely used operation in image processing, suffers from being computationally intensive. In this letter, we propose a method called additive Hough transform (AHT) to accelerate Hough transform computation. It is based on parallel processing of k2points, obtained by dividing the edge map into uniform blocks using a k times k grid. It has been shown that AHT on verification gives the same Hough profile as the conventional Hough transform. Further, AHT reduces the total computation time by at least k2times as compared to existing Hough transform architectures. Ravi Kumar Satzoda, Suchitra Sathyanarayana, Thambipillai Srikanthan |
IEEE Signal Process. Lett. | 3 |
| 2007 | Estimating Area Costs of Custom Instructions for FPGA-based Reconfigurable ProcessorsabstractFPGA (field programmable gate array) based reconfigurable processor has been shown to meet the increasingly challenging performance targets and shorter time-to-market pressures. In this paper, we propose a method to rapidly estimate the FPGA area costs of custom instructions without the need for hardware synthesis. The proposed estimation technique relies on a novel approach to partition the custom instruction data-paths into a set of clusters, where each cluster can be realized using an FPGA logic element or a coarse-grained arithmetic unit. Experiments based on 20 custom instructions reveal that the estimation results show an average of only 8% increase in the area costs when compared with the corresponding hardware synthesized results. In addition, we show that the maximum FPGA area utilized by custom instructions of each of the seven applications examined is equivalent to about 1000 Xilinx FPGA logic elements. Siew-Kei Lam, Thambipillai Srikanthan |
ASAP | 2 |
| 2007 | Temperature-Aware Submesh Allocation Scheme for Heat Balancing on Chip-MultiprocessorsabstractThis paper explores the thermal problems in future CMPs in multiprogrammed environment for heat balancing. We first give the observation of the temperature variation of cores in this scenario. Then we propose a temperature-aware submesh allocation scheme to manage cores with submeshes and allocate submeshes of cores to jobs under temperature-aware policies to balance heat chip-wide. Several scheduling policies are suggested and a HotSpot-based thermal simulator is used to evaluate the scheme and its policies under the workloads of benchmark programs. Simulation results show that our proposed scheme with global coolest policy and global neighbor-aware policy can lead to lower peak temperatures and effectively reduce the temporal variance and spatial variance of temperatures of cores to achieve better heat balance. Xiongfei Liao, Jigang Wu, Thambipillai Srikanthan |
ASAP | 3 |
| 2007 | An Embedded Systems graduate education for Singapore
Ian McLoughlin 0001, Douglas L. Maskell, Thambipillai Srikanthan, Wooi-Boon Goh |
ICPADS | 3 |
| 2007 | One-dimensional Search Algorithms for Hardware/Software PartitioningabstractHardware/software (HW/SW) partitioning is one of the key challenges in HW/SW co-design. This paper presents a new formulation to handle the HW/SW partitioning problem, which has been proved to be NP-hard. The proposed formulation transforms the partitioning problem into an extended 0-1 knapsack problem that is approximately solved in this paper by scanning a one-dimensional search space, instead of scanning a two-dimensional search space as presented in the literature cited in this paper. Two heuristic algorithms are proposed to explore the feasible partitions to meet the given constraints. The time complexity of the latest heuristic algorithm is significantly reduced from O(dxldr dyldr n3) to O(n log n + d ldr (n + m)) for the given graphs with n nodes and m edges, where dxldr dyis the number of the fragments of the scanned two-dimensional search space, and d is that of the scanned one-dimensional search space. Empirical results show that the proposed algorithms run extremely fast and still produce better or similar solutions in comparison with the latest algorithm. Jigang Wu, Thambipillai Srikanthan |
MEMOCODE | 2 |
| 2007 | Integrated Row and Column Rerouting for Reconfiguration of VLSI Arrays with Four-Port SwitchesabstractThis paper deals with the issue of developing efficient algorithms for reconfiguring two-dimensional VLSI arrays linked by four-port switches in the presence of faulty processing elements (PEs). The proposed algorithm reroutes the arrays with faults in both row and column directions at the same time. Unlike previous work, the compensation technique to replace the faulty PE is not restricted to the adjacent rows of the excluded row. Instead, we consider the neighbor rows of any faulty PE for compensation purposes. The nonfaulty PEs lying in the excluded rows are also effectively utilized to form the maximal target arrays, making the proposed algorithm more efficient in terms of both the percentages of harvest and the degradation of VLSI arrays for random and clustered faults. Empirical study shows that the improvement in harvest increases with increasing fault size and is more notable for maximal square target arrays than for maximal target arrays. Our investigations show that the improvement can be up to 8 percent and 23 percent for a 256 times 256 VLSI array with random faults of size 25 percent for maximal target arrays and for maximal square target arrays, respectively. Jigang Wu, Thambipillai Srikanthan |
IEEE Trans. Computers | 2 |
| 2006 | Efficient algorithm for functional scheduling in hardware/software co-designabstractTask scheduling is one of the crucial steps during functional hardware/software co-design. Due to the possibly concurrent execution of the tasks implemented in hardware, the NP-hard scheduling problem becomes more difficult to solve optimally. In this paper an efficient algorithm is proposed for task scheduling in functional hadware/software co-design. The proposed algorithm assigns the priority for each task combining the information both in the communication penalty and the hardware-only critical path, to enhance the parallelism of the tasks. A large body of experimental results confirm that the proposed algorithm is superior to the most widely used approaches first-come first-schedule(FCFS) and level-by-level schedule (LBLS) in hardware/software scheduling, both for random graphs and some realistic application graphs, without large increase in running time. The improvement over FCFS and LBLS is up to 10% for some random graphs, and it is more significant for FFT application graphs, according to the simulation results on the same types of graphs (under the same assumptions) as in the literature where LBLS is employed Jigang Wu, Thambipillai Srikanthan, Tao Jiao |
FPT | 2 |
| 2006 | Efficient management of custom instructions for run-time reconfigurable instruction set processorsabstractThe instruction set extension capability of RISPs (reconfigurable instruction set processors) provides an attractive means to meet the flexibility, performance, and cost demands of ubiquitous computing devices. Run-time reconfiguration can further increase the cost efficiency and hardware specialization of these processors by dynamically changing the configuration of the reconfigurable logic to the required functionality. In this paper, we propose the use of a heuristic that leads to the selection of large custom instructions for increased performance gain. Result analysis of six applications from the MiBench embedded benchmark suite show that efficient data-path merging can be applied to the custom instructions to reduce the average number of configurations to less than 8 in a run-time RISP. In addition, there is only a small difference in the average number of configurations when compared to a custom instruction selection strategy that results in lower performance Siew-Kei Lam, Bharathi N. Krishnan, Thambipillai Srikanthan |
FPT | 3 |
| 2006 | Low Area-time Complexity Averaging Scheme for Thumbnail GenerationabstractImage preprocessing is an important step in most real-time neural-network based pattern recognition systems. Thumbnail generation is a commonly used preprocessing method to reduce the hardware complexity of the input layer of a neural network. A direct implementation of the averaging method to create thumbnails incurs high hardware cost. In this paper we propose an optimized averaging method that outperforms standard averaging method in terms of device utilization, maximum operating frequency and latency, when implemented using FPGA design methodology. It eliminates the use of accumulators and dividers and offers hardware savings of as much as 96% over standard averaging method. Also, the proposed averaging is performed in a single cycle making it more suitable for real-time implementations Ravi Kumar Satzoda, Suchitra Sathyanarayana, Thambipillai Srikanthan |
ICARCV | 3 |
| 2006 | Area and delay estimation for FPGA implementation of coarse-grained reconfigurable architecturesabstractReconfigurable architecture is one solution to the increasing computational requirement that often cannot be met by the low-end embedded processors. Compiling applications to such architectures involves hardware/software partitioning. To partition the applications, a set of parameters, such as the hardware execution time and hardware area consumption, is required for each application block. Quick derivation of the parameters for all the blocks is essential. Previous research has shown that the coarse-grained reconfigurable architectures are able to accelerate the applications. However, no research effort has been made to find the area and time for application blocks implemented on such architectures. In this paper we present an estimation model for the coarse-grained reconfigurable architectures implemented on FPGA platforms. The estimation model is able to quickly produce an area-time graph, which shows the area and time relationship, for each application block. The accuracy of the estimation model has been verified on real applications. Experiment shows that the estimation error for the area consumption is within 13% and the estimation error for the time is within 8%. Leipo Yan, Thambipillai Srikanthan, Niu Gang |
LCTES | 2 |
| 2006 | Low-complex dynamic programming algorithm for hardware/software partitioning
Jigang Wu, Thambipillai Srikanthan |
Inf. Process. Lett. | 2 |
| 2006 | An efficient algorithm for the collapsing knapsack problem
Jigang Wu, Thambipillai Srikanthan |
Inf. Sci. | 2 |
| 2006 | Reconfiguration Algorithms for Power Efficient VLSI Subarrays with Four-Port SwitchesabstractTechniques to determine subarrays when processing elements of VLSI arrays become faulty have been investigated extensively. These tend to identify the largest subarray that is possible without concentrating on the power efficiency of the resulting subarray. In this paper, we propose new techniques, based on heuristic strategy and dynamic programming, to minimize the interconnect length in an attempt to reduce power dissipation without performance penalty. Our algorithms show that notable improvements in the reduction of the number of long interconnects could be realized in linear time and without sacrificing the size of the subarray. Our evaluations show that, for a VLSI array of size 256/spl times/256, the number of long interconnects in the subarray can be reduced by up to 95 percent for clustered faults and up to 50 percent and 73 percent for a random fault with density of 10 percent and 0.1 percent, respectively, when compared with the most efficient implementation cited in the literature. The interconnect power saving for a VLSI array of size 512/spl times/512 is by up to 11 percent for a random fault. We have also shown that interconnect power savings of up to 14 percent are possible for the cases investigated. Simulations based on several random and clustered fault scenarios clearly reveal the superiority of the proposed techniques for power efficient realizations. In addition, the lower bound of the performance has been proposed to demonstrate that the proposed algorithms are nearly optimal for the cases considered in this paper. Jigang Wu, Thambipillai Srikanthan |
IEEE Trans. Computers | 2 |
| 2006 | Algorithmic aspects of area-efficient hardware/software partitioning
Jigang Wu, Thambipillai Srikanthan |
J. Supercomput. | 2 |
| 2005 | Efficient Techniques and Hardware Analysis for Mesh-Connected Processors
Jigang Wu, Thambipillai Srikanthan |
ICA3PP | 2 |
| 2005 | Power Efficient Sub-Array in Reconfigurable VLSI Meshes
Jigang Wu, Thambipillai Srikanthan |
J. Comput. Sci. Technol. | 2 |
| 2005 | New adaptive color quantization method based on self-organizing mapsabstractColor quantization (CQ) is an image processing task popularly used to convert true color images to palletized images for limited color display devices. To minimize the contouring artifacts introduced by the reduction of colors, a new competitive learning (CL) based scheme called the frequency sensitive self-organizing maps (FS-SOMs) is proposed to optimize the color palette design for CQ. FS-SOM harmonically blends the neighborhood adaptation of the well-known self-organizing maps (SOMs) with the neuron dependent frequency sensitive learning model, the global butterfly permutation sequence for input randomization, and the reinitialization of dead neurons to harness effective utilization of neurons. The net effect is an improvement in adaptation, a well-ordered color palette, and the alleviation of underutilization problem, which is the main cause of visually perceivable artifacts of CQ. Extensive simulations have been performed to analyze and compare the learning behavior and performance of FS-SOM against other vector quantization (VQ) algorithms. The results show that the proposed FS-SOM outperforms classical CL, Linde, Buzo, and Gray (LBG), and SOM algorithms. More importantly, FS-SOM achieves its superiority in reconstruction quality and topological ordering with a much greater robustness against variations in network parameters than the current art SOM algorithm for CQ. A most significant bit (MSB) biased encoding scheme is also introduced to reduce the number of parallel processing units. By mapping the pixel values as sign-magnitude numbers and biasing the magnitudes according to their sign bits, eight lattice points in the color space are condensed into one common point density function. Consequently, the same processing element can be used to map several color clusters and the entire FS-SOM network can be substantially scaled down without severely scarifying the quality of the displayed image. The drawback of this encoding scheme is the additional storage overhead, which can be cut down by leveraging on existing encoder in an overall lossy compression scheme. Chip-Hong Chang, Pengfei Xu 0009, Thambipillai Srikanthan |
IEEE Trans. Neural Networks | 4 |
| 2005 | Efficient reconfigurable techniques for VLSI arrays with 6-port switchesabstractThis paper proposes an efficient techniques to reconfigure a two-dimensional degradable very large scale integration/wafer scale integration (VLSI/WSI) array under the row and column routing constraints, which has been shown to be NP-complete. The proposed VLSI/WSI array consists of identical processing elements such as processors or memory cells embedded in a 6-port switch lattice in the form of a rectangular grid. It has been shown that the proposed VLSI structure with 6-port switches eliminates the need to incorporate internal bypass within processing elements and leads to notable increase in the harvest when compared with the one using 4-port switches. A new greedy rerouting algorithm and compensation approaches are also proposed to maximize harvest through reconfiguration. Experimental results show that the proposed VLSI array with 6-port switches consistently outperforms the most efficient alternative proposed in literature, toward maximizing the harvest in the presence of fault processing elements. Jigang Wu, Thambipillai Srikanthan |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2004 | Decode filter cache for energy efficient instruction cache hierarchy in super scalar architectures
Kugan Vivekanandarajah, Thambipillai Srikanthan, Saurav Bhattacharyya |
ASP-DAC | 2 |
| 2004 | Dynamic Filter Cache for Low Power Instruction Memory HierarchyabstractFilter cache (FC) is effective in achieving energy saving at the expense of some performance degradation. The energy savings, here, comes from repeated execution of tiny loops from energy efficient FC. The absence of cacheable loops leads to performance degradation in such FC structures. Therefore, we propose a simple dynamic FC scheme, which detects the opportunity for use of the FC and enables (or disables) it dynamically. Thus providing (slightly reduced) energy savings at minimal performance degradation. A combination of the predictive filter cache with the above schemes reduces the performance and energy penalty. For the benchmarks simulated with 256 byte FC, the average performance degradation is 1.13% with the proposed scheme compared to 2.47% with just the predictive filter cache. With the same configuration, the resulting energy reduction is 42.77%. Finally, the proposed dynamic filter cache scheme is inherently simple and hence it lends well for VLSI efficient implementation. Kugan Vivekanandarajah, Thambipillai Srikanthan, Saurav Bhattacharyya |
DSD | 2 |
| 2004 | RTOS acceleration on soft-core processors using instruction set customizationabstractConfigurable system-on-chip often uses a soft-core programmable processor, running a real-time operating system (RTOS). However, the CPU overheads imposed by the RTOS increase rapidly with the number of tasks in the system. In This work, we propose instruction set customization as the means to contain RTOS-imposed CPU overheads. We present the design of custom instructions for frequently used RTOS primitives and present our area-time results for a variety of configurations. We show that complex embedded systems stand to benefit significantly from our approach. Individual RTOS routines show performance improvements in the range of 50-90%. Frequently used RTOS primitives showed an improvement of 10-35% and the Dhrystone mark of the system improved by as much as 13%. Jin Zhenyu, Mohit Sindhwani, Thambipillai Srikanthan |
FPT | 3 |
| 2004 | High-throughput image rotation using sign-prediction based redundant cordic algorithm
Suchitra Sathyanarayana, Siew-Kei Lam, Thambipillai Srikanthan |
ICIP | 3 |
| 2004 | An efficient data structure for branch-and-bound algorithm
Jigang Wu, Thambipillai Srikanthan |
Inf. Sci. | 2 |
| 2004 | Area-Time Efficient Sign Detection Technique for Binary Signed-Digit Number SystemabstractComputer arithmetic operations based on the BSD (binary signed-digit) number representation system lend themselves well to high-speed computations due to the facilitation of limited carry addition/subtraction. We propose an area-time efficient method for sign detection in a BSD number system based on optimized reverse tree structure. When compared to other popular approaches, such as the most significant carry detection-based CLA (carry look-ahead) and MRC (multilevel reverse carry) implementations, the proposed method is superior to both area and time costs in VLSI. Synthesis results for different word lengths show that the proposed approach to sign detection in the BSD number system continues to maintain its advantage over area and time measures. Thambipillai Srikanthan, Siew-Kei Lam, Mishra Suman |
IEEE Trans. Computers | 1 |
| 2003 | An adaptive initialization technique for color quantization by self organizing feature mapabstractAn unsupervised learning network, such as the self organizing feature map (SOFM), has been applied successfully to color classification for image compression and pattern recognition. Like other vector quantization algorithms, the reconstruction quality and adaptation rate of the SOFM are sensitive to the neuron initialization. We propose an efficient new initialization method, whereby an excess number of neurons is defined and the neurons are adaptively pruned, merged and split within their lattice according to the spatial distribution of the input color pixels. Comparisons with conventional gray scale initialization using subsampling and butterfly jumping sequences show that the proposed method obtains good initial code vectors that can accelerate the convergence of the SOFM and improve the reconstructed image quality significantly. Chip-Hong Chang, Thambipillai Srikanthan |
ICASSP (3) | 3 |
| 2003 | Vector quantization techniques for GMM based speaker verificationabstractThis paper explores the novel application of two vector quantization algorithms, namely Linde, Buzo, Gray (1980) and K-means algorithm for efficient speaker verification. Automatic speaker verification (ASV) is a memory and compute intensive process, giving rise to area and latency concerns in the way of its implementation for real-time efficient embedded systems. The training schemes for computing the speaker models, such as the expectation maximization are highly iterative and contribute significantly to the overall complexity in the implementation of the system. We demonstrate the use of the LBG and the K-means algorithm to realize compute efficient training method. Models trained with the LBG algorithm achieves as much as 99.88% of EM accuracy, whilst K-means achieves as much as 99.91% of EM accuracy. Moreover, the EM computational complexity is almost twice that of LBG or K-means. Thus, using LBG and K-means algorithms for training Gaussian mixture speaker models for text-independent speaker verification, we show that, that they deliver comparable performance as the EM algorithm at significantly reduced computational complexity. Thus making them an ideal choice for low-cost applications. Ashish Panda, Saurav Bhattacharyya, Thambipillai Srikanthan |
ICASSP (2) | 4 |
| 2003 | An improved reconfiguration algorithm for degradable VLSI/WSI arrays
Jigang Wu, Thambipillai Srikanthan |
J. Syst. Archit. | 2 |
| 2003 | A VLSI architecture for 3-D self-organizing map based color quantization and its FPGA implementation
N. Sudha, Thambipillai Srikanthan, Babu Mailachalam |
J. Syst. Archit. | 2 |
| 2002 | New Architecture and Algorithms for Degradable VLSI/WSI Arrays
Jigang Wu, Thambipillai Srikanthan |
COCOON | 3 |
| 2002 | Segmenting endoscopic images using adaptive progressive thresholding: a hardware perspective
Vijayan K. Asari, Thambipillai Srikanthan |
J. Syst. Archit. | 2 |
| 2002 | Heuristic techniques for accelerating hierarchical routing on road networksabstractThe route computation module is one of the most important functional blocks in a dynamic route guidance system. Although various algorithms exist for finding the shortest path, their performance tends to deteriorate as the network size increases. We present an efficient hierarchical routing algorithm that finds a near-optimal route and evaluate it on a large city road network. Solutions provided by the hierarchical routing algorithm are compared with the optimal solutions to analyze and quantify the loss of accuracy. We propose a novel yet simple heuristic to substantially improve the performance of the hierarchical routing algorithm with acceptable loss of accuracy. A network pruning technique has been incorporated into the algorithm to reduce the search space and the correctness of the results is evaluated. The improved hierarchical routing algorithm that incorporates the heuristic techniques has been found to be over 50 times faster than a typical shortest path algorithm. George Rosario Jagadeesh, Thambipillai Srikanthan, K. H. Quek |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2000 | Dynamic multicast routing in VLSI
Siew-Kei Lam, Thambipillai Srikanthan |
Comput. Commun. | 2 |
| 1999 | A Reverse Converter for the 4-moduli Superset {2n-1, 2n, 2n+1, 2n+1+1}abstractThe authors propose an extension to the popular {2/sup n/-1, 2/sup n/, 2/sup n/+1} moduli set by adding a fourth modulus "2/sup n+1/+1. This extension leads to higher parallelism while keeping the forward conversion and modular arithmetic units simple. The main challenge of efficient reverse conversion is met by three techniques described for the first time. Firstly, we reverse convert linear combinations of moduli hence reducing the number of non-zero bits in the Booth encoded multiplicands from n to merely 2. Secondly, it is shown that division by 3, if introduced at the right stage, can be implemented very efficiently and can, in turn, reduce the cost of the converter. To implement VLSI efficient modulo reduction, we propose two techniques-multiple split tables (MST) and a modified division algorithm (MDA). It is shown that the MST can reduce exponential ROM requirements to quadratic ROM requirements while the MDA can reduce these further to linear requirements. As a result of these innovations, the proposed reverse converter uses simple shift and add operations and needs a lookup with only 6 entries. The delay of the converter is approximately 10n+13 full adder delays and the area cost is quadratic in n. Manish Bhardwaj, Thambipillai Srikanthan, Christopher T. Clarke |
IEEE Symposium on Computer Arithmetic | 2 |
| 1999 | VLSI Costs of Arithmetic Parallelism: A Residue Reverse Conversion PerspectivabstractThis paper reports how VLSI cost metrics (area, delay, power) of residue reverse converters scale with the cardinality and dynamic range of moduli sets. The study uses CMAC reverse converters, reported previously by the authors to be the most efficient known to date in terms of area and delay. In all, 134 reverse converters with dynamic ranges from 32 to 120 bits and set cardinalities ranging from 4 to 20 are actually constructed and analyzed. It is seen that area, delay and power costs are cardinality insensitive once the cardinality exceeds a threshold (usually between five to eight). For cardinalities beyond this threshold, conversion costs are essentially dynamic range dependent. This insensitivity is explained in detail by noting the counterbalancing effects of the various sub-units of a CMAC reverse converter. Since practical implementations of RNS usually employ cardinalities beyond the abovementioned thresholds, the significance of this study is its conclusion that increasing the set cardinality in most implementations will have a marginal, if any, effect on VLSI reverse conversion costs. Manish Bhardwaj, Thambipillai Srikanthan, Christopher T. Clarke |
IEEE Symposium on Computer Arithmetic | 2 |
| 1998 | Designing efficient residue arithmetic based VLSI correlatorsabstractThe most important reason for the lack of commercial residue arithmetic (RA) based systems is not the "slow" and area consuming reverse conversion, but the absence of research that explores the system-level trade-offs of such arithmetic in actual VLSI implementations. Such system-level issues are-choice of the moduli set, effect of moduli imbalance on resulting VLSI implementation, choice of the reverse and forward convertors, use of lookup versus computation for modular operations, system characteristics that indicate RA suitability and finally, typical VLSI area and performance figures. This paper explains these concerns by presenting novel RA architectures for VLSI correlators employed in radioastronomy and ultrasonic blood flow measurement. A state-of-the-art, high performance (80-100 MHz), RA-based correlator ASIC was successfully fabricated as a result of this research. Aniruddha A. Deodhar, Manish Bhardwaj, Christopher T. Clarke, Thambipillai Srikanthan |
ICASSP | 4 |
| 1998 | An Internet application for on-line banking
S. K. Leong, Thambipillai Srikanthan, Gurdeep S. Hura 0001 |
Comput. Commun. | 2 |
| 1997 | An efficient adaptive routing algorithm for a network management system
Thambipillai Srikanthan, Gurdeep S. Hura 0001 |
Comput. Commun. | 1 |