VLDB 2026 Research / reviewers in the wild / expert
Mohammed Shoaib
dblp:65/2967
· DBLP profile ↗
10ranked-venue papers
5as first author
0since 2021 · last 2016
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 5 first-authorSoftware engineering, systems software and programming languages · 4 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Hardware accelerators and domain-specific architectures · 58% Energy-efficient computing · 21% Integrated circuit design · 21% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 100% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.2 | 1 | 2015 | Scalable-effort classifiers for energy-efficient machine learning · DAC 2015 |
Hardware accelerators and domain-specific architectures › bioinformatics accelerator
biomedical accelerator |
0.1 | 1 | 2011 | A low-energy computation platform for data-driven biomedical monitoring algorithms · DAC 2011 |
Energy-efficient computing
low-power design |
0.1 | 1 | 2011 | A low-energy computation platform for data-driven biomedical monitoring algorithms · DAC 2011 |
Integrated circuit design › low-power circuit design
subthreshold circuit design |
0.1 | 1 | 2011 | A low-energy computation platform for data-driven biomedical monitoring algorithms · DAC 2011 |
Machine learning › Efficient and distributed learning
model compression |
0.1 | 1 | 2015 | Scalable-effort classifiers for energy-efficient machine learning · DAC 2015 |
Medical and health informatics › electrocardiogram analysis
cardiac arrhythmia detection |
0.0 | 1 | 2011 | A low-energy computation platform for data-driven biomedical monitoring algorithms · DAC 2011 |
Methods — techniques the papers use, named apart from their topics
ensemble of biased classifiers · 0.4chain of classifiers · 0.4voltage scaling · 0.2parallelism · 0.2hardware reconfiguration · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | Creating Soft Heterogeneity in Clusters Through Firmware Re-configurationabstractCustomizing server hardware to adapt to its workload has the potential to improve both runtime and energy efficiency. In a cluster that caters to diverse workloads, employing servers with customized hardware components leads to heterogeneity, which is not scalable. In this paper, we seek to create soft heterogeneity from existing servers with homogenous hardware components through customizing the firmware configuration. We demonstrate that firmware configurations have a large impact on runtime, power, and energy efficiency of workloads. Since finding the firmware configuration that minimizes runtime and/or energy efficiency grows exponentially as a function of the number of firmware settings, we propose a methodology called FXplore that helps complete the exploration with a quadratic time complexity. Furthermore, FXplore enables system administrators to manage the degree of the heterogeneity by deriving firmware configurations for sub-clusters that can cater to multiple workloads with similar characteristics. Thus, during online operation, incoming workloads to the cluster can be mapped to appropriate sub-clusters with pre-configured firmware settings. FXplore also finds the best firmware settings in case of co-runners on the same server. We validate our methodology on a fully-instrumented cluster under a large range of parallel workloads that are representative of both high-performance compute clusters and datacenters. Compared to enabling all firmware options, our method improves average runtime and energy consumption by 11% and 15%, respectively. Xin Zhan, Mohammed Shoaib, Sherief Reda |
CCGrid | 2 |
| 2016 | X-Mem: A cross-platform and extensible memory characterization tool for the cloudabstractEffective use of the memory hierarchy is crucial to cloud computing. Platform memory subsystems must be carefully provisioned and configured to minimize overall cost and energy for cloud providers. For cloud subscribers, the diversity of available platforms complicates comparisons and the optimization of performance. To address these needs, we present X-Mem, a new open-source software tool that characterizes the memory hierarchy for cloud computing. Mark Gottscho, Sriram Govindan, Bikash Sharma, Mohammed Shoaib, Puneet Gupta 0001 |
ISPASS | 4 |
| 2015 | Rate-adaptive compressed-sensing and sparsity variance of biomedical signalsabstractBiomedical signals exhibit substantial variance in their sparsity, preventing conventional a-priori open-loop setting of the compressed sensing (CS) compression factor. In this work, we propose, analyze, and experimentally verify a rate-adaptive compressed-sensing system where the compression factor is modified automatically, based upon the sparsity of the input signal. Experimental results based on an embedded sensor platform exhibit a 16.2% improvement in power consumption for the proposed rate-adaptive CS versus traditional CS with a fixed compression factor. We also demonstrate the potential to improve this number to 24% through the use of an ultra low power processor in our embedded system. Vahid Behravan, Neil E. Glover, Rutger Farry, Patrick Chiang 0001, Mohammed Shoaib |
BSN | 5 |
| 2015 | Scalable-effort classifiers for energy-efficient machine learningabstractSupervised machine-learning algorithms are used to solve classification problems across the entire spectrum of computing platforms, from data centers to wearable devices, and place significant demand on their computational capabilities. In this paper, we propose scalable-effort classifiers, a new approach to optimizing the energy efficiency of supervised machine-learning classifiers. We observe that the inherent classification difficulty varies widely across inputs in real-world datasets; only a small fraction of the inputs truly require the full computational effort of the classifier, while the large majority can be classified correctly with very low effort. Yet, state-of-the-art classification algorithms expend equal effort on all inputs, irrespective of their difficulty. To address this inefficiency, we introduce the concept of scalable-effort classifiers, or classifiers that dynamically adjust their computational effort depending on the difficulty of the input data, while maintaining the same level of accuracy. Scalable effort classifiers are constructed by utilizing a chain of classifiers with increasing levels of complexity (and accuracy). Scalable effort execution is achieved by modulating the number of stages used for classifying a given input. Every stage in the chain contains an ensemble of biased classifiers, where each biased classifier is trained to detect a single class more accurately. The degree of consensus between the biased classifiers' outputs is used to decide whether classification can be terminated at the current stage or not. Our methodology thus allows us to transform any given classification algorithm into a scalable-effort chain. We build scalable-effort versions of 8 popular recognition applications using 3 different classification algorithms. Our experiments demonstrate that scalable-effort classifiers yield 2.79x reduction in average operations per input, which translates to 2.3x and 1.5x improvement in energy for hardware and software implementations, respectively. Swagath Venkataramani, Anand Raghunathan, Jie Liu 0001, Mohammed Shoaib |
DAC | 4 |
| 2015 | SAPPHIRE: an always-on context-aware computer vision system for portable devices
Swagath Venkataramani, Paramvir Bahl, Xian-Sheng Hua 0001, Jie Liu 0001, Jin Li 0001, Matthai Philipose, Bodhi Priyantha, Mohammed Shoaib |
DATE | 8 |
| 2015 | Signal Processing With Direct Computations on Compressively Sensed DataabstractSparsity is characteristic of a signal that potentially allows us to represent information efficiently. We present an approach that enables efficient representations based on sparsity to be utilized throughout a signal processing system, with the aim of reducing the energy and/or resources required for computation, communication, and storage. The representation we focus on is compressive sensing. Its benefit is that compression is achieved with minimal computational cost through the use of random projections; however, a key drawback is that reconstruction is expensive. We focus on inference frameworks for signal analysis. We show that reconstruction can be avoided entirely by transforming signal processing operations (e.g., wavelet transforms, finite impulse response filters, etc.) such that they can be applied directly to the compressed representations. We present a methodology and a mathematical framework that achieve this goal and also enable significant computational-energy savings through operations over fewer input samples. This enables explicit energy-versus-accuracy tradeoffs that are under the control of the designer. We demonstrate the approach through two case studies. First, we consider a system for neural prosthesis that extracts wavelet features directly from compressively sensed spikes. Through simulations, we show that spike sorting can be achieved with 54× fewer samples, providing an accuracy of 98.63% in spike count, 98.56% in firing-rate estimation, and 96.51% in determining the coefficient of variation; this compares with a baseline Nyquist-domain detector with corresponding performance of 98.97%, 99.69%, and 97.09%, respectively. Second, we consider a system for detecting epileptic seizures by extracting spectral-energy features directly from compressively sensed electroencephalogram. Through simulations of the end-to-end algorithm, we show that detection can be achieved with 21× fewer samples, providing a sensitivity of 94.43%, false alarm rate of 0.1543/h, and latency of 4.70 s; this compares with a baseline Nyquist-domain detector with corresponding performance of 96.03%, 0.1471/h, and 4.59 s, respectively. Mohammed Shoaib, Niraj K. Jha, Naveen Verma |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2013 | Algorithm-Driven Architectural Design Space Exploration of Domain-Specific Medical-Sensor ProcessorsabstractData-driven machine-learning techniques enable the modeling and interpretation of complex physiological signals. The energy consumption of these techniques, however, can be excessive, due to the complexity of the models required. In this paper, we study the tradeoffs and limitations imposed by the energy consumption of high-order detection models implemented in devices designed for intelligent biomedical sensing. Based on the flexibility and efficiency needs at various processing stages in data-driven biomedical algorithms, we explore options for hardware specialization through architectures based on custom instruction and coprocessor computations. We identify the limitations in the former, and propose a coprocessor-based platform that exploits parallelism in computation as well as voltage scaling to operate at a subthreshold minimum-energy point. We present results from post-layout simulation of cardiac arrhythmia detection with patient data from the MIT-BIH database. After wavelet-based feature extraction, which consumes 12.28 μJ, we demonstrate classification computations in the 12.00-120.05 μJ range using 10000-100000 support vectors. This represents 1170× lower energy than that of a low-power processor with custom instructions alone. After morphological feature extraction, which consumes 8.65 μJ of energy, the corresponding energy numbers are 10.24-24.51 μJ, which is 1548× smaller than one based on a custom-instruction design. Results correspond to Vdd=0.4 V and a data precision of 8 b. Mohammed Shoaib, Niraj K. Jha, Naveen Verma |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2012 | Enabling advanced inference on sensor nodes through direct use of compressively-sensed signalsabstractNowadays, sensor networks are being used to monitor increasingly complex physical systems, necessitating advanced signal analysis capabilities as well as the ability to handle large amounts of network data. For the first time, we present a methodology to enable advanced decision support on a low-power sensor node through the direct use of compressively-sensed signals in a supervised-learning framework; such signals provide a highly efficient means of representing data in the network, and their direct use overcomes the need for energy-intensive signal reconstruction. Sensor networks for advanced patient monitoring are representative of the complexities involved. We demonstrate our technique on a patient-specific seizure detection algorithm based on electroencephalograph (EEG) sensing. Using data from 21 patients in the CHB-MIT database, our approach demonstrates an overall detection sensitivity, latency, and false alarm rate of 94.70%, 5.83 seconds, and 0.199 per hour, respectively, while achieving data compression by a factor of 10x. This compares well with the state-of-the-art baseline detector with corresponding results being 96.02%, 4.59 seconds, and 0.145 per hour, respectively. Mohammed Shoaib, Niraj K. Jha, Naveen Verma |
DATE | 1 |
| 2012 | A closed-loop system for artifact mitigation in ambulatory electrocardiogram monitoringabstractMotion artifacts interfere with electrocardiogram (ECG) detection and information processing. In this paper, we present an independent component analysis based technique to mitigate these signal artifacts. We propose a new statistical measure to enable an automatic identification and removal of independent components, which correspond to the sources of noise. For the first time, we also present a signal-dependent closed-loop system for the quality assessment of the denoised ECG. In one experiment, noisy data is obtained by the addition of calibrated amounts of noise from the MIT-BIH NST database to the AHA ECG database. Arrhythmia classification based on a state-of-the-art algorithm with the direct use of noisy data thus obtained shows sensitivity and positive predictivity values of 87.7% and 90.0%, respectively, at an input signal SNR of -9 dB. Detection with the use of ECG data denoised by the proposed approach exhibits significant improvement in the performance of the classifier with the corresponding results being 96.5% and 99.1%, respectively. In a related lab trial, we demonstrate a reduction in RMS error of instantaneous heart rate estimates from 47.2% to 7.0% with the use of 56 minutes of denoised ECG from four physically active subjects. To validate our experiments, we develop a closed-loop, ambulatory ECG monitoring platform, which consumes 2.17 mW of power and delivers a data rate of 33 kbps over a dedicated UWB link. Mohammed Shoaib, Gene Marsh, Harinath Garudadri, Somdeb Majumdar |
DATE | 1 |
| 2011 | A low-energy computation platform for data-driven biomedical monitoring algorithmsabstractA key challenge in closed-loop chronic biomedical systems is the ability to detect complex physiological states from patient signals within a constrained power budget. Data-driven machine-learning techniques are major enablers for the modeling and interpretation of such states. Their computational energy, however, scales with the complexity of the required models. In this paper, we propose a low-energy, biomedical computation platform optimized through the use of an accelerator for data-driven classification. The accelerator retains selective flexibility through hardware reconfiguration and exploits voltage scaling and parallelism to operate at a sub-threshold minimum-energy point. Using cardiac arrhythmia detection algorithms with patient data from the MIT-BIH database, classification is achieved in 2.96 μJ (at Vdd = 0.4 V), over four orders of magnitude smaller than that on a low-power general-purpose processor. The energy of feature extraction is 148 μJ while retaining flexibility for a range of possible biomarkers. Mohammed Shoaib, Niraj K. Jha, Naveen Verma |
DAC | 1 |