EDBT 2026 Demo / reviewers in the wild / expert
Eugene John
dblp:20/1517 · also Eugene B. John
· DBLP profile ↗
22ranked-venue papers
0as first author
7since 2021 · last 2026
0000-0001-9494-4894ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 2Software engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | POSTER: Hardware Acceleration for Graph Neural Networks
Cory Davis, Patrick M. Stockton, Eugene John, Jeeho Ryoo, Ebod Shojaei |
CF | 3 |
| 2026 | Weightless Neural Networks on Flexible Substrates: A Novel Approach to Wearable Machine LearningabstractIn this article, we present a novel approach that seamlessly integrates machine learning (ML) algorithms into wearable technology through the use of weightless neural networks (WNNs) and flexible integrated circuits (FlexICs). Our methodology employs combinational intelligent networks (COIN) for edge inference on resource-constrained devices, highlighting the advantages of WNNs in terms of power efficiency and minimal hardware requirements. We propose an automated design flow for implementing COIN as FlexICs aimed at developing scalable, cost-effective, and environmentally sustainable wearable monitoring solutions. As a proof-of-concept demonstrator, an arrhythmia detection FlexIC was fabricated using COIN to meet the stringent requirements of medium-complexity wearable applications, offering a promising path toward personalized and accessible healthcare solutions. Igor D. S. Miranda, Velu Pillai, Tejas Musale, Mugdha P. Jadhao, Paulo C. R. Souza Neto, Zachary Susskind, Alan T. L. Bacellar, Mael Lhostis, Priscila M. V. Lima, Diego Leonel Cadette Dutra, Eugene John, Maurício Breternitz, Felipe M. G. França, Emre Ozer 0001, Lizy Kurian John |
IEEE Trans. Very Large Scale Integr. Syst. | 11 |
| 2024 | Differentiable Weightless Neural NetworksabstractWe introduce the Differentiable Weightless Neural Network (DWN), a model based on interconnected lookup tables. Training of DWNs is enabled by a novel Extended Finite Difference technique for approximate differentiation of binary values. We propose Learnable Mapping, Learnable Reduction, and Spectral Regularization to further improve the accuracy and efficiency of these models. We evaluate DWNs in three edge computing contexts: (1) an FPGA-based hardware accelerator, where they demonstrate superior latency, throughput, energy efficiency, and model area compared to state-of-the-art solutions, (2) a low-power microcontroller, where they achieve preferable accuracy to XGBoost while subject to stringent memory constraints, and (3) ultra-low-cost chips, where they consistently outperform small models in both accuracy and projected hardware area. DWNs also compare favorably against leading approaches for tabular datasets, with higher average rank. Overall, our work positions DWNs as a pioneering solution for edge-compatible high-throughput neural networks. Alan T. L. Bacellar, Zachary Susskind, Maurício Breternitz, Eugene John, Lizy Kurian John, Priscila M. V. Lima, Felipe M. G. França |
ICML | 4 |
| 2023 | Profiling Analysis for Enhancing the Performance of Graph Neural NetworksabstractSince the turn of the century, Graph Neural Networks (GNNs) have come into the forefront for collaborative filtering (CF). While not immediately recognized by the average user, GNNs are providing users CF recommendations across all aspects of life. The Graph Convolutional Network (GCN) has provided a strong foundation for CF and node classification in recent years. A great number of models are based on the design of the GCN. GNN models are highly complex and require large labeled datasets making training costly. Analyzing the performance characteristics of GNNs provides insight into opportunities for model enhancement and hardware acceleration. In this paper, we describe and analyze two GNN models: LightGCN and ExpressGNN. In LightGCN we find an abundance of elementwise operations yielding to high levels of pipeline stalls. With ExpressGNN, we find most kernels are based in GEMM (General Matrix Multiplication) and GEMV (General Matrix-Vector Multiplication) operations. ExpressGNN has the possibility to greatly benefit from accelerators, including the use of tensor cores in Nvidia's GPUs. Cory Davis, Patrick M. Stockton, Eugene John |
ICMLA | 3 |
| 2023 | Instruction Profiling Based Predictive Throttling for Power and PerformanceabstractTechnology scaling has long been the driving force for reducing power consumption in microprocessor design. As scaling has reached its limits, new techniques are being adopted to address the power problem. Throttling is an architectural mechanism that slows down the pipeline stages to reduce instant dynamic power. However, power savings due to throttling is achieved at the expense of performance degradation. Throttling is commonly applied at fetch, issue, or commit stages where slowing down a particular stage may reduce dynamic power. A balanced bandwidth of fetch, issue, and commit are maintained in designing a pipeline such that execution can flow seamlessly. However, case studies have shown that bottleneck exists at different pipeline stages that result in performance loss. The loss of performance also indicates wasted dynamic power due to pipeline flush. In this paper, we use instruction profiling based on benchmark traces to identify instructions that are causing a significant bottleneck in the targeted processor architecture. Knowledge of the probable residence of congested instructions at different pipeline stages can enable effective CPU throttling. This paper shows that instruction profiling based predictive throttling at fetch and commit stages can save dynamic power at minimal performance loss. Abdullah A. Owahid, Eugene John |
IEEE Trans. Computers | 2 |
| 2021 | Quickest Joint Detection and Classification of Faults in Statistically Periodic ProcessesabstractAn algorithm is proposed to detect and classify a change in the distribution of a stochastic process that has periodic statistical behavior. The problem is posed in the framework of independent and periodically identically distributed (i.p.i.d.) processes, a recently introduced class of processes to model statistically periodic data. It is shown that the proposed algorithm is asymptotically optimal as the rate of false alarms and the probability of misclassification goes to zero. This problem has applications in anomaly detection in traffic data, social network data, ECG data, and neural data, where periodic statistical behavior has been observed. The effectiveness of the algorithm is demonstrated by application to real and simulated data. Taposh Banerjee, Smruti Padhy, Ahmad F. Taha, Eugene John |
ICASSP | 4 |
| 2021 | Robust Quickest Change Detection in Statistically Periodic ProcessesabstractThe problem of detecting a change in the distribution of a statistically periodic process is investigated. The problem is posed in the framework of independent and periodically identically distributed (i.p.i.d.) processes, a recently introduced class of processes to model statistically periodic data. An algorithm is proposed that is shown to be robust against an uncertainty in the post-change law. The motivation for the problem comes from event detection problems in traffic data, social network data, electrocardiogram data, and neural data, where periodic statistical behavior has been observed. Taposh Banerjee, Ahmad F. Taha, Eugene John |
ISIT | 3 |
| 2020 | Sequential Methods for Detecting a Change in the Distribution of an Episodic ProcessabstractA new class of stochastic processes called episodic processes is introduced to model the statistical regularity of data observed in several applications in cyberphysical systems, neuroscience, and medicine. Algorithms are proposed to detect a change in the distribution of episodic processes. The algorithms can be computed recursively using finite memory and are shown to be asymptotically optimal for well-defined Bayesian or minimax stochastic optimization formulations. The application of the developed algorithms to detect a change in waveform patterns is also discussed. Taposh Banerjee, Edmond Adib, Ahmad F. Taha, Eugene John |
ICASSP | 4 |
| 2020 | Demystifying the MLPerf Training Benchmark SuiteabstractMLPerf, an emerging machine learning benchmark suite, strives to cover a broad range of machine learning applications. We present a study on the characteristics of MLPerf benchmarks and how they differ from previous deep learning benchmarks such as DAWNBench and DeepBench. MLPerf benchmarks are seen to exhibit moderately high memory transactions per second and moderately high compute rates, while DAWNBench creates a high-compute benchmark with low memory transaction rate, and DeepBench provides low compute rate benchmarks. We also observe that the various MLPerf benchmarks possess unique features that allow unveiling various bottlenecks in systems. We also observe variation in scaling efficiency across the MLPerf models. The variation exhibited by the different models highlight the importance of smart scheduling strategies for multi-GPU training. Another observation is that dedicated low latency interconnect between GPUs in multi-GPU systems is crucial for optimal distributed deep learning training. Furthermore, host CPU utilization increases with an increase in the number of GPUs used for training. Corroborating prior work, we also observe and quantify improvements possible by mixed-precision training using Tensor Cores. Snehil Verma, Qinzhe Wu, Bagus Hanindhito, Gunjan Jha, Eugene John, Ramesh Radhakrishnan, Lizy Kurian John |
ISPASS | 5 |
| 2019 | Instruction Profiling Based Fetch Throttling for Wasted Dynamic Power ReductionabstractIn superscalar processors the throughput of the early pipeline stages will impose an upper bound on the throughput of all the subsequent stages. Therefore, to achieve high performance, maximum instruction fetch bandwidth is maintained. This leads to higher power dissipation that often is wasted due to aggressive fetching of instruction cache. To address this problem, this paper implements fetch throttling based on instruction profiling. Instruction profile indicates how likely the instructions are to be stalled in each pipeline stages. The knowledge of the probable residence of the instructions at different stages can enable throttling mechanism to reduce wasted dynamic power with minimal performance penalty. In this paper, we propose fetch throttling based on instruction profiling for top 10 instructions that contribute to frequent pipeline stalls. Simulation results show that instruction profiling based fetch throttling improves average energy efficiency in the range of 24.36-39.70% for an average performance degradation of 5.3412.20% at different levels of fetch throttling. Abdullah A. Owahid, Eugene John |
SBAC-PAD | 2 |
| 2019 | A Study of Core Utilization and Residency in Heterogeneous Smart Phone ArchitecturesabstractIn recent years, the smart phone platform has seen a rise in the number of cores and the use of heterogeneous clusters as in the Qualcomm Snapdragon, Apple A10 and the Samsung Exynos processors. This paper attempts to understand characteristics of mobile workloads, with measurements on heterogeneous multicore phone platforms with big and little cores. It answers questions such as the following: (i) Do smart phones need multiple cores of different types (eg: big or little)? (ii) Is it energy-efficient to operate with more cores (with less time) or fewer cores even if it might take longer? (iii)What are the best frequencies to operate the cores considering energy efficiency? (iv) Do mobile applications need out-of-order speculative execution cores with complex branch prediction? (v) Is IPC a good performance indicator for early design tradeoff evaluation while working on mobile processor design? Joseph Whitehouse, Qinzhe Wu, Shuang Song 0007, Eugene John, Andreas Gerstlauer, Lizy Kurian John |
ICPE | 4 |
| 2019 | Wasted dynamic power and correlation to instruction set architecture for CPU throttling
Abdullah A. Owahid, Eugene John |
J. Supercomput. | 2 |
| 2019 | Resource Shared Galois Field Computation for Energy Efficient AES/CRC in IoT ApplicationsabstractEnd-to-end encryption and reliability of the transmitted data are essential requirements in the present era of internet enabled smart devices. Adhering to current industry standards, the Advanced Encryption Standard (AES) and Cyclic Redundancy Check (CRC) are the two most utilized methods for ensuring security and reliability. To integrate AES and CRC functionality in ultralow-power embedded System on Chips (SoCs), dedicated computation engines/co-processors are often used, consuming valuable silicon area and additional battery power. This paper presents the design of an energy-efficient multipurpose encryption engine capable of processing both AES and CRC algorithms using a shared Galois Field Computation Unit (GFCU). By decomposing the necessary Galois Field operations of AES and CRC to their fundamental binary steps, it was possible to identify shared operations in these two algorithms. This approach allowed the development of a resource shared system architecture capable of computing AES-128 and CRC-32 using a single computation unit. The GFCU based design was implemented in an area of 151μm x 151μm in 90nm technology node. The energy consumption of the design operating at 0.8 V supply voltage for a 25.6 Mbps throughput was less than 280pJ and 140pJ for AES-128 encryption and CRC-32, respectively. Safwat Mostafa, Eugene John |
IEEE Trans. Sustain. Comput. | 2 |
| 2019 | Design and Implementation of an Ultralow-Energy FFT ASIC for Processing ECG in Cardiac PacemakersabstractIn embedded biomedical applications, spectrum analysis algorithms such as Fast Fourier Transform (FFT) are crucial for pattern detection and has been the focus of continued research. In deeply embedded systems such as cardiac pacemakers, FFT based signal processing is typically computed by Application Specific Integrated Circuits (ASIC) to achieve low power operation. This research proposes a data driven design approach for an FFT ASIC solution which exploits the limited range of data encountered by these embedded systems. The optimizations proposed in this paper uses the simple concept of Hashing and Look-Up Tables (LUT) to effectively reduce the number of arithmetic operations required to perform the FFT of an electrocardiogram (ECG) signal. By reducing the dynamic power consumption and overall energy footprint of FFT computation, the proposed design aims to achieve longer battery life for a Cardiac Pacemaker. The design is synthesized using a 90nm standard cell library, and gate level switching activity is simulated to obtain accurate power consumption results. The proposed optimizations achieved a low energy consumption of 27.72nJ per FFT, which is 14.22% lower than a standard 128-point radix-2 FFT when tested with actual ECG data collected from PhysioNet. Safwat Mostafa, Eugene John, Manoj Panday |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | FlowPaP and FlowReR: Improving Energy Efficiency and Performance for STT-MRAM-Based Handheld Devices under Read DisturbanceabstractHandheld devices, such as smartphones and tablets, currently dominate the semiconductor market. The memory access patterns of CPU and IP cores are dramatically different in a handheld device, making the main memory a critical bottleneck of the entire system. As a result, non-volatile memories, such as spin transfer torque magnetoresistive random-access memory (STT-MRAM), are emerging as a replacement for the existing DRAM-based main memory, achieving a wide variety of advantages. However, replacing DRAM with STT-MRAM also results in new design challenges including read disturbance. A simple read-and-restore scheme preserves data integrity under read disturbance, but incurs significant performance and energy overheads. Consequently, by utilizing unique characteristics of mobile applications, we propose FlowPaP, a flow pattern prediction scheme to dynamically predict the write-to-last-read distances for data frames running on a handheld device. FlowPaP identifies and removes unnecessary memory restores originally required for preventing read disturbance, significantly improving energy efficiency and performance for STT-MRAM-based handheld devices. In addition, we propose a flow-based data retention time reduction scheme named FlowReR to further lower energy consumption of STT-MRAM at the expense of reducing its data retention time. FlowReR imposes a second step that marginally trades off the already improved energy efficiency for performance improvements. Experimental results show that, compared to the original read-and-restore scheme, the application of FlowPaP and FlowReR together can simultaneously improve energy efficiency by 34% and performance by 17% for a set of commonly used Android applications. Lei Jiang 0001, Lide Duan, Wei-Ming Lin, Eugene John |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2016 | Building trust in 3PIP using asset-based security property verificationabstractDetecting vulnerabilities, including hardware Trojans, in third-party intellectual property (3PIP) is a rising security challenge for modern day System-on-Chips (SoCs). This work proposes a new approach for detecting security vulnerabilities in a SoC containing 3PIP by breaking complex security flows into fine-grained security properties which are subject to formal verification. Juan Portillo, Eugene John, Seetharam Narasimhan |
VTS | 2 |
| 2014 | Performance enhancement in shared-memory multiprocessors using dynamically classified sharing informationabstractAdvances in process technology has enabled the integration of many cores on a single die. The advent of many core systems has led to a commensurate increase in cache coherence complexity. As a solution to this problem, researches have proposed directory based protocols, which are scalable alternatives to snoop-based protocols. Although write-invalidation based directory protocols enhance the performance of large-scale multiprocessors, coherence misses are intrinsic impediments in such systems. Write-update protocols were proposed as a means to reduce these coherence misses. However, previous researches have shown that pure write-update protocol is highly undesirable because of the heavy traffic caused by the aggressive updates. In order to remedy these limitations, we propose a performance-aware mechanism which dynamically classifies the sharers of each cache block, either as a weak-sharing-group or an efficient-sharing-group and exploit this dynamic classification as a metric for seamless dynamic adaptation between write-invalidate and write-update strategy on a per block basis. Exploitation of the dynamic adaptation of the protocol, based on the sharing-group speculation, reduces unnecessary accesses to the shared last level cache and hence reduces the traffic caused by coherence misses and directory accesses. Simulation results on a 64-core CMP show that our proposed method can achieve 15 % (average) speedup over the baseline directory-based MOESI cache coherence protocol with PARSEC workloads. Our proposed work also reduces the L1 cache miss rate by 17 %(average). The network traffic caused by directory accesses and L1 read misses are also reduced by 16% and 17% respectively. Nilufar Ferdous, Byeong Kil Lee, Eugene John |
IPCCC | 3 |
| 2012 | Performance-sensitivity and performance-similarity based workload reductionabstractIn computer design area, pre-silicon early-stage design exploration requires the detailed simulation which is running applications on a cycle-level microprocessor simulator. Main objectives of simulation-level design exploration include understanding the architectural behaviors of target applications and finding optimal configurations to cover wide range of applications in terms of performance and power. However, full simulation of an industry standard benchmark suite takes several weeks to several months. Among many techniques for reducing simulation time, a tool called SimPoint is popularly used. Even though, simulation load with the reduced workloads by SimPoint is still heavy considering design complexity of modern microprocessors. Basic motivation of this research is started from how design exploration is actually performed. Designers will observe the performance impact from resource variations or configuration changes. If a simulation point shows low sensitivity to resource variations, designers would eliminate those simulation points from the simulation setup procedure. In this paper, we focus on identifying those simulation points which have high sensitivity or low sensitivity, by which overall simulation methodology can be effectively improved. We also performed the performance-sensitivity-based similarity analysis for workload reduction by using statistical technique. Our experiment results show that the proposed scheme provides around 50% reduction with 0.2%-3.5% error rate over original SimPoint method. Kathlene Morales, Byeong Kil Lee, Eugene John |
IPCCC | 4 |
| 2008 | Caches for Multimedia Workloads: Power and Energy TradeoffsabstractOne of the significant workloads in current generation desktop processors and mobile devices is multimedia processing. Large on-chip caches are common in modern processors, but large caches will result in increased power consumption and increased access delays. Regular data access patterns in streaming multimedia applications and video processing applications can provide high hit-rates, but due to issues associated with access time, power and energy, caches cannot be made very large. Characterizing and optimizing the memory system is conducive for designing power and performance efficient multimedia application processors. Performance tradeoffs for multimedia applications have been studied in the past, however, power and energy tradeoffs for caches for multimedia processing have not been adequately studied in the past. In this paper, we characterize multimedia applications for I-cache and D-cache power and energy using a multilevel cache hierarchy. Both dynamic and static power increase with increasing cache sizes, however, the increase in dynamic power is small. The increase in static power is significant, and becomes increasingly relevant for smaller feature sizes. There is significant static power dissipation, ~ 45%, in L1 & L2 caches at 70 nm technology sizes, emphasizing the fact that future multimedia systems must be designed by taking leakage power reduction techniques into account. The energy consumption of on-chip L2 caches is seen to be very sensitive to cache size variations. Sizes larger than 16 k for I-caches and 32 k for D-caches will not be efficient choices to maintain power and performance balance. Since multimedia applications spend significant amounts of time in integer operations, to improve the performance, we propose implementing low power full adders and hybrid multipliers in the data path, which results in 9% to 21% savings in the overall power consumption. Dhireesha Kudithipudi, Stefan Petko, Eugene John |
IEEE Trans. Multim. | 3 |
| 2006 | Architectural enhancements for network congestion control applicationsabstractComplex network protocols and various network services require significant processing capability for modern network applications. One of the important features in modern networks is differentiated service. Along with differentiated service, rapidly changing network environments result in congestion problems. In this paper, we analyze the characteristics of representative congestion control applications-scheduling and queue management algorithms, and we propose application-specific acceleration techniques that use instruction-level parallelism (ILP) and packet-level parallelism (PLP) in these applications. From the PLP perspective, we propose a hardware acceleration model based on detailed analysis of congestion control applications. In order to get large throughputs, a large number of processing elements (PEs) and a parallel comparator are designed. Such hardware accelerators provide large parallelism proportional to the number of processing elements added. A 32-PE enhancement yields 24/spl times/ speedup for weighted fair queueing (WFQ) and 27/spl times/ speedup for random early detection (RED). For ILP, new instruction set extensions for fast conditional operations are applied for congestion control applications. Based on our experiments, proposed architectural extensions show 10%-12% improvement in performance for instruction set enhancements. As the performance of general-purpose processors rapidly increases, defining architectural extensions (e.g., multi-media extensions (MMX) as in multimedia applications) for general-purpose processors could be an alternative solution for a wide range of network applications. Byeong Kil Lee, Lizy Kurian John, Eugene John |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2005 | Architectural Support for Accelerating Congestion Control Applications in Network ProcessorsabstractComplex network protocols and various network services require significant processing capability for modern intelligent network applications. One of the significant features in modern networks is differentiated service. Along with differentiated service, rapidly changing network environments result in congestion problems. In this paper, we analyze the characteristics of representative congestion control applications - scheduling and queue management algorithms, and we propose application-specific acceleration technique using ILP (instruction level parallelism) and PLP (packet level parallelism). From the ILP perspective, new instruction set extensions for fast conditional operations are applied for congestion control applications. Based on our experiments, proposed architectural extensions show 10-12% improvement in performance for instruction set enhancements. For PLP, we propose a hardware acceleration model based on detailed analysis of congestion control applications. In order to get large throughputs, large number of processing elements and a parallel comparator are designed. Such hardware accelerators provide large parallelism proportional to the number of processing elements added. As the performance of general purpose processors rapidly increases, defining architectural extensions (e.g., MMX as in multimedia applications) for general purpose processors could be an alternative solution for wide range of network applications. Byeong Kil Lee, Lizy Kurian John, Eugene John |
ASAP | 3 |
| 1999 | A Novel Low Power Energy Recovery Full Adder CellabstractA novel low power and low transistor count static energy recovery full adder (SERF) is presented in this paper. The power consumption and general characteristics of the SERF adder are then compared against three low powerful adders; the transmission function adder (TFA) the dual value logic (DVL) adder and the fourteen transistor (14 T) full adder. The proposed SERF adder design was proven to be superior to the other three designs in power dissipation and area, and second in propagation delay only to the DVL adder. The combination of low power and low transistor count makes the new SERF cell a viable option for low power design. R. Shalem, Lizy Kurian John, Eugene John |
Great Lakes Symposium on VLSI | 3 |