James Myers

dblp:04/6622 · DBLP profile ↗
← Back
16ranked-venue papers
1as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 9 since 2021Artificial intelligence and machine learning · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author
YearPublicationVenuePosition
2026 Evaluating Cross-Architecture Performance Modeling of Distributed ML Workloads Using StableHLO
abstract
Predicting the performance of large-scale distributed machine learning (ML) workloads across multiple accelerator architectures remains a central challenge in ML system design. Existing GPU and TPU focused simulators are typically architecture-specific, while distributed training simulators rely on workload-specific analytical models or costly post-execution traces, limiting portability and cross-platform comparison. This work evaluates whether MLIR’s StableHLO dialect can serve as a unified workload representation for cross-architecture and crossfidelity performance modeling of distributed ML workloads. The study establishes a StableHLO-based simulation methodology that maps a single workload representation onto multiple performance models, spanning analytical, profiling-based, and simulator-driven predictors. Using this methodology, workloads are evaluated across GPUs and TPUs without requiring access to scaled-out physical systems, enabling systematic comparison across modeling fidelities. An empirical evaluation covering distributed GEMM kernels, ResNet, and large language model training workloads demonstrates that StableHLO preserves relative performance trends across architectures and fidelities, while exposing accuracy trade-offs and simulator limitations. Across evaluated scenarios, prediction errors remain within practical bounds for early-stage design exploration, and the methodology reveals fidelity-dependent limitations in existing GPU simulators. These results indicate that StableHLO provides a viable foundation for unified, distributed ML performance modeling across accelerator architectures and simulators, supporting reusable evaluation workflows and crossvalidation throughout the ML system design process.
Jonas Svedas, Nathan Laubeuf, Ryan Harvey, Changhai Man, Abubakr Nada, Tushar Krishna, James Myers, Debjyoti Bhattacharjee
ISPASS8
2026 Thermal Insights of 3-D BS-PDN in Cloud Server SoC Using TCAD Modeling
abstract
In this brief, the thermal performance of a large-scale cloud server system-on-chip (SoC) with the backside power delivery network (BS-PDN) and 3-D integration in memory-on-logic (MoL)/logic-on-memory (LoM) configuration with 2.5-D packaging is analyzed in advanced A10 nanosheet technology node using Sentaurus TCAD platform. The results show a 45.6% (~20.3 K) thermal penalty for the 80-core SoC in MoL with BS-PDN compared with the 2-D-baseline frontside PDN (FS-PDN), using a heatsink with forced cooling. A nonuniform power map further aggravates thermal concerns, which can be mitigated using an LoM configuration with BS-PDN, reducing the penalty to 22% (~15 K). Extending the study to a 320-core SoC, in conjunction with an advanced cooling system, LoM with BS-PDN shows 45.3% (~29 K) lower temperature than conventional MoL BS-PDN. The modeling results provide valuable insights and motivate future research into packaging and cooling techniques for BS-PDN integration.
Subrat Mishra, Herman Oprins, James Myers, Julien Ryckaert, Pieter Woltgens, Dwaipayan Biswas
IEEE Trans. Very Large Scale Integr. Syst.4
2025 TAXI: Traveling Salesman Problem Accelerator with X-bar-based Ising Macros Powered by SOT-MRAMs and Hierarchical Clustering
abstract
Ising solvers with hierarchical clustering have shown promise for large-scale Traveling Salesman Problems (TSPs), in terms of latency and energy. However, most of these methods still face unacceptable quality degradation as the problem size increases beyond a certain extent. Additionally, their hardwareagnostic adoptions limit their ability to fully exploit available hardware resources. In this work, we introduce TAXI – an inmemory computing-based TSP accelerator with crossbar(Xbar)-based Ising macros. Each macro independently solves a TSP subproblem, obtained by hierarchical clustering, without the need for any off-macro data movement, leading to massive parallelism. Within the macro, Spin-Orbit-Torque (SOT) devices serve as compact energy-efficient random number generators enabling rapid “natural annealing”. By leveraging hardware-algorithm co-design, TAXI offers improvements in solution quality, speed, and energy-efficiency on TSPs up to $\mathbf{8 5, 9 0 0}$ cities (the largest TSPLIB instance). TAXI produces solutions that are only $22 \%$ and $20 \%$ longer than the Concorde solver’s exact solution on $\mathbf{3 3, 8 1 0}$ and $\mathbf{8 5, 9 0 0}$ city TSPs, respectively. TAXI outperforms a current state-of-the-art clustering-based Ising solver, being $8 \times$ faster on average across 20 benchmark problems from TSPLib.
Sangmin Yoo, Amod Holla, Sourav Sanyal, Dong Eun Kim, Francesca Iacopi, Dwaipayan Biswas, James Myers, Kaushik Roy 0001
DAC7
2025 Late Breaking Results: Thermal Feasibility of Backside Integrated LDOs in 2.5D/3D System-in-Package Using Nanosheet Technology
abstract
Digital Low Dropout Regulators (LDOs) are an excellent candidate for area-efficient fine-grain power management in heterogeneous systems, leveraging integrated power switches. Relocating the power switches to the backside of the wafer in conjunction with the Backside Power Delivery Network (BSPDN) layer is envisaged as a System Technology Co-Optimization (STCO) booster for finer grain power management and reduced area/cost. We perform a detailed thermal analysis using power-switch-based LDOs enabling per-core DVFS for a high-performance server 3D computing chiplet in a Nanosheet CMOS (A10) technology node with BSPDN. While BSPDN introduces thermal penalties due to a lack of lateral heat spreading, our high-resolution thermal simulations explore the feasibility of moving LDOs to the backside. Increasing the LDO area from 5% to 50% of the backside die area effectively lowers the 2.5/3D System-in-Package (SiP) peak temperature, confirming that thermal concerns do not impede backside LDO integration. This study supports the cost-effective design of next-generation SiPs by demonstrating no adverse thermal impact for relocating power switches to the wafer backside in the nanosheet era.
Yukai Chen, Subrat Mishra, Julien Ryckaert, Dwaipayan Biswas, James Myers
DATE5
2025 Framework for Augmenting Main Memory with CXL-connected Emerging Memory Alternatives
abstract
The rapid evolution of memory technologies and the advent of Compute Express Link (CXL) have opened up new possibilities for scaling main memory by enabling hybrid memory systems with pooled and shared content. System-level evaluation of new memory systems during the early development stage is important for the enablement and further integration of new memory and interconnect technologies. However, existing solutions do not offer a framework neither for emerging memory protocols nor for novel memory technologies. This paper introduces CXL-HMEM to evaluate emerging CXL-based hybrid main memory architectures by applying the System Technology Co-Optimization (STCO) technique. The framework provides flexible performance metrics, workload simulation, and memory traffic analysis to assess system performance under various hybrid memory configurations, including DRAM and tiered memory hierarchies. Key features include support for memory technologies such as IGZO-based DRAM (IGZO) and FeRAM, workload scalability, and an integrated model of the CXL behavior. CXL-HMEM shows that emerging memories can improve system bandwidth and energy consumption by 7%, while having potential to further mitigate particular bottlenecks. CXL-based hybrid main memory can speed up the memory access time by >2× compared to conventional approaches of main memory extension.
Khakim Akhunov, Dwaipayan Biswas, Emil Karimov, Arvind Sharma, Hyungrock Oh, Maarten Rosmeulen, Julien Ryckaert, James Myers
ISCAS8
2025 3D SRAM Disaggregation in Advanced CMOS Nodes using Hybrid Bonding Technology
abstract
This paper studies the potential of hybrid bonded Array-under-CMOS (AuC) technology to partition logic and high-performance L1 cache in advanced technology nodes. By decoupling the SRAM bitcells from the logic tier, we achieve independent optimization of both SRAM and logic devices, as well as back-end-of-line (BEOL) interconnects. Heterogeneous integration and BEOL aspect ratio optimization is implemented with different technology nodes to mitigate the delay penalty due to hybrid bond pad staggering. Addressing the performance degradation associated with scaled technology nodes, we investigate the impact of word-line (WL) and bit-line (BL) resistance on SRAM performance. Leveraging the flexibility of decoupled SRAM BEOL and within the AuC technology framework, we explore the sensitivity to WL and BL metal aspect ratios, comparing their performance against a 2D baseline. Our results demonstrate substantial performance improvements in AuC integration through two key approaches: (1) 5% enhancement via Back-End-of-Line (BEOL) optimization, and (2) 25% improvement enabled by heterogeneous integration, achieved by decoupling memory and logic tiers.
Bhawana Kumari, Anurag Swarnkar, Dawit Burusie Abdi, Fernando García-Redondo, James Myers, Julien Ryckaert, Jaydeep P. Kulkarni, Dwaipayan Biswas
ISCAS5
2025 Bandwidth-Latency-Thermal Co-Optimization of Interconnect-Dominated Many-Core 3D-IC
abstract
The ongoing integration of advanced functionalities in contemporary system-on-chips (SoCs) poses significant challenges related to memory bandwidth, capacity, and thermal stability. These challenges are further amplified with the advancement of artificial intelligence (AI), necessitating enhanced memory and interconnect bandwidth and latency. This article presents a comprehensive study encompassing architectural modifications of an interconnect-dominated many-core SoC targeting the significant increase of intermediate, on-chip cache memory bandwidth and access latency tuning. The proposed SoC has been implemented in 3-D using A10 nanosheet technology and early thermal analysis has been performed. Our workload simulations reveal, respectively, up to 12- and 2.5-fold acceleration in the 64-core and 16-core versions of the SoC. Such speed-up comes at 40% increase in die-area and a 60% rise in power dissipation when implemented in 2-D. In contrast, the 3-D counterpart not only minimizes the footprint but also yields 20% power savings, attributable to a 40% reduction in wirelength. The article further highlights the importance of pipeline restructuring to leverage the potential of 3-D technology for achieving lower latency and more efficient memory access. Finally, we discuss the thermal implications of various 3-D partitioning schemes in High Performance Computing (HPC) and mobile applications. Our analysis reveals that, unlike high-power density HPC cases, 3-D mobile case increases$T_{\max }$only by$2~^{\circ } $C–$3~^{\circ } $C compared to 2-D, while the HPC scenario analysis requires multiconstrained efficient partitioning for 3-D implementations.
Sudipta Das, Samuel Riedel, Mohamed Naeim, Moritz Brunion, Marco Bertuletti, Luca Benini, Julien Ryckaert, James Myers, Dwaipayan Biswas, Dragomir Milojevic
IEEE Trans. Very Large Scale Integr. Syst.8
2024 3D Partitioning with Pipeline Optimization for Low-Latency Memory Access in Many-Core SoCs
abstract
This paper presents an investigation of System-on-Chip (SoC) communication latency optimization for 3D system integration and highlights the role of architectural modifications to maximize the Power, Performance, & Area (PPA) benefits. An instance of a highly configurable RISC-V SoC is implemented using ∼2nm nanosheet technology and different 3D stacking options using design flow from sign-off tools. The proposed implementation targets performance optimization for different 3D partitioning scenarios: Memory-on-Logic (MoL) & Logic-on-Logic (LoL). We target 2-die 3D Integrated Circuits (3D-IC) with high density 3D interconnect using Face-to-Face (F2F) hybrid bonding (∼1µm), and 3-die stack, as Face-to-Back (F2B) on top of F2F. Our analysis of the 16-core SoC instance shows that the proposed architectural optimizations bring a significant reduction of 4 pipeline stages in the design hierarchy at a marginal cost of 9% effective frequency loss when implemented in 3D in comparison to the baseline 2D architecture. Further, going from 2D to 3D allows more than 40% total system wire-length reduction & 10% less cell area, resulting in 20% power savings. These findings hold promise for further explorations on many-core SoC instances (256 & more) facing system interconnect challenges.
Sudipta Das, Samuel Riedel, Marco Bertuletti, Luca Benini, Moritz Brunion, Julien Ryckaert, James Myers, Dwaipayan Biswas, Dragomir Milojevic
ISCAS7
2024 Ultra-Scaled E-Tree-Based SRAM Design and Optimization With Interconnect Focus
abstract
SRAM performance is highly dominated by interconnects as technology scales down because of the significant parasitic resistance and capacitance in the interconnect. This paper introduces a framework for the co-design of technology, interconnect, and cache memory with tag array overhead, to optimize the performance of cache memory using a variety of emerging interconnect technologies. In addition, we introduce an innovative E-Tree interconnect aimed at further decreasing the average interconnect length with the consideration of realistic workloads and benchmark against its traditional H-Tree counterparts in terms of various performance metrics, such as energy-delay-area product (EDAP) or energy-delay product (EDP) in the SRAM cache memory system. A comprehensive investigation of design space is conducted, employing realistic, deeply scaled subarray designs across a range of cutting-edge technology nodes. Furthermore, the case study examines various cache memory system design parameters to assess the true potential of emerging interconnect technologies in achieving optimal performance at the cache memory system.
Zhenlin Pei, Hsiao-Hsuan Liu, Mahta Mayahinia, Mehdi Baradaran Tahoori, Francky Catthoor, Zsolt Tokei, Dawit Burusie Abdi, James Myers, Chenyun Pan
IEEE Trans. Circuits Syst. I Regul. Pap.8
2024 Multidie 3-D Stacking of Memory Dominated Neuromorphic Architectures
abstract
Event-driven neuromorphic processors for artificial intelligence (AI) inference on edge/IoT devices require largeon-chip memory capacity, for efficient execution of spiking neural networks (NNs). In this work, we evaluate 3-D stacking benefits on SENECA, a digital neuromorphic accelerator core, sweeping itson-chip memory capacity from 2 up to 32 Mb in both legacy planar and advanced nanosheet CMOS logic nodes. In a planar CMOS node (GF-22 nm), two-die memory-on-logic (MoL) partitioning enables$8\times $moreon-chip memory, and it boosts operating frequency by 7% with 26% less power than the 2-D. Moving to an advanced nanosheet technology (imec A10), multidie (up to 7 dies) MoL stacking enables a performance increase of up to 29% and power savings up to 31%. Furthermore, a core folding (CF) partitioning in A10 shows up to 16% performance improvement with 12% total power savings with respect to the 2-D implementation on the same technology. We also demonstrate no thermal overhead for multidie stacking at advanced nodes for designs exhibiting low power density. These physical design explorations lay the foundation for system technology co-optimization studies for edge devices.
Leandro M. G. Rocha, Refik Bilgic, Mohamed Naeim, Sudipta Das, Herman Oprins, Amirreza Yousefzadeh, Mario Konijnenburg, Dragomir Milojevic, James Myers, Julien Ryckaert, Dwaipayan Biswas
IEEE Trans. Very Large Scale Integr. Syst.9
2019 A 65nm switched source line sub-threshold ROM using data encoding, with 0.3V Vmin and 47fJ/b access energy
abstract
Battery-operated sensing systems have very low activity rates and power down most blocks during inactive periods to save power. During active periods, these energy-constrained systems typically operate at near or sub-threshold voltage for maintaining energy efficiency. Therefore, these ultra-low energy systems require a read-only, non-volatile code memory that can be accessed at sub-threshold voltage. ROM is the most energy efficient read-only non-volatile memory. A conventional compiler ROM cannot work reliably in near threshold regions of operation due to lack of read margin. As a result, for energy constrained IoT systems, the ROM design needs to support reliable read at low operating voltages, and a speed degradation at low voltages which is at par with logic circuits, to avoid the ROM from limiting the system performance. This paper proposes a sub-threshold ROM that addresses both reliability and performance issues by 1) Using a switched source-line (SSL) to nullify Ioff and improve read margin, 2) Compensating performance degradation caused by SSL using a novel data encoding scheme and by using unused `0' bit-cell transistors to provide local pull-down paths, 3) Tackle variation-based performance degradation by optimizing bit-cell sizing for sub-threshold, and 4) Providing a fine-grained threshold voltage adjustment technique for trading off performance improvement with leakage degradation by using a novel INverse Width Effect (INWE) bit-cell layout. The proposed SSL sub-threshold ROM is fabricated in 65nm technology, with an 8kB capacity, and a Vminof 0.3V. At 0.4V, it has a read speed of 1.1MHz and an access energy of 47fJ/bit.
Supreet Jeloka, Pranay Prabhat, Graham Knight, James Myers
ISLPED4
2018 Communicative Efficiency in Child Mandarin
Jane S. Tsay, James Myers
PACLIC2
2014 Clock-modulation based watermark for protection of embedded processors
abstract
This paper presents a novel watermark generation technique for the protection of embedded processors. In previous work, a load circuit is used to generate detectable watermark patterns in the ASIC power supply. This approach leads to hardware area overheads. We propose removing the dedicated load circuit entirely, instead to compensate the reduced power consumption the watermark power pattern is emulated by reusing existing clock gated sequential logic as a zero-overhead load circuit and modulating the clock-gating enable signal with the watermark sequence. The proposed technique has been validated through experiments using two ASICs in 65nm CMOS, one with an ARM Cortex-M0 microcontroller and one with a Cortex-A5 microprocessor. Silicon measurement results verify the viability of the technique for embedded processors. Furthermore, the proposed clock modulation technique demonstrates a significant area reduction, without compromising the detection performance. In our experiments an area overhead reduction of 98% was achieved. Through reuse of existing logic and reduction of watermark hardware implementation costs, the proposed clock modulation technique offers an improved robustness against removal attacks.
Jedrzej Kufel, Peter R. Wilson, Stephen Hill, Bashir M. Al-Hashimi, Paul N. Whatmough, James Myers
DATE6
2014 Active Mode Subclock Power Gating
abstract
This paper presents a technique, called subclock power gating, for reducing leakage power during the active mode in low performance, energy-constrained applications. The proposed technique achieves power reduction through two mechanisms: 1) power gating the combinational logic within the clock period (subclock) and 2) reducing the virtual supply to less than Vth rather than shutting down completely as is the case in conventional power gating. To achieve this reduced voltage, a pair of nMOS and pMOS transistors are used at the head and foot of the power gated logic for symmetric virtual rail clamping of the power and ground supplies. The subclock power gating technique has been validated by incorporating it with an ARM Cortex-M0 microprocessor, which was fabricated in a 65-nm process. Two sets of experiments are done: the first experimentally validates the functionality of the proposed technique in the fabricated test chip and the second investigates the utility of the proposed technique in example applications. Measured results from the fabricated chip show 27% power saving during the active mode for an example wireless sensor node application when compared with the same microprocessor without subclock power gating.
Jatin N. Mistry, James Myers, Bashir M. Al-Hashimi, David Flynn, John Biggs, Geoff V. Merrett
IEEE Trans. Very Large Scale Integr. Syst.2
2012 Cognitive Styles in Two Cognitive Sciences
James Myers
CogSci1
2012 Grammatical Approaches to Written and Graphical Communication
Colin Wilson, Neil Cohn, James Myers, Stephen Goldberg, Ariel Cohen-Goldberg
CogSci3