VLDB 2026 Research / reviewers in the wild / expert
Brett H. Meyer
dblp:16/5069
· DBLP profile ↗
47ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0002-6650-3298ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 41 · 6 first-author · 10 since 2021Software engineering, systems software and programming languages · 8 · 2 first-authorArtificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A High-Radix Switching Architecture for Machine Learning Datacenters
Rezvan Mohammadrezaee, David Rolston, Brett H. Meyer |
HPSR | 3 |
| 2026 | Latency-Aware Pruning and Quantization of Self-Supervised Speech Transformers for Edge DevicesabstractThe growing adoption of self-supervised learning transformers for speech (speech SSL) is constrained by their significant computational and memory demands, making deployment on resource-constrained edge devices challenging. We propose a latency-aware compression framework that integrates structured pruning and quantization to address these challenges. Guided by a latency model that considers the combined effects of pruning and quantization, our method dynamically identifies and removes less critical blocks while maintaining task performance, avoiding the inefficiencies of over-pruning and under-pruning seen in prior approaches. Unlike prior methods specialized in either post-training compression without fine-tuning data or in cases where fine-tuning data is available, our method is effective in both settings. Experimental results show that, in task-agnostic compression, our method achieves a 4.2× speedup on the Hikey970 edge development platform, outperforming previous task-agnostic pruning methods in most tasks, while requiring only 21–24 GPU hours—a 3× reduction compared to prior methods. Additionally, our method achieves a lower word error rate of 7.8% using task-specific pruning, while reducing computational overhead by approximately 19.4% in terms of GFLOPs compared to previous task-specific methods. Finally, our method consistently achieves higher accuracy than the state-of-the-art post-training compression approach across various latency speedup constraints, even without fine-tuning data. Seyed Milad Ebrahimipour, Seyyed Hasan Mozafari, James J. Clark, Warren J. Gross, Brett H. Meyer |
ACM Trans. Embed. Comput. Syst. | 5 |
| 2023 | Efficient 1D Grouped Convolution for PyTorch a Case Study: Fast On-Device Fine-Tuning for SqueezeBERTabstractGrouped convolution has been observed to be an effective approximation for convolution in many DNN applications. For example, SqueezeBERT, which is a light and fast BERT language processing model, utilizes 1D grouped convolutions. Though SqueezeBERT is well-optimized for inference on edge devices, it suffers from poor memory management during fine-tuning (training). This results in longer fine-tuning time on resource-limited GPUs compared to the original BERT model, BERT-base, despite being specifically designed for edge devices. We study this behavior and show that this poor memory management originates from the use of 1D grouped convolutions in SqueezeBERT. We re-implement 1D grouped convolutions using fully-connected layers, addressing the poor memory allocation and data locality of 1D grouped convolutions. We show that our method is well-suited for edge devices with limited memory; further, it has a negligible effect on inference speed. When utilizing our method, we observe a 42 % reduction in fine-tuning time for SqueezeBERT on edge devices. Seyyed Hasan Mozafari, James J. Clark, Warren J. Gross, Brett H. Meyer |
ASAP | 4 |
| 2023 | High-Throughput Edge Inference for BERT Models via Neural Architecture Search and PipelineabstractThere has been growing interest in improving the BERT inference throughput on resource-constrained edge devices for a satisfactory user experience. One methodology is to employ heterogeneous computing, which utilizes multiple processing elements to accelerate inference. Another methodology is to deploy Neural Architecture Search (NAS) to find optimal solutions in accuracy-throughput design space. In this paper, for the first time, we incorporate NAS with pipelining for BERT models. We show that performing NAS with pipelining achieves on average 53% higher throughput, compared to NAS with a homogeneous system. Hung-Yang Chang, Seyyed Hasan Mozafari, James J. Clark, Brett H. Meyer, Warren J. Gross |
ACM Great Lakes Symposium on VLSI | 4 |
| 2023 | Training Acceleration of Frequency Domain CNNs Using Activation CompressionabstractReducing the complexity of training convolutional neural networks results in lower energy consumption expended during training, or higher accuracy by admitting a greater number of training epochs within a training time budget. During backpropagation, a considerable amount of temporary data is offloaded from GPU memory to CPU memory, increasing training time. In this paper, we address this training time overhead by introducing an activation compression technique for frequency domain convolutional neural networks. Applying this compression technique on frequency domain AlexNet results in activation compression of 57.7%, and a reduction of training time by 23%, with a negligible effect on classification accuracy. Seyyed Hasan Mozafari, James J. Clark, Warren J. Gross, Brett H. Meyer |
ISCAS | 4 |
| 2023 | DT-DS: CAN Intrusion Detection with Decision Tree EnsemblesabstractThe controller area network (CAN) protocol, used in many modern vehicles for real-time inter-device communications, is known to have cybersecurity vulnerabilities, putting passengers at risk for data exfiltration and control system sabotage. To address this issue, researchers have proposed to utilize security measures based on cryptography and message authentication; unfortunately, such approaches are often too computationally expensive to be deployed in real time on CAN devices. Additionally, they have developed machine learning (ML) techniques to detect anomalies in CAN traffic and thereby prevent attacks. The main disadvantage of existing ML-based techniques is that they either depend on additional computational hardware or they heuristically assume that all communication anomalies are malicious. In this article, we show that tree-based learning ensembles outperform anomaly-based techniques like AutoRegressive Integrated Moving Average (ARIMA) and Z-Score when used to detect attacks that result in increased bus utilization. We evaluated the detection capacity of three tree-based ensembles, Adaboost, gradient boosting, and random forests, and collectively refer to these as DT-DS. We conclude that the decision tree ensemble with Adaboost performs best with an area under curve (AUC) score of 0.999, closely followed by gradient boosting and random forests with 0.997 and 0.991 AUC scores, respectively, when trained using message profiles. We observe that with an increase in the observation window, the DT-DS models present an average AUC score of 0.999, and offer a nearly perfect detection of attacks, at the cost of increased latency in detection of attacked messages. We evaluate the performance of the IDS for Aeronautical Radio, Incorporated– (ARINC) encoded CAN communication traffic in avionic systems, generated using an aerospace testbench, ARINC-825TBv2. The IDS has been evaluated against the active attacks of a state-of-the-art predictive attacker model. Additionally, we observed that the performance of IDS approaches such as ARIMA and Z-Score degrade considerably with a decrease in the size of the observation time window. In contrast, the performance of DT-DS models is consistent, with only an average drop of 0.005 in the AUC score. Jarul Mehta, Guillaume Richard, Loren Lugosch, Derek Yu, Brett H. Meyer |
ACM Trans. Cyber Phys. Syst. | 5 |
| 2022 | Fast Heterogeneous Task Mapping for Reducing Edge DNN LatencyabstractTo meet DNN inference latency constraints on resource-constrained edge devices, we employ heterogeneous computing, utilizing multiple processing elements (e.g. CPU + GPU) accelerate inference. This leads to the challenge of efficiently mapping DNN operations to heterogeneous processing elements. For this task, we introduce a novel genetic algorithm (GA) optimizer. Through intelligent initialization and a customized mutation operation, we are able to evaluate 20x fewer generations while finding superior configurations compared with a baseline GA. Using our mapping optimizer, we find device placement configurations that achieve 15%, 24%, and 31% inference speed-up for BERT, SqueezeBERT, and InceptionV3,respectively. Murray L. Kornelsen, Seyyed Hasan Mozafari, James J. Clark, Brett H. Meyer, Warren J. Gross |
ASAP | 4 |
| 2022 | Work-in-Progress: SuperNAS: Fast Multi-Objective SuperNet Architecture Search for Semantic SegmentationabstractWe present SuperNAS, a fast multi-objective neural architecture search framework for semantic segmentation. SuperNAS subsamples the structure and pre-trained parameters of DeepLabV3+, without fine-tuning, dramatically reducing training time during search. To further reduce candidate evaluation time, we use a subset of the validation dataset during search. Only the final, Pareto-dominant, candidates are ultimately fine-tuned using the complete training set. We evaluate SuperNAS by searching for models that effectively trade accuracy and computational cost on the PASCAL VOC 2012 dataset. SuperNAS finds competitive designs quickly, e.g., taking just 0.5 GPU days to discover a DeepLabV3+ variant that reduces FLOPs and parameters by 10% and 20% respectively, for less than 3% increased error. Marihan Amein, Zhuoran Xiong, Olivier Therrien, Brett H. Meyer, Warren J. Gross |
CASES | 4 |
| 2022 | Work-in-Progress: Utilizing latency and accuracy predictors for efficient hardware-aware NASabstractWith the increased size and complexity of state-of-the-art language models such as BERT, deploying them on resource-constrained devices has become challenging. Latency-aware Neural Architecture Search (NAS) is an effective solution for finding an efficient implementation of complex models that satisfy hardware limitations. However, collecting on-device accuracy and latency feedback would significantly slow down the search process, making NAS impractical. To address this, we propose a low-cost method that models both accuracy and latency of BERT-based models on the target device, NVIDIA Jetson TX2, and removes the hardware-related delays from the search loop. Using a Random Forest regressor, our predictors outperform the state-of-the-art and achieve up to 57x speedup while finding a set of near-optimal models. Negin Firouzian, Seyyed Hasan Mozafari, James J. Clark, Warren J. Gross, Brett H. Meyer |
CODES+ISSS | 5 |
| 2022 | CES-KD: Curriculum-based Expert Selection for Guided Knowledge DistillationabstractKnowledge distillation (KD) is an effective tool for compressing deep classification models for edge devices. However, the performance of KD is affected by the large capacity gap between the teacher and student networks. Recent methods have resorted to a multiple teacher assistant (TA) setting for KD, which sequentially decreases the size of the teacher model to relatively bridge the size gap between these models. This paper proposes a new technique called Curriculum Expert Selection for Knowledge Distillation (CES-KD) to efficiently enhance the learning of a compact student under the capacity gap problem. This technique is built upon the hypothesis that a student network should be guided gradually using stratified teaching curriculum as it learns easy (hard) data samples better and faster from a lower (higher) capacity teacher network. Specifically, our method is a gradual TA-based KD technique that selects a single teacher per input image based on a curriculum driven by the difficulty in classifying the image. In this work, we empirically verify our hypothesis and rigorously experiment with CIFAR-10, CIFAR-100, CINIC-10, and ImageNet datasets and show improved accuracy on VGG-like models, ResNets, and WideResNets architectures. Ibtihel Amara, Maryam Ziaeefard, Brett H. Meyer, Warren J. Gross, James J. Clark |
ICPR | 3 |
| 2022 | Efficient Fine-Tuning of BERT Models on the EdgeabstractResource-constrained devices are increasingly the deployment targets of machine learning applications. Static models, however, do not always suffice for dynamic environments. On-device training of models allows for quick adaptability to new scenarios. With the increasing size of deep neural networks, as noted with the likes of BERT and other natural language processing models, comes increased resource requirements, namely memory, computation, energy, and time. Furthermore, training is far more resource intensive than inference. Resource-constrained on-device learning is thus doubly difficult, especially with large BERT-like models. By reducing the memory usage of fine-tuning, pre-trained BERT models can become efficient enough to fine-tune on resource-constrained devices. We propose Freeze And ReconFigure (FAR), a memory-efficient training regime for BERT-like models that reduces the memory usage of activation maps during fine-tuning by avoiding unnecessary parameter updates. FAR reduces fine-tuning time on the DistilBERT model and CoLA dataset by 30 %, and time spent on memory operations by 47%. More broadly, reductions in metric performance on the GLUE and SQuAD datasets are around 1% on average. Danilo Vucetic, Mohammadreza Tayaranian, Maryam Ziaeefard, James J. Clark, Brett H. Meyer, Warren J. Gross |
ISCAS | 5 |
| 2021 | A Design Framework for Invertible LogicabstractInvertible logic using a probabilistic magnetoresistive device model has been recently presented that can compute functions in bidirectional ways and solve several problems quickly, such as factorization and combinational optimization. In this article, we present a design framework for invertible logic circuits. Our approach makes use of linear programming to create a Hamiltonian library with the minimum number of nodes for small invertible-logic functions. In addition, as the device model is approximated based on stochastic computing in synthesizable SystemVerilog, a faster simulation using the compiled SystemC binary is realized than a conventional SPICE-level simulation and is verified using field-programmable gate array (FPGA) as prototyping. Using our design framework, several invertible-logic circuits are designed and emulated (verified) in SystemC, exhibiting five order-of-magnitude faster simulation than conventional work. Naoya Onizawa, Kaito Nishino, Sean C. Smithson, Brett H. Meyer, Warren J. Gross, Hitoshi Yamagata, Hiroyuki Fujita, Takahiro Hanyu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | Worst-case Execution Time Calculation for Query-based Monitors by Witness GenerationabstractRuntime monitoring plays a key role in the assurance of modern intelligent cyber-physical systems, which are frequently data-intensive and safety-critical. While graph queries can serve as an expressive yet formally precise specification language to capture the safety properties of interest, there are no timeliness guarantees for such auto-generated runtime monitoring programs, which prevents their use in a real-time setting. While worst-case execution time (WCET) bounds derived by existing static WCET estimation techniques are safe, they may not be tight as they are unable to exploit domain-specific (semantic) information about the input models. This article presents a semantic-aware WCET analysis method for data-driven monitoring programs derived from graph queries. The method incorporates results obtained from low-level timing analysis into the objective function of a modern graph solver. This allows the systematic generation of input graph models up to a specified size (referred to as witness models ) for which the monitor is expected to take the most time to complete. Hence, the estimated execution time of the monitors on these graphs can be considered as safe and tight WCET. Additionally, we perform a set of experiments with query-based programs running on a real-time platform over a set of generated models to investigate the relationship between execution times and their estimates, and we compare WCET estimates produced by our approach with results from two well-known timing analyzers, aiT and OTAWA. Márton Búr, Kristóf Marussy, Brett H. Meyer, Dániel Varró |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2020 | Probabilistic Sequential Multi-Objective Optimization of Convolutional Neural NetworksabstractWith the advent of deeper, larger and more complex convolutional neural networks (CNN), manual design has become a daunting task, especially when hardware performance must be optimized. Sequential model-based optimization (SMBO) is an efficient method for hyperparameter optimization on highly parameterized machine learning (ML) algorithms, able to find good configurations with a limited number of evaluations by predicting the performance of candidates before evaluation. A case study on MNIST shows that SMBO regression model prediction error significantly impedes search performance in multi-objective optimization. To address this issue, we propose probabilistic SMBO, which selects candidates based on probabilistic estimation of their Pareto efficiency. With a formulation that incorporates error in accuracy prediction and uncertainty in latency measurement, probabilistic Pareto efficiency quantifies a candidate's quality in two ways: its likelihood of being Pareto optimal, and the expected number of current Pareto optimal solutions that it will dominate. We evaluate our proposed method on four image classification problems. Compared to a deterministic approach, probabilistic SMBO consistently generates Pareto optimal solutions that perform better, and that are competitive with state-of-the-art efficient CNN models, offering tremendous speedup in inference latency while maintaining comparable accuracy. Zixuan Yin, Warren J. Gross, Brett H. Meyer |
DATE | 3 |
| 2020 | Using Speech Synthesis to Train End-To-End Spoken Language Understanding ModelsabstractEnd-to-end models are an attractive new approach to spoken language understanding (SLU) in which the meaning of an utterance is inferred directly from the raw audio, without employing the standard pipeline composed of a separately trained speech recognizer and natural language understanding module. The downside of end-to-end SLU is that in-domain speech data must be recorded to train the model. We propose a strategy to overcome this requirement in which speech synthesis is used to generate a large synthetic training dataset from several artificial speakers. We confirm the effectiveness of our approach with experiments on two open-source SLU datasets, where synthesized speech is used both as a sole source of training data and as a form of data augmentation. Loren Lugosch, Brett H. Meyer, Derek Nowrouzezahrai, Mirco Ravanelli |
ICASSP | 2 |
| 2019 | Learning Recurrent Binary/Ternary Weights
Arash Ardakani, Zhengyun Ji, Sean C. Smithson, Brett H. Meyer, Warren J. Gross |
ICLR (Poster) | 4 |
| 2019 | ARINC-825TBv2: A Hardware-in-the-Ioop Simulation Platform for Aerospace Security ResearchabstractCAN sees wide use in the aerospace industry, and is proposed as the link-layer implementation of the proposed Aeronautics Radio, Incorporated 825 protocol standard (ARINC-825). Unfortunately, CAN receives far less attention in aerospace than in the automotive domain; consequently, the aerospace security research community has been unable to assess the vulnerabilities of this protocol, due to a lack of access to data and real-time simulation infrastructure. We introduce ARINC-825TBv2, a hardware-in-the-loop link-layer simulation platform that allows aerospace security researchers to observe ARINC-encoded CAN traffic with realistic data. We present a novel, predictive attacker Gaslighter that we integrate into the testbench, and demonstrate its effectiveness. Furthermore, we implement an intrusion detection algorithm from the literature, Z-Score, and validate it using Gaslighter and realistic CAN data. We show that as the number of messages in the simulation increases, the accuracy of Z-Score drops to below that of a random predictor. Derek Yu, Michael Vaquier, Evan Laflamme, Gabrielle Doucette-Poirier, Justin Tremblay, Brett H. Meyer |
RSP | 6 |
| 2019 | Characterizing the Effectiveness of Hot Sparing on Cost and Performance-per-Watt in Application Specific SIMT
Seyyed Hasan Mozafari, Brett H. Meyer |
Integr. | 2 |
| 2019 | Partitioning and Selection of Data Consistency Mechanisms for Multicore Real-Time SystemsabstractMulticore platforms are becoming increasingly popular in real-time systems. One of the major challenges in designing multicore real-time systems is ensuring consistent and timely access to shared resources. Lock-based protection mechanisms such as MPCP and MSRP have been proposed to guarantee mutually exclusive access in multicore systems at the expense of blocking. In this article, we consider partitioning and scheduling in multicore real-time systems with resource sharing. We first propose a resource-aware task partitioning algorithm for systems with lock-based protection. Wait-free methods, which ensure consistent access to shared memory resources with negligible blocking at the expense of additional memory space, are a suitable alternative when the shared resource is a communication buffer. We propose several approaches to solve the joint problem of task partitioning and the selection of a data consistency mechanism (lock-based or wait-free). The problem is first formulated as an Integer Linear Programming (ILP). For large systems where an ILP solution is not scalable, we propose two heuristic algorithms. Experimental results compare the effectiveness of the proposed approaches in finding schedulable systems with low memory cost and show how the use of wait-free methods can significantly improve schedulability. Zaid Al-bayati, Youcheng Sun, Haibo Zeng 0001, Marco Di Natale, Qi Zhu 0002, Brett H. Meyer |
ACM Trans. Embed. Comput. Syst. | 6 |
| 2018 | Mapping and Scheduling Mixed-Criticality Systems with On-Demand RedundancyabstractEmbedded systems in several domains such as avionics and automotive are subject to inspection from certification authorities. These authorities are interested in verifying the safety-critical aspects of a system and, typically, do not certify non-critical parts. The design of such Mixed-Criticality Systems (MCS) has received increasing attention in recent years. However, although MCS must be designed to overcome transient faults, their susceptibility to transient faults is often overlooked. In this paper, we consider the problem of mapping and scheduling efficient, certifiable MCS that can survive transient faults. We generalize previous MCS models and analysis to support On-Demand Redundancy (ODR). A task set transformation is proposed to generate a modified task set that supports various forms of ODR while satisfying reliability and certification requirements. The analysis is incorporated into a design space exploration algorithm that supports a wide range of fault-tolerance mechanisms and heterogeneous platforms. Experiments show that ODR can improve Quality of Service (QoS) provided to non-critical tasks by 29 percent on average, compared to lockstep execution. Moreover, combining several fault-tolerance mechanisms can lead to additional improvements in schedulability and QoS. Jonah Caplan, Zaid Al-bayati, Haibo Zeng 0001, Brett H. Meyer |
IEEE Trans. Computers | 4 |
| 2017 | Multi-armed bandits for efficient lifetime estimation in MPSoC designabstractReliability in integrated circuits is becoming a critical issue with the miniaturization of electronics. Smaller process technologies have led to higher power densities, resulting in higher temperatures and earlier device wear-out. One way to mitigate failure is by over-provisioning resources and remapping tasks from failed components to components with spare capacity, or slack. Since the slack allocation design space is large, finding the optimal is difficult, as brute-force approaches are impractical. During design space exploration, device lifetimes are typically evaluated using Monte-Carlo Simulation (MCS) by sampling each design equally; this method is inefficient since poor designs are evaluated as accurately as good designs. A better method will focus sampling time on the designs that are difficult to distinguish, reducing the time required to evaluate a set of designs; this can be accomplished using Multi-armed Bandit (MAB) Algorithms. This work demonstrates that MAB achieve the same level of accuracy as MCS in 1.45 to 5.26 times fewer samples. Calvin Ma, Aditya Mahajan, Brett H. Meyer |
DATE | 3 |
| 2017 | Area, Throughput, and Power Trade-Offs for FPGA- and ASIC-Based Execution Stream CompressionabstractAn emerging trend in safety-critical computer system design is the use of compression—for example, using cyclic redundancy check (CRC) or Fletcher checksum (FC)—to reduce the state that must be compared to verify correct redundant execution. We examine the costs and performance of CRC and FC as compression algorithms when implemented in hardware for embedded safety-critical systems. To do so, we have developed parameterizable hardware-generation tools targeting CRC and two novel FC implementations. We evaluate the resulting designs implemented for FPGA and ASIC and analyze their efficiency. While CRC is often best, FC dominates when high throughput is needed. Maria Isabel Mera, Jonah Caplan, Seyyed Hasan Mozafari, Brett H. Meyer, Peter A. Milder |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2016 | A four-mode model for efficient fault-tolerant mixed-criticality systems
Zaid Al-bayati, Jonah Caplan, Brett H. Meyer, Haibo Zeng 0001 |
DATE | 3 |
| 2016 | Capturing True Workload Dependency of BTI-induced Degradation in CPU ComponentsabstractAtomistic-based approaches accurately model Bias Temperature Instability phenomena, but they suffer from prolonged execution times, preventing their seamless integration in system-level analysis flows. In this paper we present a comprehensive flow that combines the accuracy of Capture Emission Time (CET) maps with the efficiency of the Compact Digital Waveform (CDW) representation. That way, we capture the true workload-dependent BTI-induced degradation of selected CPU components. First, we show that existing works that assume constant stress patterns fail to account for workload dependency leading to fundamental estimation errors. Second, we evaluate the impact of different real workloads on selected CPU sub-blocks from a commercial processor design. To the best of our knowledge, this is the first work that combines atomistic property and true workload-dependency for variability analysis. Dimitrios Stamoulis, Simone Corbetta, Dimitrios Rodopoulos, Pieter Weckx, Peter Debacker, Brett H. Meyer, Ben Kaczer, Praveen Raghavan, Dimitrios Soudris, Francky Catthoor, Zeljko Zilic |
ACM Great Lakes Symposium on VLSI | 6 |
| 2016 | Neural networks designing neural networks: multi-objective hyper-parameter optimizationabstractArtificial neural networks have gone through a recent rise in popularity, achieving state-of-the-art results in various fields, including image classification, speech recognition, and automated control. Both the performance and computational complexity of such models are heavily dependant on the design of characteristic hyper-parameters (e.g., number of hidden layers, nodes per layer, or choice of activation functions), which have traditionally been optimized manually. With machine learning penetrating low-power mobile and embedded areas, the need to optimize not only for performance (accuracy), but also for implementation complexity, becomes paramount. In this work, we present a multi-objective design space exploration method that reduces the number of solution networks trained and evaluated through response surface modelling. Given spaces which can easily exceed 1020 solutions, manually designing a near-optimal architecture is unlikely as opportunities to reduce network complexity, while maintaining performance, may be overlooked. This problem is exacerbated by the fact that hyper-parameters which perform well on specific datasets may yield sub-par results on others, and must therefore be designed on a per-application basis. In our work, machine learning is leveraged by training an artificial neural network to predict the performance of future candidate networks. The method is evaluated on the MNIST and CIFAR-10 image datasets, optimizing for both recognition accuracy and computational complexity. Experimental results demonstrate that the proposed method can closely approximate the Pareto-optimal front, while only exploring a small fraction of the design space. Sean C. Smithson, Warren J. Gross, Brett H. Meyer |
ICCAD | 4 |
| 2016 | Tolerating the Consequences of Multiple EM-Induced C4 Bump FailuresabstractWith ever-increasing on-chip current density, technology scaling is pushing the electromigration (EM)-induced robustness of silicon chips' controlled collapse chip connection (C4) bump array to its limit. Since the density of C4 bumps is projected to be constant in the future, it is increasingly becoming challenging to guarantee EM-failure free for all power-supply bumps without increasing chip packaging cost or encroaching on bumps sites needed for I/O. In this paper, we develop a statistical simulation framework to analyze the mechanism and consequences of multiple power-bump wearout. Our analysis shows that the penalty of a moderate number of EM-induced power-bump failures is fairly small. A mild increase in on-chip supply voltage noise guardband can tolerate these bump failures and significantly increase a system mean-time-to-failure (MTTF). As a result, the targeted system MTTF can be achieved with significantly reduced power-bump count (e.g., 43% less) and a small extra noise margin (e.g., 0.5% VddIR drop). Runjie Zhang, Brett H. Meyer, Ke Wang 0011, Mircea R. Stan, Kevin Skadron |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2015 | A cross-layer design exploration of charge-recycled power-delivery in many-layer 3d-ICabstract3D-IC technology brings both the opportunities to continue the historical trend of integration-level scaling and the challenges to deliver power reliably and efficiently. Voltage-stacking (V-S), a charge-recycled power delivery scheme that connects the different layers' supply/ground nets into a series stack, provides a scalable solution to the 3D-IC power delivery wall. While prior work has extensively discussed the implementations of V-S at circuit-level, a cross-layer study that examines its system-level implications is missing. In this paper, we start with a circuit implementation of a charge-recycled voltage regulator and build an architecture-level model to study the costs and benefits of utilizing V-S in 3D-IC. Our study shows that by significantly improving the EM-lifetime of C4 and TSV array (e.g., up to 5x) while only marginally increasing the average-case voltage noise (e.g., 0.75% Vdd IR drop), V-S provides a scalable solution for many-layer 3D-IC's power delivery challenge. Runjie Zhang, Kaushik Mazumdar, Brett H. Meyer, Ke Wang 0011, Kevin Skadron, Mircea R. Stan |
DAC | 3 |
| 2015 | Yield-aware Performance-Cost Characterization for Multi-Core SIMTabstractRedundancy is now routinely allocated in circuits, microarchitectural structures, or at the system level, to mitigate mounting manufacturing yield losses. In this paper, we propose spare lane sharing, which reduces the cost of multi-core SIMT systems by allowing one of two neighboring cores to make use of a redundant lane if necessary. We have evaluated the performance-cost trade-offs of core-, lane-, and shared-lane-sparing under a variety of benchmarks, and found that for nearly all applications shared-lane-sparing outperforms lane-sparing, reducing cost by up to 20%. Seyyed Hasan Mozafari, Brett H. Meyer, Kevin Skadron |
ACM Great Lakes Symposium on VLSI | 2 |
| 2015 | Efficient Reliability Analysis of Processor Datapath using Atomistic BTI Variability ModelsabstractIn this paper, we propose EDA methodologies for efficient, datapath-wide reliability analysis under Bias Temperature Instability (BTI). The proposed EDA flow combines the efficiency of atomistic, pseudo-transient BTI modeling with the accuracy of commercial Static Timing Analysis (STA) tools. In order to reduce the transistor inventory that needs to be tracked by the STA solver, we develop a threshold-pruning methodology to identify the variation-critical part of a design. That way, we accelerate variation-aware STA iterations, with a maximum speedup of 6.82x achieved for representative benchmark circuits. We substantiate the efficiency of the proposed framework for realistic designs. For a CPU datapath, our threshold-pruning technique outperforms built-in pruning commands of the STA solver by 16.87% in terms of runtime improvement. We demonstrate the impact of BTI after three years of operation, with clock frequency degradation up to 24% and functional yield reduction below 90% for higher frequencies. Dimitrios Stamoulis, Dimitrios Rodopoulos, Brett H. Meyer, Dimitrios Soudris, Francky Catthoor, Zeljko Zilic |
ACM Great Lakes Symposium on VLSI | 3 |
| 2015 | Transient voltage noise in charge-recycled power delivery networks for many-layer 3D-ICabstractAside from the benefits it brings, 3D-IC technology inevitably exacerbates the difficulty of power delivery with volumetrically increasing power consumption. Recent work managed to “recycle” current within the 3D stack by linking the different layers' supply/ground nets into a series connection. This charge-recycled (also known as voltage-stacked, or V-S) scheme provides a scalable solution for 3D-IC's power delivery because it supports an arbitrary number of layers with a constant off-chip current demand. Although prior work has studied the circuit implementation of a V-S power delivery network (PDN) and its current-reduction benefits, a whole-system evaluation of V-S PDNs' transient voltage noise and a noise comparison between the V-S PDN and the traditional PDN are missing. In this paper, we build a system-level model to examine voltage-stacked 3D-ICs' transient noise and explore the impact of different PDN design parameters and workload behaviors. Our results show that compared with the traditional PDN scheme, V-S provides stronger isolation for cross-layer noise interference, which in turn grants higher performance benefits for run-time noise mitigation techniques, such as dynamic margin adaptation. We observe that, compared with traditional PDNs, V-S PDNs provide up to 60% lower transient noise in the worst-case scenario. Furthermore, we show that V-S PDNs significantly reduce the packaging cost, because their noise is almost insensitive to the package impedance (e.g., a 300% impedance increase only raises worst-case noise by less than 0.3% Vdd). Runjie Zhang, Kaushik Mazumdar, Brett H. Meyer, Ke Wang 0011, Kevin Skadron, Mircea R. Stan |
ISLPED | 3 |
| 2015 | Task placement and selection of data consistency mechanisms for real-time multicore applicationsabstractMulticores are today used in automotive, controls and avionics systems supporting real-time functionality. When real-time tasks allocated on different cores cooperate through the use of shared communication resources, they need to be protected by mechanisms that guarantee access in a mutual exclusive way with bounded worst-case blocking time. Lock-based mechanisms such as MPCP and MSRP have been developed to fulfill this demand, and research papers are today tackling the problem of finding the optimal task placement in multicores while trying to meet the deadlines against blocking times. In this paper, we propose a resource-aware task allocation algorithm for systems that use MSRP to protect shared resources. Furthermore, we leverage the additional opportunity provided by wait-free methods as an alternative data consistency mechanism for the case that the shared resource is communication or state memory. An algorithm that performs both task allocation and data consistency mechanism (MSRP or wait-free) selection is proposed. The selective use of wait-free methods can significantly extend the range of schedulable systems at the cost of memory. Zaid Al-bayati, Youcheng Sun, Haibo Zeng 0001, Marco Di Natale, Qi Zhu 0002, Brett H. Meyer |
RTAS | 6 |
| 2015 | Techniques for on-demand structural redundancy for massively parallel processor arrays
Vahid Lari, Jürgen Teich, Alexandru Tanase, Michael Witterauf, Faramarz Khosravi, Brett H. Meyer |
J. Syst. Archit. | 6 |
| 2014 | Walking pads: Fast power-supply pad-placement optimizationabstractWe propose a novel C4 pad placement optimization framework for 2D power delivery grids: Walking Pads (WP). WP optimizes pad locations by moving pads according to the “virtual forces” exerted on them by other pads and current sources in the system. WP algorithms achieve the same IR drop as state-of-the-art techniques, but are up to 634X faster. We further propose an analytical model relating pad count and IR drop for determining the optimal pad count for a given IR drop budget. Ke Wang 0011, Brett H. Meyer, Runjie Zhang, Kevin Skadron, Mircea R. Stan |
ASP-DAC | 2 |
| 2014 | Walking Pads: Managing C4 Placement for Transient Voltage Noise MinimizationabstractTransient voltage noise, including resistive and reactive noise, causes timing errors at runtime. We introduce a heuristic framework---Walking Pads---to minimize transient voltage violations by optimizing power supply pad placement. We show that the steady-state optimal design point differs from the transient optimum, and further noise reduction can be achieved with transient optimization. Our methodology significantly reduces voltage violations by balancing the average transient voltage noise of the four branches at each pad site. When we optimize pad placement using a representative stressmark, voltage violations are reduced 46-80% across 11 Parsec benchmarks with respect to the results from IR-drop-optimized pad placement. We also show that the allocation of on-chip decoupling capacitance significantly influences the optimal locations of pads. Ke Wang 0011, Brett H. Meyer, Runjie Zhang, Mircea R. Stan, Kevin Skadron |
DAC | 2 |
| 2014 | Trade-offs in execution signature compression for reliable processor systemsabstractAs semiconductor processes scale, making transistors more vulnerable to transient upset, a wide variety of microarchitectural and system-level strategies are emerging to perform efficient error detection and correction computer systems. While these approaches often target various application domains and address error detection and correction at different granularities and with different overheads, an emerging trend is the use of state compression, e.g., cyclic redundancy check (CRC), to reduce the cost of redundancy checking. Prior work in the literature has shown that Fletcher's checksum (FC), while less effective where error detection probability is concerned, is less computationally complex when implemented in software than the more-effective CRC. In this paper, we reexamine the suitability of CRC and FC as compression algorithms when implemented in hardware for embedded safety-critical systems. We have developed and evaluated parameterizable implementations of CRC and FC in FPGA, and we observe that what was true for software implementations does not hold in hardware: CRC is more efficient than FC across a wide variety of target input bandwidths and compression strengths. Jonah Caplan, Maria Isabel Mera, Peter A. Milder, Brett H. Meyer |
DATE | 4 |
| 2014 | Flexibility and Circuit Overheads in Reconfigurable SIMD/MIMD SystemsabstractDynamically reconfigurable SIMD/MIMD architectures made from simple cores have emerged to exploit diverse forms of parallelism in applications [1,2]. In this work, we investigate the circuit-level overhead and flexibility tradeoffs of such architectures through the design of a custom reconfigurable SIMD/MIMD system. Saad Arrabi, Kevin Skadron, Benton H. Calhoun, John C. Lach, Brett H. Meyer |
FCCM | 7 |
| 2014 | MB-FICA: multi-bit fault injection and coverage analysisabstractRecent studies have shown a dramatic increase in multi-bit upset (MBU) events and related errors as transistors continue to shrink. Mojing Liu, Brett H. Meyer |
ACM Great Lakes Symposium on VLSI | 3 |
| 2014 | Architecture implications of pads as a scarce resourceabstractDue to non-ideal technology scaling, delivering a stable supply voltage is increasingly challenging. Furthermore, competition for limited chip interface resources (i.e., C4 pads) between power supply and I/O, and the loss of such resources to electromigration, means that constructing a power delivery network (PDN) that satisfies noise margins without compromising performance is and will remain a critical problem for architects and circuit designers alike. Simple guardbanding will no longer work, as the consequent performance penalty will grow with technology scaling. In this paper, we develop a pre-RTL PDN model, VoltSpot, for the purpose of studying the performance and noise tradeoffs among power supply and I/O pad allocation, the effectiveness of noise mitigation techniques, and the consequent implications of electromigration-induced PDN pad failure. Our simulations demonstrate that, despite their integral role in the PDN, power/ground pads can be aggressively reduced (by conversion into I/O pads) to their electromigration limit with minimal performance impact from extra voltage noise - provided the system implements a suitable noise-mitigation strategy. The key observation is that even though reducing power/ground pads significantly increases the number of voltage emergencies, the average noise amplitude increase is small. Overall, we can triple I/O bandwidth while maintaining target lifetimes and incurring only 1.5% slowdown. Runjie Zhang, Ke Wang 0011, Brett H. Meyer, Mircea R. Stan, Kevin Skadron |
ISCA | 3 |
| 2014 | Cost-effective lifetime and yield optimization for NoC-based MPSoCsabstractAs manufacturing processes scale, designers are increasingly dependent on techniques to mitigate manufacturing defect and permanent failure. In embedded systems-on-chip, system lifetime and yield can be increased using slack —under-utilization in execution and storage resources—so that when components are defective, data and tasks can be remapped and rescheduled. For any given system, the design space of possible slack allocations is both large and complex, consisting of every possible way to replace each component in the initial system with another from the component library. Based on the observation that useful slack is often quantized, we have developed Critical Quantity Slack Allocation (CQSA), an approach that effectively and efficiently allocates execution and storage slack to jointly optimize system yield and cost. While exploring less than 1.4% of the slack allocation design space, our approach consistently outperforms alternative slack allocation techniques to find sets of designs within 1.4% of the lifetime-cost Pareto-optimal front. When applied to yield-cost optimization, our approach again outperforms alternative techniques, exploring less than 1.62% of the design space to find sets of designs within 4.27% of the yield-cost Pareto-optimal front. One advantage of managing failure at the system level is that the same techniques that improve lifetime often also improve yield. As a result, with little modification, CQSA is further able to perform effective joint optimization of lifetime and yield, finding designs within 1.6% of the Pareto-optimal front. Brett H. Meyer, Adam S. Hartman, Donald E. Thomas |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2013 | Architectural implications of spatial thermal filtering
Karthik Sankaranarayanan, Brett H. Meyer, Wei Huang 0004, Robert J. Ribando, Hossein Haj-Hariri, Mircea R. Stan, Kevin Skadron |
Integr. | 2 |
| 2012 | ArchFP: Rapid prototyping of pre-RTL floorplans
Gregory G. Faust, Runjie Zhang, Kevin Skadron, Mircea R. Stan, Brett H. Meyer |
VLSI-SoC | 5 |
| 2011 | Cost-effective safety and fault localization using distributed temporal redundancyabstractCost pressure is driving vendors of safety-critical systems to integrate previously distributed systems. One natural approach we have previous introduced is On-Demand Redundancy (ODR), which allows safety-critical and non-critical tasks, traditionally isolated to limit interference, to execute on shared resources. Our prior work has shown that relaxed dedication (RD), one ODR strategy which allows non-critical tasks (NCTs) to execute on idle critical task resources (CTRs), significantly increases NCT throughput. Unfortunately, there are circumstances under which, in spite of this opportunity, it is difficult to effectively schedule NCTs. Brett H. Meyer, Benton H. Calhoun, John C. Lach, Kevin Skadron |
CASES | 1 |
| 2011 | Reducing the cost of redundant execution in safety-critical systems using relaxed dedicationabstractWe introduce on-demand redundancy, a set of architectural techniques that leverage the tightly-coupled nature of components in systems-on-chip to reduce the cost of safety-critical systems. On-demand redundancy eases the assumptions that traditionally segregate the execution of critical and non-critical tasks (NCTs), making resources available for critical tasks at potentially arbitrary points in both space and time, and otherwise freeing resources to execute non-critical tasks when critical tasks are not executing. Relaxed dedication is one such technique that allows non-critical tasks to execute on critical task resources. Our results demonstrate that for a wide variety of applications and architectures, relaxed dedication is more cost-effective than a traditional approach that employs dedicated resources executing in lockstep. Applied to dual-modular redundancy (DMR), relaxed dedication exposes 73% more NCT cycles than traditional DMR on average, across a wide variety of usage scenarios. Brett H. Meyer, Nishant George, Benton H. Calhoun, John C. Lach, Kevin Skadron |
DATE | 1 |
| 2010 | Cost-effective slack allocation for lifetime improvement in NoC-based MPSoCsabstractWear-out related permanent faults are projected to make system lifetime a critical issue for all designs. In embedded systems, lifetime can be increased using slack, underutilization in execution and storage resources, so that when components fail, data and tasks can be re-mapped and re-scheduled. The design space of possible slack allocation is both large and complex. However, based on the observation that useful slack is often quantized, we have developed an approach that effectively and efficiently allocates execution and storage slack to jointly optimize system lifetime and cost. While exploring less than 1.4% of the slack allocation design space, our approach consistently outperforms alternative slack allocation techniques to find sets of designs within 1.4% of the lifetime-cost Pareto-optimal front. Brett H. Meyer, Adam S. Hartman, Donald E. Thomas |
DATE | 1 |
| 2010 | Temperature-to-power mappingabstractAccurate power maps are useful for power model validation, process variation characterization, leakage estimation, and power optimization, but are hard to measure directly. Deriving power maps from measured thermal maps is the inverse problem of the power-to-temperature mapping, extensively studied through thermal simulation. Until recently this inverse heat conduction problem has received little attention in the microarchitecture research community. This paper first identifies the source of difficulties for the problem. The inverse mapping is then performed by applying constraints from microarchitecture-level observations. The inherent large sensitivity of the resultant power map is minimized through thermal map-filtering and constrained least-squares optimization. Choices of filter parameters and optimization constraints are investigated and their effects are evaluated. Furthermore, the paper highlights the differences between the grid and block modeling in the inverse mapping which were often ignored by previous schemes. The proposed methods reduce the mapping error by more than 10× compared to unoptimized solutions. To our best knowledge this is the first work to quantitatively evaluate and minimize the noise effect in the temperature to power mapping problem at the microarchitecture level for both grid and block mode, and for the steady and transient case. Zhenyu Qi 0001, Brett H. Meyer, Wei Huang 0004, Robert J. Ribando, Kevin Skadron, Mircea R. Stan |
ICCD | 2 |
| 2007 | Rethinking Automated Synthesis of MPSoC ArchitecturesabstractEmerging heterogeneous multiprocessors have custom memory and bus architectures that must balance resource sharing and system partitioning to meet cost constraints. We propose an augmented simulated annealing synthesis tool that uses system performance and layout evaluation to drive simultaneous data mapping, memory allocation and bus synthesis. A detailed look at the resulting automated design process reveals an approach that, contrary to prior approaches, optimizes bus topology first rather than last, providing design insight for the development of future tools. Brett H. Meyer, Donald E. Thomas |
IPDPS | 1 |
| 2005 | Power-Performance Simulation and Design Strategies for Single-Chip Heterogeneous MultiprocessorsabstractSingle chip heterogeneous multiprocessors (SCHMs) are becoming more commonplace, especially in portable devices where reduced energy consumption is a priority. The use of coordinated collections of processors which are simpler or which execute at lower clock frequencies is widely recognized as a means of reducing power while maintaining latency and throughput. A primary limitation of using this approach to reduce power at the system level has been the time to develop and simulate models of many processors at the instruction set simulator level. High-level models, simulators, and design strategies for SCHMs are required to enable designers to think in terms of collections of cooperating, heterogeneous processors in order to reduce power. Toward this end, this paper has two contributions. The first is to extend a unique, preexisting high-level performance simulator, the Modeling Environment for Software and Hardware (MESH), to include power annotations. MESH can be thought of as a thread-level simulator instead of an instruction-level simulator. Thus, the problem is to understand how power might be calibrated and annotated with program fragments instead of at the instruction level. Program fragments are finer-grained than threads and coarser-grained than instructions. Our experimentation found that compilers produce instruction patterns that allow power to be annotated at this level using a single number over all compiler-generated fragments executing on a processor. Since energy is power*time, this makes system runtime (i.e., performance) the dominant factor to be dynamically calculated at this level of simulation. The second contribution arises from the observation that high-level modeling is most beneficial when it opens up new possibilities for organizing designs. Thus, we introduce a design strategy, enabled by the high-level performance power-simulation, which we refer to as spatial voltage scaling. The strategy both reduces overall system power consumption and improves performance in our example. The design space for this design strategy could not be explored without high-level SCHM power-performance simulation. Brett H. Meyer, Joshua J. Pieper, JoAnn M. Paul, Jeffrey E. Nelson, Sean M. Pieper, Anthony Rowe 0001 |
IEEE Trans. Computers | 1 |