EDBT 2026 Demo / reviewers in the wild / expert
Bradley McDanel
dblp:163/7366
· DBLP profile ↗
21ranked-venue papers
8as first author
9since 2021 · last 2026
0000-0001-6684-8918ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 7 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 2 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Individualized Quizzes From Student Code with LLMsabstractThe rapid integration of AI coding assistants into student workflows has fundamentally challenged the role of traditional programming assignments as a measure of individual understanding. Rather than attempting to prohibit these tools, we propose a pedagogical shift: from assessing the process of code creation to verifying a student's comprehension and ownership of their final submitted work. We present a fully automated pipeline that leverages a Large Language Model (LLM) to generate individualized, in-class quizzes with questions targeting specific segments of each student's own code. This approach compels students to engage deeply with their submissions, as they must be prepared to explain their logic and implementation choices. While students might rely too heavily on AI tools to complete out-of-class assignments, this method places the responsibility of understanding the code on the student. The instructor remains central to the process, reviewing the quizzes before distributing to the class, and hand-grading the quizzes to provide meaningful, nuanced feedback. This ensures both fairness and a human connection in the assessment loop. Edmund Novak, Bradley McDanel |
SIGCSE (2) | 2 |
| 2025 | AMUSD: Asynchronous Multi-Device Speculative Decoding for LLM AccelerationabstractLarge language models typically generate tokens autoregressively, using each token as input for the next. Recent work on Speculative Decoding has sought to accelerate this process by employing a smaller, faster draft model to more quickly generate candidate tokens. These candidates are then verified in parallel by the larger (original) verify model, resulting in overall speedup compared to using the larger model by itself in an autoregressive fashion. In this work, we introduce AMUSD (Asynchronous Multi-device Speculative Decoding), a system that further accelerates generation by decoupling the draft and verify phases into a continuous, asynchronous approach. Unlike conventional speculative decoding, where only one model (draft or verify) performs token generation at a time, AMUSD enables both models to perform predictions independently on separate devices (e.g., GPUs). We evaluate our approach over multiple datasets and show that AMUSD achieves an average 29% improvement over speculative decoding and up to 1.96× speedup over conventional autoregressive decoding, while achieving identical output quality. Our system is open-source and available at https://github.com/BradMcDanel/AMUSD/. Bradley McDanel |
ISCAS | 1 |
| 2025 | Designing LLM-Resistant Programming Assignments: Insights and Strategies for CS EducatorsabstractThe rapid advancement of Large Language Models (LLMs) like ChatGPT has raised concerns among computer science educators about how programming assignments should be adapted. This paper explores the capabilities of LLMs (GPT-3.5, GPT-4, and Claude Sonnet) in solving complete, multi-part CS homework assignments from the SIGCSE Nifty Assignments list. Through qualitative and quantitative analysis, we found that LLM performance varied significantly across different assignments and models, with Claude Sonnet consistently outperforming the others. The presence of starter code and test cases improved performance for advanced LLMs, while certain assignments, particularly those involving visual elements, proved challenging for all models. LLMs often disregarded assignment requirements, produced subtly incorrect code, and struggled with context-specific tasks. Based on these findings, we propose strategies for designing LLM-resistant assignments. Our work provides insights for instructors to evaluate and adapt their assignments in the age of AI, balancing the potential benefits of LLMs as learning tools with the need to ensure genuine student engagement and learning. Bradley McDanel, Edmund Novak |
SIGCSE (1) | 1 |
| 2023 | Dynamic Patch Sampling for Efficient Training and Dynamic Inference in Vision TransformersabstractWe introduce the notion of a Patch Sampling Schedule (PSS), that varies the number of Vision Transformer (ViT) patches used per batch during training. Since all patches are not equally important for most vision objectives (e.g., classification), we argue that less important patches can be used in fewer training iterations, leading to shorter training time with minimal impact on performance. Additionally, we observe that training with a PSS makes a ViT more robust to a wider patch sampling range during inference. This allows for a fine-grained, dynamic trade-off between throughput and accuracy during inference. We evaluate using PSSs on ViTs for ImageNet both trained from scratch and pre-trained using a reconstruction loss function. For the pre-trained model, we achieve a 0.26% reduction in classification accuracy for a 31% reduction in training time (from 25 to 17 hours) compared to using all patches each iteration. Bradley McDanel, Chi Phuong Huynh |
ICMLA | 1 |
| 2023 | StitchNet: Composing Neural Networks from Pre-Trained FragmentsabstractWe propose StitchNet, a novel neural network cre-ation paradigm that stitches together fragments (one or more consecutive network layers) from multiple pre-trained neural networks. StitchNet allows the creation of high-performing neural networks without the large compute and data requirements needed under traditional model creation processes via backprop-agation training. We leverage Centered Kernel Alignment (CKA) as a compatibility measure to efficiently guide the selection of these fragments in composing a network for a given task tailored to specific accuracy needs and computing resource constraints. We then show that these fragments can be stitched together to create neural networks with accuracy comparable to that of traditionally trained networks at a fraction of computing resource and data requirements. Finally, we explore a novel on-the-fly personalized model creation and inference application enabled by this new paradigm. The code is available at https://github.com/steerapi/stitchnet. Surat Teerapittayanon, Marcus Z. Comiter, Bradley McDanel, H. T. Kung 0001 |
ICMLA | 3 |
| 2022 | FAST: DNN Training Under Variable Precision Block Floating Point with Stochastic RoundingabstractBlock Floating Point (BFP) can efficiently support quantization for Deep Neural Network (DNN) training by providing a wide dynamic range via a shared exponent across a group of values. In this paper, we propose a Fast First, Accurate Second Training (FAST) system for DNNs, where the weights, activations, and gradients are represented in BFP. FAST supports matrix multiplication with variable precision BFP input operands, enabling incremental increases in DNN precision throughout training. By increasing the BFP precision across both training iterations and DNN layers, FAST can greatly shorten the training time while reducing overall hardware resource usage. Our FAST Multipler-Accumulator (fMAC) supports dot product computations under multiple BFP precisions. We validate our FAST system on multiple DNNs with different datasets, demonstrating a 2-6× speedup in training on a single-chip platform over prior work based on mixed-precision or block floating point number systems while achieving similar performance in validation accuracy. Sai Qian Zhang, Bradley McDanel, H. T. Kung 0001 |
HPCA | 2 |
| 2022 | Accelerating DNN Training with Structured Data Gradient PruningabstractWeight pruning is a technique to make Deep Neural Network (DNN) inference more computationally efficient by reducing the number of model parameters over the course of training. However, most weight pruning techniques generally does not speed up DNN training and can even require more iterations to reach model convergence. In this work, we propose a novel Structured Data Gradient Pruning (SDGP) method that can speed up training without impacting model convergence. This approach enforces a specific sparsity structure, where only N out of every M elements in a matrix can be nonzero, making it amenable to hardware acceleration. Modern accelerators such as the Nvidia A100 GPU support this type of structured sparsity for 2 nonzeros per 4 elements in a reduction. Assuming hardware support for 2:4 sparsity, our approach can achieve a 15-25% reduction in total training time without significant impact to performance. Source code and pre-trained models are available at https://github.com/BradMcDanel/sdgp. Bradley McDanel, Helia Dinh, John Magallanes |
ICPR | 1 |
| 2021 | Training for multi-resolution inference using reusable quantization termsabstractLow-resolution uniform quantization (e.g., 4-bit bitwidth) for both Deep Neural Network (DNN) weights and data has emerged as an important technique for efficient inference. Departing from conventional quantization, we describe a novel training approach to support inference at multiple resolutions by reusing a single set of quantization terms (the same set of nonzero bits in values). The proposed approach streamlines the training and supports dynamic selection of resolution levels during inference. We evaluate the method on a diverse range of applications including multiple CNNs on ImageNet, an LSTM on Wikitext-2, and YOLO-v5 on COCO. We show that models resulting from our multi-resolution training can support up to 10 resolutions with only a moderate performance reduction (e.g., ≤ 1%) compared to training them individually. Lastly, using an FPGA, we compare our multi-resolution multiplier-accumulator (mMAC) against other conventional MAC designs and evaluate the inference performance. We show that the mMAC design broadens the choices in trading off cost, efficiency, and latency across a range of computational budgets. Sai Qian Zhang, Bradley McDanel, H. T. Kung 0001, Xin Dong 0009 |
ASPLOS | 2 |
| 2021 | Saturation RRAM Leveraging Bit-Level Sparsity Resulting from Term QuantizationabstractThe proposed saturation RRAM for in-memory computing of a pre-trained Convolutional Neural Network (CNN) inference imposes a limit on the maximum analog value output from each bitline in order to reduce analog-to-digital (A/D) conversion costs. The proposed scheme uses term quantization (TQ) to enable flexible bit annihilation at any position for a value in the context of a group of weights values in RRAM. This enables a drastic reduction in the required ADC resolution while still maintaining CNN model accuracy. Specifically, we show that the A/D conversion errors after TQ have a minimum impact on the classification accuracy of the inference task. For instance, for a 64×64 RRAM, reducing the ADC resolution from 6 bits to 4 bits enables a 1.58× reduction in the total system power, without a significant impact to classification accuracy. Bradley McDanel, Sai Qian Zhang, H. T. Kung 0001 |
ISCAS | 1 |
| 2020 | Term quantization: furthering quantization at run timeabstractWe present a novel technique, called Term Quantization (TQ), for furthering quantization at run time for improved computational efficiency of deep neural networks (DNNs) already quantized with conventional quantization methods. TQ operates on power-of-two terms in expressions of values. In computing a dot-product computation, TQ dynamically selects a fixed number of largest terms to use from values of the two vectors. By exploiting weight and data distributions typically present in DNNs, TQ has a minimal impact on DNN model performance (e.g., accuracy or perplexity). We use TQ to facilitate tightly synchronized processor arrays, such as systolic arrays, for efficient parallel processing. We evaluate TQ on an MLP for MNIST, multiple CNNs for ImageNet and an LSTM for Wikitext-2. We demonstrate significant reductions in inference computation costs (between 3-10×) compared to conventional uniform quantization for the same level of model performance. H. T. Kung 0001, Bradley McDanel, Sai Qian Zhang |
SC | 2 |
| 2019 | Maestro: A Memory-on-Logic Architecture for Coordinated Parallel Use of Many Systolic ArraysabstractWe present the Maestro memory-on-logic 3D-IC architecture for coordinated parallel use of a plurality of systolic arrays (SAs) in performing deep neural network (DNN) inference. Maestro reduces under-utilization common for a single large SA by allowing parallel use of many smaller SAs on DNN weight matrices of varying shapes and sizes. In order to buffer immediate results in memory blocks (MBs) and provide coordinated high-bandwidth communication between SAs and MBs in transferring weights and results Maestro employs three innovations. (1) An SA on the logic die can access its corresponding MB on the memory die in short distance using 3D-IC interconnects, (2) through an efficient switch based on H-trees, an SA can access any MB with low latency, and (3) the switch can combine partial results from SAs in an elementwise fashion before writing back to a destination MB. We describe the Maestro architecture, including a circuit and layout design, detail scheduling of the switch, analyze system performance for real-time inference applications using input with batch size equal to one, and showcase applications for deep learning inference, with ShiftNet for computer vision and recent Transformer models for natural language processing. For the same total number of systolic cells, Maestro, with multiple smaller SAs, leads to 16x and 12x latency improvements over a single large SA on ShiftNet and Transformer, respectively. Compared to a floating-point GPU implementation of ShiftNet and Transform, a baseline Maestro system with 4,096 SAs (each with 8x8 systolic cells) provides significant latency improvements of 30x and 47x, respectively. H. T. Kung 0001, Bradley McDanel, Sai Qian Zhang, Xin Dong 0009, Chih-Chiang Chen |
ASAP | 2 |
| 2019 | Packing Sparse Convolutional Neural Networks for Efficient Systolic Array Implementations: Column Combining Under Joint OptimizationabstractThis paper describes a novel approach of packing sparse convolutional neural networks into a denser format for efficient implementations using systolic arrays. By combining multiple sparse columns of a convolutional filter matrix into a single dense column stored in the systolic array, the utilization efficiency of the systolic array can be substantially increased (e.g., 8x) due to the increased density of nonzero weights in the resulting packed filter matrix. In combining columns, for each row, all filter weights but the one with the largest magnitude are pruned. The remaining weights are retrained to preserve high accuracy. We study the effectiveness of this joint optimization for both high utilization efficiency and classification accuracy with ASIC and FPGA designs based on efficient bit-serial implementations of multiplier-accumulators. We demonstrate that in mitigating data privacy concerns the retraining can be accomplished with only fractions of the original dataset (e.g., 10% for CIFAR-10). We present analysis and empirical evidence on the superior performance of our column combining approach against prior arts under metrics such as energy efficiency (3x) and inference latency (12x). H. T. Kung 0001, Bradley McDanel, Sai Qian Zhang |
ASPLOS | 2 |
| 2019 | Full-stack optimization for accelerating CNNs using powers-of-two weights with FPGA validationabstractWe present a full-stack optimization framework for accelerating inference of CNNs (Convolutional Neural Networks) and validate the approach with a field-programmable gate array (FPGA) implementation. By jointly optimizing CNN models, computing architectures, and hardware implementations, our full-stack approach achieves unprecedented performance in the trade-off space characterized by inference latency, energy efficiency, hardware utilization, and inference accuracy. An FPGA implementation is used as the validation vehicle for our design, achieving a 2.28ms inference latency for the ImageNet benchmark. Our implementation shines in that it has 9x higher energy efficiency compared to other implementations while achieving comparable latency. A highlight of our approach which contributes to the achieved high energy efficiency is an efficient Selector-Accumulator (SAC) architecture for implementing CNNs with powers-of-two weights. Compared to an FPGA implementation for a traditional 8-bit MAC, SAC substantially reduces required hardware resources (4.85x fewer lookup tables) and power consumption (2.48x). Bradley McDanel, Sai Qian Zhang, H. T. Kung 0001, Xin Dong 0009 |
ICS | 1 |
| 2019 | Systolic Building Block for Logic-on-Logic 3D-IC Implementations of Convolutional Neural NetworksabstractWe present a building block architecture for systolic array 3D-IC implementations of convolutional neural network (CNN) inference. The building block can be part of a library offered by a chip design service provider to support efficient CNN implementations. We describe how the building block can form systolic arrays for implementing low-latency, energy-efficient CNN inference for models of any size, while incorporating advanced packaging features such as “logic-on-logic” 3D-IC (micro-bump/TSV, monolithic 3D or other 3D technology). We present delay and power analysis for 2D and 3D implementations, and argue that as systolic arrays scale in size, 3D implementations based on, e.g., micro-bump/TSV, lead to significant performance improvements over 2D implementations. H. T. Kung 0001, Bradley McDanel, Sai Qian Zhang, C. T. Wang, Jin Cai, Victor C. Y. Chang, M. F. Chen, Jack Yuan-Chen Sun, Douglas Yu |
ISCAS | 2 |
| 2018 | Adaptive Tiling: Applying Fixed-size Systolic Arrays To Sparse Convolutional Neural NetworksabstractWe introduce adaptive tiling, a method of partitioning layers in a sparse convolutional neural network (CNN) into blocks of filters and channels, called tiles, each implementable with a fixed-size systolic array. By allowing a tile to adapt its size so that it can cover a large sparse area, we minimize the total number of tiles, or equivalently, the number of systolic array calls required to perform CNN inference. The proposed scheme resolves a challenge of applying systolic array architectures, traditionally designed for dense matrices, to sparse CNNs. To validate the approach, we construct a highly sparse Lasso-Mobile network by pruning MobileNet trained with an l1 regularization penalty, and demonstrate that adaptive tiling can lead to a 2- 3x reduction in systolic array calls, on Lasso-Mobile, for several benchmark datasets. H. T. Kung 0001, Bradley McDanel, Sai Qian Zhang |
ICPR | 2 |
| 2017 | Embedded Binarized Neural Networks
Bradley McDanel, Surat Teerapittayanon, H. T. Kung 0001 |
EWSN | 1 |
| 2017 | Distributed Deep Neural Networks Over the Cloud, the Edge and End DevicesabstractWe propose distributed deep neural networks (DDNNs) over distributed computing hierarchies, consisting of the cloud, the edge (fog) and end devices. While being able to accommodate inference of a deep neural network (DNN) in the cloud, a DDNN also allows fast and localized inference using shallow portions of the neural network at the edge and end devices. When supported by a scalable distributed computing hierarchy, a DDNN can scale up in neural network size and scale out in geographical span. Due to its distributed nature, DDNNs enhance sensor fusion, system fault tolerance and data privacy for DNN applications. In implementing a DDNN, we map sections of a DNN onto a distributed computing hierarchy. By jointly training these sections, we minimize communication and resource usage for devices and maximize usefulness of extracted features which are utilized in the cloud. The resulting system has built-in support for automatic sensor fusion and fault tolerance. As a proof of concept, we show a DDNN can exploit geographical diversity of sensors to improve object recognition accuracy and reduce communication cost. In our experiment, compared with the traditional method of offloading raw sensor data to be processed in the cloud, DDNN locally processes most sensor data on end devices while achieving high accuracy and is able to reduce the communication cost by a factor of over 20x. Surat Teerapittayanon, Bradley McDanel, H. T. Kung 0001 |
ICDCS | 2 |
| 2017 | Incomplete Dot Products for Dynamic Computation Scaling in Neural Network InferenceabstractWe propose the use of incomplete dot products (IDP) to dynamically adjust the number of input channels used in each layer of a convolutional neural network during feedforward inference. IDP adds monotonically non-increasing coefficients, referred to as a “profile”, to the channels during training. The profile orders the contribution of each channel in non-increasing order. At inference time, the number of channels used can be dynamically adjusted to trade off accuracy for lowered power consumption and reduced latency by selecting only a beginning subset of channels. This approach allows for a single network to dynamically scale over a computation range, as opposed to training and deploying multiple networks to support different levels of computation scaling. Additionally, we extend the notion to multiple profiles, each optimized for some specific range of computation scaling. We present experiments on the computation and accuracy trade-offs of IDP for popular image classification models and datasets. We demonstrate that, for MNIST and CIFAR-10, IDP reduces computation significantly, e.g., by 75%, without significantly compromising accuracy. We argue that IDP provides a convenient and effective means for devices to lower computation costs dynamically to reflect the current computation budget of the system. For example, VGG-16 with 50% IDP (using only the first 50% of channels) achieves 70% in accuracy on the CIFAR-10 dataset compared to the standard network which achieves only 35% accuracy when using the reduced channel set. Bradley McDanel, Surat Teerapittayanon, H. T. Kung 0001 |
ICMLA | 1 |
| 2016 | BranchyNet: Fast inference via early exiting from deep neural networksabstractDeep neural networks are state of the art methods for many learning tasks due to their ability to extract increasingly better features at each network layer. However, the improved performance of additional layers in a deep network comes at the cost of added latency and energy usage in feedforward inference. As networks continue to get deeper and larger, these costs become more prohibitive for real-time and energy-sensitive applications. To address this issue, we present BranchyNet, a novel deep network architecture that is augmented with additional side branch classifiers. The architecture allows prediction results for a large portion of test samples to exit the network early via these branches when samples can already be inferred with high confidence. BranchyNet exploits the observation that features learned at an early layer of a network may often be sufficient for the classification of many data points. For more difficult samples, which are expected less frequently, BranchyNet will use further or all network layers to provide the best likelihood of correct prediction. We study the BranchyNet architecture using several well-known networks (LeNet, AlexNet, ResNet) and datasets (MNIST, CIFAR10) and show that it can both improve accuracy and significantly reduce the inference time of the network. Surat Teerapittayanon, Bradley McDanel, H. T. Kung 0001 |
ICPR | 2 |
| 2015 | Outlier detection for large scale manufacturing processesabstractIntegrated circuit manufacturing consists of tests at various stages to ensure functionality and performance using numerous test metrics for each system on chip (SoC) captured as part of assessment. At a later stage, functional units are evaluated in terms of multiple performance characteristics. In this paper, we propose a system that uses test metrics as features for machine learning models to predict the performance characteristics of each SoC. We show that these models are robust against erroneous or noisy signal in test metrics and provide accurate prediction. Given accurate models, we build a system that automatically detects systematic changes in the manufacturing process from week to week and identifies wafers, a grouping of patterned dies in the fabrication process, which have significantly higher than average prediction error and label them as outliers. These outliers are analyzed in order to determine the cause of the discrepancy and to assess potential problems in the manufacturing process. The system has been proven applicable across multiple products and process technologies. Abhinav Jauhri, Bradley McDanel, Chris Connor |
IEEE BigData | 2 |
| 2015 | Taming Wireless Fluctuations by Predictive Queuing Using a Sparse-Coding Link-State ModelabstractWe introduce State-Informed Link-Layer Queuing (SILQ), a system that models, predicts, and avoids packet delivery failures caused by temporary wireless outages in everyday scenarios. By stabilizing connections in adverse link conditions, SILQ boosts throughput and reduces performance variation for network applications, for example by preventing unnecessary TCP timeouts due to dead zones, elevators, and subway tunnels. SILQ makes predictions in real-time by actively probing links, matching measurements to an overcomplete dictionary of patterns learned offline, and classifying the resulting sparse feature vectors to identify those that precede outages. We use a clustering method called sparse coding to build our data-driven link model, and show that it produces more variation-tolerant predictions than traditional loss-rate, location-based, or Markov chain techniques. Stephen J. Tarsa, Marcus Z. Comiter, Michael B. Crouse, Bradley McDanel, H. T. Kung 0001 |
MobiHoc | 4 |