EDBT 2026 Demo / reviewers in the wild / expert
Daniel H. Noronha
dblp:214/1044 · also Daniel Holanda Noronha
· DBLP profile ↗
8ranked-venue papers
3as first author
4since 2021 · last 2022
0000-0003-1043-0920ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 first-author · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Boosting Domain-Specific Debug Through Inter-frame CompressionabstractAcceleration of machine learning models is proving to be an important application for FPGAs. Unfortunately, debugging such models during training or inference is difficult. Software simulations of a machine learning system may be of insufficient detail to provide meaningful debug insight, or may require infeasibly long run-times. Thus, it is often desirable to debug the accelerated model while it is running on real hardware. Effective on-chip debug often requires instrumenting a design with additional circuitry to store run-time data, consuming valuable chip resources. Previous work has developed methods to perform lossy compression of signals by exploiting machine learning specific knowledge, thereby increasing the amount of debug context that can be stored in an on-chip trace buffer. However, all prior work compresses each successive element in a signal of interest independently. Since debug signals may have temporal similarity in many machine learning applications there is an opportunity to further increase trace buffer utilization. In this paper, we present an architecture to perform lossless temporal compression in addition to the existing lossy element-wise compression. We show that, when applied to a typical machine learning algorithm in realistic debug scenarios, we are able to store twice as much information in an on-chip buffer while increasing the total area of the debug instrument by approximately 25%. The impact is that, for a given instrumentation budget, a significantly larger trace window is available during debug, possibly allowing a designer to narrow down the root cause of a bug faster. Zakary Nafziger, Martin Chua, Daniel H. Noronha, Steve Wilton |
FPT | 3 |
| 2022 | Adaptive Clock Management of HLS-generated Circuits on FPGAsabstractIn this article, we present Syncopation , a performance-boosting fine-grained timing analysis and adaptive clock management technique for High-Level Synthesis-generated circuits implemented on Field-Programmable Gate Arrays. The key idea is to use the HLS scheduling information along with the placement and routing results to determine the worst-case timing path for individual clock cycles. By adjusting the clock period on a cycle-by-cycle basis, we can increase performance of an HLS-generated circuit. Our experiments show that Syncopation improves performance by 3.2% (geomean) across all benchmarks (up to 47%). In addition, by employing targeted synthesis techniques along with Syncopation, we can achieve 10.3% performance improvement (geomean) across all benchmarks (up to 50%). Syncopation instrumentation is implemented entirely in soft logic without requiring alterations to the HLS-synthesis toolchain or changes to the FPGA, and has been validated on real hardware. Kahlan Gibson, Esther Roorda, Daniel H. Noronha, Steve Wilton |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2021 | Flexible Instrumentation for Live On-Chip Debug of Machine Learning Training on FPGAsabstractFPGAs have recently shown promise for accelerating machine learning training. This has led to research into the co-design of narrow-precision accelerator architectures and the investigation of novel machine learning models. Such research can be extremely expensive, as the steep cost of training a model can increase several-fold due to the need of performing hyper-parameter tuning and adjustments to the model to ensure acceptable convergence speed and accuracy. In this scenario, monitoring key data on-chip is essential to more quickly understand and diagnose problems, significantly reducing training costs.Previous work has proposed on-chip debug instrumentation to monitor key signals for both general-purpose circuits and inference algorithms. This instrumentation either performs limited on-chip compression, or is extremely restricted in the amount of run-time customization that may occur. We argue that for training applications, the extremely long and expensive training runs warrant significantly more flexibility in the on-chip instrumentation, even at the expense of some chip area.In this paper, we propose flexible debug instrumentation that allows for the live debugging of machine learning systems during training. Different from previous debug instrumentation, our instrumentation offers firmware programmability, allowing the researcher to gather data in a large variety of ways that would likely not be anticipated at compile time. Daniel H. Noronha, Zhiqiang Que, Wayne Luk, Steve Wilton |
FCCM | 1 |
| 2021 | In-circuit tuning of deep learning designs
Zhiqiang Que, Daniel H. Noronha, Ruizhe Zhao, Xinyu Niu, Steve Wilton, Wayne Luk |
J. Syst. Archit. | 2 |
| 2020 | Syncopation: Adaptive Clock Management for High-Level Synthesis Generated Circuits on FPGAsabstractHigh-level synthesis (HLS) tools improve hardware designer productivity by enabling software design techniques during hardware development. During HLS the delay of paths can only be estimated, so the resulting circuit may suffer from unbalanced computational path delays across clock cycles. Since the maximum operating frequency of circuits is determined statically using the worst-case timing path, unbalanced paths may lead to reduced performance compared to circuits designed at the hardware level. In this paper, we address this using Syncopation, a performance-boosting fine-grained timing analysis and adaptive clock management technique for HLS circuits. The key idea is to use the HLS scheduling information along with the results from placement and routing to determine the worst-case timing path for individual clock cycles. By then adjusting the clock period on a cycle-to-cycle basis, we can increase circuit performance. Our experiments show that Syncopation and fine-grained timing analysis can improve performance without altering the HLS-synthesis toolchain. Kahlan Gibson, Esther Roorda, Daniel H. Noronha, Steve Wilton |
FPL | 3 |
| 2019 | On-chip FPGA Debug Instrumentation for Machine Learning ApplicationsabstractFPGAs provide a promising implementation option for many machine learning applications. Although simulations or software models can be used to explore the design space of these applications, often the final behaviour can not be evaluated until the design is mapped to the FPGA and integrated into the target system. This may be because long run-times are required, or because the environment can not be adequately described using a software model. Once unexpected behaviour is observed, on-chip debug is notoriously difficult; typically a design is instrumented with on-chip trace buffers that record the run-time behaviour for later interrogation. In this paper, we describe instrumentation that can accelerate the process of debugging machine learning applications implemented on an FPGA. Unlike previous work, our instrumentation is optimized to take advantage of characteristics of this application domain. Our instruments gather useful domain-specific information about the observed variables instead of recording the raw values of those elements. Results show that the proposed instruments provide at least 17.8x longer visibility in the most conservative of our experiments at a low area and latency cost. Daniel H. Noronha, Ruizhe Zhao, Jeffrey B. Goeders, Wayne Luk, Steve Wilton |
FPGA | 1 |
| 2019 | Towards In-Circuit Tuning of Deep Learning DesignsabstractThis paper presents InTune, a novel approach for in-circuit tuning of deep learning designs targeting implementations in field-programmable gate array technology. This approach combines two promising techniques: domain-specific adaptation and in-circuit tuning. Domain-specific adaptation exploits domain-specific information in adapting pre-trained models to specific application domains, replacing standard convolution layers with efficient convolution blocks; the effects of such adaptation are then assessed by in-circuit tuning instruments to provide information to application builders for tuning the design. This approach is illustrated by its deployment in tuning deep neural networks, and its potential for a new generation of domain-specific tools with tight integration of synthesis and in-circuit tuning is explored. Zhiqiang Que, Daniel H. Noronha, Ruizhe Zhao, Steve Wilton, Wayne Luk |
ICCAD | 2 |
| 2018 | LeFlow: Automatic Compilation of TensorFlow Machine Learning Applications to FPGAsabstractAcceleration of Machine Learning applications on Field-Programmable Gate Arrays (FPGAs) has shown to have advantages over other computing platforms in recent work. However, since machine learning code is often specified in a high-level software language such as Python, the manual translation of the algorithm to either C code for high-level synthesis or to Register Transfer Level (RTL) code for synthesis is time consuming and requires the designer to have expertise in designing hardware. In order to show how we can make FPGAs more accessible to software developers, we present a demonstration of LeFlow: an open-source tool which maps numerical computation models written in TensorFlow to synthesizable RTL. This demonstration includes two examples which begin with a model written in TensorFlow and show how a designer would use the LeFlow tool to generate Verilog, simulate the result, and synthesize the design to target FPGAs. Daniel H. Noronha, Kahlan Gibson, Bahar Salehpour, Steve Wilton |
FPT | 1 |