Jackson Melchert

dblp:210/1782 · DBLP profile ↗
← Back
14ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0002-8232-1603ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 4 first-author · 6 since 2021Software engineering, systems software and programming languages · 4 · 2 first-author · 4 since 2021Theory of computation · 3 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 PEak: A Single Source of Truth for Hardware Design and Verification
abstract
Domain-specific languages for hardware can significantly enhance designer productivity, but sometimes at the cost of ease of verification. On the other hand, ISA specification languages are too static to be used during early stage design space exploration. We present PEak, an open-source hardware design and specification language, which aims at improving both design productivity and verification capability. PEak does this by providing a single source of truth for functional models, formal specifications, and RTL. PEak has been used in several academic projects, and PEak-generated RTL has been included in three fabricated hardware accelerators. In these projects, the formal capabilities of PEak were crucial for enabling both novel design space exploration techniques and automated compiler synthesis.
Caleb Donovick, Jackson Melchert, Ross Daly, Leonard Truong, Priyanka Raina, Pat Hanrahan, Clark W. Barrett
ACM Trans. Embed. Comput. Syst.2
2025 Automated Translation Validation of a Compiler for Statically Scheduled Accelerators
Jackson Melchert, Caleb Terrill, Aron Ricardo Perez-Lopez, Clark W. Barrett, Priyanka Raina
FMCAD1
2024 Efficiently Synthesizing Lowest Cost Rewrite Rules for Instruction Selection
Ross Daly, Caleb Donovick, Caleb Terrill, Jackson Melchert, Priyanka Raina, Clark W. Barrett, Pat Hanrahan
FMCAD4
2024 Onyx: A Programmable Accelerator for Sparse Tensor Algebra
abstract
•Applications ranging from scientific computing to machine learning can have extremely sparse inputs
Kalhan Koul, Maxwell Strange, Jackson Melchert, Alex Carsello, Yuchen Mei, Olivia Hsu, Taeyoung Kong, Huifeng Ke, Keyi Zhang, Qiaoyi Liu, Gedeon Nyengele, Akhilesh Balasingam, Jayashree Adivarahan, Ritvik Sharma, Zhouhua Xie, Christopher Torng, Joel S. Emer, Fredrik Kjolstad, Mark Horowitz, Priyanka Raina
HCS3
2024 Cascade: An Application Pipelining Toolkit for Coarse-Grained Reconfigurable Arrays
abstract
While coarse-grained reconfigurable arrays (CGRAs) have emerged as promising programmable accelerator architectures, they require automatic pipelining of applications during their compilation flow to achieve high performance. Current CGRA compilers either lack pipelining altogether resulting in low application performance, or perform exhaustive pipelining resulting in high power and resource consumption. We address these challenges by proposing Cascade, an end-to-end open-source application compiler for CGRAs that achieves both state-of-the-art performance and fast compilation times. The contributions of this work are: (1) a novel post place-and-route (PnR) application pipelining technique for CGRAs that accounts for interconnect hop delays during pipelining but in a unique way that avoids cyclic scheduling and place-and-route, (2) a register resource usage optimization technique that leverages the scheduling logic in CGRA memory tiles to minimize the number of register resources used during pipelining, and (3) an automated CGRA timing model generator, an application timing analysis tool, and a large set of existing and novel application pipelining techniques integrated into an end-to-end compilation flow. Cascade achieves 8 -34× lower critical path delay and 7 -190× lower energy-delay product (EDP) across a variety of dense image processing and machine learning workloads, and 3 -5.2× lower critical path delay and 2.5 -5.2× lower EDP on sparse workloads, compared to a compiler without pipelining. Cascade mitigates the performance and energy-efficiency drawbacks of existing CGRA compilers, and enables further research into CGRAs as flexible, yet competitive accelerator architectures.
Jackson Melchert, Yuchen Mei, Kalhan Koul, Qiaoyi Liu, Mark Horowitz, Priyanka Raina
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 APEX: A Framework for Automated Processing Element Design Space Exploration using Frequent Subgraph Analysis
abstract
The architecture of a coarse-grained reconfigurable array (CGRA) processing element (PE) has a significant effect on the performance and energy-efficiency of an application running on the CGRA. This paper presents APEX, an automated approach for generating specialized PE architectures for an application or an application domain. APEX first analyzes application domain benchmarks using frequent subgraph mining to extract commonly occurring computational subgraphs. APEX then generates specialized PEs by merging subgraphs using a datapath graph merging algorithm. The merged datapath graphs are translated into a PE specification from which we automatically generate the PE hardware description in Verilog along with a compiler that maps applications to the PE. The PE hardware and compiler are inserted into a flexible CGRA generation and compilation toolchain that allows for agile evaluation of CGRAs. We evaluate APEX for two domains, machine learning and image processing. For image processing applications, our automatically generated CGRAs with specialized PEs achieve from 5% to 30% less area and from 22% to 46% less energy compared to a general-purpose CGRA. For machine learning applications, our automatically generated CGRAs consume 16% to 59% less energy and 22% to 39% less area than a general-purpose CGRA. This work paves the way for creation of application domain-driven design-space exploration frameworks that automatically generate efficient programmable accelerators, with a much lower design effort for both hardware and compiler generation.
Jackson Melchert, Kathleen Feng, Caleb Donovick, Ross Daly, Ritvik Sharma, Clark W. Barrett, Mark Horowitz, Pat Hanrahan, Priyanka Raina
ASPLOS (3)1
2023 AHA: An Agile Approach to the Design of Coarse-Grained Reconfigurable Accelerators and Compilers
abstract
With the slowing of Moore’s law, computer architects have turned to domain-specific hardware specialization to continue improving the performance and efficiency of computing systems. However, specialization typically entails significant modifications to the software stack to properly leverage the updated hardware. The lack of a structured approach for updating the compiler and the accelerator in tandem has impeded many attempts to systematize this procedure. We propose a new approach to enable flexible and evolvable domain-specific hardware specialization based on coarse-grained reconfigurable arrays (CGRAs). Our agile methodology employs a combination of new programming languages and formal methods to automatically generate the accelerator hardware and its compiler from a single source of truth. This enables the creation of design-space exploration frameworks that automatically generate accelerator architectures that approach the efficiencies of hand-designed accelerators, with a significantly lower design effort for both hardware and compiler generation. Our current system accelerates dense linear algebra applications but is modular and can be extended to support other domains. Our methodology has the potential to significantly improve the productivity of hardware-software engineering teams and enable quicker customization and deployment of complex accelerator-rich computing systems.
Kalhan Koul, Jackson Melchert, Kavya Sreedhar, Leonard Truong, Gedeon Nyengele, Keyi Zhang, Qiaoyi Liu, Jeff Setter, Yuchen Mei, Maxwell Strange, Ross Daly, Caleb Donovick, Alex Carsello, Taeyoung Kong, Kathleen Feng, Dillon Huff, Ankita Nayak, Rajsekhar Setaluri, James Thomas 0003, Nikhil Bhagdikar, David Durst, Zachary A. Myers, Nestan Tsiskaridze, Stephen Richardson, Rick Bahr, Kayvon Fatahalian, Pat Hanrahan, Clark W. Barrett, Mark Horowitz, Christopher Torng, Fredrik Kjolstad, Priyanka Raina
ACM Trans. Embed. Comput. Syst.2
2022 Synthesizing Instruction Selection Rewrite Rules from RTL using SMT
Ross Daly, Caleb Donovick, Jackson Melchert, Rajsekhar Setaluri, Nestan Tsiskaridze, Priyanka Raina, Clark W. Barrett, Pat Hanrahan
FMCAD3
2022 Amber: Coarse-Grained Reconfigurable Array-Based SoC for Dense Linear Algebra Acceleration
abstract
Dedicated hardware accelerators popular for imaging, vision, and machine learning (ML) applications
Kathleen Feng, Alex Carsello, Taeyoung Kong, Kalhan Koul, Qiaoyi Liu, Jackson Melchert, Gedeon Nyengele, Maxwell Strange, Keyi Zhang, Ankita Nayak, Jeff Setter, James Thomas 0003, Kavya Sreedhar, Nikhil Bhagdikar, Zachary A. Myers, Brandon D'Agostino, Pranil Joshi, Stephen Richardson, Rick Bahr, Christopher Torng, Mark Horowitz, Priyanka Raina
HCS6
2020 Creating an Agile Hardware Design Flow
abstract
Although an agile approach is standard for software design, how to properly adapt this method to hardware is still an open question. This work addresses this question while building a system on chip (SoC) with specialized accelerators. Rather than using a traditional waterfall design flow, which starts by studying the application to be accelerated, we begin by constructing a complete flow from an application expressed in a high-level domain-specific language (DSL), in our case Halide, to a generic coarse-grained reconfigurable array (CGRA). As our under-standing of the application grows, the CGRA design evolves, and we have developed a suite of tools that tune application code, the compiler, and the CGRA to increase the efficiency of the resulting implementation. To meet our continued need to update parts of the system while maintaining the end-to-end flow, we have created DSL-based hardware generators that not only provide the Verilog needed for the implementation of the CGRA, but also create the collateral that the compiler/mapper/place and route system needs to configure its operation. This work provides a systematic approach for desiging and evolving high-performance and energy-efficient hardware-software systems for any application domain.
Rick Bahr, Clark W. Barrett, Nikhil Bhagdikar, Alex Carsello, Ross Daly, Caleb Donovick, David Durst, Kayvon Fatahalian, Kathleen Feng, Pat Hanrahan, Teguh Hofstee, Mark Horowitz, Dillon Huff, Fredrik Kjolstad, Taeyoung Kong, Qiaoyi Liu, Makai Mann, Jackson Melchert, Ankita Nayak, Aina Niemetz, Gedeon Nyengele, Priyanka Raina, Stephen Richardson, Rajsekhar Setaluri, Jeff Setter, Kavya Sreedhar, Maxwell Strange, James Thomas 0003, Christopher Torng, Leonard Truong, Nestan Tsiskaridze, Keyi Zhang
DAC18
2019 SAADI: a scalable accuracy approximate divider for dynamic energy-quality scaling
abstract
Approximate computing can significantly improve the energy efficiency of arithmetic operations in error-resilient applications. In this paper, we propose an approximate divider design that facilitates dynamic energy-quality scaling. Conventional approximate dividers lack runtime energy-quality scalability, which is the key to maximizing the energy efficiency while meeting dynamically varying accuracy requirements. Our divider design, named SAADI, makes an approximation to the reciprocal of the divisor in an incremental manner, thus the division speed and energy efficiency can be dynamically traded for accuracy by controlling the number of iterations. For the approximate 8-bit division of 32-bit/16-bit division, the average accuracy of SAADI can be adjusted in between 92.5% and 99.0% by varying latency up to 7x. We evaluate the accuracy and energy consumption of SAADI for various design parameters and demonstrate its efficacy for low-power signal processing applications.
Setareh Behroozi, Jackson Melchert, Younghyun Kim 0001
ASP-DAC3
2019 SAADI-EC: A Quality-Configurable Approximate Divider for Energy Efficiency
abstract
Energy efficiency is one of the most crucial constraints that dictate performance, lifetime, form factor, and cost in modern computing system design. However, the energy efficiency improvement driven by semiconductor technology scaling is coming to an end with the prediction of the end of Moore's law in the near future. Approximate computing is a new paradigm to accomplish energy-efficient computing in this twilight of Moore's law by relaxing exactness requirement of computation results for intrinsically error-resilient applications, such as some machine learning and signal processing. In this paper, we propose an approximate binary divider design that features dynamic configurability of accuracy and energy consumption. Conventional approximate binary dividers lack runtime energy-quality scalability, which is the key to maximizing energy efficiency while meeting the dynamically varying accuracy requirements of the application. Our divider, named Scalable Accuracy Approximate Divider with Error Compensation (SAADI-EC), supports dynamic energy-quality scalability by the incremental approximation of the reciprocal of the divisor using Taylor series expansion. As a result, the speed and energy efficiency of division can be dynamically traded for accuracy by controlling the number of iterations for the approximation. In addition, SAADI-EC corrects the approximation error using simple, yet effective error compensation hardware to greatly improve the accuracy compared to the base implementation, SAADI. For the 8-bit approximation of 32-bit/16-bit division, the average accuracy of SAADI-EC can be adjusted from 94.2% to 99.6% by varying latency and energy 7×. In terms of energy × delay cost, our design costs up to 87% less than other approximate binary dividers for the same accuracy level. We evaluate the accuracy and energy consumption of SAADI-EC for various design parameters and demonstrate its efficacy for low-power signal processing applications including k-means color quantization, JPEG image compression, and image division for video sequences.
Jackson Melchert, Setareh Behroozi, Younghyun Kim 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2018 A Comparative Study of Local Net Modeling Using Machine Learning
abstract
Local nets are by default ignored during global routing but can contribute to a high percentage (up to 30%) of total number of nets in the design. Prior work proposed simple models for how local nets are routed and showed benefits such as better congestion analysis post-placement, integration with global routing, and better track assignment. In this work we study local net modeling using machine learning. Our model predicts utilization by local routes inside each global cell. We model this as a regression problem and as reference use local route utilization data from the detailed routing stage using a commercial tool. To solve the problem we identify suitable machine learning algorithms. Within our modeling, we study and rank different features which utilize various layout attributes. We identify the most beneficial features and show our model performs superior to prior work which were based on pin density and Steiner tree models. Our model also performs better for the subset of local nets which are routed in more than one global cell.
Jackson Melchert, Boyu Zhang 0001, Azadeh Davoodi
ACM Great Lakes Symposium on VLSI1
2017 Are Proximity Attacks a Threat to the Security of Split Manufacturing of Integrated Circuits?
abstract
Split manufacturing is a technique that allows manufacturing the transistor-level and lower metal layers of an integrated circuit (IC) at a high-end, untrusted foundry, while manufacturing only the higher metal layers at a smaller, trusted foundry. Using split manufacturing is only viable if the untrusted foundry cannot reverse engineer the higher metal layer connections (and thus the overall IC design) from the lower layers. This paper studies the effectiveness of proximity attack as a key step to reverse engineer a design at the untrusted foundry. We propose and study different proximity attacks based on how a set of candidates are defined for each broken connection. The attacks use both placement and routing information along with factors which capture the router's behavior such as per-layer routing congestion. Our studies are based on designs having millions of nets routed across nine metal layers and significant layer-by-layer wire size variation. Our results show that a common, Hamming distance-based proximity attack seldom achieves a match rate over 5%. But our proposed attack yields a relatively small list of candidates which often contains the correct match. Finally, we propose a procedure to artificially insert routing blockages in a design at a desired split level, without causing any area overhead, in order to trick the router to make proximity-based reverse engineering significantly more challenging.
Jonathon Magaña, Daohang Shi, Jackson Melchert, Azadeh Davoodi
IEEE Trans. Very Large Scale Integr. Syst.3