VLDB 2026 Research / reviewers in the wild / expert
Tzung-Han Juang
dblp:246/9455
· DBLP profile ↗
9ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0003-1629-8617ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SkeleShare: Algorithmic Skeletons and Equality Saturation for Hardware Resource SharingabstractCompiling functional programs into efficient Field Programmable Gate Array (FPGA) designs is difficult. Hardware resources must be explicitly allocated and shared to maximize resource efficiency. This requires careful orchestration of several transformations to expose and exploit sharing opportunities.This paper introduces SkeleShare, a novel approach that automates the problem of resource allocation and sharing. It leverages equality saturation and algorithmic skeletons to expose sharing opportunities across abstraction levels. A solver-based extractor then selects a design that consolidates computations, meeting resource constraints while maintaining performance.This approach is evaluated on neural networks and image processing targeting a real FPGA. The paper shows how SkeleShare is used to express the various algorithmic patterns and transformation rules inherent in neural network operators. The experimental evaluation demonstrates that SkeleShare’s fully automated resource allocation and sharing matches and exceeds the performance of prior work, which involves expert manual extraction of sharing opportunities. Jonathan Van der Cruysse, Tzung-Han Juang, Shakiba Bolbolian Khah, Christophe Dubach |
CGO | 2 |
| 2026 | A Functional Approach to Synthesizing Routable Programmable Accelerators for Neural NetworksabstractProducing optimized accelerators is tedious, as even modern HDLs (Hardware Description Languages) such as Chisel, require reasoning about low-level concepts. Recent functional approaches, such as Aetherling and SHIR, treat hardware as composition of pure operators. This raises the abstraction level, allowing for systematic optimizations through rewriterules for FPGAs (Field Programmable Gate Arrays). Tzung-Han Juang, Paul Teng, Christophe Dubach |
LCTES | 1 |
| 2025 | Scalable multi-agent reinforcement learning for factory-wide dynamic scheduling in semiconductor manufacturingabstractReal-time dynamic scheduling in modern manufacturing is highly complex due to frequent disturbances and intricate operational constraints. While reinforcement learning (RL) has shown promise, existing approaches often rely on extensive dispatching rules and struggle to scale to factory-wide settings. This study introduces a scalable multi-agent RL (MARL) framework with a leader–follower structure, enabling decentralized agents to handle sub-problems while maintaining global coordination through abstract goals. To further enhance robustness, a limited rule-based conversion algorithm is proposed to mitigate performance degradation from poor agent decisions. Our experimental results demonstrate that the proposed model outperforms the state-of-the-art deep RL-based scheduling methods in various aspects. Additionally, the proposed model provides the most robust scheduling performance to demand changes. Overall, the proposed MARL-based scheduling model presents a promising solution to the real-time scheduling problem, with potential applications in various manufacturing industries. Jaeyeon Jang, Diego Klabjan, Han Liu 0001, Nital S. Patel, Xiuqi Li, Balakrishnan Ananthanarayanan, Husam Dauod, Tzung-Han Juang |
Eng. Appl. Artif. Intell. | 8 |
| 2025 | Learning multiple coordinated agents under directed acyclic graph constraints
Jaeyeon Jang, Diego Klabjan, Han Liu 0001, Nital S. Patel, Xiuqi Li, Balakrishnan Ananthanarayanan, Husam Dauod, Tzung-Han Juang |
Expert Syst. Appl. | 8 |
| 2025 | Maximizing Data and Hardware Reuse for HLS with Early-Stage Symbolic PartitioningabstractWhile traditional High-Level Synthesis (HLS) converts “high-level” C-like programs into hardware automatically, producing high-performance designs still requires hardware expertise. Optimizations such as data partitioning can have a large impact on performance since they directly affect data reuse patterns and the ability to reuse hardware. However, optimizing partitioning is a difficult process since minor changes in the parameter choices can lead to totally unpredictable performance. Functional array-based languages have been proposed instead of C-based approaches, as they offer stronger performance guarantees. This article proposes to follow a similar approach and exposes a divide-and-conquer primitive at the algorithmic level to let users partition any arbitrary computation. The compiler is then free to explore different partition shapes to maximize both data and hardware reuse automatically. The main challenge remains that the impact of partitioning is only known much later in the compilation flow. This is due to the hard-to-predict effects of the many optimizations applied during compilation. To solve this problem, the partitioning is expressed using a set of symbolic tunable parameters, introduced early in the compilation pipeline. A symbolic performance model is then used in the last compilation stage to predict performance based on the possible values of the tunable parameters. Using this approach, a design space exploration is conducted on an Intel Arria 10 Field Programmable Gate Arrays (FPGAs), and competitive performance is achieved on the classical VGG and TinyYolo neural networks. Tzung-Han Juang, Christophe Dubach |
ACM Trans. Archit. Code Optim. | 1 |
| 2023 | Let Coarse-Grained Resources Be Shared: Mapping Entire Neural Networks on FPGAsabstractTraditional High-Level Synthesis (HLS) provides rapid prototyping of hardware accelerators without coding with Hardware Description Languages (HDLs). However, such an approach does not well support allocating large applications like entire deep neural networks on a single Field Programmable Gate Array (FPGA) device. The approach leads to designs that are inefficient or do not fit into FPGAs due to resource constraints. This work proposes to shrink generated designs by coarse-grained resource control based on function sharing in functional Intermediate Representations (IRs). The proposed compiler passes and rewrite system aim at producing valid design points and removing redundant hardware. Such optimizations make fitting entire neural networks on FPGAs feasible and produce competitive performance compared to running specialized kernels for each layer. Tzung-Han Juang, Christof Schlaak, Christophe Dubach |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2022 | Optimizing data reshaping operations in functional IRs for high-level synthesisabstractFPGAs (Field Programmable Gate Arrays) have become the substrate of choice to implement accelerators. They deliver high performance with low power consumption, while offering the flexibility of being re-programmable. But they are notoriously hard to program directly using HDLs (Hardware Description Languages). Traditional HLS (High-Level Synthesis) methods are addressing some of these issues but are far from being perfect. Programmers are still required to write hardware-specific code and existing HLS tools often produce sub-optimal designs. Christof Schlaak, Tzung-Han Juang, Christophe Dubach |
LCTES | 2 |
| 2022 | Memory-Aware Functional IR for Higher-Level Synthesis of AcceleratorsabstractSpecialized accelerators deliver orders of a magnitude of higher performance than general-purpose processors. The ever-changing nature of modern workloads is pushing the adoption of Field Programmable Gate Arrays (FPGAs) as the substrate of choice. However, FPGAs are hard to program directly using Hardware Description Languages (HDLs). Even modern high-level HDLs, e.g., Spatial and Chisel, still require hardware expertise. This article adopts functional programming concepts to provide a hardware-agnostic higher-level programming abstraction. During synthesis, these abstractions are mechanically lowered into a functional Intermediate Representation (IR) that defines a specific hardware design point. This novel IR expresses different forms of parallelism and standard memory features such as asynchronous off-chip memories or synchronous on-chip buffers. Exposing such features at the IR level is essential for achieving high performance. The viability of this approach is demonstrated on two stencil computations and by exploring the optimization space of matrix-matrix multiplication. Starting from a high-level representation for these algorithms, our compiler produces low-level VHSIC Hardware Description Language (VHDL) code automatically. Several design points are evaluated on an Intel Arria 10 FPGA, demonstrating the ability of the IR to exploit different hardware features. This article also shows that the designs produced are competitive with highly tuned OpenCL implementations and outperform hardware-agnostic OpenCL code. Christof Schlaak, Tzung-Han Juang, Christophe Dubach |
ACM Trans. Archit. Code Optim. | 2 |
| 2020 | Collision-free Navigation of Human-centered Robots via Markov GamesabstractWe exploit Markov games as a framework for collision-free navigation of human-centered robots. Unlike the classical methods which formulate robot navigation as a single-agent Markov decision process with a static environment, our framework of Markov games adopts a multi-agent formulation with one primary agent representing the robot and the remaining auxiliary agents form a dynamic or even competing environment. Such a framework allows us to develop a path-following type adversarial training strategy to learn a robust decentralized collision avoidance policy. Through thorough experiments on both simulated and real-world mobile robots, we show that the learnt policy outperforms the state-of-the-art algorithms in both sample complexity and runtime robustness. Guo Ye, Qinjie Lin, Tzung-Han Juang, Han Liu 0001 |
ICRA | 3 |