EDBT 2026 Demo / reviewers in the wild / expert
Richard Sun
dblp:78/4772
· DBLP profile ↗
9ranked-venue papers
0as first author
4since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ParSCo: Performance-Driven Partitioning and Scheduling Co-optimization Framework for Processor-based EmulationabstractAs the scale and complexity of designs increase, functional verification becomes a critical part of the very-large-scale integration (VLSI) design flow. However, existing processor-based emulation systems suffer from inefficiencies due to the misalignment objective between partitioning and scheduling, which are traditionally treated as separate and independent stages during compilation. To address this issue, we propose ParSCo , a partitioning and scheduling co-optimization framework that explicitly aligns the objectives of both stages by jointly considering cut minimization and topological order balancing (TOB) under multiple constraints. To integrate these objectives and constraints into our framework, we incorporate them into all partitioning and scheduling stages and further develop a set of novel techniques, including TOB-aware coarsening with multiple constraints , global growing initial partitioning with fixed nodes , TopoRefinement , and partitioning-aware scheduling , which collectively enhance the co-optimization process in emulation compilation. Furthermore, we establish theorems that reduce the time complexity of gain calculation and update to O (1), significantly improving the computational efficiency of the whole process. Furthermore, we evaluate the proposed method on the public and open-source chip design benchmarks, which have up to nearly 10 million cells. ParSCo significantly extends ideas and algorithms that first appeared in our previous work TopoOrderPart and achieves a 15% improvement. Extensive experimental results demonstrate the effectiveness of ParSCo , achieving an average improvement of 22.5% in time step reduction, 72% enhancement in TOB, and 55% acceleration in CPU time compared to the state-of-the-art (SOTA) two-stage partitioning and scheduling approach. Shunyang Bi, Hailong You, Cong Li 0023, Richard Sun |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2025 | OLTA: Optimizing bait seLection for TArgeted sequencingabstractMOTIVATION: Targeted enrichment via capture probes, also known as baits, is a promising complementary procedure for next-generation sequencing methods. This technique uses short biotinylated oligonucleotide probes that hybridize with complementary genetic material in a sample. Following hybridization, the target fragments can be easily isolated and processed with minimal contamination from irrelevant material. Designing an efficient set of baits for a set of target sequences, however, is an NP-hard problem. RESULTS: We develop a novel heuristic algorithm that leverages the similarities between the characteristics of the Minimum Bait Cover and the Closest String problems to reduce the number of baits to cover a given target sequence. Our results on real and synthetic datasets demonstrate that our algorithm, OLTA produces fewest baits for nearly all experimental settings and datasets. On average, it produces 6% and 11% fewer baits than the next best state-of-the-art methods for two major real datasets, AIV and MEGARES. Also, its bait set has the highest utilization and the minimum redundancy. AVAILABILITY AND IMPLEMENTATION: Our algorithm is available at github.com/FuelTheBurn/OLTA-Optimizing-bait-seLection-for-TArgeted-sequencing. Test data and other software are archived at doi.org/10.5281/zenodo.15086636. Mete Orhun Minbay, Richard Sun, Vijay Ramachandran, Ahmet Ay, Tamer Kahveci |
Bioinform. | 2 |
| 2024 | TopoOrderPart: a Multi-level Scheduling-Driven Partitioning Framework for Processor-Based EmulationabstractIn a compilation flow of processor-based emulation (PBE), partitioning involves dividing a large netlist into smaller pieces and assigning them to different processors. Furthermore, the scheduling process must adhere to the levels of the netlist, which are determined by topological ordering, and the logic gates in the same level can be emulated in parallel. However, during the netlist partitioning stage, assigning most gates at the same level to one processor would undermine the benefits of parallelization in scheduling, leading to overall performance degradation. This paper proposes the TopoOrderPart, the first scheduling-driven partitioning framework for simultaneous balancing topological order and minimizing the cut size, which holds significant value in reducing time steps of scheduling. In particular, the topological order balancing and cut size are considered throughout the multilevel paradigm, and balance-aware coarsening achieves balancing between clusters in the early stage, with super-far root growing initial partitioning obtaining the better partition by selecting those root nodes in distant relationship within the connection space and two novel TopoRefine algorithms further enhancing the solution. Experimental results show TopoOrderPart can improve 69% topological order balancing and 0.53× run time while maintaining comparable cut size, compared to the state-of-the-art partitioner. Shunyang Bi, Hailong You, Cong Li 0023, Richard Sun |
ICCAD | 6 |
| 2024 | MaPart: An Efficient Multi-FPGA System-Aware Hypergraph Partitioning FrameworkabstractMulti-FPGA systems (MFSs) are increasingly important in addressing VLSI circuit emulation and prototyping. However, the limitations of I/O resources between FPGAs have driven the usage of TDM and FPGA-hop technologies, which complicate the partitioning problem. Consequently, designing a suitable partitioning process for MFS has emerged as a critical research question affecting overall system performance. This paper proposes MaPart, a novel hypergraph partitioning framework, which aims to minimize the maximum path delay in MFS. MaPart combines binary search with a non-hop partitioner, TopoPart+, to minimize the maximum hop count during the partitioning process. Compared to previous non-hop partitioner, TopoPart+ provides enhanced problem-solving capabilities and achieves a remarkable 96% reduction in cut-size. Furthermore, the framework incorporates two successive local refinement algorithms that optimize the time-division multiplexing ratio, reduce total hop count, and alleviate congestion on critical paths. Additionally, MaPart includes a system-level router based on layered graphs, enabling flexible control of the hop count based on the timing criticality of each path. Experimental results demonstrate that the proposed framework achieves a significant 37% reduction in delay compared to baseline algorithms when evaluated using publicly available benchmarks. Benzheng Li, Shunyang Bi, Hailong You, Zhongdong Qi, Richard Sun |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2020 | Timing Driven Partition for Multi-FPGA Systems with TDM AwarenessabstractMulti-FPGA system is a popular approach to achieve hardware acceleration with the scalability to accommodate large designs. To overcome the connectivity constraint between each pair of FPGAs, Time-division multiplexing (TDM) is adopted with the expense of additional delay that dominates the performance on multi-FPGA system based emulator. To the best of our knowledge, there is no prior work on partitioning for multi-FPGA system considering hardware configuration and the impact of TDM. This work proposes a partition methodology to improve timing performance for multi-FPGA system. Delay introduced by TDM is estimated and optimized using look-up table for better efficiency. Our experimental result shows 43% improvement in maximum delay while considering both hardware configuration and impact of TDM compared with cut driven partition approach. Sin-Hong Liou, Sean S.-Y. Liu, Richard Sun, Hung-Ming Chen |
ISPD | 3 |
| 2019 | 2019 CAD Contest: System-level FPGA Routing with Timing Division Multiplexing TechniqueabstractThe time division multiplexing technique overcomes the bandwidth limitation by allowing FPGA chips to transmit multiple signals the maximum clocking frequency. With the additional multiplexers, this technique dramatically increases system-level routing capability in the FPGA-based emulator. However, the large number of virtual wires in the chip interconnection may impact emulation performance. The system-level FPGA routing tends to connect all virtual wires (signals) and considers emulation performance. At the same time, the challenge for system-level FPGA routing using time division multiplexing lies in the emulation performance. Yu-Hsuan Su, Richard Sun, Pei-Hsin Ho |
ICCAD | 2 |
| 2018 | Simultaneous partitioning and signals grouping for time-division multiplexing in 2.5D FPGA-based systemsabstractThe 2.5D FPGA is a promising technology to accommodate a large design in one FPGA chip, but the limited number of inter-die connections in a 2.5D FPGA may cause routing failures. To resolve the failures, input/output time-division multiplexing is adopted by grouping cross-die signals to go through one routing channel with a timing penalty after netlist partitioning. However, grouping signals after partitioning might lead to a suboptimal solution. Consequently, it is desirable to consider simultaneous partitioning and signal grouping although the optimization objectives of partitioning and grouping are different, and the time complexity of such simultaneous optimization is usually high. In this paper, we propose a simultaneous partitioning and grouping algorithm that can not only integrate the two objectives smoothly, but also reduce the time complexity to linear time per partitioning iteration. Experimental results show that our proposed algorithm outperforms the state-of-the-arts flow in both cross-die signal timing criticality and system-clock periods. Shih-Chun Chen, Richard Sun, Yao-Wen Chang |
ICCAD | 2 |
| 2018 | Challenges in Large FPGA-based Logic Emulation SystemsabstractFunctional verification is an important aspect of electronic design automation. Traditionally, simulation at the register transfer-level has been the mainstream functional verification approach. Formal verification and various static analysis checkers have been used to complement specific corners of logic simulation. However, as the size of IC designs grow exponentially, all the above approaches fail to scale with the design growth. In recent years, logic emulation have gained popularity in functional verification, partly due to their performance and scalability benefits. There are two main approaches to logic emulation: ASIC and commercial field-programmable gate array (FPGA). In this paper, we focus on commercial FPGA based logic emulation and present various challenging problems in this area for the academic community. William N. N. Hung, Richard Sun |
ISPD | 2 |
| 2018 | Pin Assignment Optimization for Multi-2.5D FPGA-based SystemsabstractAdvanced 2.5D FPGAs with larger logic capacity and higher pin counts compared to conventional FPGAs are commercially available. Some multi-FPGA systems have already utilized 2.5D FPGAs. Commercial 2.5D FPGA consists of multiple dies connected through an interposer. The interposer provides a fraction of the amount of interconnect resources with increased delay compared to that within individual dies. A recent study has shown the benefits of reducing signal crossings between dies on routability and timing when a circuit is mapped to a 2.5D FPGA. In a multi-2.5D FPGA system with multiplexed hardwired inter-FPGA connections, there can be tens of thousands of inter-FPGA signals incident with each FPGA and their pin assignment can greatly affect the amount of signal crossings between dies. In this paper, we formulate the pin assignment problem for such system with the objective of minimizing signal crossings between dies within the individual FPGAs. Taking into consideration of the multi-die structure of 2.5D FPGA, we propose an effective and efficient iterative improvement algorithm based on integer linear programming to the pin assignment problem. Experimental results show that our algorithm can reduce signal crossings between dies in the individual FPGAs by over 30% on average compared to two heuristic approaches. Wan-Sin Kuo, Shi-Han Zhang, Wai-Kei Mak, Richard Sun, Yoon Kah Leow |
ISPD | 4 |