EDBT 2026 Demo / reviewers in the wild / expert
Tianji Liu
dblp:327/1768
· DBLP profile ↗
11ranked-venue papers
7as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 6 first-author · 10 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient and Effective E-graph-based Logic OptimizationabstractRecent efforts of applying e-graphs in logic synthesis have shown promising results. Nevertheless, e-graph-based gate-level logic optimization suffers from inefficiency and limited extraction quality. In this article, we propose a fast parallel e-matching algorithm for speeding up e-graph rewriting, and an efficient netlist extraction framework with high quality of results in both area and delay. Experiments show that e-graph rewriting can be accelerated by up to 8.3× over a high-performance e-graph library, and our extraction framework achieves 11.0% and 1.0% improvements in size and level on average, compared to the best results of the state-of-the-art netlist extraction method. Tianji Liu, Nutdranai Jaruthikorn, Shiju Lin, Bentian Jiang, Guannan Guo, Weihua Sheng, Evangeline F. Y. Young |
ACM Trans. Design Autom. Electr. Syst. | 1 |
| 2025 | Simulation-based Parallel Sweeping: A New Perspective on Combinational Equivalence CheckingabstractCombinational equivalence checking (CEC) is a fundamental task in the realization of digital designs which is unlikely to have universally efficient algorithms due to its co-NP-completeness. Recent researches of CEC have been focusing on SAT sweeping. This paper provides a new perspective other than SAT for tackling CEC, namely exhaustive simulation, and presents a simulation-based CEC engine constructed with fast GPU-parallel algorithms. The proposed engine can solve 4 out of the 9 large cases in the experiments on its own, with up to $88.11 \times$ speed-up compared with the checker in ABC. Moreover, a combination of the proposed engine with the ABC checker achieves averaged accelerations of $4.89 \times$ and $4.88 \times$ over the standalone ABC checker and a commercial checker, respectively. Tianji Liu, Evangeline F. Y. Young |
DAC | 1 |
| 2025 | ExactMap: Enhancing Delay Optimization in Parallel ASIC Technology MappingabstractASIC technology mapping consists of mapping a technology-independent Boolean network into an equivalent circuit utilizing cells from a specified library, a process that is vital in electronic design automation (EDA). However, existing algorithms in the literature are sequential in nature and often neglect to account for the actual delay of the cells during the mapping process, resulting in significant discrepancy between estimated and actual delay. In this paper, we propose an ASIC technology mapper that considers load information of intermediate solutions. This approach enables the identification of critical nodes and a better selection of delay-oriented cells for those nodes. Furthermore, we introduce a dual recovery method for delay and area to enhance performance. Finally, these innovations are integrated within a configuration-level parallelism framework on GPU. Experimental results on different technology libraries demonstrate that our method achieves on average 32% reduction in delay with 3% area penalty compared to the public synthesis tool ABC. Additionally, our approach provides a significant speedup of 65.33×. Zhenxuan Xie, Tianji Liu, Evangeline F. Y. Young |
ICCAD | 3 |
| 2025 | A Unified Parallel Framework for LUT Mapping and Logic OptimizationabstractLookup-table (LUT) mapping has been extensively utilized in logic synthesis, including being an indispensable step in FPGA design, serving as a building block in high-effort synthesis flows, and providing an algorithmic framework for logic optimization. Hence, a fast mapping algorithm is vital to satisfying the demand for synthesizing high-quality, large-scale modern VLSI designs. This article proposes two efficient GPU-parallel algorithms, namely LUT mapping and and-inverter graph (AIG) optimization using a precomputed database, which rely on a common parallel mapping framework that consists of novel fine-grained parallel mapping passes with high degree of parallelism. The mapping pass is enhanced by specifically tailored cut evaluation and memory management methods for GPUs that enable fast mapping of large circuits with limited GPU memory. Parallel timing analysis passes and parallel cut expansion passes are also proposed for constructing a fully GPU-accelerated LUT mapping flow. The core of parallel AIG optimization is a plugin of the mapping framework, which contains a self-adaptive parallel candidate structure evaluation procedure with high time efficiency and low hardware resource usage. Experiments show that on average, GPU LUT mapping and AIG optimization achieve$34.6\times $and$99.9\times $speedup with similar result quality, compared with the high-performance LUT mapper and AIG optimization algorithm with a database implemented in ABC, respectively, on large benchmarks. When combining the two algorithms with other GPU logic optimization algorithms, a GPU-based sequence targeting LUT network synthesis achieves$46.7\times $speedup with 4.7% smaller area and 0.2% smaller delay over ABC. Tianji Liu, Lei Chen 0031, Xing Li 0023, Mingxuan Yuan, Evangeline F. Y. Young |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | FineMap: A Fine-grained GPU-parallel LUT Mapping EngineabstractLookup-table (LUT) mapping is an indispensable step in FPGA design flows, and also serves as a building block in many technology-independent optimization algorithms. Therefore, it is crucial to accelerate LUT mapping in order to satisfy the demand for synthesizing high-quality, large-scale VLSI designs. Previous work on GPU LUT mapping suffers from low speedup due to limited degree of parallelism. In this paper, we propose an ultra-fast GPU-parallel LUT mapping engine named FineMap, which is composed of a novel fine-grained mapping phase with a high degree of parallelism, a parallel cut expansion phase and a parallel timing analysis pass. The mapping phase is enhanced by specifically tailored cut evaluation and memory management algorithms for GPUs that enable fast mapping of large circuits with limited GPU memory. Experiments show that compared with the high-performance mapper implemented in ABC, FineMap achieves 128.7× speedup with better quality in terms of area on large benchmarks. Tianji Liu, Lei Chen 0031, Xing Li 0023, Mingxuan Yuan, Evangeline F. Y. Young |
ASPDAC | 1 |
| 2024 | Massively Parallel AIG ResubstitutionabstractResubstitution is a flexible algorithmic framework for circuit restructuring that has been incorporated into many high-effort logic optimization flows. It is thus important to speed up resubstitution in order to obtain high-quality realizations of large-scale designs. This paper proposes a massively parallel AIG resubstitution algorithm targeting GPUs, with effective approaches to addressing cyclic dependencies and restructuring conflicts. Compared with ABC and mockturtle, our algorithm achieves 41.9× and 50.3× acceleration on average without quality degradation. When combining our resubstitution with other GPU algorithms, a GPU-based resyn2rs sequence obtains 46.4× speedup over ABC with 0.8% and 5.8% smaller area and delay respectively. Tianji Liu, Martin D. F. Wong, Evangeline F. Y. Young |
DAC | 2 |
| 2024 | On Advanced Methodologies for Microarchitecture Design Space ExplorationabstractWith the ever-increasing complexity of microprocessors, microarchitectural design becomes over-challenging. Design space exploration (DSE) of microarchitecture configurations to obtain high-quality designs with different PPA trade-offs is time-consuming, due to the huge configuration space and inefficient VLSI verification flow. Many DSE frameworks proposed in previous works failed to systematically analyze the contribution of each algorithmic component to the full flow. This paper provides a novel methodology for designing DSE frameworks by separating DSE flow into stages, and discussing algorithmic instantiations in each stage with theoretical and experimental analyses. Newly formulated DSE frameworks guided by this methodology achieve state-of-the-art results in ICCAD’22 DSE contest evaluation environments. Tianji Liu, Qijing Wang, Evangeline F. Y. Young |
ACM Great Lakes Symposium on VLSI | 1 |
| 2024 | Improving Physical Layer Security with RIS-Assisted Symbiotic RadioabstractReconfigurable intelligent surface (RIS) has been widely exploited for secure communications in physical layer security (PLS) by destructing the eavesdropper's channel via reflect beamforming. In this paper, we investigate RIS-aided secure communications with a novel RIS design scheme. The proposed design leverages RIS to increase the achievable secrecy rate via transmitting the artificial noise (AN) instead. To do so, RIS modulates its information over the incident signal and reflects it to the legitimate receiver and eavesdropper. The RIS modulation scheme is a prior knowledge available at the legitimate receiver, but not available at the eavesdropper. Thus, the reflected signal through the RIS is naturally an additional multi-path component for the legitimate user but a type of AN for the eavesdropper, yielding a mutualistic symbiosis between the RIS and the legitimate user but a parasitic symbiosis between the RIS and the eavesdropper as demonstrated in symbiotic radio (SR). From this SR perspective, we consider an achievable secrecy rate maximization problem by optimizing the RIS reflect beamforming. To this end, we use the path-following algorithm to solve the problem iteratively. Further-more, the comparison of the conventional destruct-channel (DC)- RIS design and the proposed AN - RIS design is conducted. Finally, simulation results show that the proposed AN - RIS design outperforms the DC- RIS design in general cases where the reflecting link is weaker than the direct link of eavesdropper. Tianji Liu, Hu Zhou 0001, Ruizhe Long, Ying-Chang Liang |
ICC | 1 |
| 2024 | Parmesan: Efficient Partitioning and Mapping Flow for DNN Training on General Device TopologyabstractRecently, various pipeline parallelism strategies are proposed to tackle the scalability problem of training a large DNN model on a distributed system. However, most of the works focus on pipeline scheduling while lacking a general methodology to handle network partitioning and mapping to distributed systems with heterogeneous interconnection. In this work, we propose an efficient design flow, named Parmesan, to map the training of a large DNN onto a system with general device topology to maximize the throughput. Parmesan works in an end-to-end manner and solves the whole optimization problem in two phases. The first phase aims at producing well-balanced partitions, and the second phase works towards placing the DNN on devices connected by an arbitrary topology network, considering the heterogeneity of the interconnection bandwidth. We show that Parmesan speeds up the pipeline training throughput on systems with different GPU topologies and is able to handle the mapping problem for heterogeneously interconnected architectures. We believe our proposed general device topology mapping algorithm will provide valuable information for architecture designers and assist them in designing a more DNN-friendly architecture. Tianji Liu, Bentian Jiang, Evangeline F. Y. Young |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | Rethinking AIG Resynthesis in ParallelabstractThe efficiency issue of logic optimization becomes critical as the scale of VLSI designs grows. Since various algorithms are interleaved during optimization to ensure quality, it is necessary to accelerate those commonly used algorithms for obtaining substantial total speed-up. This paper proposes novel parallel algorithms for AIG refactoring and AND-balancing. Equipped with delicately designed parallel-friendly, data-race-free frameworks and GPU data structures, our algorithms obtain significant speed-up and enable the resyn2 sequence to be fully GPU-parallelized when combined with GPU rewriting. Experiments show that on large AIGs, we achieve average accelerations up to 45.9×over ABC with comparable or better qualities. Tianji Liu, Evangeline F. Y. Young |
DAC | 1 |
| 2022 | NovelRewrite: node-level parallel AIG rewritingabstractLogic rewriting is an important part in logic optimization. It rewrites a circuit by replacing local subgraphs with logically equivalent ones, so that the area and the delay of the circuit can be optimized. This paper introduces a parallel AIG rewriting algorithm with a new concept of logical cuts. Experiments show that this algorithm implemented with one GPU can be on average 32X faster than the logic rewriting in the logic synthesis tool ABC on large benchmarks. Compared with other logic rewriting acceleration works, ours has the best quality and the shortest running time. Shiju Lin, Tianji Liu, Martin D. F. Wong, Evangeline F. Y. Young |
DAC | 3 |