EDBT 2026 Demo / reviewers in the wild / expert
Shunbin Li
dblp:176/0878
· DBLP profile ↗
7ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0002-9350-0287ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-author · 3 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Architectural Exploration for Waferscale Switching SystemabstractWith the end of Moore’s law and Dennard scaling, waferscale systems or processors that integrate multiple pre-tested known good dies (KGDs) on a waferscale-interposer are new approaches to further improve the chiplet-based system’s performance. This article explores the network on wafer (NoW) architecture of waferscale switching system under several physical constraints. A software-based approach is proposed to redefine the topological property. A five-level butterfly fat-tree (BFT)-like logical topology with 8.96-Tb/s (896 ports$\times 10$Gb/s/port) switching bandwidth is achieved based on 2-D-mesh-like physical topology. We show that the proposed BFT-like topology with breadth-first-search (BFS) based traffic balanced routing algorithm reduces 55.6% hops, 41.4% transmission delay, and improves 24.2% throughput compared to 2-D-mesh-like topology under different traffic distributions. This BFT-like waferscale switching system is suitable for high-performance computing and data centers. In addition, the numerical analysis shows that the waferscale package can provide significant power efficiency and latency advantages compared to the typical single-chip package, which mainly benefits from the short-reach IO requirements. Note that the proposed waferscale switching system is compatible with high-switch-capacity dies with advanced process technology, which can further improve system performance. Finally, we present the physical implementations for the waferscale system with heterogeneous dies. Zhiquan Wan, Zhipeng Cao 0001, Shunbin Li, Peijie Li, Qingwen Deng, Kun Zhang 0037, Guandong Liu, Ruyun Zhang 0001, Qinrang Liu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2024 | Modeling and Analysis of Waferscale Switching Network with Multiple System FaultsabstractWith the end of Moore's Law and Dennard scaling, waferscale systems that integrate multiple pre-tested known good dies (KGDs) on a waferscale-interposer are new approaches to further improve chiplets-based systems' performance. This paper explores the manufacturing defects of waferscale systems and the effect to network on wafer (NoW). A traffic-balanced routing algorithm is proposed for the irregular network with multiple faults of NoW. The results show that the routing algorithm reduces transmission delay by 57% and improves throughput by 36.4% compared to the normal breadth-first-search algorithm. Besides, we build a throughput model by using the nonlinear least squares method (NLLS) with the parameters of faults number, location and concentration ratio. The results show that the model can reach 0.98 goodness of fit and accurately predict the throughput performance of different irregular NoW topologies. The model can be used to predict the system performance and avoid critical faults of the physical waferscale switching system. Zhiquan Wan, Zhipeng Cao 0001, Shunbin Li, Dehao Ye |
ISCAS | 3 |
| 2023 | Mangling Rules Generation With Density-Based Clustering for Password GuessingabstractRule-based password generation is one of the most effective and often employed techniques in the highly compute-intensive password recovery process. However, it is challenging to design and maintain a practical password mangling ruleset, which is a time-consuming task requiring specialized expertise. This paper therefore introduced MDBSCAN (Modified Density-Based Spatial Clustering of Applications with Noise), a novel density-based cluster approach in machine learning, to build an automatic password mangling rule generator. To evaluate the proposed method, cross-checks across 4 different real-world password datasets leaked from popular Internet services and applications are adopted. The results indicate that the proposed generator could produce high-quality mangling rules with a better hit rate and enhance current mangling rules by identifying hidden or omitted rules. The proposed approach also shows strong interpretability and computational efficiency. When examining the RockYou password dataset with the top 77 rules, the hit rate may rise by 11% to 104% proportionally to other well-known solutions. Furthermore, by combining the top 77 rules generated by MDBSCAN with those from other rulesets, 3–12.67% more real-world passwords can be retrieved. Shunbin Li, Ruyun Zhang 0001, Chunming Wu 0001, Hanguang Luo |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2022 | Secured Data Transmission Over Insecure Networks-on-Chip by Modulating Inter-Packet DelaysabstractAs the network-on-chip (NoC) integrated into an SoC design can come from an untrusted third party, there is a growing risk that data integrity and security get compromised when supposedly sensitive data flows through such an untrusted NoC. We thus introduce a new method that can ensure secure and secret data transmission over such an untrusted NoC. Essentially, the proposed scheme relies on encoding binary data as delays between packets travelling across the source and destination pair. The maximum data transmission rate of this inter-packet-delay (IPD)-based communication channel can be determined from the analytical model developed in this article. To further improve the undetectability and robustness of the proposed data transmission scheme, a new block coding method and communication protocol are also proposed. Experimental results show that the proposed IPD-based method can achieve a packet error rate (PER) of as low as 0.3% and an effective throughput of$\boldsymbol {2.3\times 10^{5}}$b/s, outperforming the methods of thermal covert channel, cache covert channel, and circuit-based encryption and, thus, is suitable for secure data transmission in unsecure systems. Jiaen Xu, Xiaohang Wang 0001, Yingtao Jiang, Amit Kumar Singh 0002, Chongyan Gu, Letian Huang, Mei Yang 0001, Shunbin Li |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2019 | Energy-Efficient RAR3 Password Recovery with Dual-Granularity Data Path StrategyabstractPassword recovery tools are used to recover lost passwords and regain access to precious data. Due to the extremely large time and energy consumption of password recovery, efficient hardware accelerators are demanded to accelerate the recovery process. However, simple and regular data interconnect paths between data sources and non-blocking hash pipelines are hard to construct for Roshal ARchive version 3 (RAR3) algorithm based on field programmable gate array (FPGA) devices. The difficulty comes from the fact that the message format of the hash pipeline inputs vary with the password length and the secure hash algorithm 1 (SHA-1) iteration phase. To attack this problem, a dual-granularity data path adjustment strategy is proposed to eliminate the randomness of message block formats caused by the irregularity of password length and to efficiently schedule the data through the regular data interconnect paths. Experimental results show that the proposed hardware accelerator for RAR3 password recovery is 3.3 × more energy-efficient than a state-of-the-art implementation Hashcat on NVIDIA GTX 1060 GPU. Qingyuan Ding, Shunbin Li, Peng Liu 0016 |
ISCAS | 3 |
| 2019 | An Energy-Efficient Accelerator Based on Hybrid CPU-FPGA Devices for Password RecoveryabstractPassword recovery tools are needed to recover lost and forgotten passwords so as to regain access to valuable information. As the process of password recovery can be extremely compute-intensive, hardware accelerators are often needed to expedite the recovery process. This paper thus presents a high performance, energy-efficient accelerator built upon modern hybrid CPU-FPGA SoC devices. The proposed password recovery accelerator relies on the development of a set of intellectual property (IP) cores for implementing variety of encryption algorithms with vastly different characteristics and complexities. To keep the resource requirements of each IP core running on a resource-strapped FPGA to the minimum, while achieving the highest throughput possible, the most performance critical computational hash functions are mapped to the FPGA with two specific optimization techniques, namely the fixed message padding for hashing and loop transformation for deep pipelining. The proposed password recovery accelerator implements a non-blocking deep pipeline design that does not incur any data and structural hazards, which is made possible by applying a task scheduling scheme through the use of block RAMs. Synchronization between tasks that are mapped to run separately on CPU and FPGA is achieved through task reordering and a communication protocol for maximum parallelism and low overhead. The proposed design is evaluated on Xilinx XC7Z030-3 device, and it is compared much favorably with other known implementations. The proposed hardware accelerator design is found 12.5 and 3.1 times more resource-efficient than the pure FPGA-based password recovery accelerators for TrueCrypt and WPA-2, respectively. The proposed implementation also shows more than 200 percent improvement in energy efficiency over a state-of-the-art implementation on NVIDIA GTX 750 Ti GPU. Peng Liu 0016, Shunbin Li, Qingyuan Ding |
IEEE Trans. Computers | 2 |
| 2017 | An Adaptive PAM-4 Analog Equalizer With Boosting-State Detection in the Time DomainabstractThis paper introduces an improved adaptive analog equalizer that is required in high speed serial receivers using four-level pulse amplitude modulation signaling. By performing boosting-state detection in the time domain, the proposed adaptive analog equalizer can effectively overcome a serious problem that the received signal's eye-opening tends to be compromised by the convergence accuracy of the adaptive control loop. To suppress the pattern-dependent jitters (PDJs), an inductor-less, cross-stage feedback structure is employed in the proposed analog equalizer to help broaden its effective tuning bandwidth. Multichannel simulations have confirmed that the proposed equalizer is able to achieve a 42% improvement in eye-height opening when compared with the equalizers employing popular spectrum-comparing schemes. Trellis diagram analyses under different date rates have revealed that the bandwidth of the proposed equalizer can be extended by as much as 60%, thus effectively bringing the PDJs from 44% down to 27.5%. Shunbin Li, Yingtao Jiang, Peng Liu 0016 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |