Shixin Chen

dblp:131/7874 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 5 first-author · 8 since 2021Security and privacy · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 IP-Matcher: An Efficient One-to-Many Matching Framework for Analog Circuit Design and Reusing
abstract
The design efficiency of analog circuits is generally lower than that of digital circuits, presenting a significant bottleneck in the current integrated circuit industry. One promising method to accelerate design processes is the modular design philosophy adapted from digital methodologies. However, there is a lack of an efficient framework for reusing mature analog circuit topologies and the corresponding layout designs. To achieve a rapid design iteration while utilizing specialized expertise in design, we propose IP-Matcher, an efficient IP-based analog circuit matching and reusing framework. The framework consists of three components: Analog Graph Converter, Analog IP Manager, and IP-based Matcher, which collaborate to enhance both matching accuracy and speed, thereby improving analog IP reusability. We leverage the unique characteristics of analog circuits to significantly prune the matching space, overcoming the limitations of traditional circuit matching strategies. Experimental results show that our work not only outperforms the state-of-the-art method by 32% in accuracy but also achieves a 16× speedup.
Shixin Chen, Peng Xu 0052, Tinghuan Chen, Bei Yu 0001
DATE1
2026 CHASE: A CHiplet Architecture Simulation and Exploration Framework with Decoupled Multi-Fidelity Optimization
abstract
Chiplet-based architecture is a promising emerging technology with benefits in cost, reusability, and performance. However, designing a complicated system to fulfill the comprehensive design metrics is challenging, and designers frequently suffer from tedious evaluation iterations. We propose the CHASE framework, i.e., a CHiplet-based Architecture Simulation and Exploration framework, which jointly considers both performance metrics and manufacturing. In the framework, simulation component ChipletSIM offers holistic modeling of chiplet-based architectures, integrating critical performance metrics (e.g., latency, power) and manufacturing metrics (e.g., yield, cost) across design stages. The exploration component, ChipletDSE, adopts a decoupled multi-fidelity exploration strategy to boost design exploration efficiency and reduce resource consumption. Our framework substantially improves the probability of attaining optimal designs in the early design phase via a comprehensive simulation process and an efficient exploration approach. Compared to previous methods, the experimental results demonstrate the effectiveness of the CHASE framework in comprehensive simulation and efficient exploration.
Shixin Chen, Jianwang Zhai, Bei Yu 0001
ISPD1
2025 Bootstrappable Fully Homomorphic Attribute-Based Encryption with Unbounded Circuit Depth
Feixiang Zhao, Shixin Chen, Man Ho Au, Jian Weng 0001, Huaxiong Wang, Jian Guo 0001
ASIACRYPT (7)2
2025 The Survey of 2.5D Integrated Architecture: An EDA perspective
abstract
Enhancing performance while reducing costs is the fundamental design philosophy of integrated circuits (ICs). With advancements in packaging technology, interposer-based chiplet architecture has emerged as a promising solution. Chiplet integration, often referred to as 2.5D IC, offers significant benefits, including cost-effectiveness, reusability, and improved performance. However, realizing these advantages heavily relies on effective electronic design automation (EDA) processes. EDA plays a crucial role in optimizing architecture design, partitioning, combination, physical design, reliability analysis, etc. Currently, optimizing the automation methodologies for chiplet architecture is a popular focus; therefore, we propose a survey to summarize current methods and discuss future directions. This paper will review the research literature on design automation methods for chiplet-based architectures, highlighting current challenges and exploring opportunities in 2.5D IC from an EDA perspective. We expect this survey will provide valuable insights for the future development of EDA tools chiplet-based integrated architectures.
Shixin Chen, Zichao Ling, Jianwang Zhai, Bei Yu 0001
ASP-DAC1
2025 Breaking the Trilemma: Toward Efficient, Privacy-Preserving, and Forward-Secure Data Sharing in the Post-Quantum Era
abstract
Cloud-based data sharing has emerged as a prevailing solution for enterprises and end users, supporting various online services in our daily lives. However, the current cloud security solutions are vulnerable to the “harvest now, decrypt later” threat imposed by future quantum computers. To encounter the threat, lattice-based cryptographic solutions for supporting cloud data encryption and search have been extensively investigated by both academia and industry. Despite these efforts, existing lattice-based schemes fall into a trilemma: (1) lack of efficient access control for data retrieval; (2) inadequate protection of keyword privacy in both ciphertext and search token; and (3) difficulty in realizing forward secrecy to safeguard historical data. These limitations result in a substantial burden for lattice-based solutions to be adopted in real-world cloud data sharing. To our knowledge, no prior work has comprehensively addressed these issues at the same time, motivating us to design a more flexible, efficient, and secure lattice-based solution. In this paper, we propose an efficient, privacy-preserving, and forward-secure data sharing framework centered around a novel primitive called Forward-Secure Authenticated Searchable Encryption (FS-ASE). Specifically, we first construct an Authenticated Searchable Encryption (ASE) scheme based on ideal lattices, enabling efficient one-to-many search functionality and ensuring keyword privacy in both ciphertext and search token. On top of this primitive, we present the FS-ASE scheme, which achieves forward secrecy through a highly efficient key evolution mechanism, thereby keeping the confidentiality of historical data even if the current secret key is compromised. Finally, the security of our construction is proven under the Ring Learning With Errors (RLWE) assumption, and experimental results show that it achieves performance improvements of 158× in data retrieval and 350× in token generation over state-of-the-art approaches, indicating its practicality in real use.
Jian Weng 0001, Pengfei Wu 0003, Shixin Chen, Jianfei Sun, Guomin Yang, Robert H. Deng
IEEE Trans. Inf. Forensics Secur.4
2025 Rank-DSE: Neural Pareto Comparator of Microarchitecture Design Space Exploration
abstract
The complexity of microarchitecture design has surged due to the expanding design space and time-intensive verification processes. Existing regression-based machine learning methods struggle with inaccurate estimations because of limited training samples. To address these challenges, we propose Rank-DSE, a novel framework for microarchitecture design space exploration (DSE) that leverages a Neural Pareto Comparator (NPC) to directly model the comparative relationships between different architecture designs. Rank-DSE bypasses the inaccuracies of absolute PPA (performance, power, area) predictions by focusing on relative comparisons. The NPC computes the probability of one architecture dominating another and employs semi-supervised learning to reduce the reliance on labeled data. Additionally, a reinforcement-learning-based sampling scheme with an updating baseline Pareto set accelerates the exploration process. Experimental results on the ICCAD 2021 benchmark demonstrate that Rank-DSE achieves superior search quality and cost-efficiency compared to state-of-the-art methods. Specifically, Rank-DSE improves hypervolume by up to 7% while reducing exploration cost by 53.09% compared to cutting-edge approaches. These results highlight the advantages of Rank-DSE in terms of efficiency and effectiveness for microarchitecture DSE.
Peng Xu 0052, Su Zheng, Mingzi Wang, Ziyang Yu 0001, Shixin Chen, Tinghuan Chen, Keren Zhu 0001, Tsung-Yi Ho, Bei Yu 0001
ACM Trans. Design Autom. Electr. Syst.5
2024 SoC-Tuner: An Importance-guided Exploration Framework for DNN-targeting SoC Design
abstract
Designing a system-on-chip (SoC) for deep neural network (DNN) acceleration requires balancing multiple metrics such as latency, power, and area. However, most existing methods ignore the interactions among different SoC components and rely on inaccurate and error-prone evaluation tools, leading to inferior SoC design. In this paper, we present SoC-Tuner, a DNN-targeting exploration framework to find the Pareto optimal set of SoC configurations efficiently. Our framework constructs a thorough SoC design space of all components and divides the exploration into three phases. We propose an importance-based analysis to prune the design space, a sampling algorithm to select the most representative initialization points, and an information-guided multi-objective optimization method to balance multiple design metrics of SoC design. We validate our framework with the actual very-large-scale-integration (VLSI) flow on various DNN benchmarks and show that it outperforms previous methods. To the best of our knowledge, this is the first work to construct an exploration framework of SoCs for DNN acceleration.
Shixin Chen, Su Zheng, Wenqian Zhao 0002, Bei Yu 0001
ASPDAC1
2024 WinoGen: A Highly Configurable Winograd Convolution IP Generator for Efficient CNN Acceleration on FPGA
abstract
The convolution neural network (CNN) has been widely adopted in computer vision tasks. In the FPGA-based CNN accelerator design, Winograd convolution can effectively improve computation performance and save hardware resources. However, building efficient and highly compatible IP for arbitrary Winograd convolution on FPGA remains underexplored. To address this issue, we propose a novel and efficient reformulation of Winograd convolution, named Structured Direct Winograd Convolution (SDW). We further develop WinoGen, a Chisel-based highly configurable Winograd convolution IP generator. Given arbitrary input/output tile size and kernel size, it can generate optimized high-performance IP automatically. Meanwhile, our generated IP can be compatible with multiple kernel sizes and tile sizes. Experimental results show that the IP generated by WinoGen achieves DSP efficiency up to 3.80 GOPS/DSP and energy efficiency up to 652.77 GOPS/W while showing 2.45× and 3.10× improvements when processing a same CNN model compared with state-of-the-arts.
Pengjia Li, Shixin Chen, Beichen Li 0003, Chong Tong, Jianlei Yang 0001, Tinghuan Chen, Bei Yu 0001
DAC4
2024 WVFL: Weighted Verifiable Secure Aggregation in Federated Learning
abstract
Federated learning has shown great potential in Internet of Things (IoTs) for performing intelligent decision making. It allows IoT devices to collaboratively train a neural network upon the data they collect while separately keeping these data staying local. However, several research works have shown that such architecture still faces security challenges that adversaries could raise inference attack to the transferring model parameters to reveal data from devices. Moreover, another security risk in federated learning is that malicious devices may launch model pollution attack to reduce the quality of the aggregated model, or dishonest server may output incorrect aggregated result to the devices. Most existing privacy-preserving federated learning protocols could not deal with both problems. In this paper, we present WVFL, a secure weighted aggregation protocol in which aims to minimize the effect of wrong local models to the aggregated model, meanwhile allowing devices to verify the correctness of the aggregation result. All important intermediate values in the process are in encrypted form so that they would not be revealed to both devices and servers to guarantee privacy. At the end of this paper, we give implementation of our WVFL scheme, showing its efficiency compared with previous work.
Yijian Zhong, Wuzheng Tan, Zhifeng Xu 0003, Shixin Chen, Jia-Si Weng 0001, Jian Weng 0001
IEEE Internet Things J.4
2024 GTCO: Graph and Tensor Co-Design for Transformer-Based Image Recognition on Tensor Cores
abstract
Deep learning frameworks or compilers optimize the operators in computation graph using fixed templates via significant engineering efforts, which may miss potential optimizations such as operator fusion. Therefore, automatically implementing and optimizing the emerging new combinations of operators on a specific hardware accelerator is of importance. In this article, we introduce GTCO, a tensor compilation system designed to accelerate transformer-based vision models’ inference on GPUs. GTCO tackles the operator fusion techniques in the transformer-based model using a novel dynamic programming algorithm and proposes a search policy with new sketch generation rules for the fused batch matrix multiplication and softmax operators. Tensor programs are sampled from an effective search space, and a hardware abstraction with hierarchical mapping from tensor computation to domain-specific accelerators (Tensor Cores) is formally defined. Finally, our framework can map and transform tensor expression into efficient CUDA kernels with hardware intrinsics on GPU. Our experimental results demonstrate that GTCO improves the end-to-end execution performance by up to$1.73\times $relative to the cutting-edge deep learning library TensorRT on NVIDIA GPUs with Tensor Cores.
Xufeng Yao, Qi Sun 0002, Wenqian Zhao 0002, Shixin Chen, Zixiao Wang 0001, Bei Yu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2024 Floorplet: Performance-Aware Floorplan Framework for Chiplet Integration
abstract
A chiplet is an integrated circuit (IC) that encompasses a well-defined subset of an overall systems functionality. In contrast to traditional monolithic system-on-chips (SoCs), chipletbased architecture can reduce costs and increase reusability, representing a promising avenue for continuing Moore’s Law. Despite the advantages of multi-chiplet architectures, floorplan design in a chiplet-based architecture has received limited attention. Conflicts between cost and performance necessitate a trade-off in chiplet floorplan design since additional latency introduced by advanced packaging can decrease performance. Consequently, balancing performance, cost, area, and reliability is of paramount importance. To address this challenge, we propose Floorplet (Floorplan chiplet), a framework comprising simulation tools for performance reporting and comprehensive models for cost and reliability optimization. Our framework employs the open-source Gem5 simulator to establish the relationship between performance and floorplan for the first time, guiding the floorplan optimization of multi-chiplet architecture. The experimental results show that our method decreases inter-chiplet communication costs by 24.81%.
Shixin Chen, Shanyi Li, Zhen Zhuang, Su Zheng, Zheng Liang 0003, Tsung-Yi Ho, Bei Yu 0001, Alberto L. Sangiovanni-Vincentelli
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1