Shuting Cai

dblp:49/4159 · DBLP profile ↗
← Back
44ranked-venue papers
0as first author
37since 2021 · last 2026
0000-0002-2842-6439ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 29 · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Comprehensive Delay-Aware Net Weighting Framework for Timing-Driven Global Placement
Lixin Chen, Keyu Peng, Jinghui Zhou, Shuting Cai, Ziran Zhu
ASP-DAC6
2026 Optimizing detailed-routability for 3D global routing through dynamic resource model and routability-aware cost scheme
Juntao Jian, Shuting Cai, Xiaoming Xiong
Integr.3
2026 Routability-wirelength co-guided cell inflation with explainable multi-task learning for global placement optimization
Zicheng Deng, Shuting Cai, Xiaoming Xiong
Integr.3
2026 A novel recurrent topology-based memristor: Simplification and circuit simulation for multi-scroll attractor generation
Yifeng Diao, Shufeng Huang, Xiaoming Xiong, Shuting Cai
Inf. Sci.6
2026 High-Efficiency Bidirectional Translator between SystemC and Verilog
abstract
The SystemC language, with its higher level of abstraction, plays a critical role in facilitating hardware/software co-design and architecture exploration. However, as most hardware models are predominantly written in Verilog and translating between SystemC and Verilog remains a challenge, an efficient and reliable tool for translating between these two languages is essential to streamline system development. This article proposes SCAV, a bidirectional translator between SystemC and Verilog, which breaks these limitations. SCAV provides a fully automated solution for translating both SystemC to Verilog and Verilog to SystemC, leveraging a translation framework with front-end/back-end separation. Additionally, SCAV incorporates an Abstract Syntax Tree (AST) filter, optimizing the translation process by filtering out invalid content. The experimental results demonstrate that SCAV achieves a 100% adaptation rate for Verilog and a 98% adaptation rate for SystemC, with 100% accuracy in both directions. Furthermore, SCAV outperforms existing tools, delivering a minimum speedup of 18% across various test cases.
Xin Zheng 0001, Yongfeng Zhong, Shaofen Zeng, Huaien Gao, Shuting Cai, Xiaoming Xiong
ACM J. Emerg. Technol. Comput. Syst.6
2026 An efficient DSP packing framework for FPGA-based mixed-precision DCNN processor
Xueming Li 0001, Jinhui Pan, Hongmin Huang, Yuanmiao Lin, Xianghong Hu 0001, Shuting Cai, Xiaoming Xiong
J. Syst. Archit.6
2026 NTT-LSU: Tightly Coupled Architecture for Efficient NTT Implementation on RISC-V Processor
abstract
Polynomial multiplication is one of the most computationally intensive operations in lattice-based cryptographic systems, directly impacting overall computational efficiency. Although the Number Theoretic Transform (NTT) reduces the time complexity of polynomial multiplication fromO(n2) toO(nlogn), the varying parameter requirements of different lattice algorithms limit the generality of hardware designs. To address this issue, we propose an innovative hardware architecture that tightly couples the NTT unit with the Load and Store Unit (LSU) in the pipeline of a RISC-V processor. We also propose a hybrid width data path method that effectively reduces data transfer time. Compared to previous designs, our architecture minimizes data transfer latency while enhancing computational flexibility and scalability. Specifically, we have customized NTT-related instructions to support Inverse Number Theoretic Transform (INTT) and various NTT parameter configurations. Experimental results demonstrate that our solution significantly shortens the NTT computation cycle, achieving over 10× speedup compared to software implementations. In comparison to existing solutions, our architecture exhibits superior area-time product (ATP) performance.
Yinqiao Zhao, Zilong Xie, Ruidian Zhan, Xiaoming Xiong, Yun Chen 0004, Shuting Cai
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2026 Low Redundancy Fourth-Order Loose Complementary Coprime Array and FPGA Implementation
abstract
Fourth-order cumulants (FOC) is well known for direction of arrivals (DOA) estimation due to the high degrees of freedom (DOF), especially when combined with a sparse array. In this paper, we provide a set of FOC sparse array design strategies for a real-time mutual coupling (MC) hardware system. Data repetition is reduced by decreasing array redundancy of fourth-order coarray. The fourth-order difference coarray problem is decomposed into a two-level second-order difference coarray optimization problem. Three types of structures are proposed to reduce array redundancy. The first structure provides a redundancy reduction framework by removing antennas with respect to parity. The second structure can save more physical antennas while the DOF is still maintained at high level. The third structure extends the DOF using complementary element shifting strategy. The redundant data size and MC are substantially reduced, which provides extra benefits for hardware implementation. More than 50% antennas can be saved compared to generalized coprime array (GCA). The proposed approach enables high accuracy and low latency DOA estimation while Field Programmable Gate Array (FPGA) resource is limited.
Zhe Fu 0001, Dongsheng Lv, Shuting Cai
IEEE Trans. Circuits Syst. I Regul. Pap.4
2026 An FPGA-Efficient CNN Accelerator for Hybrid Model Compression With Scalable Bit-Serial and Bit-Parallel MAC
abstract
Mixed-precision quantization and unstructured pruning have emerged as two effective compression techniques, demonstrating great potential in reducing model size and computational cost in the deployment of convolutional neural networks (CNNs). However, their joint deployment still faces two major challenges: 1) The former introduces heterogeneous bit-widths in the bit-level, while the latter results in irregular sparsity in the value-level; their fundamentally incompatible data representations and computation patterns require two distinct types of hardware overhead to process them separately, which severely limits hardware execution efficiency. 2) Jointly applying both techniques often leads to notable accuracy degradation. In this paper, we propose a novel compression perspective that reinterprets zero-values generated by unstructured pruning as multiple consecutive 0-bits. We further introduce column-based bit-level sparsity, which provides a unified representation for weights after mixed-precision quantization and unstructured pruning, requiring only a single type of hardware overhead. Based on these techniques, we develop a hybrid compression framework that jointly optimizes model size, accuracy, and hardware implementation. Our method achieves weight/activation precision of 2.13b/4.06b on VGG16, delivering 7.40$\times $compression and 2.74$\times $speedup with 0.92% accuracy loss compared to the 8b baseline. Compared to state-of-the-art accelerators, our design achieves 1.12$\times $-6.23$\times $and 1.31$\times $-6.60$\times $improvements in energy efficiency and LUT efficiency when deploying VGG16 and ResNet50.
Yuanmiao Lin, Xueming Li 0001, Hongmin Huang, Heng Mai, Ruidian Zhan, Xianghong Hu 0001, Shuting Cai, Xiaoming Xiong
IEEE Trans. Circuits Syst. I Regul. Pap.8
2026 A Low-Cost Local Masking Radix-4 NTT Against Soft-Analytical Side-Channel Attacks
abstract
The number theoretic transform (NTT) is essential for accelerating polynomial multiplication in lattice-based cryptography. However, it is vulnerable to soft-analytical side-channel attacks (SASCAs). Although local masking countermeasure provides theoretical resistance against such attacks, its direct implementation in Radix-4 NTT architecture leads to more than a 4 times increase in modular multiplications, resulting in substantial hardware overhead. To address this challenge, we propose the modular multiplication parallel mask sharing (MMPMS) scheme, which optimizes the modular multiplication parallelism of the Radix-4 butterfly units and shares random twiddle factors, thereby achieving a balance between hardware overhead and security. Then, we construct a complete local masking NTT/INTT algorithm and efficiently implement it on the Artix-7 field-programmable gate array (FPGA). Experimental results show that compared with the state-of-the-art local masking NTT, our scheme reduces the equivalent area and ATP overhead by more than 8.24 times and 6.74 times, respectively. In addition, a nonspecifict-test analysis indicates no significant side-channel leakage.
Congwei Chen, Jinwei Pu, Jianxiong Zhang 0003, Jiaying Liao, Ruidian Zhan, Yun Chen 0004, Shuting Cai
IEEE Trans. Very Large Scale Integr. Syst.8
2026 FPUltra: An Area-Efficient Single-Precision Floating-Point Unit for Cost-Sensitive RISC-V Cores
abstract
Area efficiency is vital for floating-point units (FPUs) in resource-constrained IoT devices. However, existing designs suffer from rigid architectures and costly arithmetic units, limiting performance-area optimization. To this end, this work presents FPUltra, an area-efficient single-precision FPU for cost-sensitive RISC-V cores. FPUltra adopts a novel phase-decoupled control architecture to mitigate timing hazards and improve execution efficiency. A parallel approximate floating-point multiplier (FPM) is designed using combinational logic, based on the Mitchell algorithm with error compensation. A Newton–Raphson-based subinstruction decomposition method is presented to support floating-point division (Fdiv) and square root (Fsqrt). Compared with state-of-the-art FPUs, FPUltra achieves 9%–695% and 101%–14 186% improvements in equivalent slices efficiency (Eq.Slices Eff.) on FPGA and equivalent area efficiency (Eq.Area Eff.) on ASIC, respectively. Our code will be available athttps://github.com/LX-IC/FPUltra
Xian Lin, Jiahao Lan, Xin Zheng 0001, Huanxin Zhuang, Huaien Gao, Shuting Cai, Xiaoming Xiong
IEEE Trans. Very Large Scale Integr. Syst.6
2026 Efficient FPGA Acceleration for 4-bit CNNs via Quantization-Induced Structured Sparsity and LUT-Based Multiplication
abstract
N:M structured sparsity is key to convolutional neural network (CNN) compression and acceleration, but two challenges remain. From the algorithm perspective, prior works have mainly focused on 8-bit quantized models, where N:M sparsity yields limited hardware efficiency. From the hardware perspective, the cost differences across N:M sparsity have not been analyzed. To address these issues, we present a unified algorithm–hardware co-design framework for 4-bit CNN acceleration. We show that 4-bit quantization induces over 80% zero weights and strongly structured sparsity, with over 95% of weight groups satisfying 4:8, 8:16, or 16:32 patterns. We propose a pruning-after-quantization (PAQ) algorithm that enforces strict N:M sparsity with minimal accuracy loss. We also analyze the hardware overhead of activation fetch units (AFUs) under different N:M sparsity patterns (4:8, 8:16, 16:32), revealing that the 4:8 AFU reduces look-up table (LUT) cost by up to 66.7% compared to 16:32. Finally, we introduce a 4-bit LUT-based sign-magnitude multiplier (LBSMM) requiring only 11 LUT6 resources, outperforming existing multipliers. Integrated on a Xilinx VCU118 field-programmable gate array (FPGA), our accelerator achieves$2.51\times $–$12.89\times $improvements in equivalent LUT efficiency over SOTA designs. The implementations of the PAQ algorithm and the RTL of LBSMM are available athttps://github.com/haden-01/PAQ-and-LBSMM.git
Yuanmiao Lin, Xueming Li 0001, Hongmin Huang, Ruidian Zhan, Xianghong Hu 0001, Shuting Cai, Xiaoming Xiong
IEEE Trans. Very Large Scale Integr. Syst.8
2025 Expediting the discovery of promising photothermal cyanine molecules through a transfer learning approach
abstract
Cyanine-based molecules have gained significant attention in photothermal therapy due to their unique fluorescence brightness and tunable spectral properties. However, the development of new photothermal agents is often constrained by the complexity of the chemical landscape and the need for biocompatibility. To address these challenges, we present an innovative transfer learning approach for rapidly identifying promising photothermal agent candidates with excellent photothermal properties, high synthetic feasibility, and superior biocompatibility. Using natural language processing, our pretrained model generated a molecular library based on cyanine scaffolds. The most promising candidates were screened rigorously through a weighted analysis of chemical indicators, such as photothermal performance and synthetic accessibility and biological indicators, including bio-toxicity. From these, three molecules were selected for retrosynthetic analysis. This artificial intelligence-driven approach provides a robust solution to the traditional challenges in photothermal agent design, significantly enhancing their potential applications in cancer bioimaging, mitochondrial phototherapy, and image-guided surgery.
Siwei Wu, Liqiang He, Guining Cao, Jiacheng Tang, Zhenxing Pan, Zihui Huang, Andi Li, Shuting Cai, Xujie Liu
Briefings Bioinform.11
2025 An FPGA-based bit-level weight sparsity and mixed-bit accelerator for neural networks
Xianghong Hu 0001, Shansen Fu, Yuanmiao Lin, Xueming Li 0001, Chaoming Yang, Rongfeng Li 0001, Hongmin Huang, Shuting Cai, Xiaoming Xiong
J. Syst. Archit.8
2025 FLALM: A Flexible Low Area-Latency Montgomery Modular Multiplication on FPGA
abstract
Montgomery Modular Multiplication (MMM) is widely used in many public key cryptography systems. This paper presents a Flexible Low Area-Latency MMM (FLALM) implementation, which supports Generic Montgomery Modular Multiplication (GMM) and Square Montgomery Modular Multiplication (SMM) operations. A new SMM schedule for the Finely Integrated Product Scanning (FIPS) GMM algorithm is proposed to accelerate SMM with tiny additional design. Furthermore, a new FIPS dual-schedule is proposed to solve the data hazards of this algorithm. Finally, we explore the trade-off between area and latency, and present the FLALM to accelerate GMM and SMM. The FLALM is implemented on FPGA (Virtex-7 platform). The results show that the area*latency (AL) value of FLALM (wordsize$w$=128) is 38.1% and 44.7% better than the previous state-of-art scalable references when performing 1024-bit and 2048-bit GMM, respectively. Moreover, when computing SMM, the advantage of AL value is raised to 73.7% and 86.3% respectively.
Yujun Xie 0001, Yuan Liu 0022, Xin Zheng 0001, Bohan Lan, Dengyun Lei, Dehao Xiang, Shuting Cai, Xiaoming Xiong
IEEE Trans. Computers7
2025 GNN-Based Timing Prediction in Prerouting Stage With Multitask Learning Strategy
abstract
Static timing analysis tools are commonly used to evaluate timing performance and guide optimization during placement stage. However, traditional timing analysis struggles to fast and accurately evaluate timing violation due to absence of detailed routing information necessary for RC parasitic parameter extraction. Therefore, a timing analyzer based on graph neural network is proposed in this article. Compared to previous works, a novel representation of circuit delay model is proposed in this article, employing timing arcs and virtual pins to predict net delay and arrival time (AT) in the prerouting phase. Additionally, to our knowledge, this is the first attempt to improve the quality of timing analyzer through a strategy of multitask learning, with the proposed enhanced dynamic weight average method. The experimental results demonstrate that our model excels in predicting net delay and AT, with average correlations of 0.9540 and 0.9058, respectively, on the testing set. In comparison to the previous state-of-the-art methods, our approach maintains accuracy in net delay prediction while enhancing the overall$R^{2}$score for AT prediction by 0.0271. Additionally, our method reduces inference time by 25.8%.
Haisen Zhang, Xiaoming Xiong, Shuting Cai
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2025 A Lightweight Heterogeneous Graph Embedding Framework for Hotspot Detection
abstract
Hotspot detection is a crucial step in ensuring the manufacturability of integrated circuits, as it seeks to identify potential defects in the layout. Pattern matching methods have been widely used to accelerate the detection of these defects. However, they often struggle with complicated deviations. Image-based machine learning methods were introduced to confront this challenge, but they often involved distorted information extraction and incurred significant runtime overhead. In this article, we introduce a novel detection framework based on the modified transitive closure graph (MTCG). By applying the concept of MTCG, the layout can be accurately modeled as a heterograph. The embeddings of these heterographs are extracted using an optimized lightweight 3-hop message-passing graph neural network (GNN) and subsequently utilized for classification. Furthermore, a dynamic edge transformation method based on the properties of MTCG is proposed for data augmentation. The proposed method is evaluated with datasets from ICCAD 2012 and ICCAD 2019, demonstrating outstanding performance in recall and false alarm, along with a significantly decreased inference time.
Haopeng Yan, Yuzhe Ma, Xiaoming Xiong, Shuting Cai
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2025 A Precision-Scalable Accelerator with Sign-Magnitude Representation and Dual Adder Trees
abstract
Currently, there are two mainstream acceleration methods; one is mixed precision and the other is sparsity. Few accelerators support both mixed precision and sparsity, and most enable precision configurations across layers rather than within a single layer. Furthermore, most of accelerators adopt the traditional two’s complement (2C) data representation method, and we found that 2C brings many invalid ”1” when representing signed data, which brings more resources overhead for mixed precision and many invalid operations for bit-level sparsity. Therefore, we propose a high-efficiency accelerator featuring a precision-scalable Sign-Magnitude Processing Element (SM-PE), which adopts a data representation method of SM and can flexibly support various precision calculations (2, 4, 8 bits) and bit-level sparsity. In addition, a dynamic quantization algorithm named DoReFaLike and a bit-level column sparsity (BLCS) technique are proposed to improve the efficiency of SM-PEs. Under the same accuracy constraint, the sparsity rate of the SM scheme is 3.5× higher than that of the 2C format. The accelerator has been synthesized on a 55nm CMOS ASIC platform. When scaled to 28nm, experimental results show that the energy efficiency of the proposed accelerator reaches 15.50, 25.37, 101.54 TOPS/W with 8-bit, 4-bit, and 2-bit input activations, respectively, and weights represented in sparse 8-bit precision, operating at 400 MHz. Compared to state-of-the-art accelerators, the proposed design achieves a performance improvement of 1.1× to 3.9×.
Xianghong Hu 0001, Chaoming Yang, Xueming Li 0001, Rongfeng Li 0001, Yuanmiao Lin, Shansen Fu, Hongmin Huang, Shuting Cai, Xiaoming Xiong
ACM Trans. Embed. Comput. Syst.8
2025 HEFPA: hyperbolic embedding and fast position-aware network for point cloud registration
Wenlin Huang, Shuting Cai, Jing Guo 0007
J. Supercomput.3
2025 RTMF: Routing based on TDM for Multi-FPGA System
abstract
As modern VLSI design advances, the significance of multi-FPGA systems in prototyping and verification is steadily growing. Due to the physical I/O limitations, the Time-Division Multiplexing (TDM) and I/O assignment techniques are introduced to solve these problems. However, most multi-FPGA systems primarily focus on inter-FPGA routing while overlooking intra-FPGA routing. In this article, a comprehensive routing framework based on TDM for Multi-FPGA systems (RTMF) is hereby presented. To our knowledge, this is the first attempt to jointly optimize intra-level and inter-level routing in the work of multi-FPGA systems (MFS). The RTMF framework, tailored for system-level and intra-level routing under constrained wiring resources, integrates routing demands within and between FPGAs. Through the integration of TDM technology and adaptive optimization algorithms, RTMF effectively meets routing requirements and delivers efficient solutions. Furthermore, RTMF demonstrates remarkable adaptability, allowing for dynamic adjustments and optimizations to address diverse routing demands and constraints. In comparison to the state-of-the-art methodologies, in benchmark designs with a scale greater than 50,000, our approach on average reduces the maximum routing weight by 59.98% and 46.70%, respectively.
Shiyan Liang, Jingui Lin, Wenxiong Lin, Yuzhe Ma, Xiaoming Xiong, Shuting Cai
ACM Trans. Design Autom. Electr. Syst.9
2025 Early Stage DRC Hotspot Prediction for Mixed-Size Designs Through an Efficient Graph-Based Deep Learning
abstract
Predicting hotspot locations in the early stage of Design Rule Check (DRC) is crucial for designers to proactively prevent design rule violations. However, obtaining an accurate and efficient predictor faces significant challenges due to the influence of available information and severe data imbalance. In this study, we investigate the potential of utilizing Graph Neural networks (GNN) to address this challenge. Our focus is specifically on accurately predicting DRC hotspot locations without relying on global routing techniques. We consider the presence of macros in mixed-size designs. We propose an adaptive adjacency matrix that demonstrates superior application effectiveness compared with traditional adjacency matrices. Furthermore, experimental results on benchmark circuits show significant improvements in the true positive rate (22.38% for the RouteNet model and 26.90% for the GNN model) and accuracy (6.97% and 6.76%, respectively) compared with these models. Our proposed model also maintains a low false positive rate and outperforms other Convolutional Neural Network and GNN models. Additionally, its efficient learning capability and lower computational time contribute to its outstanding training performance, with training time being approximately 10% of that required by other models.
Jingui Lin, Shiyan Liang, Wenxiong Lin, Xiaoming Xiong, Shuting Cai
ACM Trans. Design Autom. Electr. Syst.8
2025 Optimizing FPGA Routing with Explainable Co-Learning of Congestion and Wirelength
abstract
In FPGA routing, machine learning-based optimization methods have achieved improved routing solutions by integrating traditional heuristics with predictive capabilities. However, these approaches mostly relied on single-task learning models with black-box nature and often neglected the complex trade-offs and inter-dependencies between routing metrics. To address these limitations, this paper introduces a novel multi-task learning-based routing optimization method. In the congestion-wirelength co-learning stage, the simultaneous prediction of congestion and wirelength is formulated as a multi-task learning problem. A multi-task learning model, named CWNet, is proposed to tackle this challenge effectively. During the congestion-wirelength impact interpretation, the contribution of congestion to wirelength is quantified using an XAI technique known as DeepSHAP, producing a congestion-wirelength impact map. In the congestion-wirelength co-guided routing optimization (CWRO) stage, the VTR router’s lookahead map is enhanced based on the impact map, guiding the router to avoid locations where congestion significantly affect wirelength. Experimental results demonstrate that CWNet outperforms most baseline learning models in terms of both prediction performance and computational efficiency. Additionally, the impact map visually illustrates the complex and nonlinear relationship between congestion and wirelength. Ultimately, CWRO significantly reduces congestion, wirelength, and critical path delay, while maintaining a competitive runtime compared to baseline routers.
Shuting Cai, Xiaoming Xiong
ACM Trans. Design Autom. Electr. Syst.3
2025 Layout Congestion Prediction Based on Regression-ViT
abstract
To accelerate the back-end design flow of integrated circuit (IC), numerous studies have made exploratory advancements in machine learning (ML) for electronic design automation (EDA). However, most research works are limited to deep learning (DL) models predominantly based on convolutional neural networks, and the models often suffer from poor generalization due to the scarcity of data. In this study, we propose the Double generative adversarial networks (D-GAN) model to enrich the dataset and propose the Regression Vision Transformer (R-ViT) model to predict layout congestion information. Compared with the baseline model, experimental results show improvements of 3.03% and 2.64% in Receiver Operating Characteristic-Area under Curve (ROC-AUC) and Precision-Recall Curve-Area under Curve (PRC-AUC) respectively. To further enhance the prediction accuracy of the model, an adaptive Huber loss function is designed to optimize the training process, resulting in an improvement of up to 11.03% in ROC-AUC compared with the baseline model. Lastly, extended experiments are conducted to study the effects of parameters and convolutional kernel size on performance, which find a better configuration.
Guiqi Mo, Yimin Xia, Jianhong Ou, Shuting Cai, Xiaoming Xiong
ACM Trans. Design Autom. Electr. Syst.4
2025 PSCaps: High-Performance Pose-Sensitive Layout Hotspot Detector based on CapsNet
abstract
Advanced technology nodes face challenges with Design Rule Violations (DRVs), primarily due to the possibility of nm-level small variations that can lead to the occurrence of DRVs. Various Machine Learning (ML) techniques have been introduced to detect whether the layout design conforms to manufacturing rules, thus alleviating the time-consuming challenge associated with traditional lithography simulations. However, existing ML models still face challenges in detecting layout violations where there are pose variations among geometric shapes in the layout. In this study, we propose a hotspot detector called PSCaps based on the CapsNet, which considers the pose information of geometric shapes in the layout. The method effectively captures spatial information and hierarchical structures between geometric shapes in the layout. Through the dynamic routing mechanism, the model adaptively learns the relationships and weight allocations between different capsules. Additionally, we employ multiple data augmentation methods to alleviate the problem of imbalanced hotspot and non-hotspot data in the open-source dataset. The benchmarks of ICCAD-2012 and ICCAD-2019 are used to validate our method. The experimental results demonstrate that our proposed hotspot detector outperforms other state-of-the-art works.
Haopeng Yan, Xiaoming Xiong, Shuting Cai
ACM Trans. Design Autom. Electr. Syst.7
2025 Concurrent Prediction of Timing and wire Length Using A Multi-Task Graph Neural Network
abstract
Traditional supervised single-task learning models are used in timing-driven placement exploration to improve both effectiveness and efficiency by predicting wire length, wire delay, and cell delay separately. However, these metrics are interdependent, with the two delays being timing-based and wire length non-timing, which makes it difficult for single-task models to capture their complex relationships. Moreover, the limited existing multi-task learning methods can only predict either multiple timing or non-timing metrics. To address these limitations, this article introduces DLGNN, a novel multi-task graph learning model that simultaneously predicts these three metrics through an embedder-predictor architecture featuring two residual connections, a combination of both soft and hard parameter sharing, and a geometric loss strategy. Cross-design experimental results on the Nangate 45nm library demonstrate that DLGNN outperforms baseline models in terms of both predictive performance and time efficiency. Additionally, ablation studies emphasize the critical roles of the residual connections, the combination of soft and hard parameter sharing, and the geometric loss strategy in improving DLGNN’s predictive performance. The generalization experiment on the ASAP 7nm library further confirms DLGNN’s advantages for more advanced technology nodes.
Shuting Cai, Xiaoming Xiong
ACM Trans. Design Autom. Electr. Syst.4
2025 Efficient Design Space Exploration for the BOOM Using SAC-Based Reinforcement Learning
abstract
Design space exploration (DSE) is crucial for optimizing the performance, power, and area (PPA) of CPU microarchitectures ($\mu $-archs). While various machine learning (ML) algorithms have been applied to the$\mu $-arch DSE problem, the potential of reinforcement learning (RL) remains underexplored. In this article, we propose a novel RL-based approach to address the reduced instruction set computer V (RISC-V) CPU$\mu $-arch DSE problem. This approach enables dynamic selection and optimization of$\mu $-arch parameters without relying on predefined modification sequences, thus significantly enhancing exploration flexibility. To address the challenges posed by high-dimensional action spaces and sparse rewards, we use a discrete soft actor-critic (SAC) framework with entropy maximization to promote efficient exploration. In addition, we integrate multistep temporal-difference (TD) learning, an experience replay (ER) buffer, and return normalization to improve sample efficiency and learning stability during training. Our method further aligns optimization with user-defined preferences by normalizing PPA metrics relative to baseline designs. Experimental results on the Berkeley out-of-order machine (BOOM) demonstrate that the proposed approach achieves superior performance compared with state-of-the-art methods, showcasing its effectiveness and efficiency for$\mu $-arch DSE. Our code is available athttps://github.com/exhaust-create/SAC-DSE.
Mingjun Cheng, Xin Zheng 0001, Xian Lin, Huaien Gao, Shuting Cai, Xiaoming Xiong, Bei Yu 0001
IEEE Trans. Very Large Scale Integr. Syst.6
2024 Segmentation-aware prior assisted joint global information aggregated 3D building reconstruction
Hongxin Peng, Yongjian Liao, Chuanyu Fu, Ziquan Ding, Qiku Cao, Shuting Cai
Adv. Eng. Informatics9
2024 Designing higher-dimensional digital chaotic systems via reverse derivation of iterative function from strongly connected graph and its application
Shufeng Huang, Qianxue Wang, Xiaoming Xiong, Shuting Cai, Christophe Guyeux
Expert Syst. Appl.4
2024 An Efficient Method of DRC Violation Prediction with a Serial Deep Learning Model
abstract
In VLSI design, the utilization of Design Rule Check (DRC) tools in the early stage is crucial for predicting and resolving violations, thereby expediting the physical design process. In our study, we present an efficient model that predicts DRC violations prior to the routing stage. Additionally, our model incorporates a sliding-window technique to enhance the feature extraction process. We extract structural features using Graph Convolutional Networks and utilize feature reuse techniques to fully recover the lost information in neural layers, which serves as input to the Convolutional Neural Network model, resulting in more accurate hotspot prediction. The experimental results demonstrate that our model successfully identifies 95.78% of DRC violations, with a mere 4.17% false-alarm rate. Not only does our method deliver improved feature preprocessing results, but it also enhances prediction accuracy compared to alternative approaches.
Jingui Lin, Wenxiong Lin, Shiyan Liang, Xiaoming Xiong, Shuting Cai
ACM Trans. Design Autom. Electr. Syst.8
2024 WCPNet: Jointly Predicting Wirelength, Congestion and Power for FPGA Using Multi-Task Learning
abstract
To speed up the design closure and improve the QoR of FPGA, supervised single-task machine learning techniques have been used to predict individual design metric based on placement results. However, the design objective is to achieve optimal performance while considering multiple conflicting metrics. The single-task approaches predict each metric in isolation and neglect the potential correlations or dependencies among them. To address the limitations, this article proposes a multi-task learning approach to jointly predict wirelength, congestion and power. By sharing the common feature representations and adopting the joint optimization strategy, the novel WCPNet models (including WCPNet-HS and WCPNet-SS) cannot only predict the three metrics of different scales simultaneously, but also outperform the majority of single-task models in terms of both prediction performance and time cost, which are demonstrated by the results of the cross design experiment. By adopting the cross-stitch structure in the encoder, WCPNet-SS outperforms WCPNet-HS in prediction performance, but WCPNet-HS is faster because of the simpler parameters sharing structure. The significance of the feature image pinUtilization on predicting power and wirelength are demonstrated by the ablation experiment.
Juming Xian, Shuting Cai, Xiaoming Xiong, Zhengfa Hu
ACM Trans. Design Autom. Electr. Syst.3
2024 FPUx: High-Performance Floating-Point Support for Cost-Constrained RISC-V Cores
abstract
In the Internet of Things (IoT) field, cloud and fog computing dramatically increase the complexity of floating-point (FP) calculations. Cost-constrained microcontrollers (MCUs) urgently need more efficient FP computing methods, such as integrated FP units (FPUs). To this end, this brief proposes FPUx, a high-performance FPU designed through a hybrid pipeline and state-machine approach. The FPUx is integrated into E203 for implementation (E203-FPUx). Furthermore, the Easy-lite is proposed to reduce handshake delay and a range of single-precision FP (FP32) arithmetic IPs are designed to customize FPUs. Compared with E203-FPnew and E203, the performance of E203-FPUx is improved by$1.5\times $and$36\times $, and the total energy consumption is saved by 36% and 1430% on average, respectively.
Xian Lin, Heming Liu, Xin Zheng 0001, Huaien Gao, Shuting Cai, Xiaoming Xiong
IEEE Trans. Very Large Scale Integr. Syst.5
2024 BSSE: Design Space Exploration on the BOOM With Semi-Supervised Learning
abstract
With the rising prominence of RISC-V-based microprocessors in processor design, the challenge of exploring the vast and complex RISC-V microarchitecture design space has become increasingly apparent. We propose the Berkeley Out-of-Order Machine Semi-Supervised Explorer (BSSE)—a novel framework leveraging the semi-supervised learning method and parallel emulation to speed up and make tradeoffs on the RISC-V microarchitecture design space exploration (DSE). BSSE constructs the initial training dataset with the microarchitecture experimental design sampling (MEDS) method and then employs the cotraining-style k-nearest neighbors (Co-KNN) model to fit the microarchitecture features to the architectural metric value space. The trained Co-KNN model assists in searching a Pareto-optimal set with parallel emulation. Finally, a distance-based method is proposed to select a designer-preferred microarchitecture from the identified Pareto-optimal set. Extensive experiments on the Berkeley Out-of-Order Machine (BOOM) show that our proposed BSSE method can search for a better Pareto-optimal set with less time consumption compared to the state-of-the-art methods and can find microarchitectures that are equivalent to or even better than the existing manually designed BOOM microarchitectures.
Xin Zheng 0001, Mingjun Cheng, Jiasong Chen, Huaien Gao, Xiaoming Xiong, Shuting Cai
IEEE Trans. Very Large Scale Integr. Syst.6
2023 An Efficient and Scalable RFID Anti-Collision Algorithm on Optimal Partition and Collided Block Bit-Mapping
abstract
In RFID systems, many anti-collision algorithms, driven by the concept of rescheduling the response sequence between the reader and unidentified tags, have been put forward to solve tag collision problem, including ALOHA-based, tree-based and hybrid algorithms. In this paper, we propose a novel RFID anti-collision algorithm called EAQ-CBB, which adopts three main approaches: tag population estimation based on collided bit detection method, optimal partitions and trimmed query tree based on the strategy of collided block bit-mapping (QTCBB). The relatively accurate estimation of tag backlog and optimal partition ensure a great reduction of collisions in the initial phase. For each collided partition, a QTCBB process is introduced immediately, which eliminates all the empty slots and significantly reduces the collided slots. Simulation results show that EAQ-CBB performs good stability and scalability when the key parameters change. Compared with the existing algorithms, such as DFSA, QTI, T-GDFSA and CT, EAQ-CBB outperforms the others with high system throughput, low normalized latency and low normalized overhead at a low cost of energy, which makes it easier to be used widely in the efficient-aware and energy-aware applications.
Jian Yang 0008, Yonghua Wang 0001, Shuting Cai
Int. J. Pattern Recognit. Artif. Intell.3
2023 Safe-LBP: A visually meaningful image encryption scheme based on LBP and compressive sensing
Zhanwei Yuan, Shufeng Huang, Linqing Huang, Yuxiao Du, Shuting Cai, Xiaoming Xiong
J. Inf. Secur. Appl.5
2023 Sequential Routing-based Time-division Multiplexing Optimization for Multi-FPGA Systems
abstract
Multi-field programming gate array (FPGA) systems are widely used in various circuit design-related areas, such as hardware emulation, virtual prototypes, and chiplet design methodologies. However, a physical resource clash between inter-FPGA signals and I/O pins can create a bottleneck in a multi-FPGA system. Specifically, inter-FPGA signals often outnumber I/O pins in a multi-FPGA system. To solve this problem, time-division multiplexing (TDM) is introduced. However, undue time delay caused by TDM may impair the performance of a multi-FPGA system. Therefore, a more efficient TDM solution is needed. In this work, we propose a new routing sequence strategy to improve the efficiency of TDM. Our strategy consists of two parts: a weighted routing algorithm and TDM assignment optimization. The algorithm takes into account the weight of the net to generate a high-quality routing topology. Then, a net-based TDM assignment is performed to obtain a lower TDM ratio for the multi-FPGA system. Experiments on the public dataset of CAD Contest 2019 at ICCAD showed that our routing sequence strategy achieved good results. Especially in those testcases of unbalanced designs, the performance of multi-FPGA systems was improved up to 2.63. Moreover, we outperformed the top two contest finalists as to TDM results in most of the testcases.
Wenxiong Lin, Wenjun Luo, Shuting Cai, Xiaoming Xiong
ACM Trans. Design Autom. Electr. Syst.5
2022 Novel and secure plaintext-related image encryption algorithm based on compressive sensing and tent-sine system
abstract
Abstract In this paper, a secure plaintext‐related image encryption scheme based on compressive sensing and a tent‐sine system is proposed. First, the discrete wavelet transform (DWT) is used to transform the plain image to get a coefficient matrix. Second, several chaotic sequences generated by the tent‐sine chaotic map are used to scramble the coefficient matrix and construct a measurement matrix. Afterward, compressive sensing is performed on the coefficient matrix to obtain a small‐sized encrypted image. Finally, an image encryption scheme related to plaintext is designed. In particular, the proposed system uses the original image information to participate in the encryption process, ensuring the high sensitivity of the cryptosystem to minor differences in the plain image and good performance on resisting known/selected plaintext attacks. Furthermore, to convey the plaintext‐related parameters to the receiver, the dimension of the ciphertext image is expanded, and the parameters are embedded into the ciphertext image. Simulation results and security analysis show that the proposed image encryption system has strong plaintext sensitivity and robustness for effectively resisting various typical attacks such as brute‐force attacks, statistical attacks, and differential attacks.
Shufeng Huang, Linqing Huang, Shuting Cai, Xiaoming Xiong, Yuan Liu 0022
IET Image Process.3
2022 Research on intelligent slice planning method for free-form surfaces of shaped workpieces
abstract
A surface slicing planning method based on the K-means clustering algorithm and improved fruit fly optimization algorithm (FOA) is proposed to address the efficiency problem in the surface processing of special shaped workpieces. The NURBS surface reconstruction is performed on the complex surface of the selected workpiece, and the K-means clustering algorithm with the curvature-distance factor is used to subdivide the surface, and the FOA with the hopping strategy is used to plan the optimal connection path for each subdivision with different starting points. This improves the local search capability of the FOA by improving the mixing of different classes that occurs when the surface is subdivided. Finally, the proposed method is demonstrated to be effective for surface slicing by simulating the slicing plan on the constructed heterogeneous workpiece surface, and the shortest connectivity paths are obtained for different starting points.
Yuxiao Du, Shuting Cai, Kongyang Chen, Xianghuan Li
Int. J. Intell. Syst.3
2018 The Robotic Fish Strategy Based on The Extreme Learning Machine Optimized by Particle Swarm Optimization Algorithm
abstract
Aiming at URWPGSim2D platform, in order to realize rapid and accurate adjustment of the robotic fish, action decision strategy based on the extreme learning machine optimized by particle swarm algorithm is put forward. According to the current environmental information of robotic fish, use the extreme learning machine to choose the optimal hitting point independently, and to determine the optimal combination of velocity and angular velocity of robotic fish. At the same time, the particle swarm optimization algorithm is introduced to optimize the extreme learning machine, which can improve the accuracy and robustness of the extreme learning machine. Verified by URWPGSim2D platform show that: the robotic fish path can be adjusted according to the strategy, realize combinatorial optimization of speed and direction, and find the target in the shortest time and distance. This shows that action decision-making strategy based on extreme learning machine can fully consider the real-time information of robotic fish and water polo, choose a different strategy in different cases, have a strong ability to adapt, meet the requirements of robotic fish for the action decisions.
Xuexi Zhang, Shuibiao Chen, Zhiguang Cao, Shuting Cai, Zerong Peng, Xiaoming Xiong
ICIS4
2018 Integrate and Conquer: Double-Sided Two-Dimensional k-Means Via Integrating of Projection and Manifold Construction
abstract
In this article, we introduce a novel, general methodology, called integrate and conquer, for simultaneously accomplishing the tasks of feature extraction, manifold construction, and clustering, which is taken to be superior to building a clustering method as a single task. When the proposed novel methodology is used on two-dimensional (2D) data, it naturally induces a new clustering method highly effective on 2D data. Existing clustering algorithms usually need to convert 2D data to vectors in a preprocessing step, which, unfortunately, severely damages 2D spatial information and omits inherent structures and correlations in the original data. The induced new clustering method can overcome the matrix-vectorization-related issues to enhance the clustering performance on 2D matrices. More specifically, the proposed methodology mutually enhances three tasks of finding subspaces, learning manifolds, and constructing data representation in a seamlessly integrated fashion. When used on 2D data, we seek two projection matrices with optimal numbers of directions to project the data into low-rank, noise-mitigated, and the most expressive subspaces, in which manifolds are adaptively updated according to the projections, and new data representation is built with respect to the projected data by accounting for nonlinearity via adaptive manifolds. Consequently, the learned subspaces and manifolds are clean and intrinsic, and the new data representation is discriminative and robust. Extensive experiments have been conducted and the results confirm the effectiveness of the proposed methodology and algorithm.
Chong Peng 0001, Zhao Kang 0001, Shuting Cai, Qiang Shawn Cheng
ACM Trans. Intell. Syst. Technol.3
2015 Image super-resolution via 2D tensor regression learning
Ming Yin 0002, Junbin Gao, Shuting Cai
Comput. Vis. Image Underst.3
2015 Design and ARM-Embedded Implementation of a Chaotic Map-Based Real-Time Secure Video Communication System
abstract
A systematic methodology is proposed for a chaotic map-based real-time video encryption and decryption system with advanced Reduced Instruction Set Computer machine (ARM)-embedded hardware implementation. According to the anticontrol principle of dynamical systems, first, an 8-D discrete-time chaotic map-based system is constructed, which possesses the required property of 1-1 surjection in the integer range$[{0,\,N-1}]$, where$N$is the number of frame pixels, suitable for position scrambling of each video frame. Then, an 8-D discrete-time hyperchaotic system is designed for encryption–decryption of red, green, and blue (RGB) tricolor pixel values. Using the ARM-embedded platform super4412 model with Cortex-A9 processor, together with the standard QT cross-platform, an integrated chaotic map-based real-time secure video communication system is designed, implemented, and evaluated. In addition, the security performance of the designed system is tested using criteria from the National Institute of Standards and Technology statistical test suite. The main feature of this method is that, both scrambling–antiscrambling of RGB tricolor pixel positions and encryption–decryption of pixel values are realized simultaneously for enhancing the security. As is well known, compared with numerical simulations, hardware implementation for such a secure video communication system is very difficult to achieve, but we successfully implemented and tested in a real-world network environment. Both theoretical analysis and experimental results validate the feasibility and real-time performance of the new secure video communication system.
Zhuosheng Lin, Simin Yu, Jinhu Lü 0001, Shuting Cai, Guanrong Chen
IEEE Trans. Circuits Syst. Video Technol.4
2014 Blocky artifact removal with low-rank matrix recovery
abstract
In this paper, a novel image blocky artifact removal scheme based on low-rank matrix recovery is proposed. The problem of suppressing blocky artifacts is formulated as recovering a low-rank matrix from corrupted observations. During the deblocking processing, we do not directly recover the whole clean image but only its high-frequency component and then synthesize the clean image by incorporating the low-frequency component of blocky image. To take advantage of the low-rank matrix recovery paradigm, we first cluster the similar patches of the high-frequency component of image via local pixel clustering, then the clean high-frequency component of image is recovered by formulating an optimization problem of the nuclear norm and ℓ1-norm. The experimental results show that the proposed algorithm can achieve competitive performance in terms of both quantitative and subjective quality.
Ming Yin 0002, Junbin Gao, Shuting Cai
ICASSP4
2014 Blind image deblurring via coupled sparse representation
Ming Yin 0002, Junbin Gao, David Tien, Shuting Cai
J. Vis. Commun. Image Represent.4
2013 Robust face recognition via double low-rank matrix recovery for feature extraction
abstract
Feature extraction is one of the most fundamental problems in face recognition tasks. In this paper, motivated by low-rank representation (LRR) model on exploring the multiple subspace structures of observation data, we propose a double low-rank matrix recovery method to learn low-rank subspaces from face images, where it takes into account the recovery of row space and column space information simultaneously. Applying Augmented Lagrangian Multiplier (ALM), the optimization problem on minimization of nuclear norm is resolved efficiently. By evaluating on public face databases, experimental results show that our proposed method works much better than existing face recognition methods based on feature extraction. It is more robust to outliers, varying illumination and occlusion.
Ming Yin 0002, Shuting Cai, Junbin Gao
ICIP2