EDBT 2026 Demo / reviewers in the wild / expert
Xiaoming Xiong
dblp:54/5449
· DBLP profile ↗
38ranked-venue papers
0as first author
35since 2021 · last 2026
0000-0002-2421-7621ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 29 · 28 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Optimizing detailed-routability for 3D global routing through dynamic resource model and routability-aware cost scheme
Juntao Jian, Shuting Cai, Xiaoming Xiong |
Integr. | 5 |
| 2026 | Routability-wirelength co-guided cell inflation with explainable multi-task learning for global placement optimization
Zicheng Deng, Shuting Cai, Xiaoming Xiong |
Integr. | 5 |
| 2026 | A novel recurrent topology-based memristor: Simplification and circuit simulation for multi-scroll attractor generation
Yifeng Diao, Shufeng Huang, Xiaoming Xiong, Shuting Cai |
Inf. Sci. | 5 |
| 2026 | High-Efficiency Bidirectional Translator between SystemC and VerilogabstractThe SystemC language, with its higher level of abstraction, plays a critical role in facilitating hardware/software co-design and architecture exploration. However, as most hardware models are predominantly written in Verilog and translating between SystemC and Verilog remains a challenge, an efficient and reliable tool for translating between these two languages is essential to streamline system development. This article proposes SCAV, a bidirectional translator between SystemC and Verilog, which breaks these limitations. SCAV provides a fully automated solution for translating both SystemC to Verilog and Verilog to SystemC, leveraging a translation framework with front-end/back-end separation. Additionally, SCAV incorporates an Abstract Syntax Tree (AST) filter, optimizing the translation process by filtering out invalid content. The experimental results demonstrate that SCAV achieves a 100% adaptation rate for Verilog and a 98% adaptation rate for SystemC, with 100% accuracy in both directions. Furthermore, SCAV outperforms existing tools, delivering a minimum speedup of 18% across various test cases. Xin Zheng 0001, Yongfeng Zhong, Shaofen Zeng, Huaien Gao, Shuting Cai, Xiaoming Xiong |
ACM J. Emerg. Technol. Comput. Syst. | 7 |
| 2026 | An efficient DSP packing framework for FPGA-based mixed-precision DCNN processor
Xueming Li 0001, Jinhui Pan, Hongmin Huang, Yuanmiao Lin, Xianghong Hu 0001, Shuting Cai, Xiaoming Xiong |
J. Syst. Archit. | 7 |
| 2026 | NTT-LSU: Tightly Coupled Architecture for Efficient NTT Implementation on RISC-V ProcessorabstractPolynomial multiplication is one of the most computationally intensive operations in lattice-based cryptographic systems, directly impacting overall computational efficiency. Although the Number Theoretic Transform (NTT) reduces the time complexity of polynomial multiplication fromO(n2) toO(nlogn), the varying parameter requirements of different lattice algorithms limit the generality of hardware designs. To address this issue, we propose an innovative hardware architecture that tightly couples the NTT unit with the Load and Store Unit (LSU) in the pipeline of a RISC-V processor. We also propose a hybrid width data path method that effectively reduces data transfer time. Compared to previous designs, our architecture minimizes data transfer latency while enhancing computational flexibility and scalability. Specifically, we have customized NTT-related instructions to support Inverse Number Theoretic Transform (INTT) and various NTT parameter configurations. Experimental results demonstrate that our solution significantly shortens the NTT computation cycle, achieving over 10× speedup compared to software implementations. In comparison to existing solutions, our architecture exhibits superior area-time product (ATP) performance. Yinqiao Zhao, Zilong Xie, Ruidian Zhan, Xiaoming Xiong, Yun Chen 0004, Shuting Cai |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2026 | An FPGA-Efficient CNN Accelerator for Hybrid Model Compression With Scalable Bit-Serial and Bit-Parallel MACabstractMixed-precision quantization and unstructured pruning have emerged as two effective compression techniques, demonstrating great potential in reducing model size and computational cost in the deployment of convolutional neural networks (CNNs). However, their joint deployment still faces two major challenges: 1) The former introduces heterogeneous bit-widths in the bit-level, while the latter results in irregular sparsity in the value-level; their fundamentally incompatible data representations and computation patterns require two distinct types of hardware overhead to process them separately, which severely limits hardware execution efficiency. 2) Jointly applying both techniques often leads to notable accuracy degradation. In this paper, we propose a novel compression perspective that reinterprets zero-values generated by unstructured pruning as multiple consecutive 0-bits. We further introduce column-based bit-level sparsity, which provides a unified representation for weights after mixed-precision quantization and unstructured pruning, requiring only a single type of hardware overhead. Based on these techniques, we develop a hybrid compression framework that jointly optimizes model size, accuracy, and hardware implementation. Our method achieves weight/activation precision of 2.13b/4.06b on VGG16, delivering 7.40$\times $compression and 2.74$\times $speedup with 0.92% accuracy loss compared to the 8b baseline. Compared to state-of-the-art accelerators, our design achieves 1.12$\times $-6.23$\times $and 1.31$\times $-6.60$\times $improvements in energy efficiency and LUT efficiency when deploying VGG16 and ResNet50. Yuanmiao Lin, Xueming Li 0001, Hongmin Huang, Heng Mai, Ruidian Zhan, Xianghong Hu 0001, Shuting Cai, Xiaoming Xiong |
IEEE Trans. Circuits Syst. I Regul. Pap. | 9 |
| 2026 | FPUltra: An Area-Efficient Single-Precision Floating-Point Unit for Cost-Sensitive RISC-V CoresabstractArea efficiency is vital for floating-point units (FPUs) in resource-constrained IoT devices. However, existing designs suffer from rigid architectures and costly arithmetic units, limiting performance-area optimization. To this end, this work presents FPUltra, an area-efficient single-precision FPU for cost-sensitive RISC-V cores. FPUltra adopts a novel phase-decoupled control architecture to mitigate timing hazards and improve execution efficiency. A parallel approximate floating-point multiplier (FPM) is designed using combinational logic, based on the Mitchell algorithm with error compensation. A Newton–Raphson-based subinstruction decomposition method is presented to support floating-point division (Fdiv) and square root (Fsqrt). Compared with state-of-the-art FPUs, FPUltra achieves 9%–695% and 101%–14 186% improvements in equivalent slices efficiency (Eq.Slices Eff.) on FPGA and equivalent area efficiency (Eq.Area Eff.) on ASIC, respectively. Our code will be available athttps://github.com/LX-IC/FPUltra Xian Lin, Jiahao Lan, Xin Zheng 0001, Huanxin Zhuang, Huaien Gao, Shuting Cai, Xiaoming Xiong |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2026 | Efficient FPGA Acceleration for 4-bit CNNs via Quantization-Induced Structured Sparsity and LUT-Based MultiplicationabstractN:M structured sparsity is key to convolutional neural network (CNN) compression and acceleration, but two challenges remain. From the algorithm perspective, prior works have mainly focused on 8-bit quantized models, where N:M sparsity yields limited hardware efficiency. From the hardware perspective, the cost differences across N:M sparsity have not been analyzed. To address these issues, we present a unified algorithm–hardware co-design framework for 4-bit CNN acceleration. We show that 4-bit quantization induces over 80% zero weights and strongly structured sparsity, with over 95% of weight groups satisfying 4:8, 8:16, or 16:32 patterns. We propose a pruning-after-quantization (PAQ) algorithm that enforces strict N:M sparsity with minimal accuracy loss. We also analyze the hardware overhead of activation fetch units (AFUs) under different N:M sparsity patterns (4:8, 8:16, 16:32), revealing that the 4:8 AFU reduces look-up table (LUT) cost by up to 66.7% compared to 16:32. Finally, we introduce a 4-bit LUT-based sign-magnitude multiplier (LBSMM) requiring only 11 LUT6 resources, outperforming existing multipliers. Integrated on a Xilinx VCU118 field-programmable gate array (FPGA), our accelerator achieves$2.51\times $–$12.89\times $improvements in equivalent LUT efficiency over SOTA designs. The implementations of the PAQ algorithm and the RTL of LBSMM are available athttps://github.com/haden-01/PAQ-and-LBSMM.git Yuanmiao Lin, Xueming Li 0001, Hongmin Huang, Ruidian Zhan, Xianghong Hu 0001, Shuting Cai, Xiaoming Xiong |
IEEE Trans. Very Large Scale Integr. Syst. | 9 |
| 2025 | Framework design and empirical analysis of intelligent scheduling system for high-altitude photovoltaic power generation based on mixed optimization of long-nosed raccoon optimization algorithm and black winged kite optimization algorithm (COA-BKA)abstractThis study proposes an intelligent scheduling system for high-altitude photovoltaic power generation, utilizing a hybrid optimization approach that combines the Long-nosed Raccoon Optimization Algorithm (COA) and the Black-winged Kite Optimization Algorithm (BKA) (COA-BKA). The goal is to enhance scheduling accuracy, stability, and response speed under the unique environmental conditions of high-altitude regions, such as fluctuating light intensity, extreme temperatures, and dynamic load demands. In experimental comparisons with traditional algorithms like Particle Swarm Optimization (PSO) and Genetic Algorithm (GA), COA-BKA achieved a scheduling accuracy of 0.98, outperforming PSO (0.92) and GA (0.90). COA-BKA also demonstrated superior convergence speed, reaching the optimal solution by the 50th iteration, while PSO and GA required more iterations (80 and 100, respectively). Additionally, COA-BKA completed scheduling in just 4.5 s, significantly faster than PSO (6.3 s) and GA (7.2 s). The system effectively handled fluctuating light intensity and load demand changes, showcasing its robust adaptability. These results suggest that COA-BKA provides a highly efficient and stable solution for intelligent scheduling in high-altitude photovoltaic power systems, improving operational efficiency and reducing costs, while offering significant advancements for real-time optimization in smart grids. Heng Hu, Xiaoming Xiong, Taidong Yan, Yuancheng Zhang, Shuang Gan |
Discov. Comput. | 2 |
| 2025 | An FPGA-based bit-level weight sparsity and mixed-bit accelerator for neural networks
Xianghong Hu 0001, Shansen Fu, Yuanmiao Lin, Xueming Li 0001, Chaoming Yang, Rongfeng Li 0001, Hongmin Huang, Shuting Cai, Xiaoming Xiong |
J. Syst. Archit. | 9 |
| 2025 | FLALM: A Flexible Low Area-Latency Montgomery Modular Multiplication on FPGAabstractMontgomery Modular Multiplication (MMM) is widely used in many public key cryptography systems. This paper presents a Flexible Low Area-Latency MMM (FLALM) implementation, which supports Generic Montgomery Modular Multiplication (GMM) and Square Montgomery Modular Multiplication (SMM) operations. A new SMM schedule for the Finely Integrated Product Scanning (FIPS) GMM algorithm is proposed to accelerate SMM with tiny additional design. Furthermore, a new FIPS dual-schedule is proposed to solve the data hazards of this algorithm. Finally, we explore the trade-off between area and latency, and present the FLALM to accelerate GMM and SMM. The FLALM is implemented on FPGA (Virtex-7 platform). The results show that the area*latency (AL) value of FLALM (wordsize$w$=128) is 38.1% and 44.7% better than the previous state-of-art scalable references when performing 1024-bit and 2048-bit GMM, respectively. Moreover, when computing SMM, the advantage of AL value is raised to 73.7% and 86.3% respectively. Yujun Xie 0001, Yuan Liu 0022, Xin Zheng 0001, Bohan Lan, Dengyun Lei, Dehao Xiang, Shuting Cai, Xiaoming Xiong |
IEEE Trans. Computers | 8 |
| 2025 | GNN-Based Timing Prediction in Prerouting Stage With Multitask Learning StrategyabstractStatic timing analysis tools are commonly used to evaluate timing performance and guide optimization during placement stage. However, traditional timing analysis struggles to fast and accurately evaluate timing violation due to absence of detailed routing information necessary for RC parasitic parameter extraction. Therefore, a timing analyzer based on graph neural network is proposed in this article. Compared to previous works, a novel representation of circuit delay model is proposed in this article, employing timing arcs and virtual pins to predict net delay and arrival time (AT) in the prerouting phase. Additionally, to our knowledge, this is the first attempt to improve the quality of timing analyzer through a strategy of multitask learning, with the proposed enhanced dynamic weight average method. The experimental results demonstrate that our model excels in predicting net delay and AT, with average correlations of 0.9540 and 0.9058, respectively, on the testing set. In comparison to the previous state-of-the-art methods, our approach maintains accuracy in net delay prediction while enhancing the overall$R^{2}$score for AT prediction by 0.0271. Additionally, our method reduces inference time by 25.8%. Haisen Zhang, Xiaoming Xiong, Shuting Cai |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2025 | A Lightweight Heterogeneous Graph Embedding Framework for Hotspot DetectionabstractHotspot detection is a crucial step in ensuring the manufacturability of integrated circuits, as it seeks to identify potential defects in the layout. Pattern matching methods have been widely used to accelerate the detection of these defects. However, they often struggle with complicated deviations. Image-based machine learning methods were introduced to confront this challenge, but they often involved distorted information extraction and incurred significant runtime overhead. In this article, we introduce a novel detection framework based on the modified transitive closure graph (MTCG). By applying the concept of MTCG, the layout can be accurately modeled as a heterograph. The embeddings of these heterographs are extracted using an optimized lightweight 3-hop message-passing graph neural network (GNN) and subsequently utilized for classification. Furthermore, a dynamic edge transformation method based on the properties of MTCG is proposed for data augmentation. The proposed method is evaluated with datasets from ICCAD 2012 and ICCAD 2019, demonstrating outstanding performance in recall and false alarm, along with a significantly decreased inference time. Haopeng Yan, Yuzhe Ma, Xiaoming Xiong, Shuting Cai |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2025 | A Precision-Scalable Accelerator with Sign-Magnitude Representation and Dual Adder TreesabstractCurrently, there are two mainstream acceleration methods; one is mixed precision and the other is sparsity. Few accelerators support both mixed precision and sparsity, and most enable precision configurations across layers rather than within a single layer. Furthermore, most of accelerators adopt the traditional two’s complement (2C) data representation method, and we found that 2C brings many invalid ”1” when representing signed data, which brings more resources overhead for mixed precision and many invalid operations for bit-level sparsity. Therefore, we propose a high-efficiency accelerator featuring a precision-scalable Sign-Magnitude Processing Element (SM-PE), which adopts a data representation method of SM and can flexibly support various precision calculations (2, 4, 8 bits) and bit-level sparsity. In addition, a dynamic quantization algorithm named DoReFaLike and a bit-level column sparsity (BLCS) technique are proposed to improve the efficiency of SM-PEs. Under the same accuracy constraint, the sparsity rate of the SM scheme is 3.5× higher than that of the 2C format. The accelerator has been synthesized on a 55nm CMOS ASIC platform. When scaled to 28nm, experimental results show that the energy efficiency of the proposed accelerator reaches 15.50, 25.37, 101.54 TOPS/W with 8-bit, 4-bit, and 2-bit input activations, respectively, and weights represented in sparse 8-bit precision, operating at 400 MHz. Compared to state-of-the-art accelerators, the proposed design achieves a performance improvement of 1.1× to 3.9×. Xianghong Hu 0001, Chaoming Yang, Xueming Li 0001, Rongfeng Li 0001, Yuanmiao Lin, Shansen Fu, Hongmin Huang, Shuting Cai, Xiaoming Xiong |
ACM Trans. Embed. Comput. Syst. | 9 |
| 2025 | RTMF: Routing based on TDM for Multi-FPGA SystemabstractAs modern VLSI design advances, the significance of multi-FPGA systems in prototyping and verification is steadily growing. Due to the physical I/O limitations, the Time-Division Multiplexing (TDM) and I/O assignment techniques are introduced to solve these problems. However, most multi-FPGA systems primarily focus on inter-FPGA routing while overlooking intra-FPGA routing. In this article, a comprehensive routing framework based on TDM for Multi-FPGA systems (RTMF) is hereby presented. To our knowledge, this is the first attempt to jointly optimize intra-level and inter-level routing in the work of multi-FPGA systems (MFS). The RTMF framework, tailored for system-level and intra-level routing under constrained wiring resources, integrates routing demands within and between FPGAs. Through the integration of TDM technology and adaptive optimization algorithms, RTMF effectively meets routing requirements and delivers efficient solutions. Furthermore, RTMF demonstrates remarkable adaptability, allowing for dynamic adjustments and optimizations to address diverse routing demands and constraints. In comparison to the state-of-the-art methodologies, in benchmark designs with a scale greater than 50,000, our approach on average reduces the maximum routing weight by 59.98% and 46.70%, respectively. Shiyan Liang, Jingui Lin, Wenxiong Lin, Yuzhe Ma, Xiaoming Xiong, Shuting Cai |
ACM Trans. Design Autom. Electr. Syst. | 8 |
| 2025 | Early Stage DRC Hotspot Prediction for Mixed-Size Designs Through an Efficient Graph-Based Deep LearningabstractPredicting hotspot locations in the early stage of Design Rule Check (DRC) is crucial for designers to proactively prevent design rule violations. However, obtaining an accurate and efficient predictor faces significant challenges due to the influence of available information and severe data imbalance. In this study, we investigate the potential of utilizing Graph Neural networks (GNN) to address this challenge. Our focus is specifically on accurately predicting DRC hotspot locations without relying on global routing techniques. We consider the presence of macros in mixed-size designs. We propose an adaptive adjacency matrix that demonstrates superior application effectiveness compared with traditional adjacency matrices. Furthermore, experimental results on benchmark circuits show significant improvements in the true positive rate (22.38% for the RouteNet model and 26.90% for the GNN model) and accuracy (6.97% and 6.76%, respectively) compared with these models. Our proposed model also maintains a low false positive rate and outperforms other Convolutional Neural Network and GNN models. Additionally, its efficient learning capability and lower computational time contribute to its outstanding training performance, with training time being approximately 10% of that required by other models. Jingui Lin, Shiyan Liang, Wenxiong Lin, Xiaoming Xiong, Shuting Cai |
ACM Trans. Design Autom. Electr. Syst. | 7 |
| 2025 | Optimizing FPGA Routing with Explainable Co-Learning of Congestion and WirelengthabstractIn FPGA routing, machine learning-based optimization methods have achieved improved routing solutions by integrating traditional heuristics with predictive capabilities. However, these approaches mostly relied on single-task learning models with black-box nature and often neglected the complex trade-offs and inter-dependencies between routing metrics. To address these limitations, this paper introduces a novel multi-task learning-based routing optimization method. In the congestion-wirelength co-learning stage, the simultaneous prediction of congestion and wirelength is formulated as a multi-task learning problem. A multi-task learning model, named CWNet, is proposed to tackle this challenge effectively. During the congestion-wirelength impact interpretation, the contribution of congestion to wirelength is quantified using an XAI technique known as DeepSHAP, producing a congestion-wirelength impact map. In the congestion-wirelength co-guided routing optimization (CWRO) stage, the VTR router’s lookahead map is enhanced based on the impact map, guiding the router to avoid locations where congestion significantly affect wirelength. Experimental results demonstrate that CWNet outperforms most baseline learning models in terms of both prediction performance and computational efficiency. Additionally, the impact map visually illustrates the complex and nonlinear relationship between congestion and wirelength. Ultimately, CWRO significantly reduces congestion, wirelength, and critical path delay, while maintaining a competitive runtime compared to baseline routers. Shuting Cai, Xiaoming Xiong |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2025 | Layout Congestion Prediction Based on Regression-ViTabstractTo accelerate the back-end design flow of integrated circuit (IC), numerous studies have made exploratory advancements in machine learning (ML) for electronic design automation (EDA). However, most research works are limited to deep learning (DL) models predominantly based on convolutional neural networks, and the models often suffer from poor generalization due to the scarcity of data. In this study, we propose the Double generative adversarial networks (D-GAN) model to enrich the dataset and propose the Regression Vision Transformer (R-ViT) model to predict layout congestion information. Compared with the baseline model, experimental results show improvements of 3.03% and 2.64% in Receiver Operating Characteristic-Area under Curve (ROC-AUC) and Precision-Recall Curve-Area under Curve (PRC-AUC) respectively. To further enhance the prediction accuracy of the model, an adaptive Huber loss function is designed to optimize the training process, resulting in an improvement of up to 11.03% in ROC-AUC compared with the baseline model. Lastly, extended experiments are conducted to study the effects of parameters and convolutional kernel size on performance, which find a better configuration. Guiqi Mo, Yimin Xia, Jianhong Ou, Shuting Cai, Xiaoming Xiong |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2025 | PSCaps: High-Performance Pose-Sensitive Layout Hotspot Detector based on CapsNetabstractAdvanced technology nodes face challenges with Design Rule Violations (DRVs), primarily due to the possibility of nm-level small variations that can lead to the occurrence of DRVs. Various Machine Learning (ML) techniques have been introduced to detect whether the layout design conforms to manufacturing rules, thus alleviating the time-consuming challenge associated with traditional lithography simulations. However, existing ML models still face challenges in detecting layout violations where there are pose variations among geometric shapes in the layout. In this study, we propose a hotspot detector called PSCaps based on the CapsNet, which considers the pose information of geometric shapes in the layout. The method effectively captures spatial information and hierarchical structures between geometric shapes in the layout. Through the dynamic routing mechanism, the model adaptively learns the relationships and weight allocations between different capsules. Additionally, we employ multiple data augmentation methods to alleviate the problem of imbalanced hotspot and non-hotspot data in the open-source dataset. The benchmarks of ICCAD-2012 and ICCAD-2019 are used to validate our method. The experimental results demonstrate that our proposed hotspot detector outperforms other state-of-the-art works. Haopeng Yan, Xiaoming Xiong, Shuting Cai |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2025 | Concurrent Prediction of Timing and wire Length Using A Multi-Task Graph Neural NetworkabstractTraditional supervised single-task learning models are used in timing-driven placement exploration to improve both effectiveness and efficiency by predicting wire length, wire delay, and cell delay separately. However, these metrics are interdependent, with the two delays being timing-based and wire length non-timing, which makes it difficult for single-task models to capture their complex relationships. Moreover, the limited existing multi-task learning methods can only predict either multiple timing or non-timing metrics. To address these limitations, this article introduces DLGNN, a novel multi-task graph learning model that simultaneously predicts these three metrics through an embedder-predictor architecture featuring two residual connections, a combination of both soft and hard parameter sharing, and a geometric loss strategy. Cross-design experimental results on the Nangate 45nm library demonstrate that DLGNN outperforms baseline models in terms of both predictive performance and time efficiency. Additionally, ablation studies emphasize the critical roles of the residual connections, the combination of soft and hard parameter sharing, and the geometric loss strategy in improving DLGNN’s predictive performance. The generalization experiment on the ASAP 7nm library further confirms DLGNN’s advantages for more advanced technology nodes. Shuting Cai, Xiaoming Xiong |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2025 | Efficient Design Space Exploration for the BOOM Using SAC-Based Reinforcement LearningabstractDesign space exploration (DSE) is crucial for optimizing the performance, power, and area (PPA) of CPU microarchitectures ($\mu $-archs). While various machine learning (ML) algorithms have been applied to the$\mu $-arch DSE problem, the potential of reinforcement learning (RL) remains underexplored. In this article, we propose a novel RL-based approach to address the reduced instruction set computer V (RISC-V) CPU$\mu $-arch DSE problem. This approach enables dynamic selection and optimization of$\mu $-arch parameters without relying on predefined modification sequences, thus significantly enhancing exploration flexibility. To address the challenges posed by high-dimensional action spaces and sparse rewards, we use a discrete soft actor-critic (SAC) framework with entropy maximization to promote efficient exploration. In addition, we integrate multistep temporal-difference (TD) learning, an experience replay (ER) buffer, and return normalization to improve sample efficiency and learning stability during training. Our method further aligns optimization with user-defined preferences by normalizing PPA metrics relative to baseline designs. Experimental results on the Berkeley out-of-order machine (BOOM) demonstrate that the proposed approach achieves superior performance compared with state-of-the-art methods, showcasing its effectiveness and efficiency for$\mu $-arch DSE. Our code is available athttps://github.com/exhaust-create/SAC-DSE. Mingjun Cheng, Xin Zheng 0001, Xian Lin, Huaien Gao, Shuting Cai, Xiaoming Xiong, Bei Yu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2024 | Quantization-aware Optimization Approach for CNNs Inference on CPUsabstractData movements through the memory hierarchy are a fundamental bottleneck in the majority of convolutional neural network (CNN) deployments on CPUs. Loop-level optimization and hybrid bitwidth quantization are two representative optimization approaches for memory access reduction. However, they were carried out independently because of the significantly increased complexity of design space exploration. We present QAOpt, a quantization-aware optimization approach that can reduce the high complexity when combining both for CNN deployments on CPUs. We develop a bitwidth-sensitive quantization strategy that can perform the trade-off between model accuracy and data movements when deploying both loop-level optimization and mixed precision quantization. Also, we provide a quantization-aware pruning process that can reduce the design space for high efficiency. Evaluation results demonstrate that our work can achieve better energy efficiency under acceptable accuracy loss. Jiasong Chen, Zeming Xie, Weipeng Liang, Bosheng Liu, Xin Zheng 0001, Jigang Wu, Xiaoming Xiong |
ASPDAC | 7 |
| 2024 | iEDA: An Open-source infrastructure of EDAabstractBy leveraging the power of open-source software, the EDA tool offers a cost-effective and flexible solution for designers, researchers, and hobbyists alike. Open-source EDA promotes collaboration, innovation, and knowledge sharing within the EDA community. It emphasizes the role of the toolchain in accelerating the development of electronic systems, reducing design costs, and improving design quality. This paper presents an open-source EDA project, iEDA, aiming to build a basic infrastructure for EDA technology evolution and closing the industrial-academic gap in the EDA area. As the foundation for developing EDA tools and researching EDA algorithms and technologies, iEDA is mainly composed of file system, database, manager, operator and interface. To demonstrate the effectiveness of iEDA, we implement and tape out four chips of different scales (from 700k to 500M gates) on different process nodes (110nm and 28nm) with iEDA. iEDA is publicly available on the project home page https://github.com/OSCC-Project/iEDA. Zengrong Huang, Simin Tao, Zhipeng Huang 0009, Chunan Zhuang, Yihang Qiu, Guojie Luo, Huawei Li 0001, Haihua Shen, Mingyu Chen 0001, Dongbo Bu, Wenxing Zhu, Ye Cai 0001, Xiaoming Xiong, Yi Heng, Peng Zhang 0007, Bei Yu 0001, Biwei Xie, Yungang Bao |
ASPDAC | 16 |
| 2024 | Designing higher-dimensional digital chaotic systems via reverse derivation of iterative function from strongly connected graph and its application
Shufeng Huang, Qianxue Wang, Xiaoming Xiong, Shuting Cai, Christophe Guyeux |
Expert Syst. Appl. | 3 |
| 2024 | An Efficient Method of DRC Violation Prediction with a Serial Deep Learning ModelabstractIn VLSI design, the utilization of Design Rule Check (DRC) tools in the early stage is crucial for predicting and resolving violations, thereby expediting the physical design process. In our study, we present an efficient model that predicts DRC violations prior to the routing stage. Additionally, our model incorporates a sliding-window technique to enhance the feature extraction process. We extract structural features using Graph Convolutional Networks and utilize feature reuse techniques to fully recover the lost information in neural layers, which serves as input to the Convolutional Neural Network model, resulting in more accurate hotspot prediction. The experimental results demonstrate that our model successfully identifies 95.78% of DRC violations, with a mere 4.17% false-alarm rate. Not only does our method deliver improved feature preprocessing results, but it also enhances prediction accuracy compared to alternative approaches. Jingui Lin, Wenxiong Lin, Shiyan Liang, Xiaoming Xiong, Shuting Cai |
ACM Trans. Design Autom. Electr. Syst. | 7 |
| 2024 | WCPNet: Jointly Predicting Wirelength, Congestion and Power for FPGA Using Multi-Task LearningabstractTo speed up the design closure and improve the QoR of FPGA, supervised single-task machine learning techniques have been used to predict individual design metric based on placement results. However, the design objective is to achieve optimal performance while considering multiple conflicting metrics. The single-task approaches predict each metric in isolation and neglect the potential correlations or dependencies among them. To address the limitations, this article proposes a multi-task learning approach to jointly predict wirelength, congestion and power. By sharing the common feature representations and adopting the joint optimization strategy, the novel WCPNet models (including WCPNet-HS and WCPNet-SS) cannot only predict the three metrics of different scales simultaneously, but also outperform the majority of single-task models in terms of both prediction performance and time cost, which are demonstrated by the results of the cross design experiment. By adopting the cross-stitch structure in the encoder, WCPNet-SS outperforms WCPNet-HS in prediction performance, but WCPNet-HS is faster because of the simpler parameters sharing structure. The significance of the feature image pinUtilization on predicting power and wirelength are demonstrated by the ablation experiment. Juming Xian, Shuting Cai, Xiaoming Xiong, Zhengfa Hu |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2024 | FPUx: High-Performance Floating-Point Support for Cost-Constrained RISC-V CoresabstractIn the Internet of Things (IoT) field, cloud and fog computing dramatically increase the complexity of floating-point (FP) calculations. Cost-constrained microcontrollers (MCUs) urgently need more efficient FP computing methods, such as integrated FP units (FPUs). To this end, this brief proposes FPUx, a high-performance FPU designed through a hybrid pipeline and state-machine approach. The FPUx is integrated into E203 for implementation (E203-FPUx). Furthermore, the Easy-lite is proposed to reduce handshake delay and a range of single-precision FP (FP32) arithmetic IPs are designed to customize FPUs. Compared with E203-FPnew and E203, the performance of E203-FPUx is improved by$1.5\times $and$36\times $, and the total energy consumption is saved by 36% and 1430% on average, respectively. Xian Lin, Heming Liu, Xin Zheng 0001, Huaien Gao, Shuting Cai, Xiaoming Xiong |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2024 | BSSE: Design Space Exploration on the BOOM With Semi-Supervised LearningabstractWith the rising prominence of RISC-V-based microprocessors in processor design, the challenge of exploring the vast and complex RISC-V microarchitecture design space has become increasingly apparent. We propose the Berkeley Out-of-Order Machine Semi-Supervised Explorer (BSSE)—a novel framework leveraging the semi-supervised learning method and parallel emulation to speed up and make tradeoffs on the RISC-V microarchitecture design space exploration (DSE). BSSE constructs the initial training dataset with the microarchitecture experimental design sampling (MEDS) method and then employs the cotraining-style k-nearest neighbors (Co-KNN) model to fit the microarchitecture features to the architectural metric value space. The trained Co-KNN model assists in searching a Pareto-optimal set with parallel emulation. Finally, a distance-based method is proposed to select a designer-preferred microarchitecture from the identified Pareto-optimal set. Extensive experiments on the Berkeley Out-of-Order Machine (BOOM) show that our proposed BSSE method can search for a better Pareto-optimal set with less time consumption compared to the state-of-the-art methods and can find microarchitectures that are equivalent to or even better than the existing manually designed BOOM microarchitectures. Xin Zheng 0001, Mingjun Cheng, Jiasong Chen, Huaien Gao, Xiaoming Xiong, Shuting Cai |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2023 | Safe-LBP: A visually meaningful image encryption scheme based on LBP and compressive sensing
Zhanwei Yuan, Shufeng Huang, Linqing Huang, Yuxiao Du, Shuting Cai, Xiaoming Xiong |
J. Inf. Secur. Appl. | 6 |
| 2023 | High-performance Reconfigurable DNN Accelerator on a Bandwidth-limited Embedded SystemabstractDeep convolutional neural networks (DNNs) have been widely used in many applications, particularly in machine vision. It is challenging to accelerate DNNs on embedded systems because real-world machine vision applications should reserve a lot of external memory bandwidth for other tasks, such as video capture and display, while leaving little bandwidth for accelerating DNNs. In order to solve this issue, in this study, we propose a high-throughput accelerator, called reconfigurable tiny neural network accelerator (ReTiNNA), for the bandwidth-limited system and present a real-time object detection system for the high-resolution video image. We first present a dedicated computation engine that takes different data mapping methods for various filter types to improve data reuse and reduce hardware resources. We then propose an adaptive layer-wise tiling strategy that tiles the feature maps into strips to reduce the control complexity of data transmission dramatically and to improve the efficiency of data transmission. Finally, a design space exploration (DSE) approach is presented to explore design space more accurately in the case of insufficient bandwidth to improve the performance of the low-bandwidth accelerator. With a low bandwidth of 2.23 GB/s and a low hardware consumption of 90.261K LUTs and 448 DSPs, ReTiNNA can still achieve a high performance of 155.86 GOPS on VGG16 and 68.20 GOPS on ResNet50, which is better than other state-of-the-art designs implemented on FPGA devices. Furthermore, the real-time object detection system can achieve a high object detection speed of 19 fps for high-resolution video. Xianghong Hu 0001, Hongmin Huang, Xueming Li 0001, Xin Zheng 0001, Qinyuan Ren, Jingyu He, Xiaoming Xiong |
ACM Trans. Embed. Comput. Syst. | 7 |
| 2023 | Sequential Routing-based Time-division Multiplexing Optimization for Multi-FPGA SystemsabstractMulti-field programming gate array (FPGA) systems are widely used in various circuit design-related areas, such as hardware emulation, virtual prototypes, and chiplet design methodologies. However, a physical resource clash between inter-FPGA signals and I/O pins can create a bottleneck in a multi-FPGA system. Specifically, inter-FPGA signals often outnumber I/O pins in a multi-FPGA system. To solve this problem, time-division multiplexing (TDM) is introduced. However, undue time delay caused by TDM may impair the performance of a multi-FPGA system. Therefore, a more efficient TDM solution is needed. In this work, we propose a new routing sequence strategy to improve the efficiency of TDM. Our strategy consists of two parts: a weighted routing algorithm and TDM assignment optimization. The algorithm takes into account the weight of the net to generate a high-quality routing topology. Then, a net-based TDM assignment is performed to obtain a lower TDM ratio for the multi-FPGA system. Experiments on the public dataset of CAD Contest 2019 at ICCAD showed that our routing sequence strategy achieved good results. Especially in those testcases of unbalanced designs, the performance of multi-FPGA systems was improved up to 2.63. Moreover, we outperformed the top two contest finalists as to TDM results in most of the testcases. Wenxiong Lin, Wenjun Luo, Shuting Cai, Xiaoming Xiong |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2022 | Novel and secure plaintext-related image encryption algorithm based on compressive sensing and tent-sine systemabstractAbstract In this paper, a secure plaintext‐related image encryption scheme based on compressive sensing and a tent‐sine system is proposed. First, the discrete wavelet transform (DWT) is used to transform the plain image to get a coefficient matrix. Second, several chaotic sequences generated by the tent‐sine chaotic map are used to scramble the coefficient matrix and construct a measurement matrix. Afterward, compressive sensing is performed on the coefficient matrix to obtain a small‐sized encrypted image. Finally, an image encryption scheme related to plaintext is designed. In particular, the proposed system uses the original image information to participate in the encryption process, ensuring the high sensitivity of the cryptosystem to minor differences in the plain image and good performance on resisting known/selected plaintext attacks. Furthermore, to convey the plaintext‐related parameters to the receiver, the dimension of the ciphertext image is expanded, and the parameters are embedded into the ciphertext image. Simulation results and security analysis show that the proposed image encryption system has strong plaintext sensitivity and robustness for effectively resisting various typical attacks such as brute‐force attacks, statistical attacks, and differential attacks. Shufeng Huang, Linqing Huang, Shuting Cai, Xiaoming Xiong, Yuan Liu 0022 |
IET Image Process. | 4 |
| 2022 | Large-Size Data Distribution in IoV Based on 5G/6G Compatible Heterogeneous NetworkabstractThe distribution of large-size data block in the Internet of Vehicles (IoV), especially in the urban IoV with dense vehicles, is still a challenge issue. Though the methods based on 5G-cellular network commonly can efficiently distribute the large-size file in IoVs, they also have obvious defects, such as occupying scarce 5G resources, generating communication fees and limited service coverage. This paper focuses on distributing large-size data block in IoVs based on dedicated vehicular ad hoc network to reduce the relying on cellular (5G/6G) resources. A heterogeneous vehicular network (HetVNET), in which a short-range OFDM wideband communication (SOWC) with ultra-high data rate and the original IEEE 802.11p protocol are included, is first proposed. To match with the proposed HetVNET, a content-centric data distribution scheme based on edge caching is designed. The content-centric data distribution architecture for distributing large-size file, the cache node selection and the process of data distribution and control based on the HetVNET are studied. The 5G/6G communication can be enabled in the scheme to further enhance the engineering stability of the massive infrastructure-to-vehicle (I2V) broadcast in IoVs. The evaluation results show that the proposed approach has low delivery delay and high penetration ratio. Xiuwen Yin, Jianqi Liu, Xiaochun Cheng, Xiaoming Xiong |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2021 | Subgraph feature extraction based on multi-view dictionary learning for graph classification
Xin Zheng 0001, Shouzhi Liang, Bo Liu 0002, Xiaoming Xiong, Xianghong Hu 0001, Yuan Liu 0022 |
Knowl. Based Syst. | 4 |
| 2020 | A multi-task transfer learning method with dictionary learning
Xin Zheng 0001, Luyue Lin, Bo Liu 0002, Yanshan Xiao, Xiaoming Xiong |
Knowl. Based Syst. | 5 |
| 2020 | The Software/Hardware Co-Design and Implementation of SM2/3/4 Encryption/Decryption and Digital Signature SystemabstractThe security of smart devices is facing great challenges. This article presents a new hybrid cipher framework suitable for such devices. Using software/hardware (SW/HW) co-design method, an efficient encryption/decryption and digital signature scheme based on SM2, SM3, and SM4 algorithms is implemented. First, the framework is partitioned into software and hardware parts based on the analysis result of a pure software solution and the hardware overhead under the conditions of design constraints. Second, the flow of the whole scheme and embedded algorithms are realized by SW/HW co-design. Finally, an improved implementation of SM2/3/4 algorithms is proposed to achieve higher efficiency. In this implementation, some SW/HW modules are parallelized to reduce the running time and enhance the performance. The proposed design is safe and can resist simple power analysis (SPA) attacks. Especially, the AHB bus interface IP and software scheduling approach are adopted in data transfer and manipulation. The design is taped-out on a silicon chip with SMIC 110-nm technology process. The chip uses about 199K logic gates and 1-mm2areas. The operating frequency of the design is 36 MHz and the chip consumes 23-mW power. The comparison with similar previous works shows our proposed design is more efficient with speed increasing of more than 10%. Xin Zheng 0001, Chongyao Xu, Xianghong Hu 0001, Yun Zhang 0001, Xiaoming Xiong |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2018 | The Robotic Fish Strategy Based on The Extreme Learning Machine Optimized by Particle Swarm Optimization AlgorithmabstractAiming at URWPGSim2D platform, in order to realize rapid and accurate adjustment of the robotic fish, action decision strategy based on the extreme learning machine optimized by particle swarm algorithm is put forward. According to the current environmental information of robotic fish, use the extreme learning machine to choose the optimal hitting point independently, and to determine the optimal combination of velocity and angular velocity of robotic fish. At the same time, the particle swarm optimization algorithm is introduced to optimize the extreme learning machine, which can improve the accuracy and robustness of the extreme learning machine. Verified by URWPGSim2D platform show that: the robotic fish path can be adjusted according to the strategy, realize combinatorial optimization of speed and direction, and find the target in the shortest time and distance. This shows that action decision-making strategy based on extreme learning machine can fully consider the real-time information of robotic fish and water polo, choose a different strategy in different cases, have a strong ability to adapt, meet the requirements of robotic fish for the action decisions. Xuexi Zhang, Shuibiao Chen, Zhiguang Cao, Shuting Cai, Zerong Peng, Xiaoming Xiong |
ICIS | 6 |