EDBT 2026 Demo / reviewers in the wild / expert
Junying Huang
dblp:161/4629
· DBLP profile ↗
29ranked-venue papers
6as first author
25since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 26 · 3 first-author · 22 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DCLOG: Don't Cares-based Logic Optimization using Pre-training Graph Neural NetworksabstractLogic rewriting serves as a robust optimization technique that enhances Boolean networks by substituting small segments with more effective implementations. The incorporation of don’t cares in this process often yields superior optimization results. Nevertheless, the calculation of don’t cares within a Boolean network can be resourceintensive. Therefore, it is crucial to develop effective strategies that mitigate the computational costs associated with don’t cares while simultaneously facilitating the exploration of improved optimization outcomes. To address these challenges, this paper proposes DCLOG, a don’t cares-based logic optimization framework, to efficiently and effectively optimize a given Boolean network. DCLOG leverages a pretrained graph neural network model to filter out cuts without don’t cares and then performs an incremental window simulation to calculate don’t cares for each cut. Experimental results demonstrate the effectiveness and efficiency of DCLOG on large Boolean networks, specifically average size reductions of 15.64 % and 1.44 % while requiring less than 23.84 % and $44.70 \%$ of the average runtime compared with state-of-the-art methods for the majority-inverter graph (MIG), respectively. Rongliang Fu, Libo Shen, Ziyi Wang 0010, Zhengxing Lei, Zixiao Wang 0001, Junying Huang, Bei Yu 0001, Tsung-Yi Ho |
ASP-DAC | 6 |
| 2026 | THLR: A Top-down Hierarchical Logic Rewrite Framework for Xor-Majority-Inverter GraphsabstractWith the increasing complexity of integrated circuits, multiple Boolean network types have been developed to support efficient logic rewriting methods. Although the Xor-Majority-Inverter Graphs (XMG) have a relatively compact expressive power, due to the characteristics of the rewriting itself and the inherent properties of XMG, the rewriting does not always perform optimally in terms of optimization performance on XMG. In this paper, we propose a novel top-down hierarchical logic rewriting framework for XMG that exploits the complementary expressive capabilities of multiple Boolean network types. To more effectively leverage the rewriting potential of hierarchical Boolean network types, we propose a type-aware partitioning strategy that decomposes the network into structurally meaningful sub-circuits. This enables targeted optimizations tailored to the structural characteristics of each sub-circuit, effectively balancing rewriting quality with computational efficiency. Experimental results demonstrate that our framework significantly improves circuit quality, achieving an approximate 4.31% reduction in node-depth product (NDP) compared to state-of-the-art rewriting methods, while also reducing the runtime by about 13.62%. Moreover, after ASIC mapping, THLR delivers a 3.10% improvement in area-delay product (ADP) over the state-of-the-art approaches. Rongliang Fu, Shuo Ren 0001, Wenxing Li, Xiaochun Ye, Tsung-Yi Ho, Junying Huang |
ACM Great Lakes Symposium on VLSI | 9 |
| 2026 | RECALLS: Reinforcement Learning Enhanced Generative Model for Logic Synthesis Optimization
Xinda Chen, Rongliang Fu, Chunyang He, Tsung-Yi Ho, Junying Huang |
ISCAS | 8 |
| 2026 | JPnR: A Length-Matching Placement and Routing Framework for Single-Flux-Quantum CircuitsabstractSuperconducting rapid single-flux-quantum (RSFQ) logic is a promising candidate for advancing future computing technologies due to its low-energy consumption and high-frequency capabilities. However, precise timing alignment is crucial for its physical design, posing significant challenges in length-matching placement and routing. This paper introduces JPnR, a physical design framework tailored for RSFQ circuits, featuring a clock-aware length-matching placer and a length-matching multi-terminal router. The placer simultaneously considers both clock distribution and timing constraints, distributing clock pulses heuristically and transforming the placement problem into a single-source shortest-path problem. This allows it to minimize vertical wirelength using dynamic programming and iteratively optimize placement via a barycenter-like reordering method. The router tackles challenges related to splitter placement and length-matching multi-terminal routing using a two-layer planar Manhattan routing model. Initial routing assigns tracks based on the left-edge algorithm to minimize routing width while employing the dogleg algorithm to resolve cycles in the vertical constraint graph. Length-matching is achieved via a splitter tree-based hierarchical approach with maximum-flow-based detour insertion. Finally, a PTL region expansion strategy is employed for unsatisfied connections. Experimental results on RSFQ benchmarks demonstrate the effectiveness and efficiency of JPnR. Rongliang Fu, Minglei Zhou, Xinda Chen, Junying Huang, Xiaochun Ye, Zhimin Zhang 0004, Tsung-Yi Ho |
IEEE Trans. Computers | 5 |
| 2026 | WindScatter: An Ultra-Low-Power, Long-Range, Large-Scale Wind Speed Monitoring SystemabstractWind speed monitoring is crucial for environmental management and forecasting. However, current solutions often struggle with high power consumption, especially at the end device, which typically has a sensor and wireless radios with limited battery capacity. To this end, we present WindScatter, an ultra-low-power, long-range, and large-scale wind speed monitoring system. WindScatter adopts the Integrated Sensing and Communication (ISAC) paradigm to enable low-power operation. It reuses the sensed data for communication by leveraging a TMR (Tunnel Magneto-Resistance) switch sensor to measure the wind speed information and control the backscatter communication simultaneously, thus avoiding the need for analog-to-digital conversion and a microcontroller for communication control. Our hardware-software co-design enables accurate measurements and stable concurrent transmission. We implement WindScatter and conduct extensive experiments and case studies to evaluate its performance. Results show that WindScatter supports measurements of all wind speed levels on the Extended Beaufort scale, from 1.5 m/s to 60 m/s, with an average error rate of 0.78%. WindScatter can sense and transmit wind speed data at a distance of 800 m with a power consumption of 136.5$\mu$W. Compared with commodity devices, WindScatter achieves comparable measurement range and accuracy while reducing cost by$91.5\times$and power consumption by$8,791\times$. Junying Huang, Chaojie Gu, Xiuzhen Guo, Shibo He, Yuanchao Shu, Jiming Chen 0001 |
IEEE Trans. Mob. Comput. | 1 |
| 2025 | Late Breaking Results: Hybrid Logic Optimization with Predictive Self-SupervisionabstractHybrid optimization is an emerging approach in logic synthesis, focusing on applying diverse optimization methods to different parts of a logic circuit. This paper analyzes the relationship between each vertex and its corresponding optimization method. We extract a subgraph centered on each vertex and quantify the logic optimization results of these subgraphs as vertex features. Based on these features, we propose a circuit partitioning method to cluster the logic circuit, enabling the final optimized circuit to be constructed by merging clusters optimized with their respective methods. Additionally, we introduce a self-supervised prediction model to efficiently obtain vertex features. The experimental results targeting LUT mapping demonstrate that our method achieves improvements of $8.48 \%$ in area and 9.81% in delay compared to the state-of-the-art. Rongliang Fu, Zhengyuan Shi, Yuan Pu 0001, Junying Huang, Qiang Xu 0001, Tsung-Yi Ho |
DAC | 6 |
| 2025 | An Optimal DFF-Oriented Technology Legalization Algorithm for Rapid Single-Flux-Quantum Circuits
Minglei Zhou, Rongliang Fu, Xiaochun Ye, Tsung-Yi Ho, Junying Huang |
ACM Great Lakes Symposium on VLSI | 6 |
| 2025 | J2Place: A Multiphase Clocking-Oriented Length-Matching Placement for Rapid Single-Flux-Quantum CircuitsabstractSuperconducting Rapid Single-Flux-Quantum (RSFQ) logic, characterized by low power consumption and high-frequency operation, has broad application prospects and holds substantial potential for future computing technologies. However, ensuring the correct operation of RSFQ circuits requires inserting numerous D flip-flops (DFFs), which substantially increase circuit area and energy dissipation. Recent studies have demonstrated that the multiphase clocking scheme can effectively reduce the number of required DFFs. Despite these advantages, existing placement tools do not support multiphase clocking RSFQ circuits. To address this limitation, this paper introduces J2Place, a novel multiphase clocking-oriented length-matching placement framework for RSFQ circuits. Our approach introduces two new RSFQ cells, TFFDO and TFFDE, to simplify the clock network in two-phase clocking designs. We propose a maximum flow-based method to generate the clock distribution column by column and utilize dynamic programming to minimize the total vertical wirelength while maintaining fixed placement orders. Additionally, to expand the solution space, we propose a length-aware reordering method to reduce the wirelength further. Experimental results on ISCAS85 and EPFL benchmarks demonstrate the effectiveness and efficiency of J2Place compared with state-of-the-art methods. Rongliang Fu, Minglei Zhou, Huilong Jiang, Junying Huang, Xiaochun Ye, Tsung-Yi Ho |
ICCAD | 4 |
| 2025 | JBSA: A Bit-Serial Accelerator for Deep Neural Networks Using Superconducting SFQ LogicabstractThe potential of superconducting single flux quantum (SFQ) devices in accelerating deep neural networks (DNNs) has garnered significant attention due to their ultra-fast and lowpower switching capabilities.However, existing SFQ-based DNN accelerators face limitations in scaling up to larger-scale instances due to the stringent area constraints and complex architectures.Additionally, another challenge in SFQ-based DNN acceleration lies in bridging the gap between the ultrahigh computing speed offered by SFQ technology and the relatively low memory bandwidth.To address these challenges, we propose JBSA, an SFQ-based bit-serial accelerator for DNN inference acceleration.JBSA leverages bit-serial computing to alleviate area constraints and reduce bandwidth requirements.A bit-serial processing element is designed to implement multiply-accumulate operations using SFQ logic cells. Huilong Jiang, Haofei Yin, Rongliang Fu, Junying Huang, Xiaochun Ye, Zhimin Zhang 0004, Tsung-Yi Ho, Dongrui Fan |
ICS | 6 |
| 2025 | A fast test compaction method using dedicated Pure MaxSAT solver embedded in DFT flow
Zhiteng Chao, Xindi Zhang 0001, Junying Huang, Zizhen Liu, Jing Ye 0001, Shaowei Cai 0001, Huawei Li 0001, Xiaowei Li 0001 |
Integr. | 3 |
| 2025 | CGCGraph: Efficient CPU-GPU Co-execution for Concurrent Dynamic Graph ProcessingabstractWith the continuous growth of user scale and application data, the demand for large-scale concurrent graph processing is increasing. Typically, large-scale concurrent graph processing jobs need to process corresponding snapshots of dynamically changing graph data to obtain information at different time points. To enhance the throughput of such applications, current solutions concurrently process multiple graph snapshots on the GPU. However, when dealing with rapidly changing graph data, transferring multiple snapshots of concurrent jobs to the GPU results in high data transfer overhead between CPU and GPU. Additionally, the execution mode of existing work suffers from underutilization of GPU computational resources. In this work, we introduce CGCGraph, which can be integrated into existing GPU graph processing systems like Subway, to enable efficient concurrent graph snapshot processing jobs and enhance overall system resource utilization. The key idea is to offload unshared graph data of multiple concurrent snapshots to the CPU, reducing CPU-GPU transfer overhead. By implementing CPU-GPU co-execution, there is potential for enhanced utilization of GPU computing resources. Specifically, CGCGraph leverages kernel fusion to process shared graph data concurrently on the GPU, while executing all snapshots in parallel on the CPU, with each snapshot assigned a dedicated thread. This approach enables efficient concurrent processing within a novel CPU-GPU co-execution model, incorporating three optimization strategies targeting storage, computation, and synchronization. We integrate CGCGraph with Subway, an existing system designed for out-of-GPU-memory static graph processing. Experimental results show that the integration of CGCGraph with current GPU-based systems obtains performance improvements ranging from 1.7 to 4.5 times. Jie Zhang 0130, Huawei Cao, Yuan Zhang 0031, Xuejun An, Junying Huang, Xiaochun Ye |
ACM Trans. Archit. Code Optim. | 6 |
| 2025 | Memory-Efficient and Adaptive Heterogeneous Framework for Gate-Level Fault SimulationabstractGate-level fault simulation is essential for automatic test pattern generation (ATPG). The traditional event-driven simulation is time-consuming due to the large number of faults. While parallel fault simulation with GPGPUs shows promise, it faces reduced parallel efficiency on large circuits. This is mainly due to the increased space required to store fault values, limiting the number of faults that can be processed in parallel and preventing full utilization of the GPU’s capabilities. In this study, we propose a memory-efficient fault machine implementation FM gpu based on a circular vector, which is tailored for GPU fault simulation with some sacrifices of time efficiency and a variable length limit. We also propose a fully adaptive parallel fault simulation framework based on the CPU-GPU heterogeneous system, which includes two stages on the GPU and performs CPU simulation at the same time. All parameters related to GPU memory optimization and workload balancing in the framework can be adjusted adaptively. The experimental results demonstrate that our method achieves better memory efficiency and speedup compared to the previous GPU fault simulation methods, a maximum speedup of 137.48× compared to the baseline open-source simulator with 32 threads, and a maximum speedup of 2.52× compared to a 32-thread commercial tool. Zhiteng Chao, Junying Huang, Wenjie Li 0004, Jing Ye 0001, Huawei Li 0001, Xiaowei Li 0001 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2024 | A Fast Test Compaction Method for Commercial DFT Flow Using Dedicated Pure-MaxSAT SolverabstractMinimizing the testing cost is crucial in the context of the design for test (DFT) flow. In our observation, the test patterns generated by commercial ATPG tools in test compression mode still contain redundancy. To tackle this obstacle, we propose a post-flow static test compaction method that utilizes a partial fault dictionary instead of a full fault dictionary, and leverages a dedicated Pure-MaxSAT solver to re-compact the test patterns generated by commercial ATPG tools. We also observe that commercial ATPG tools offer a more comprehensive selection of candidate patterns for compaction in the “n-detect” mode, leading to superior compaction efficacy. In experiments on ISCAS89, ITC99, and open-source RISC-V CPU benchmarks, our method achieves an average reduction of 21.58% and a maximum of 29.93% in test cycles evaluated by commercial tools while maintaining fault coverage. Furthermore, our approach demonstrates improved performance compared with existing methods. Zhiteng Chao, Xindi Zhang 0001, Junying Huang, Jing Ye 0001, Shaowei Cai 0001, Huawei Li 0001, Xiaowei Li 0001 |
ASPDAC | 3 |
| 2024 | JPlace: A Clock-Aware Length-Matching Placement for Rapid Single-Flux-Quantum CircuitsabstractSuperconducting rapid single-flux-quantum (RSFQ) logic has emerged as a promising candidate for future computing technology, owing to its low power consumption and high frequency characteristics. Given its ultra-high frequency operation, achieving precise timing alignment is crucial for RSFQ circuit physical design. To address the timing issue, this paper introduces JPlace, a clock-aware length-matching placement framework for RSFQ circuits. JPlace simultaneously addresses data and clock signal length matching, effectively ensuring accurate timing alignment and mitigating timing alignment challenges during the routing phase. We propose a heuristic method for constructing the clock distribution and a dynamic programming-based approach for minimizing the total vertical wirelength while maintaining fixed placement orders. Additionally, we introduce a barycenter-based reordering method to further explore the solution space and reduce wirelength. Experimental results on the RSFQ benchmark demonstrate the effectiveness and efficiency of JPlace. Rongliang Fu, Junying Huang, Zhimin Zhang 0004, Xiaochun Ye, Tsung-Yi Ho, Dongrui Fan |
DATE | 3 |
| 2024 | A Fully Pipelined High-Performance Elliptic Curve Cryptography Processor for NIST P-256abstractElliptic curve cryptography (ECC) is widely used in public key encryption, but its high-speed deployment faces challenges due to algorithmic and arithmetic complexity. In this paper, we present a high-performance ECC processor for the elliptic curve point multiplication (ECPM) of NIST P-256. Our approach employs a fully pipelined architecture featuring a 7-stage, 256-bit multiplier operating at a high frequency. To manage the data flow of the ECPM operation process, we devise a controller equipped with configurable instructions, which provides ECPM operations with higher flexibility to meet diverse contextual requirements. Additionally, we introduce a compact pipeline schedule to reduce ECPM computation clock cycles. The proposed LUT-based design achieves ECPM computation in 0.039 ms on FPGA (Virtex-7 platform) and 0.037 ms on ASIC (90nm technology), requiring only 10712 clock cycles. Junying Huang, Jing Ye 0001, Huawei Li 0001, Xiaowei Li 0001 |
ETS | 3 |
| 2023 | A Template Attack on Reduction Without Reference Device on KyberabstractIn July 2022, the National Institute of Standards and Technology (NIST) announced its selection of four algorithms for post-quantum cryptography standardization in advance. Among these algorithms, Kyber was chosen as the only key encapsulation mechanism (KEM). In the Kyber KEM, the modular reduction function is utilized in numerous areas. We have discovered that by modeling controllable modular reduction functions, unknown modular reduction functions can be targeted. And attacks can then be constructed. Henceforth, profiling can be mounted on the target device. In this paper, we present a machine-learning-based key recovery attack on Kyber, without needing a reference device. We have effectively attacked the modular reduction function. Furthermore, this vulnerability that enables the reuse of the same function could be utilized in other attacks. Yipei Yang, Junying Huang, Zongyue Wang, Jing Ye 0001, Junfeng Fan, Huawei Li 0001, Xiaowei Li 0001, Yuan Cao 0003 |
ATS | 2 |
| 2023 | BOMIG: A Majority Logic Synthesis Framework for AQFP LogicabstractAdiabatic quantum-flux-parametron (AQFP) logic, an energy-efficient superconductor logic with no static power consumption and ultra-low switching energy, is a promising candidate for energy-efficient computing systems. Due to the native majority function in AQFP logic, which can represent more complex logic with the same cost as the AND/OR function, the design of AQFP circuits differs from AND-OR-inverter-based logic circuits. Besides, AQFP logic has the path balancing requirement and fan-out limitation, making traditional majority-based logic optimization methods not applicable. This paper proposes a global optimization method over the majority-inverter graph (MIG) to minimize the JJ number and circuit depth of AQFP circuits. MIG-based transformation methods are first illustrated to construct the feasible domain. The normalized energy-delay-product (EDP), the product of the JJ number and circuit depth of AQFP circuits, is used as the objective function. Then, Bayesian optimization is used to explore the global optimal transformation sequence applied to AQFP MIG-based logic optimization. Experimental results show that the proposed method has a significant improvement in the JJ number and circuit depth compared with the state-of-the-art. Rongliang Fu, Junying Huang, Mengmeng Wang 0006, Nobuyuki Yoshikawa, Bei Yu 0001, Tsung-Yi Ho, Olivia Chen |
DATE | 2 |
| 2023 | JRouter: A Multi-Terminal Hierarchical Length-Matching Router under Planar Manhattan Routing Model for RSFQ CircuitsabstractSuperconducting rapid single-flux-quantum (RSFQ) logic has shown great potential for high-energy-efficient computing systems. To ensure correct operations at ultra-high frequencies, it is necessary to incorporate length-matching constraints into the routing problem. Existing routing algorithms, however, can only address 2-pin connections or support the conventional horizontal/vertical routing model, which substantially limits the optimization space for routing solutions. This paper presents JRouter, an RSFQ router that considers the two-layer planar Manhattan routing model while simultaneously coping with splitter (SPL) placement and length-matching multi-terminal routing. JRouter contains a track-assignment-based initial routing that minimizes the initial routing width while avoiding conflicts in the horizontal constraint graph. Moreover, JRouter implements an SPL-tree-based hierarchical routing with an iterative maximum-flow-based formulation to insert the detours for multi-terminal routing. A routing region extension algorithm is also developed to insert the detours for unsatisfied connections. According to the experimental results, JRouter achieves an average routing width reduction of 35.71% and 22.46% on a 16-bit RSFQ Sklansky adder compared to Kito's and Kou's routing algorithms. For randomly generated benchmarks, JRouter reduces the routing width by an average of 38.77%, 38.20%, 21.65%, and 7.01% compared to Kito's, Kou's, and two of Yan's routing algorithms, respectively, while maintaining reasonable runtime. Xinda Chen, Rongliang Fu, Junying Huang, Huawei Cao, Zhimin Zhang 0004, Xiaochun Ye, Tsung-Yi Ho, Dongrui Fan |
ACM Great Lakes Symposium on VLSI | 3 |
| 2023 | Alleviating Transfer Latency in DataFlow Accelerator for DSP ApplicationsabstractTowards multiple domains, dataflow accelerators show superiority for their flexible programmability and high efficiency. This efficiency relies highly on data communication between processing elements (PEs), which is sensitive to PE location, array scale and workload size. Laying out instructions as a dataflow graph on the PE array creates more instruction-level parallelism. However, the farther distance between remote PEs and memory banks introduces extra transfer latency, bringing performance degradation to high real-time applications. This paper examines the workloads of digital signal processing across different data scales and classifies latency problems related to data transfers and kernel switching. Specifically, we propose a novel forwarding network on chip to alleviate transfer latency and improve multi-destination sharing in the dataflow execution. Moreover, we devise bandwidth reusing mechanism to speedup kernel switching. The experiment results show that our scalable design achieves up to 2.19× (1.45× on average) speedup while reducing switching overhead by 9.85×, with an area overhead of 10.82% over the conventional dataflow accelerator. Zhihua Fan, Zhen Wang 0045, Tianyu Liu 0007, Junying Huang, Shengzhong Tang, Yanhuan Liu, Kunming Zhang, Xiaochun Ye, Dongrui Fan |
ICCD | 6 |
| 2023 | FSGraph: fast and scalable implementation of graph traversal on GPUs
Yuan Zhang 0031, Huawei Cao, Jie Zhang 0130, Junying Huang, Xiaochun Ye, Xuejun An |
CCF Trans. High Perform. Comput. | 5 |
| 2023 | IRA-FSOD: Instant-Response and Accurate Few-Shot Object DetectorabstractAiming at recognizing and localizing objects of novel categories with just a few reference samples, few-shot object detection (FSOD) is quite a challenging task. Previous works rely heavily on the fine-tuning process to transfer their models to the novel categories. They are flawed in the real application since the fine-tuning process is time-consuming and it suffers from serious deterioration on the low-quality support set. Based on the observation, this paper proposes an instant-response and accurate few-shot object detector (IRA-FSOD) that can detect the objects from novel categories without fine-tuning. We carefully analyze the limitations of widely-used Faster R-CNN and transform it to IRA-FSOD. Specifically, we first propose a novel semi-supervised Region Proposal Network (SS-RPN) module and a switch classifier module to precisely recognize the potential foreground instances from novel categories without fine-tuning. Moreover, we introduce two explicit inference strategies into the localization module, including explicit localization score and semi-explicit box regression, to alleviate over-fitting towards the base categories. Extensive experiments demonstrates that the proposed IRA-FSOD not only accomplish few-shot object detection with the instant-response, but also reaches state-of-the-art performance under various FSOD protocols and settings. Junying Huang, Junhao Cao, Liang Lin 0004, Dongyu Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Enhancing Prototypical Few-Shot Learning By Leveraging The Local-Level StrategyabstractAiming at recognizing the samples from novel categories with few reference samples, few-shot learning (FSL) is a challenging problem. We found that the existing works often build their few-shot model based on the image-level feature by mixing all local-level features, which leads to the discriminative location bias and information loss in local details. To tackle the problem, this paper returns the perspective to the local-level feature and proposes a series of local-level strategies. Specifically, we present (a) a local-agnostic training strategy to avoid the discriminative location bias between the base and novel categories, (b) a novel local-level similarity measure to capture the accurate comparison between local-level features, and (c) a local-level knowledge transfer that can synthesize different knowledge transfers from the base category according to different location features. Extensive experiments justify that our proposed local-level strategies can significantly boost the performance and achieve 2.8%–7.2% improvements over the baseline across different benchmark datasets, which also achieves the state-of-the-art accuracy. Junying Huang, Keze Wang, Liang Lin 0004, Dongyu Zhang 0002 |
ICASSP | 1 |
| 2022 | A survey on superconducting computing technology: circuits, architectures and design tools
Junying Huang, Rongliang Fu, Xiaochun Ye, Dongrui Fan |
CCF Trans. High Perform. Comput. | 1 |
| 2022 | JBNN: A Hardware Design for Binarized Neural Networks Using Single-Flux-Quantum CircuitsabstractAs a high-performance application of low-temperature superconductivity, superconducting single-flux-quantum (SFQ) circuits have high speed and low-power consumption characteristics, which have recently received extensive attention, especially in the field of neural network inference accelerations. Despite these promising advantages, they are still limited by storage capacity and manufacture reliability, making them unfriendly for feedback loops and very large-scale circuits. The Binarized Neural Network (BNN), with minimal memory requirements and no reliance on multiplication, is undoubtedly an attractive candidate for implementing inference hardware using SFQ circuits. This work presents the first SFQ-based Binarized Neural Network inference accelerator, namely JBNN, with a new representation to binarize weights and activation variables. Every SFQ gate is essentially a pipeline stage, making conventional design methods of the accumulator unsuitable for SFQ circuits. So an SFQ-based accumulative parallel counter using SFQ logic cells including T1, OR, and AND is designed to realize the accumulation, where the data size is reduced to a quarter after passing the XNOR column and the AU layer, largely declining the hardware cost. Our evaluation shows that the proposed design outperforms a cryogenic CMOS-based BNN accelerator design running at 77K by 70.92 times while maintaining 97.89% accuracy on the MNIST benchmark dataset. Without the cooling cost, the power efficiency increases up to 929.18 times. Rongliang Fu, Junying Huang, Xiaochun Ye, Dongrui Fan, Tsung-Yi Ho |
IEEE Trans. Computers | 2 |
| 2021 | Equivalence Checking for Superconducting RSFQ Logic CircuitsabstractEquivalence checking is a key component of the verification methodology for digital circuit designs. In this paper, we propose an equivalence checking framework for superconducting rapid single-flux-quantum (RSFQ) logic circuits which include acyclic circuits and bit-slice-based cyclic circuits. It consists of a structure checker and a logic checker. The structure checker is used to check whether the circuit meets the design rules of superconducting RSFQ logic circuits. The logic checker can be used to check whether two RSFQ gate-level circuits have the same logic function. For the logic checker, we propose a logic equivalence checking method based on logic cone partition. The circuit network is simplified layer by layer and iteratively partitioned into logic cones, each of which is verified by the SMT solver. The experimental results show the feasibility of our approach on superconducting RSFQ logic circuits. Rongliang Fu, Junying Huang, Zhimin Zhang 0004 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2020 | Design Automation Methodology from RTL to Gate-level Netlist and Schematic for RSFQ Logic CircuitsabstractThe superconducting rapid single flux quantum (RSFQ) logic circuit has the characteristics of high speed and low power consumption, making it an attractive candidate for future supercomputers. However, computer-aided design (CAD) tools for CMOS cannot be directly applied to RSFQ logic due to their distinct properties. For instance, the RSFQ logic gate can work properly when all its fan-ins have the same logic level. This paper presents the design flow from RTL to RSFQ logic netlist and schematic. First, we implement logic synthesis for RSFQ logic circuits. It achieves path balancing while minimizing the number of DFFs. In addition, we propose an automatic schematic generator for the RSFQ logic circuits. It converts the synthesized netlist into its equivalent schematic. A layer assignment algorithm is proposed, which makes all gates layered in the order of the clock arrival time. Experimental results with ISCAS85 and EPFL benchmarks along with some Kogge-Stone adders have shown a 29.2% reduction in the number of DFFs over the breadth-first first search; moreover, 59.57% and 5.3% decrease in the number of layers of the schematic and number of edge crossings over the ELK tool. Rongliang Fu, Zhimin Zhang 0004, Guang-Ming Tang, Junying Huang, Xiaochun Ye, Dongrui Fan, Ninghui Sun |
ACM Great Lakes Symposium on VLSI | 4 |
| 2019 | iATPG: Instruction-level Automatic Test Program Generation for Vulnerabilities under DVFS attackabstractWith the growing cost of powering and cooling, the Dynamic Voltage Frequency Scaling (DVFS) technique has been adopted in many mobiles and embedded devices nowadays. However, attackers are capable of maliciously manipulating the DVFS to threaten application programs including the security related ones. This paper proposes an instruction-level Automatic Test Program Generation (iATPG) framework, which generates test programs to test the vulnerabilities of CPU instructions under the DVFS attack. The conditions that the test program needs to meet, the testability of CPU instructions, and the iATPG algorithm are proposed. It is applied to an arm CPU in a mobile phone. Typical instructions are tested, and some are found vulnerable. The application programs using these instructions are then attacked to prove the effectiveness of the proposed framework. Kuozhong Zhang, Junying Huang, Jing Ye 0001, Xiaochun Ye, Dongrui Fan, Huawei Li 0001, Xiaowei Li 0001, Zhimin Zhang 0004 |
IOLTS | 2 |
| 2019 | Instruction Vulnerability Test and Code Optimization Against DVFS AttackabstractWith the growing cost of powering and cooling, the Dynamic Voltage Frequency Scaling (DVFS) technique has been adopted in many mobiles and embedded devices nowadays. However, attackers are capable of maliciously manipulating the DVFS to threaten application programs including the security related ones. This paper first proposes a test method to test the vulnerabilities of CPU instructions under the DVFS attack. The test program feature, the testability of CPU instructions, and the Test Program Generation Algorithm (TPGA) are proposed. It is applied to an arm CPU in a mobile phone. Typical instructions are tested, and some are found vulnerable. Then, based on the test result, a method for code optimization by instruction substitution is proposed. The application program using vulnerable instructions are then attacked and optimized to prove the effectiveness of the proposed methods. Junying Huang, Jing Ye 0001, Xiaochun Ye, Dongrui Fan, Huawei Li 0001, Xiaowei Li 0001, Zhimin Zhang 0004 |
ITC-Asia | 1 |
| 2014 | Size aware placement for island style FPGAsabstractIn this paper we first examine the impact of FPGA size on overall performance and run-time of placement and routing in the context of cluster-based island-style FPGAs. Based on the observations, an FPGA placement algorithm, Min-Size, is introduced to alleviate the deterioration of performance and run-time of placement and routing when using a large FPGA to implement a circuit. We achieve this by allowing Min-Size to generate a more compact placement of logic, I/O and hard blocks. Our experimental results have shown a 3X and AX speedup in placement and routing run-time, a 38% and 41% reduction in wire length, and a 8% and 5% improvement in critical path delay when FPGA size increases 10 times. Junying Huang, Colin Yu Lin, Haigang Yang |
FPT | 1 |