VLDB 2026 Research / reviewers in the wild / expert
Weixiong Jiang
dblp:250/3724
· DBLP profile ↗
25ranked-venue papers
10as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 17 · 5 first-author · 15 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An adaptive expansion network for incremental fault diagnosis in open and dynamic industrial systems
Zongzhen Ye, Weixiong Jiang, Xuesong He, Jixian Dong, Jun Wu 0012 |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | ROFD: Event-based 5k-fps Real-Time Optical Flow Detector for Transient Radiant Expanding FlaresabstractDetecting transient radiant expanding flares is crucial for assessing the status of key devices like Hall thrusters, for which real-time detection is vital to maintain satellite stability in orbit. Event cameras are better suited for this task than high-speed cameras, offering faster perception speed, lower power consumption, and more compact sizes. While previous event-based works have significantly enhanced the visual processing speed, current approaches still fail to meet real-time requirements for transient flare detection, due to complex algorithms and inefficient hardware designs. To address the issues, we propose ROFD, an event-based 5k-fps real-time optical flow detector that includes: 1) A residual spatio-temporal-average optical flow detection algorithm that reduces the computing complexity and shortens the detection time window. 2) A tile-based interleaving memory mapping method that minimizes wasted memory access time. 3) Conflict-free data flows that eliminate data dependency and enhance parallelism. Experiments demonstrate that our FPGA-implemented ROFD operates at 5k-fps, achieving a maximum speedup of 35.4 × and a maximum accuracy improvement of 2.98×, while saving 87% DSPs, 33% BRAMs, and 58% power consumption compared to SOTA. Boyi Wei, Yibo Zhang 0008, Wenzhe Zheng, Weixiong Jiang, Chenyang Shi, Yajun Ha |
ISCAS | 7 |
| 2025 | Exemplar-free class incremental learning for rotating machinery fault diagnosis via adaptive prototype correction and separation network
Zongzhen Ye, Jun Wu 0012, Xuesong He, Lixiang Wang, Weixiong Jiang |
Adv. Eng. Informatics | 5 |
| 2025 | Manifold transfer and ensemble filter strategy for axial piston pump fault diagnosis under varied pressure pulsation
Weixiong Jiang, Jun Wu 0012, Zuoyi Chen, Haiping Zhu 0001, Yaqiong Lv |
Appl. Intell. | 1 |
| 2025 | A Gradient Alignment Federated Domain Generalization Framework for Rotating Machinery Fault DiagnosisabstractEmpowered by the huge amounts of sensor data in Industrial Internet of Things (IIOT), deep learning models have made remarkable achievements in the field of rotating machinery fault diagnosis. To improve the diagnosis performance under unknown working conditions, domain generalization technologies have been extensively studied. However, the existing methods predominantly gather the sensor data from multiple source domains together for model training, which poses a threat to data privacy in the IIOT. To address this problem, this paper proposes a novel gradient alignment federated domain generalization (GAFedDG) framework for rotating machinery fault diagnosis. In the proposed GAFedDG, an intra-domain gradient aligning mechanism is designed to minimize the gradient discrepancy between the current classifier on raw signals and augmented signals, effectively preventing the local model from overfitting the domain-specific fault knowledge. In addition, to bridge the domain shifts across multiple scattered source domains, an inter-domain gradient aligning mechanism is implemented to minimize the gradient discrepancy between the current classifier and other domain classifiers. By combining the two mechanisms above, a domain-agnostic model that can generalize well on unseen working conditions is established. Extensive experimental results on two self-built test rigs show that the GAFedDG possesses superior generalization capability in privacy-preserving scenarios. Zongzhen Ye, Jun Wu 0012, Xuesong He, Weixiong Jiang |
IEEE Internet Things J. | 4 |
| 2025 | Human-machine collaborative health estimation of industrial robot based on fuzzy self-attention network and manifold cluster
Weixiong Jiang, Jun Wu 0012, Haiping Zhu 0001 |
Knowl. Based Syst. | 1 |
| 2025 | FiDRL: Flexible Invocation-Based Deep Reinforcement Learning for DVFS Scheduling in Embedded SystemsabstractDeep Reinforcement Learning (DRL)-based Dynamic Voltage Frequency Scaling (DVFS) has shown great promise for energy conservation in embedded systems. While many works were devoted to validating its efficacy or improving its performance, few discuss the feasibility of the DRL agent deployment for embedded computing. State-of-the-art approaches focus on the miniaturization of agents’ inferential networks, such as pruning and quantization, to minimize their energy and resource consumption. However, this spatial-based paradigm still proves inadequate for resource-stringent systems. In this paper, we address the feasibility from a temporal perspective, where FiDRL, a flexible invocation-based DRL model is proposed to judiciously invoke itself to minimize the overall system energy consumption, given that the DRL agent incurs non-negligible energy overhead during invocations. Our approach is three-fold: (1) FiDRL that extends DRL by incorporating the agent's invocation interval into the action space to achieve invocation flexibility; (2) a FiDRL-based DVFS approach for both inter- and intra-task scheduling that minimizes the overall execution energy consumption; and (3) a FiDRL-based DVFS platform design and an on/off-chip hybrid algorithm specialized for training the DRL agent for embedded systems. Experiment results show that FiDRL achieves 55.1% agent invocation cost reduction, under 23.3% overall energy reduction, compared to state-of-the-art approaches. Jingjin Li, Weixiong Jiang, Yuting He 0002, Qingyu Yang 0004, Anqi Gao, Yajun Ha, Ender Özcan, Ruibin Bai, Tianxiang Cui, Heng Yu 0001 |
IEEE Trans. Computers | 2 |
| 2025 | RefSCAT: Formal Verification of Logic-Optimized Multipliers via Automated Reference Multiplier Generation and SCA-SAT SynergyabstractFormally verifying logic-optimized integer multipliers remains a crucial yet insufficiently addressed problem in both industry and academia, presenting significant verification challenges, particularly when verifying the large-scale logic-optimized multipliers with diverse architectures. Satisfiability (SAT)-based methods require structurally similar and known correct reference multipliers, which may not always be readily accessible. Symbolic computer algebra (SCA) techniques can verify multipliers without references but encounter difficulties with optimized multipliers due to unclear adder boundaries. To enable effective formal verification of the optimized multipliers, we propose the RefSCAT framework, which contains a reference multiplier generator that produces references structurally similar to the optimized multiplier with clear adder boundaries, enabling a synergistic SCA-SAT verification flow. First, we propose a reverse engineering algorithm that extracts the essential adder tree from the optimized multiplier, ensuring similarity. Second, since only a partial netlist is extractable after optimization, we propose a constraint satisfaction algorithm to complete the generation using only adders while following the extracted netlist, ensuring both similarity and clear adder boundaries. Third, leveraging the generated reference, we propose a synergized SCA-SAT verification flow that verifies the generated reference using SCA and then uses it as a correct reference for the SAT-based verification. The experiments demonstrate that RefSCAT can successfully verify logic-optimized multipliers with diverse partial-product-based architectures up to 128 bits, outperforming the state-of-the-art methods by verifying at least 29% more benchmarks. Rui Li 0095, Lin Li 0079, Heng Yu 0001, Masahiro Fujita 0004, Weixiong Jiang, Yajun Ha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2025 | RefSCAT-2.0: Formal Verification of Large-Scale Optimized Multipliers via Quantum-Inspired Ant Colony Optimization-Based Reference GenerationabstractFormal verification of large-scale optimized integer multipliers remains a critical yet insufficiently addressed challenge in industry and academia. Current methods employ reference multiplier generators to automatically construct structurally similar reference multipliers, which are then used by Satisfiability (SAT)-based techniques to verify equivalence with optimized multipliers. However, these approaches face limitations when generating references for large-scale optimized multipliers within acceptable timeframes. To address these limitations, we introduce the RefSCAT-2.0 framework, designed to rapidly produce high-quality large-scale reference multipliers. Firstly, we generate the macro-architecture to determine the number of adders required for constructing the reference multiplier. We propose a novel Integer Linear Programming (ILP)-based macro-architecture generation algorithm that minimizes the number of allocated adders, thereby reducing the overall problem complexity. Secondly, we organize the allocated adders into groups to simplify the subsequent generation process. We present a multi-level scheduler that automatically decomposes adders into groups with minimized interdependencies, ensuring both the quality of generation and a reduction in overall generation complexity. Thirdly, we generate the micro-architecture for each scheduled group, wherein we finalize the connections between adders. We present a graph-based design space representation coupled with a quantum-inspired ant colony optimization (QACO)-based generation algorithm that can efficiently explores the micro-architectures of each scheduled group. Experimental results show that RefSCAT-2.0 successfully verifies all 124 cases in a 256-bit optimized multiplier benchmark suite, outperforming SCA-based tcad22revsca and hybrid RefSCATTCAD24 methods which solve only 24 cases each. Rui Li 0095, Lin Li 0079, Heng Yu 0001, Masahiro Fujita 0004, Weixiong Jiang, Yajun Ha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2025 | A Deep Investigation on Stealthy DVFS Fault Injection Attacks at DNN Hardware AcceleratorsabstractWith increasing computation of various applications, dynamic voltage and frequency scaling (DVFS) is gradually deployed on FPGAs to improve performance and save energy. However, its reliability and security have not been sufficiently evaluated, which incurs quite many concerns. In this article, we propose an evaluation framework for deep investigation of stealthy DVFS fault injection attacks on the state-of-the-art deep neural networks (DNNs) deployed on modern FPGAs. The evaluation framework mainly consists of a DVFS attack striker and a time-to-digital converter (TDC)-based hardware profiler. Two modes of evaluation are derived, and their effectiveness is demonstrated on a platform composed of a SkyNet accelerator and three ImageNet models built on a Xilinx deep learning processor unit (DPU). Experimental results show that more than 99% detection accuracy loss can be measured targeting at all tested DNN models under prospective operation mode but without any performance degradation in frame per second (FPS). In our investigation of sensitive layer mode, more than 93% average accuracy loss with 84.7% fault probability can be measured on a single bundle of the SkyNet. We characterize the vulnerabilities of different DNN layers subject to DVFS attacks through leveraging the TDC-based hardware profiler to precisely control the timing of fault injection. Junge Xu, Fan Zhang 0010, Wenguang Jin, Kun Yang 0012, Zeke Wang, Weixiong Jiang, Yajun Ha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2025 | QuantTPM: Efficient Mixed-Precision Quantization Framework for Tractable Probabilistic ModelsabstractTractable probabilistic models (TPMs) can perform reliable probabilistic inference and enhance the reasoning capabilities of edge devices, such as aiding decision-making for autonomous vehicles. To deploy TPMs in edge scenarios with constrained hardware resources and energy, efficient quantization algorithms are necessary. However, the traditional quantization methods for neural networks are not applicable to TPMs due to the irregular model structure and highly varying data distribution. To address the issues, we propose QuantTPM, a mixed-precision quantization framework designed to enhance the energy and resource efficiency of TPM inference. First, we reformulate the irregular model structure into a unified format, as irregular structures are inefficient for hardware implementation. Second, we divide the reformulated model graph into hierarchical levels, so as to assign appropriate quantization bit-widths for different levels with varying precision requirements. Third, we decompose the entire mixed-precision quantization search into several steps with smaller search spaces, so as to reduce the algorithm complexity and save search time. Compared with state-of-the-art works, our mixed-precision quantization framework achieves, on average,$3.7\times $weight compression,$6.0\times $resource efficiency, and$4.8\times $energy consumption, while maintaining competitive accuracy. Guangyao Yan, Weixiong Jiang, Yajun Ha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2025 | An FPGA-Based Real-Time Loop Closure Detection Framework With Ultra-Fast Descriptor GeneratorabstractLoop closure detection (LCD) is crucial in LiDAR-based Simultaneous Localization and Mapping (SLAM) for smart vehicles, demanding both real-time performance and high-accuracy. Unfortunately, although the state-of-the-art LCD algorithms offer high-accuracy, the large search space in clustering, the high complexity in descriptor computation, and the slow speed in retrieval prevent them from achieving real-time performance. To address the issue, we propose three key techniques to achieve a real-time FPGA-based LCD framework. First, we reduce the clustering time by designing a Range Image-based clustering accelerator that significantly reduces the search space and achieves high-parallelism. Second, we reduce the descriptor computation time by designing an accelerator that selects only high-quality features and employs simplified operations to achieve low complexity. Third, we reduce the retrieval time by proposing a novel dual-descriptor cross-verification mechanism that uses fewer descriptors while maintaining high-accuracy. Compared to the state-of-the-art, experimental results show that our LCD accelerator demonstrates 128.8x performance improvement across multiple datasets in various scenes and achieves the required real-time performance. Shijie Meng, Weixiong Jiang, Jinjie Huang, Hao Sun 0035, Yajun Ha |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | An Energy-Efficient and Real-Time FPGA-Based Point Cloud Registration Framework with Ultra-Fast and Configurable Multi-Mode Correspondence SearchabstractPoint cloud registration is a fundamental task in LiDAR-based localization and mapping, widely employed in robotics and autonomous vehicles. However, existing registration solutions lose geometric topology continuity and lack scanline-aware, configurable correspondence search, restricting their real-time applicability. To solve these issues, we propose a fundamentally re-architected, energy-efficient FPGA framework for real-time point cloud registration, featuring a configurable, multi-mode correspondence search engine. First, we introduce a scanline-aided range-projection structure (SA-RPS) that reorganizes LiDAR points within configurable segmentation domains into contiguous memory while preserving scanline topology, enabling efficient and flexible multi-mode correspondence search. Second, we develop a deeply pipelined, ultra-fast SA-RPS-based correspondence search (SA-RPS-CS) accelerator that supports dynamic configuration of search mode and parallelism and incorporates a sliding-window cache and scanline-aware K-selection module for high-throughput, multi-mode correspondence extraction. Third, we present a co-designed registration framework that integrates the accelerator with dynamic parameter configuration, enabling adaptive, real-time processing across diverse SLAM scenarios. Experimental results demonstrate that the proposed SA-RPS-CS accelerator delivers \(2.3\times\) – \(32.4\times\) faster search and \(1.8\times\) – \(26.2\times\) higher energy efficiency than previous state-of-the-art FPGA designs, achieving real-time registration for 64-channel LiDAR at 21.5 FPS with negligible loss in accuracy. Hao Sun 0035, Yuhao Shu, Jianzhong Xiao, Weixiong Jiang, Hui Wang 0036, Yajun Ha |
ACM Trans. Reconfigurable Technol. Syst. | 5 |
| 2024 | Health assessment of wind turbine gearbox via parallel ensemble and fuzzy derivation collaboration approach
Weixiong Jiang, Jun Wu 0012, Chengjie Wang 0013, Haiping Zhu 0001, Xianbo Wang |
Adv. Eng. Informatics | 1 |
| 2024 | Multimodel Fusion Health Assessment for Multistate Industrial Robot via Fuzzy Deep Residual Shrinkage Network and Versatile ClusterabstractTo assess the health condition of industrial robots roundly and make hierarchical maintenance decisions, a multimodel fusion health assessment method is proposed for multistate industrial robots. Herein, many symptom parameters (SPs) are used to reflect the operation state of the industrial robot from aspects of vibration, temperature, and torque. Then, fuzzy deep residual shrinkage network is proposed to establish the SP-based status membership function as a single assessment model. The probabilities of robot operation states are determined and formulated as the hesitation fuzzy number (HFN). These HFNs from multiple assessment models are integrated into a collective hesitation fuzzy assessment matrix. Thus, the best worst method is adopted to estimate the confidence of each assessment model, and TOPSIS is used to judge the impact of different operation states on the industrial robot's behavior. Finally, a novel health index is defined for industrial robot, and robot health degree is identified by versatile cluster for hierarchical maintenance decisions. A self-built industrial robot test stand is adopted to validate the effectiveness of the proposed method, and sensitivity and comparison analysis results demonstrated that our method has advantages in terms of the situation adaptability and performance stability. Weixiong Jiang, Jun Wu 0012, Haiping Zhu 0001, Liang Gao 0001 |
IEEE Trans. Fuzzy Syst. | 1 |
| 2023 | AOS: An Automated Overclocking System for High-Performance CNN Accelerator Through Timing Delay Measurement on FPGAabstractWith the inherent algorithmic error resilience of conventional neural networks (CNNs) and the worst-case design methodologies of current electronic design automation tools, overclocking-based timing speculation is a promising technique to improve the performance of CNN accelerators on FPGA by removing unnecessary timing margins. To avoid potential timing errors, timing delay measurement should be used during overclocking. However, current approaches are not yet good at measuring paths with more intense variability factors such as jitter and lack an automated process for testing circuit delays. In this article, we first propose 2-dimension multiframe fusion to deal with the sampling jitter, then present a timing delay measurement-based automatic overclocking system (AOS) running on heterogeneous FPGA for high-performance CNN accelerators. On the FPGA side, AOS is composed of timing delay monitors (TDMs) that can measure all types of timing paths, a TDM controller that converts the sampled values of TDMs into timing delay in terms of the ratio of path delay to the clock period. On the CPU side, AOS converts the path delay from clock period ratio to absolute delay value and decides the frequency of the accelerator in the next iteration. We demonstrate AOS with a SkyNet accelerator on the Xilinx ZCU104 board and achieve 657 FPS at 436 MHz without accuracy degradation, which is$1.41\times $performance compared to the baseline. Weixiong Jiang, Heng Yu 0001, Fupeng Chen, Yajun Ha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | A High-Throughput Full-Dataflow MobileNetv2 Accelerator on Edge FPGAabstractFPGA accelerators for lightweight neural networks, such as MobileNetv2, are of great need in edge computing applications with high throughput requirements. Dataflow architecture has been considered a promising approach to optimize throughput since the intermediate feature map transfers can be significantly saved. However, previous MobileNetv2 accelerators only achieved a partial-dataflow architecture, and just one-third of the feature map transfers can be saved. To solve this issue, we propose a scheme to achieve a full-dataflow MobileNetv2 accelerator on FPGA. The scheme contains four techniques. First, we improve the full-integer quantization for easier deployment on hardware. Second, we propose tunable activation weight imbalance transfer for less quantization accuracy loss. Third, we present several highly optimized accelerator components whose parallelism can be flexibly adjusted and implement residual connection with deeper FIFO so that the requirements of the full-dataflow architecture can be fully met. Finally, we present a computing resource allocation strategy to balance the latency of each layer, and a memory resource allocation strategy to effectively use the on-chip memory. Compared to the state-of-the-art, experimental results show that the accelerator achieves 1910 FPS with$1.8\times $speedup when implemented on the Xilinx ZCU102 FPGA. In addition, it reaches 72.98% Top-1 accuracy with 8-bit integer quantization that outperforms all the other MobileNetv2 accelerators. Weixiong Jiang, Heng Yu 0001, Yajun Ha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2022 | WSQ-AdderNet: Efficient Weight Standardization Based Quantized AdderNet FPGA Accelerator Design with High-Density INT8 DSP-LUT Co-Packing OptimizationabstractConvolutional neural networks (CNNs) have been widely adopted for various machine intelligence tasks. Nevertheless, CNNs are still known to be computational demanding due to the convolutional kernels involving expensive Multiply-ACcumulate (MAC) operations. Recent proposals on hardware-optimal neural network architectures suggest that AdderNet with a lightweight ℓ1-norm based feature extraction kernel can be an efficient alternative to the CNN counterpart, where the expensive MAC operations are substituted with efficient Sum-of-Absolute-Difference (SAD) operations. Nevertheless, it lacks an efficient hardware implementation methodology for AdderNet as compared to the existing methodologies for CNNs, including efficient quantization, full-integer accelerator implementation, and judicious resource utilization of DSP slices of FPGA devices. In this paper, we present WSQ-AdderNet, a generic framework to quantize and optimize AdderNet-based accelerator designs on embedded FPGA devices. First, we propose a weight standardization technique to facilitate weight quantization in AdderNet. Second, we demonstrate a full-integer quantization hardware implementation strategy, including weight and activation quantization methodologies. Third, we apply DSP packing optimization to maximize the DSP utilization efficiency, where Octo-INT8 can be achieved via DSP-LUT co-packing. Finally, we implement the design using Xilinx Vitis HLS (high-level synthesis) and Vivado to Xilinx Kria KV-260 FPGA. Our experimental results of ResNet-20 using WSQ-AdderNet demonstrate that the implementations achieve 89.9% inference accuracy with INT8 implementation, which shows little performance loss as compared to the FP32 and INT8 CNN designs. At the hardware level, WSQ-AdderNet achieves up to 3.39× DSP density improvement with nearly the same throughput as compared to INT8 CNN design. The reduction in DSP utilization makes it possible to deploy large network models on resource-constrained devices. When further scaling up the PE sizes by 39.8%, WSQ-AdderNet can achieve 1.48× throughput improvement while still achieving 2.42× DSP density improvement. Weixiong Jiang, Yajun Ha |
ICCAD | 3 |
| 2022 | Quality Optimization of Adaptive Applications via Deep Reinforcement Learning in Energy Harvesting Edge DevicesabstractApplications with adaptability are widely available on the edge devices with energy harvesting capabilities. For their runtime quality optimization, however, current approaches cannot tackle the variations of quality modeling and harvested energy simultaneously. Therefore, in this article, we are the first to propose a deep reinforcement learning (DRL)-based dynamic voltage frequency scaling (DVFS) method that optimizes the application execution quality of energy harvesting edge devices to mitigate the variations. First, we propose a baseline DRL formulation that novelly migrates the objective of quality maximization into a reward function and constructs a DRL quality agent. Second, we devise a long short-term memory (LSTM)-based selector that performs DRL quality agent selection based on the energy harvesting history. Third, we further propose two optimization methods to alleviate the nonnegligible overhead of DRL computations: 1) an improved thinking-while-moving concurrent DRL scheme to compromise the “state drifting” issue during the DRL decision process and 2) a variable interstate duration decision scheme that compromises the DVFS overhead incurred in each action taken. The experiments take an adaptive stereo matching application as a case study. The results show that the proposed DRL-based DVFS method on average achieves 17.9% runtime reduction and 22.05% quality improvement compared to state-of-the-art solutions. Fupeng Chen, Heng Yu 0001, Weixiong Jiang, Yajun Ha |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | A Reliable 8T SRAM for High-Speed Searching and Logic-in-Memory OperationsabstractTo efficiently implement searching and logic functions with the SRAM-based in-memory computing (IMC), we need to perform computations on bitlines (BLs) (called compute access) via multiple wordline (WL) activations. However, this may cause prominent read disturbance when the IMC is implemented with the standard 6 T SRAM. To address this reliability issue, existing solutions adopt either auxiliary assistance circuits or alternative bitcell topologies, but they lead to substantial overheads of the access speed or array density. In this article, we propose a novel 8T compute SRAM (CSRAM) for reliable and high-speed in-memory searching and compound logic-in-memory computations. Our 8T CSRAM features a pair of pMOS access transistors and split-WLs dedicated to the compute access. A thorough circuit-level analysis reveals that the pMOS-based compute access port is essential for significantly mitigating the read disturbance. Moreover, we propose an elevated precharge voltage scheme and a low-skewed inverter-based sensing amplifier to improve the sensing speed. We have validated the proposed 8T CSRAM design in a 16 Kb array with a 28-nm CMOS technology. Compared to the state-of-the-art 8 T CSRAM, results show that our design is not only reliable but also 3.1 times faster, with a maximum operating frequency upping to 2.44 GHz. Yuqi Wang 0004, Yuhao Shu, Weixiong Jiang, Yajun Ha |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2022 | FODM: A Framework for Accurate Online Delay Measurement Supporting All Timing Paths in FPGAabstractVoltage and frequency scaling (VFS) has been widely used to improve energy efficiency, lifespan, and system reliability by converting conservative timing margins into$V_{\text {dd}}$reduction. Along these lines, to investigate the potential implementation of VFS technique in exploring the timing margins under different voltages and frequencies,in situor online circuit delay measurement is required to monitor all timing paths, which are usually ended with terminal registers. The previously reported online delay measurement approaches require the output of a terminal register to be measurable. However, some FPGA timing paths are ended with embedded hardcores such as DSPs or BRAMs. It is impossible to measure the output of the terminal register inside a hardcore. To address the issue, we propose an online delay monitor (ODM) that can accurately measure the delay of any type of timing path in real-time conditions. The ODM is mainly composed of two shadow registers and a phase-shifted clock. The shadow registers use a phase-shifted clock signal as the input and the output signal of the combinational logic as the clock. In addition, we present an automatic tool and its corresponding design flow (FODM) for inserting an ODM to monitor a path. Compared with the state-of-the-art, our experimental results indicate that the proposed method has the ability to accurately measure the delays online for all the potential timing paths, regardless of their path termination types. Moreover, we demonstrate an average measurement error of only 1.51% using eight floating-point operators at different voltages. Weixiong Jiang, Heng Yu 0001, Hongtu Zhang, Yuhao Shu, Rui Li 0095, Yajun Ha |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2021 | TAIT: One-Shot Full-Integer Lightweight DNN Quantization via Tunable Activation Imbalance TransferabstractBoth parameter quantization and depthwise convolution are essential measures to provide high-accuracy, lightweight, and resource-friendly solutions when deploying deep neural networks (DNNs) onto edge-AI devices. However, combining the two methodologies may lead to adverse effects: It either suffers from significant accuracy loss or long finetuning time. Besides, contemporary quantization methods are only selectively applied to weight and activation values but not bias and scaling factor values, making them less practical for ASIC/FPGA accelerators. To solve these issues, we propose a novel quantization framework that is effectively optimized for depthwise convolution networks. We discover that the uniformity of the value range within a tensor can serve as a predictor for the tensor’s quantization error. Under the guidance of this predictor, we develop a mechanism called Tunable Activation Imbalance Transfer (TAIT), which tunes the value range uniformity between an activated feature map and its latter weights. Moreover, TAIT fully supports full-integer quantization. We demonstrate TAIT on SkyNet and deploy it on FPGA. Compared to the state-of-the-art, our quantization framework and system design achieve 2.2%+ IoU, $2.4 \times$ speed, and $1.8 \times$ energy efficiency improvements, without any requirement of finetuning. Weixiong Jiang, Heng Yu 0001, Hao Sun 0035, Rui Li 0095, Yajun Ha |
DAC | 1 |
| 2020 | DVFS-Based Scrubbing Scheduling for Reliability Maximization on Parallel Tasks in SRAM-based FPGAsabstractTo obtain high reliability but avoiding the huge area overhead of traditional triple modular redundancy (TMR) methods in SRAM-based FPGAs, scrubbing based methods reconfigure the configuration memory of each task just before its execution. However, due to the limitation of the FPGA reconfiguration module that can only scrub one task at a time, parallel tasks may leave stringent timing requirements to schedule their scrubbing processes. Thus the scrubbing requests may be either delayed or omitted, leading to a less reliable system. To address this issue, we propose a novel optimal DVFS-based scrubbing algorithm to adjust the execution time of user tasks, thus significantly enhance the chance to schedule scrubbing successfully for parallel tasks. Besides, we develop an approximation algorithm to speed up its optimal version and develop a novel K-Means based method to reduce the memory usage of the algorithm. Compared to the state-of-the-art, experimental results show that our work achieves up to 36.11% improvement on system reliability with comparable algorithm execution time and memory consumption. Rui Li 0095, Heng Yu 0001, Weixiong Jiang, Yajun Ha |
DAC | 3 |
| 2020 | An Accurate FPGA Online Delay Monitor Supporting All Timing PathsabstractAccurate circuit delay measurement is essential for various purposes such as aging detection, health monitoring, and dynamic voltage and frequency scaling. State-of-the-art measurement techniques exhibit several limitations. For example, they are insufficiently informative by only returning binary results on the status of the circuit being normal or abnormal. More importantly, current approaches are not applicable for measuring the delay of timing paths that end with DSPs and BRAMs. To address the issues, we propose a novel online delay monitor (ODM) for modern FPGA platforms that (1) accurately returns the numerical delay values, (2) and is compatible with all types of timing paths in FPGAs. Our proposed ODM is achieved by employing a shadow register triggered by the output signal of a combinational circuit to sample a phase shifting clock. Besides, our design is capable of conveniently measuring the clock jitters, so we are able to propose an associated jitter management scheme to ensure correct ODM sampling. Experimental results show that our ODM achieves an error within 2% with respect to the ground truth. Weixiong Jiang, Rui Li 0095, Heng Yu 0001, Yajun Ha |
ISCAS | 1 |
| 2019 | Early Development of Infant Brain Complex Network
Weixiong Jiang, Han Zhang 0002, Li-Ming Hsu, Dan Hu 0004, Guoshi Li, Ye Wu 0001, Dinggang Shen |
MICCAI (2) | 1 |