EDBT 2026 Demo / reviewers in the wild / expert
Jian Dong 0010
dblp:58/3444-10
· DBLP profile ↗
24ranked-venue papers
0as first author
18since 2021 · last 2026
0000-0002-2980-9431ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 11 since 2021Software engineering, systems software and programming languages · 8 · 7 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TranSIC: A Storage in Computing Framework for Designing Efficient Transformer Accelerators in Energy-Constrained ScenariosabstractIntelligent applications such as wearable health monitors, IoT devices, nano-drones, and embodied intelligence systems are increasingly deployed in energy-constrained environments. Though Transformer models offer advanced processing capabilities, their integration is limited by energy and computational constraints. Extensive research has shown that data movement, particularly for large Transformer weights, is the primary source of energy consumption, accounting for up to 37% of total energy use. Recent techniques like processing in-memory (PIM) and processing near-memory (PNM) are hard to fundamentally reduce the number of data movements, limiting their effectiveness in lowering energy consumption. This paper investigates Storage in Computing (SIC), integrating weight matrices directly into computation circuits, thereby eliminating external memory access during inference and reducing operand data movement significantly. To realize SIC, we propose TranSIC, an innovative approach to create efficient Transformer accelerators with minimized data movement. Key advancements facilitating practical SIC implementation include a reuse strategy to reduce required SIC units and a low-overhead optimization method for this strategy to decrease data routing overhead. Experimental results demonstrate that TranSIC achieves up to 8.65× and 56.24× energy efficiency over state-of-the-art PIM and traditional approaches, respectively. Jian Dong 0010, Gang Qu 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | BPUFuzzer: Effective Fuzz Testing for Branching Transient Execution Vulnerabilities of RISC-V CPUabstractThis paper presents BPUFuzzer, a fuzz testing tool for detecting branching transient execution vulnerabilities in CPU RTL design. BPUFuzzer addresses two key challenges: generating testcases that capture complex control flows, and extracting essential data from vast hardware states to guide testcase selection. Utilizing a control flow graph-based testcase generation strategy with anomaly detection and employing fitness and coverage metrics, BPUFuzzer works on testcases that cover broader program flows and deliberately selects testcases to discover transient execution vulnerabilities effectively. When applied on RISC-V Boom v3, BPUFuzzer uncovered more Spectre types than the state-of-the-arts, including a previously unidentified variant, named Spectre-LOOP. Rihui Sun, Hanyin Liu, Zikang Tao, Gang Qu 0001, Dongsheng Wang 0002, Yongqiang Lyu 0001, Jian Dong 0010 |
DAC | 8 |
| 2025 | MCTA: A Multi-Stage Co-Optimized Transformer Accelerator with Energy-Efficient Dynamic Sparse OptimizationabstractAs Transformer-based models continue to enhance service quality across various domains, their intensive computational requirements are exacerbating the AI energy crisis. Traditional energy-efficient Transformer architectures primarily focus on optimizing the Attention stage due to its high algorithmic complexity$(O(n^{2})\overline{)}$. However, linear layers can also be significant energy consumers, sometimes accounting for over 70% of total energy usage. Although existing approaches such as sparsity have improved the Attention stage, the optimization space within linear layers is not fully exploited. In this paper, we introduce the multi-stage co-optimized Transformer accelerator (MCTA) for optimizing energy efficiency. Our approach independently enhances the Query-Key-Value generation, Attention, and Feed-forward Neural Network stages. It employs two novel techniques: Low-overhead Mask Generation (LMG) for dynamically identifying unimportant calculations with minimal energy costs, and Cascaded Mask Derivation (CMD) for streamlining the mask generation process through parallel processing. Experimental results show that MCTA achieves an average energy reduction of 1.48 × with only a 1% accuracy loss compared to state-of-the-art accelerators. This work demonstrates the potential for significant energy savings in Transformer models without the need for retraining, paving the way for more sustainable AI applications. Jian Dong 0010 |
DATE | 5 |
| 2025 | CMC: Compound Memory-Computing Architecture for Energy-Efficient CNN AcceleratorsabstractData movement contributes significantly to energy consumption in CNNs. To address this overhead, this paper presents the Compound Memory-Computing (CMC) architecture for CNN inference accelerators that integrates both computing logic and weights in look-up tables, eliminating the need to load weights from any storage sources, including external memory, on-chip memory, or on-chip cache. CMC efficiently manages weight matrix values and structures without physical weight storage. Experimental results demonstrate that CMC achieves up to$2.9 \times$and$4.5 \times$energy efficiency gains over state-of-the-art CNN accelerators, and$24 \times$compared to a typical edge GPU, significantly mitigating energy consumption challenges in CNNs. Jian Dong 0010, Gang Qu 0001 |
ICCD | 3 |
| 2025 | Cheetah: Pipelined BFT Consensus Protocol with High Throughput and Low LatencyabstractRecent blockchain consensus protocols, such as HotStuff and Jolteon, employ pipelining combined with leader rotation to achieve both efficiency and fairness. However, these pipelined protocols encounter a significant challenge: performance degradation in the presence of faulty replicas. This issue stems from the dependency of block commits on multiple consecutive honest leaders within the pipelined protocols. The inability to ascertain the honesty of subsequent leaders exacerbates this problem. To address this limitation, we introduce Cheetah, a novel protocol that implements an independent block commitment mechanism. This mechanism enables Cheetah to commit blocks successfully without relying on multiple leaders (whether consecutive or not), thus preventing performance degradation in the presence of faulty replicas. Our results demonstrate that Cheetah consistently outperforms the state-of-the-art pipelined protocols across various scenarios. Jian Dong 0010 |
SRDS | 3 |
| 2025 | Harmonia: Enhancing Liveness in Two-Phase HotStuff Without Sacrificing Core PropertiesabstractHotStuff stands out as the first Byzantine Fault Tolerant (BFT) state machine replication protocol to achieve linear communication complexity while maintaining optimistic responsiveness and fairness. However, it is required at least three phases to commit a block, which leads to high latency. Recent studies have focused on developing two-phase variants of HotStuff to reduce its block commitment latency. Nonetheless, two-phase HotStuff suffers from the liveness issue, forcing these protocols to sacrifice one of the three fundamental properties of HotStuff when addressing this problem, thereby facing a trade-off among these properties. In this paper, we introduce a novel protocol, Harmonia, which, to the best of our knowledge, is the first BFT protocol to simultaneously achieve linear communication complexity, optimistic responsiveness, fairness, and the minimum commit latency of two phases. Harmonia overcomes the liveness issue of two-phase HotStuff through its specialized block unlocking mechanism while preserving all the original properties of HotStuff. Experimental results show that Harmonia consistently outperforms HotStuff across various metrics. Songyan Ji, Jian Dong 0010 |
SRDS | 4 |
| 2025 | SEOF: Reducing Spikes for Efficient SCNN Accelerators With Approximate ComputingabstractSpiking Convolutional Neural Networks (SCNNs), a variant of Spiking Neural Networks (SNNs) based on the architecture of Convolutional Neural Networks (CNNs), have demonstrated superior performance in tasks compared to traditional SNNs. However, this enhanced capability results in higher energy consumption due to increased computation and weight access. Research has extensively demonstrated that SCNN computation is directly related to spike counts. Thus, reducing the number of spikes can lead to more efficient implementations of SCNN inference accelerators. However, arbitrarily reducing spikes often results in an uncontrollable decline in accuracy. In this paper, we present a spike-efficient optimization framework, SEOF, which integrates approximate computing principles and exploits SCNN-specific characteristics to achieve energy savings while maintaining low inference accuracy loss. SEOF incorporates a novel spike-efficient controller for spiking neurons, a spike ratio-oriented objective function designed to produce accelerator-friendly SCNN models, and a spatial mask that lowers energy consumption by reducing weight accesses. Our experimental results indicate that SEOF can achieve 31%-51% computation energy savings by reducing spikes, with an accuracy loss of less than 1%, and a speedup of up to 1.56×. Additionally, SEOF can reduce weight access by up to 54.93%, leading to overall energy savings of up to 50.77%. Gang Qu 0001, Jian Dong 0010 |
IEEE Trans. Sustain. Comput. | 6 |
| 2024 | Lightning: Leveraging DVFS-induced Transient Fault Injection to Attack Deep Learning Accelerator of GPUsabstractGraphics Processing Units (GPU) are widely used as deep learning accelerators because of its high performance and low power consumption. Additionally, it remains secure against hardware-induced transient fault injection attacks, a classic type of attacks that have been developed on other computing platforms. In this work, we demonstrate that well-trained machine learning models are robust against hardware fault injection attacks when the faults are generated randomly. However, we discover that these models have components, which we refer to as sensitive targets, that are vulnerable to faults. By exploiting this vulnerability, we propose the Lightning attack, which precisely strikes the model’s sensitive targets with hardware-induced transient faults based on the Dynamic Voltage and Frequency Scaling (DVFS). We design a sensitive targets search algorithm to find the most critical processing units of Deep Neural Network (DNN) models determining the inference results, and develop a genetic algorithm to automatically optimize the attack parameters for DVFS to induce faults. Experiments on three commodity Nvidia GPUs for four widely-used DNN models show that the proposed Lightning attack can reduce the inference accuracy by 69.1% on average for non-targeted attacks, and, more interestingly, achieve a success rate of 67.9% for targeted attacks. Rihui Sun, Pengfei Qiu, Yongqiang Lyu 0001, Jian Dong 0010, Haixia Wang 0001, Dongsheng Wang 0002, Gang Qu 0001 |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2023 | ATC: Approximate Temporal Coding for Efficient Implementations of Spiking Neural NetworksabstractSpiking Neural Networks (SNN) update their neurons' states, the most energy consuming action, only after receiving or firing spikes for energy efficiency. So reducing the number of spikes would lead to more efficient SNN implementations. We propose an approximate temporal coding (ATC) for this purpose. Because the reduction of spikes leads to more synapses being used rarely, we develop a pruning method for further energy improvement. Experimental results validate the efficiency of ATC and the pruning method. On the MNIST dataset, for example, 61% of the spikes are reduced, leading to 60% energy saving without any accuracy loss. Jian Dong 0010, Gang Qu 0001 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2023 | A Guided Mutation Strategy for Smart Contract FuzzingabstractSmart contracts manage a large number of digital assets which is attractive to attackers. There have been many attacks that have caused huge financial losses. Therefore, it is of great importance to detect vulnerabilities in smart contracts. Fuzzing is considered a promising approach to test smart contracts. However, the complexity of changing state variables and the handling of external parameters during mutation pose critical technical challenges for current smart contract fuzzers, hindering their ability to cover branches under complex constraints and leaving potential vulnerabilities for attackers to exploit. To tackle these problems, we design a guided mutation strategy combined with two novel techniques: Dynamic Dependency Learning (DDL) and Dynamic Variables Analysis (DVA). DDL learns the dependencies of sequences to provide guided transaction sequence generation for handling state variables in complex constraints, while DVA leverages variable-level dynamic taint analysis to process the external parameters and guide the mutation. We implement the proposed strategy on a fuzzer, called SeqFuzz. The experimental results show that SeqFuzz could cover more branches and detect more bugs in real-world smart contracts compared with state-of-the-art tools. Songyan Ji, Jian Dong 0010, Lishi Lu |
ICSME | 2 |
| 2023 | Effuzz: Efficient fuzzing by directed search for smart contracts
Songyan Ji, Junfu Qiu, Jian Dong 0010 |
Inf. Softw. Technol. | 4 |
| 2022 | FADATest: Fast and Adaptive Performance Regression Testing of Dynamic Binary Translation SystemsabstractDynamic binary translation (DBT) is the cornerstone of many important applications. In practice, however, it is quite difficult to maintain the performance efficiency of a DBT system due to its inherent complexity. Although performance regression testing is an effective approach to detect potential performance regression issues, it is not easy to apply performance regression testing to DBT systems, because of the natural differences between DBT systems and common software systems and the limited availability of effective test programs. In this paper, we present FADATest, which devises several novel techniques to address these challenges. Specifically, FADATest automatically generates adaptable test programs from existing real benchmark programs of DBT systems according to the runtime characteristics of the benchmarks. The test programs can then be used to achieve highly efficient and adaptive performance regression testing of DBT systems. We have implemented a prototype of FADATest. Experimental results show that FADATest can successfully uncover the same performance regression issues across the evaluated versions of two popular DBT systems, QEMU and Valgrind, as the original benchmark programs. Moreover, the testing efficiency is improved significantly on two different hardware platforms powered by x86-64 and AArch64, respectively. Jian Dong 0010, Ruili Fang, Wenwen Wang 0001, De-Cheng Zuo |
ICSE | 2 |
| 2022 | WDBT: Non-volatile memory wear characterization and mitigation for DBT systems
Jian Dong 0010, Ruili Fang, Wenwen Wang 0001, De-Cheng Zuo |
J. Syst. Softw. | 2 |
| 2022 | Double-Shift: A Low-Power DNN Weights Storage and Access Framework based on Approximate Decomposition and QuantizationabstractOne major challenge in deploying Deep Neural Network (DNN) in resource-constrained applications, such as edge nodes, mobile embedded systems, and IoT devices, is its high energy cost. The emerging approximate computing methodology can effectively reduce the energy consumption during the computing process in DNN. However, a recent study shows that the weight storage and access operations can dominate DNN's energy consumption due to the fact that the huge size of DNN weights must be stored in the high-energy-cost DRAM. In this paper, we propose Double-Shift, a low-power DNN weight storage and access framework, to solve this problem. Enabled by approximate decomposition and quantization, Double-Shift can reduce the data size of the weights effectively. By designing a novel weight storage allocation strategy, Double-Shift can boost the energy efficiency by trading the energy consuming weight storage and access operations for low-energy-cost computations. Our experimental results show that Double-Shift can reduce DNN weights to 3.96%–6.38% of the original size and achieve an energy saving of 86.47%–93.62%, while introducing a DNN classification error within 2%. Jian Dong 0010, Gang Qu 0001 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2022 | RMLIM: A Runtime Machine Learning Based Identification Model for Approximate Computing on Data Flow GraphsabstractApproximate computing (AC) is an effective energy-efficient method for applications that have intrinsic error resilience. Early research efforts select noncritical portion of the computation, operations that have little impact on the accuracy of the results, for approximation. They ignore the runtime information and result in either under-approximation, which fails to reach the full potential of energy saving, or over-approximation, which causes unacceptable errors in the computation. A recently proposed runtime approach first estimates the noncritical portion for given input values and then performs accurate computation only on the critical portion. However, its complicated estimation process brings large runtime overhead and may not be suitable for real time embedded software. In this paper, we solve this problem by proposing a Runtime Machine Learning based Identification Model (RMLIM) to locate the noncritical portion in the data flow graph representation of any software program. RMLIM is trained offline by generated training data set and then applied at runtime for each input. This reduces the runtime complexity of identifying noncritical parts. Our experiments show that, compared with the existing runtime AC method, our machine learning based approach can maintain similar energy efficiency and computation accuracy, but reduces the execution time by 40 percent–61 percent. Jian Dong 0010, Yanxin Liu, Chunpei Wang, Gang Qu 0001 |
IEEE Trans. Sustain. Comput. | 2 |
| 2021 | FTApprox: A Fault-Tolerant Approximate Arithmetic Computing Data FormatabstractApproximate computing (AC) is an effective energy-efficient method for error-resilient applications. The essence behind AC is to reduce energy consumption by slightly sacrificing computation accuracy purposefully while providing quality-acceptable results. On the other hand, soft error is a common problem during program execution and may cause unacceptable outputs or catastrophic failure to the system. As AC introduces errors and soft errors are mitigated by fault-tolerant mechanisms, they have conflict goals and contradictory approaches. To the best of our knowledge, there is no previous efforts to consider the two at the same time. In this paper, we study the problem of AC with soft errors in order to guarantee the safe execution of the program while reducing energy (by AC). More specifically, we propose FTApprox, a fault-tolerant approximate arithmetic computing data format, to enable the detection and correction of SEs. As an approximate data format, FTApprox can use 16 bits to approximate any 32-bit integers and fixed-point numbers, and will select only the most significant part of operands for AC at runtime. Energy saving is obtained by converting 32-bit arithmetic operations to 8-bit operations. Meanwhile, for soft errors such as random bit flips, FTApprox not only can detect all single bit flips and most 2-bit flips, it can also correct most of these errors. The experimental results show that FTApprox has significant resistance against soft errors while providing 66.4%-79.6% energy saving. Jian Dong 0010, Qian Xu 0022, Gang Qu 0001 |
DATE | 2 |
| 2021 | Increasing Fuzz Testing Coverage for Smart Contracts with Dynamic Taint AnalysisabstractNowadays, smart contracts manage more and more digital assets and have become an attractive target for adversaries. To prevent smart contracts from malicious attacks, a thorough test is indispensable and must be finished before deployment because smart contracts cannot be modified after being deployed. Fuzzing is an important testing approach, but most existing smart contract fuzzers can hardly solve the constraints which involve deeply nested conditional statements, resulting in low coverage. To address this problem, we propose Targy, an efficient targeted mutation strategy based on dynamic taint analysis. We obtain the taint flow by dynamic taint propagation, and generate a more accurate mutation strategy for the input parameters of functions to simultaneously satisfy all conditional statements. We implemented Targy on sFuzz with 3.6 thousand smart contracts running on Ethereum. The numbers of covered branches and detected vulnerabilities increase by 6% and 7% respectively, and the average time required for covering a branch is reduced by 11 %. Songyan Ji, Jian Dong 0010, Junfu Qiu, Bowen Gu, Tongqi Wang |
QRS | 2 |
| 2021 | Effective exploitation of SIMD resources in cross-ISA virtualizationabstractSystem virtualization is a fundamental technology that enables many important applications. However, existing virtualization techniques suffer from a critical limitation due to their limited exploitation of host SIMD hardware resources, especially when a guest application does not have inherently fine-grained data-level parallelism. To bridge this utilization gap and unleash the full potential of host SIMD resources, this paper proposes an effective and unconventional SIMD exploitation technique. The proposed exploitation takes advantage of ample host SIMD registers and powerful host SIMD instructions to generate more efficient host binary code for guest applications even without any fine-grained data-level parallelism. It also mitigates the shortage of general-purpose registers on the host platform, as well as improves the efficiency of accessing guest registers. We have implemented the exploitation in an extensively-used virtualization platform, QEMU. Experimental results on a comprehensive list of benchmarks from PARSEC, SPEC-CPU2017, and Google Octane JavaScript benchmark suite show that an average of 2.2X performance speedup can be achieved for AArch64 binaries on an x86-64 host machine. We believe the proposed technique will provide a new perspective for our community to rethink the exploitation of SIMD hardware resources. Jian Dong 0010, Ruili Fang, Xiaoli Gong, Wenwen Wang 0001, De-Cheng Zuo |
VEE | 2 |
| 2020 | A Machine Learning based Approximate Computing Approach on Data Flow Graphs: Work-in-ProgressabstractWe report our ongoing work towards a machine learning based runtime approximate computing (AC) approach that can be applied on the data flow graph representation of any software program. This approach can utilize runtime inputs together with prior information of the software to identify and approximate the noncritical portion of a computation with low runtime overhead. Some preliminary experimental results show that compared with previous runtime AC approaches, our approach can significantly reduce the time overhead with little loss on the energy efficiency and computation accuracy. Jian Dong 0010, Yanxin Liu, Chunpei Wang, Gang Qu 0001 |
EMSOFT | 2 |
| 2020 | Is It Approximate Computing or Malicious Computing?abstractApproximate computing (AC) is an attractive energy efficient technique that can be implemented at almost all the design levels including data, algorithm, and hardware. The basic idea behind AC is to deliberately control the trade-off between computation accuracy and energy efficiency. However, with the introduction of AC, traditional computing frameworks are having many potential security vulnerabilities. In this paper, we analyze these vulnerabilities and the associated attacks as well as corresponding countermeasures. More importantly, we propose the vulnerability at data level and demonstrate that without appropriate security mechanism, adversaries can modify the data and convert a secure and trusted AC process to one that produces unexpected errors in the final output. Furthermore, it is difficult to distinguish whether such errors are caused by the approximation nature of AC or from malicious modification and injection. Finally, we propose the information hiding based countermeasures to defend against both existing attacks and the proposed data level attacks, which helps to answer the question: given an error in AC, whether it comes from approximation or it is maliciously introduced. Jian Dong 0010, Qian Xu 0022, Zhaojun Lu, Gang Qu 0001 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2020 | PerfDBT: Efficient Performance Regression Testing of Dynamic Binary TranslationabstractDynamic binary translation (DBT) has been adopted in many important applications. Due to the large scale and complexity of a DBT system, a minor code change may lead to unexpected impact on the performance. Therefore, it is necessary to conduct performance regression testing for DBT systems. However, existing benchmark suites are not suitable for daily performance testing due to the extremely-long testing time. To address this challenge, we propose PerffiBT, which employs a novel approach to automatically generate test programs from existing long-running benchmarks. The execution times of the generated test programs are much shorter, which allows them to be used for daily performance regression testing of a DBT system. Experimental results demonstrate that the test programs generated by PerfDBT can achieve an average of 71X testing efficiency compared to original benchmarks. Furthermore, it can also deliver similar testing results to original benchmarks. Jian Dong 0010, Ruili Fang, Wenwen Wang 0001, De-Cheng Zuo |
ICCD | 2 |
| 2020 | Improving the Dependability of Self-Adaptive Cyber Physical System With Formal Compositional ContractabstractTo adapt to the uncertain environment smartly and timely, cyber physical systems (CPSs) have to interact with the physical world in a decentralized but rigorous, organized way. Guaranteeing the timing reliability is key to achieve consensus on the order of distributed events, as well as dependable cooperative decision processing. Based on our hierarchically decentralized compositional self-adaptive framework, we propose a formal compositional reliability-contract-based solution to guarantee the timing reliability of event observation and decision processing in a large-scale, geographically distributed CPS. As the prophetic decision may not fit the local situation well because of the uncertainties, we propose a gradual contract optimization solution to refine the dependability, timeliness, and energy consumption. Following the seven proposed composition schemes, we employ the nondominated sorting genetic algorithm II (NSGA-II) algorithm to optimize arrangement of decision. Moreover, a topology-aware time reserving solution is applied to improve the resilience of processing time and to tolerance timing failures. Both simulation results and real-world testing are introduced to evaluate the efficacy of our proposal. We believe that the formal compositional contract will be a competitive CPS solution to analyze requirements and optimize the self-adaptation decision at runtime. Peng Zhou 0005, De-Cheng Zuo, Kun Mean Hou, Zhan Zhang 0002, Jian Dong 0010 |
IEEE Trans. Reliab. | 5 |
| 2019 | Information Hiding behind Approximate ComputationabstractThere are many interesting advances in approximate computing recently targeting the energy efficiency in system design and execution. The basic idea is to trade computation accuracy for power and energy during all phases of the computation, from data to algorithm and hardware implementation. In this paper, we explore how to utilize approximate computing for security based information hiding. More specifically, we will demonstrate with examples the potential of embedding information in approximate hardware and approximate data, as well as during approximate computation. We analyze both the security vulnerabilities that this may cause and the potential security applications enabled by such information hiding. We argue that information could be hidden behind approximate computation without compromising the computation accuracy or energy efficiency. Qian Xu 0022, Gang Qu 0001, Jian Dong 0010 |
ACM Great Lakes Symposium on VLSI | 4 |
| 2019 | Experimental Analysis and Comparison of Load Prediction Algorithms in Cloud Data CenterabstractDue to the increasing scale of cloud data center, the issue of energy consumption is becoming pretty significant. To tackle this problem, an extremely effective approach is increasing the utilization of resource in data center. Researchers have found that accurate load prediction can help allocator distribute resource reasonably, so as to increase the utilization. There are a lot of traditional prediction algorithms which have been applied to cloud data center, such as linear regression. However, with the development of technologies, a number of novel prediction algorithms are brought out, for example, neural network. This paper assesses and analyzes the performance of several different prediction algorithms applying on data sets from real world. We get some meaningful and interesting conclusions from comparison among these algorithms, which may offer references for system designers of cloud data center. Yanxin Liu, Jian Dong 0010, De-Cheng Zuo, Hongwei Liu 0002 |
QRS | 2 |