EDBT 2026 Demo / reviewers in the wild / expert
Zhenmin Li
dblp:41/2867
· DBLP profile ↗
40ranked-venue papers
10as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 27 · 5 first-author · 10 since 2021Software engineering, systems software and programming languages · 6 · 3 first-authorDatabases, data management, data science and information retrieval · 4 · 1 first-authorArtificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CRA-BPNet: A time-frequency feature fusion network for cuffless blood pressure estimation featuring constrained rhythm cross-attention
Zixiang Jin, Zhenmin Li, Ruohai Hu, Yongqiang Yu |
Expert Syst. Appl. | 2 |
| 2026 | A Falcon Signature Verification Accelerator Using Area-Efficient NTT Architecture With Simplified Barrett Modular MultiplierabstractFalcon is a lattice-based post-quantum digital-signature scheme standardized by U.S. National Institute of Standards and Technology (NIST). As for its hardware implementation, signature verification is the critical operation, in which the polynomial multiplier and the hash generator are computationally intensive. Existing hardware implementations face an inherent resource–performance tradeoff in the multiplier, and the overall verification latency remains high. In this article, we present an accelerator for Falcon signature verification using area-efficient number-theoretic transform (NTT) architecture with simplified Barrett modular multiplier. The precomputed constant of Barrett modular-multiplication unit is adjusted from 21 845 to 21 840. This modification maintains computation correctness while lowering resource consumption and improving throughput, which serves as an efficient building block for NTT/inverse NTT (INTT) operations. Moreover, a fully pipelined, nonstored radix-2 multipath delay commutator (R2-MDC) architecture is adopted for the NTT/INTT, which effectively reduces the delay and achieves a more balanced resource–performance tradeoff. Both the polynomial multiplier and the hash generator are parallelized to improve performance. A complete Falcon-1024 signature verification accelerator is implemented on FPGA platform. Experimental results show that, for the two-stage NTT inside the multiplier, the proposed design attains the lowest area–time product (ATP) compared to state-of-the-art designs, and the ATP of the entire Falcon-1024 verification procedure is reduced by 50%. Hongran Hu, Yangyi Chu, Gaoming Du, Duoli Zhang, Zhenmin Li |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2026 | An Efficient Accelerator for Dehazing Neural Network Based on Physical Perception Model and Cross-Scale Pixel AttentionabstractHaze reduces visibility, hindering real-time image processing applications. Although deep learning-based dehazing algorithms can significantly enhance image quality, their high computational complexity and storage demands make them challenging to deploy on resource-constrained hardware platforms. To address this, we propose an efficient dehazing neural network featuring physical perception model and cross-scale pixel attention (EP-CSANet), together with its hardware accelerator. In terms of algorithm, we design an efficient physical perception model (EPM) based on the atmospheric scattering model (ASM), combined with adaptive channel attention (ACA), which accurately approximates atmospheric light (A) and transmittance [t(x)], and effectively adapts to various environmental conditions, significantly improving dehazing accuracy. Furthermore, to enhance multiscale information interaction and preserve fine-grained texture details, we introduce cross-scale pixel attention (CSPA), which utilizes a dual-branch approach to extract high receptive-field information while maintaining fine-grained textures. In terms of hardware, we design a dedicated FPGA-based dehazing acceleration architecture. Through an efficient 16-stage pipeline, we achieve a dehazing rate of 127 frames per second (fps) for$640\times 480$images, meeting real-time processing requirements. In addition, by incorporating sparse convolution optimization techniques, we significantly improve resource utilization: LUTs increased by 31.8%, FFs by 27.9%, and DSPs by 22.3%. Experimental results demonstrate that EP-CSANet, using only 2361 parameters on the SOTS dataset, outperforms other dehazing algorithms. The source code is publicly available athttps://github.com/netflymachine/EP-CSANet Gaoming Du, Zhenmin Li, Duoli Zhang |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2026 | ParaPM: Efficient Hardware Accelerator for Postquantum Signature With High-Performance Polynomial MultiplierabstractDuring the Institute of Standards and Technology (NIST) postquantum cryptography standardization process, the lattice-based Dilithium scheme was selected as one of the three third-round finalists for digital signature algorithms. Although numerous hardware implementations of Dilithium have been proposed, there remains substantial room for performance optimization, particularly in terms of computational speed. In this work, we target the two most time-consuming operations, namely the coefficient generation and polynomial multiplication. We propose a high-speed hardware architecture that fully exploits hardware parallelism. Our design introduces a fully pipelined radix-2 multipath delay commutator (R2MDC) structure supporting both NTT and inverse NTT (INTT) modes, a throughput-matched polynomial multiplier enabling on-the-fly pointwise multiplication (PWM) without buffering, and a low-resource Keccak. These optimizations collectively eliminate the throughput mismatch bottleneck and maximize hardware utilization through task-level parallelism. We implement coefficient generation, signing, and verification for three security levels on the A7 and Z7 platforms. Experimental results demonstrate that our current design achieves the optimal area-time product (ATP) across all three security levels. Gaoming Du, Yongsheng Yin, Duoli Zhang, Zhenmin Li |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2025 | Image encryption/decryption accelerator based on Fast Cosine Number Transform
Zhenmin Li, Gaoming Du |
Integr. | 2 |
| 2025 | HALTRAV: Design of a High-Performance and Area-Efficient Latch With Triple-Node-Upset Recovery and Algorithm-Based VerificationsabstractWith the rapid advancement of semiconductor technologies, latches become increasingly sensitive to soft errors, especially triple node upsets (TNUs), in harsh radiation environments. In this article, we first propose a high-performance and area-efficient latch, namely, HALTRAV, featuring complete TNU-recovery. The storage portion of HALTRAV consists of 28 interlocked source-drain cross-coupled inverters (SCIs) for complete TNU-recovery with area efficiency and low delay. To mitigate the issue that node-upset-recovery verifications for existing latches highly relies on electronic design automation tools, we further propose an algorithm-based verification method that can automatically verify the node-upset-recovery of latches, which greatly simplifies the reliability-verification flow. Simulation results demonstrate the TNU-recovery of HALTRAV and also show that HALTRAV achieves 40.38%, 8.17%, and 31.89% reduction in delay, area, and delay-power–area product (DPAP) on average, respectively; however; it is at the cost of power as compared to typical latches that are TNU-recoverable. Comparison results also demonstrate the moderate sensitivity of HALTRAV to the impacts of the process, voltage, and temperature (PVT) variations. Zhenmin Li, Xiaoqing Wen, Patrick Girard 0001, Aibin Yan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2024 | A Low-Latency Polynomial Multiplier Accelerator for CRYSTALS-Dilithium Digital SignatureabstractIn the post-quantum signature CRYSTALS-Dilithium algorithm, polynomial multiplication accounts for 32% of the total computation Latency. It is necessary to design a high-performance polynomial multiplication module. In this paper, we have proposed a low-Latency polynomial multiplier. First, we insert registers in the Radix-2 Multi-path Delay Commutator (R2MDC) structure to increase clock frequency. Additionally, to further enhance the computational clock frequency, we have designed a five-stage pipelined Butterfly unit. Secondly, we have proposed a fully pipelined polynomial multiplier that supports polynomial point-wise multiplication during NTT/INTT transformations to save a significant number of cycles. We also designed a configurable Polynomial Pointwise Multiplication (PPM) module that supports calculations with three different security levels. Our polynomial multiplier structure is implemented on the K7 model of FPGA. Compared to the currently fastest design in terms of computational speed, we have saved between 13% and 39% of the computation latency. Gaoming Du, Zhuo Chen 0037, Zhenmin Li, Duoli Zhang |
ACM Great Lakes Symposium on VLSI | 3 |
| 2024 | ICLTR: A Input-split Inverters and C-elements based Low-Cost Latch with Triple-Node-Upset RecoveryabstractAs the semiconductor technology continues to advance, integrated circuits (ICs) are becoming increasingly sensitive to soft errors, e.g., double-node upsets (DNUs) and triplenode upsets (TNUs), induced by harsh radiation. In this paper, a low-cost latch design, namely ICLTR, using input-split inverters (ISIs) and C-elements to provide complete TNU recovery, is proposed. ICLTR consists of seven ISIs, seven 2-input C-elements and a clock-gated inverter, and all these elements are interlocked. Simulation results show the complete TNU recovery for ICLTR. The simulation results also show that ICLTR can save 59.5% of the transmission delay, 36.1% of the power consumption and 81.6% of the delay-area-power product (DAPP) on average when compared with the same type of TNU recovery latch designs. Zhenmin Li, Gaoyang Shan, Xiaoqing Wen |
ITC-Asia | 2 |
| 2024 | A fast hardware accelerator for nighttime fog removal based on image fusion
Tianyi Lv, Gaoming Du, Zhenmin Li, Peiyi Teng |
Integr. | 3 |
| 2022 | A BNN Accelerator Based on Edge-skip-calculation Strategy and Consolidation Compressed TreeabstractBinarized neural networks (BNNs) and batch normalization (BN) have already become typical techniques in artificial intelligence today. Unfortunately, the massive accumulation and multiplication in BNN models bring challenges to field-programmable gate array (FPGA) implementations, because complex arithmetics in BN consume too much computing resources. To relax FPGA resource limitations and speed up the computing process, we propose a BNN accelerator architecture based on consolidation compressed tree scheme by combining both XNOR and accumulation operation of the low bit into a systematic one. During the compression process, we adopt 0-padding (not ±1) to achieve no-accuracy-loss from software modeling to hardware implementation. Moreover, we introduce shift-addition-BN free binarization technique to shorten the delay path and optimize on-chip storage. To sum up, we drastically cut down the hardware consumption while maintaining great speed performance with the same model complexity as the previous design. We evaluate our accelerator on MNIST and CIFAR-10 dataset and implement the whole system on the ARTIX-7 100T FPGA with speed performance of 2052.65 GOP/s and area efficiency of 70.15 GOPS/KLUT. Gaoming Du, Bangyi Chen, Zhenmin Li, Zhenxing Tu, Shenya Wang, Qinghao Zhao, Yongsheng Yin |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2021 | An abnormal event detection method based on the Riemannian manifold and LSTM network
Limin Xia, Zhenmin Li |
Neurocomputing | 2 |
| 2021 | SCNET: A Novel UGI Cancer Screening Framework Based on Semantic-Level Multimodal Data FusionabstractUpper gastrointestinal (UGI) cancer has been identified as one of the ten most common causes of cancer deaths globally. UGI cancer screening is critical to improving the survival rate of UGI cancer patients. While many approaches to UGI cancer screening rely on single-modality data such as gastroscope imaging, limited studies have been dedicated to UGI cancer screening exploiting multisource and multimodal medical data, which could potentially lead to improved screening results. In this paper, we propose semantic-level cancer-screening network (SCNET), a framework for UGI cancer screening based on semantic-level multimodal upper gastrointestinal data fusion. Specifically, the proposed SCNET consists of a gastrointestinal image recognition flow and a textual medical record processing flow. High-level features of upper gastrointestinal data are extracted by identifying effective feature channels according to the correlation between the textual features and the spatial structure of the image features. The final screening results are obtained after the data fusion step. The experimental results show that the improvement of our approach over the state-of-the-art ones reached 4.01% in average. The source code of SCNET is available at https://github.com/netflymachine/SCNET. Shuai Ding 0001, Zhenmin Li, Xiao Liu 0004, Shanlin Yang |
IEEE J. Biomed. Health Informatics | 3 |
| 2021 | A new method of abnormal behavior detection using LSTM network with temporal attention mechanism
Limin Xia, Zhenmin Li |
J. Supercomput. | 2 |
| 2021 | A Real-Time Effective Fusion-Based Image Defogging Architecture on FPGAabstractFoggy weather reduces the visibility of photographed objects, causing image distortion and decreasing overall image quality. Many approaches (e.g., image restoration, image enhancement, and fusion-based methods) have been proposed to work out the problem. However, most of these defogging algorithms are facing challenges such as algorithm complexity or real-time processing requirements. To simplify the defogging process, we propose a fusional defogging algorithm on the linear transmission of gray single-channel. This method combines gray single-channel linear transform with high-boost filtering according to different proportions. To enhance the visibility of the defogging image more effectively, we convert the RGB channel into a gray-scale single channel without decreasing the defogging results. After gray-scale fusion, the data in the gray-scale domain should be linearly transmitted. With the increasing real-time requirements for clear images, we also propose an efficient real-time FPGA defogging architecture. The architecture optimizes the data path of the guided filtering to speed up the defogging speed and save area and resources. Because the pixel reading order of mean and square value calculations are identical, the shift register in the box filter after the average and the computation of the square values is separated from the box filter and put on the input terminal for sharing, saving the storage area. What’s more, using LUTs instead of the multiplier can decrease the time delays of the square value calculation module and increase efficiency. Experimental results show that the linear transmission can save 66.7% of the total time. The architecture we proposed can defog efficiently and accurately, meeting the real-time defogging requirements on 1920 × 1080 image size. Gaoming Du, Jiting Wu, Hongfang Cao, Kun Xing, Zhenmin Li, Duoli Zhang |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2019 | Efficient Softmax Hardware Architecture for Deep Neural NetworksabstractDeep neural network (DNN) has become a pivotal machine learning and object recognition technology in the big data era. The softmax layer is one of the key component layers for completing multi-classification tasks. However, the softmax layer contains complex exponential and division operations, resulting in low accuracy and long critical paths in hardware accelerator design. In order to solve the above issues, we present a softmax hardware architecture with proper accuracy, good trade-off and strong expansibility. We summarize the classification rules of neural network and balance the calculation accuracy between resource consumption. On this basis, we proposed an exponential calculation unit based on the group lookup table, and improve a natural logarithmic calculation unit based on the Maclaurin series and the data preprocessing scheme matching them. The experimental results show that the softmax hardware architecture proposed in this paper can achieve the calculation accuracy of 3 decimal fraction and the classification accuracy of $99.01%$. Theoretically, it can accomplish the classification task of infinite categories. Gaoming Du, Zhenmin Li, Duoli Zhang, Yongsheng Yin |
ACM Great Lakes Symposium on VLSI | 3 |
| 2019 | NR-MPA: Non-Recovery Compression Based Multi-Path Packet-Connected-Circuit Architecture of Convolution Neural Networks AcceleratorabstractConvolution Neural Networks (CNNs) involve massive data to be calculated and stored. To meet the challenges above, parallel hardware accelerators consisting of hundreds of Processing Elements (PEs) arranged as a many-core systemon-chip, connected by a Network-on-Chip (NoC) are proposed, which achieve high throughput exploiting parallel PE array. However, most of existing accelerators focus on only one aspect, such as compute structure of PE and data movement overhead above NoC, which causes the throughout, area and latency of the accelerator not fully optimized. In this paper, we propose an efficient general purpose CNN accelerator including both compute based on Non-Recovery Compression (NRC) method and data movement by novel Multi-Paths Packet Connection Circuit (MP-PCC). NRC can save computation time due to zero multiplier through shift decoding in PE and improve power efficiency by saving a large number of data transmission. MPPCC, evolved from Packet Connection Circuit, supports single and multicast transmission modes at the same time, and changes the multicast (X, Y) routing algorithm to multicast Y algorithm to improve the transmission efficiency. The proposed architecture which was implemented on Xilinx FPGA achieves 17.7x faster computation speed and 2.2x fewer memory accesses compared with the state-of-the-art method. Gaoming Du, Zhenwen Yang, Zhenmin Li, Duoli Zhang, Yongsheng Yin, Zhonghai Lu |
ICCD | 3 |
| 2019 | Smart electronic gastroscope system using a cloud-edge collaborative framework
Shuai Ding 0001, Zhenmin Li, Hao Wang 0081, Yanchun Zhang |
Future Gener. Comput. Syst. | 3 |
| 2019 | Diabetic complication prediction using a similarity-enhanced latent Dirichlet allocation model
Shuai Ding 0001, Zhenmin Li, Xiao Liu 0004, Shanlin Yang |
Inf. Sci. | 2 |
| 2019 | SSS: Self-aware System-on-chip Using a Static-dynamic Hybrid MethodabstractNetwork-on-Chip (NoC) has become the de facto communication standard for multi-core or many-core System-on-Chip (SoC) due to its scalability and flexibility. However, an important factor in NoC design is temperature, which affects the overall performance of SoC—decreasing circuit frequency, increasing energy consumption, and even shortening chip lifetime. In this article, we propose SSS, a self-aware SoC using a static-dynamic hybrid method that combines dynamic mapping and static mapping to reduce the hotspot temperature for NoC-based SoCs. First, we propose monitoring and thermal modeling for self-state sensoring. Then, in static mapping stage, we calculate the optimal mapping solutions under different temperature modes using the discrete firefly algorithm to help self-decision making. Finally, in dynamic mapping stage, we achieve dynamic mapping through configuring NoC and SoC sentient units for self-optimizing. Experimental results show that SSS has substantially reduced the peak temperature by up to 37.52%. The FPGA prototype proves the effectiveness and smartness of SSS in reducing hotspot temperature. Gaoming Du, Guanyu Liu, Zhenmin Li, Duoli Zhang, Minglun Gao, Zhonghai Lu |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2019 | Allocation and scheduling of SystemJ programs on chip multiprocessors with weighted TDMA scheduling
Muhammad Nadeem 0002, Zhenmin Li, Avinash Malik, Morteza Biglari-Abhari, Zoran A. Salcic |
J. Syst. Archit. | 2 |
| 2017 | SSS: self-aware system-on-chip using static-dynamic hybrid method (work-in-progress)abstractNetwork on chip has become the de facto communication standard for multi-core or many-core system on chip, due to its scalability and flexibility. However, temperature is an important factor in NoC design, which affects the overall performance of SoC---decreasing circuit frequency, increasing energy consumption, and even shortening chip lifetime. In this paper, we propose SSS, a self-aware SoC using a static-dynamic hybrid method, which combines dynamic mapping and static mapping to reduce the hot-spots temperature for NoC based SoCs. First, we propose monitoring the thermal distribution for self-state sensoring. Then, in static mapping stage, we calculate the optimal mapping solutions under different temperature modes using discrete firefly algorithm to help self-decision making. Finally, in dynamic mapping stage, we achieve dynamic mapping through configuring NoC and SoC sentient unit for self-optimizing. Experimental results show SSS can reduce the peak temperature by up to 30.64%. FPGA prototype shows the effectiveness and smartness of SSS in reducing hot-spots temperature. Gaoming Du, Shibi Ma, Zhenmin Li, Zhonghai Lu, Minglun Gao |
CASES | 3 |
| 2017 | On the Accuracy of Stochastic Delay Bound for Network on ChipabstractDelay bound guarantee in network on chip (NoC) is important for hard real-time applications, and deterministic network calculus (DNC) is a effective tool for delay bound modeling. But for soft real-time applications, delay bound derivation using DNC is often over-pessimistic, resulting in too much chip area (e.g., router buffer) and power consumption; stochastic network calculus (SNC), on the contrary, improves the delay bound accuracy by providing stochastic service curves. Existing service models assume that contention takes place as long as there exist contention flows from different input channels requesting the same output channel. These models only consider flow paths in flows contention analyzing. We have observed that, beyond flow path contentions, the arrival rate also has deep influence on the flow contention in NoC, consequently affecting delay bound. In this paper, we further analyze the intrinsic factors affecting the flow contention, and propose a stochastic analytic model of per-flow delay bound to improve the calculation accuracy, according to both path and arrival rate. Within this model, the end-to-end delay bound is evaluated based on SNC. Experimental results show that our proposed model is both effective and accurate. Gaoming Du, Zhenmin Li, Guanyu Liu, Duoli Zhang |
NOCS | 3 |
| 2017 | Using design space exploration for finding schedules with guaranteed reaction times of synchronous programs on multi-core architecture
Zhenmin Li, HeeJong Park 0001, Avinash Malik, Kevin I-Kai Wang, Zoran A. Salcic, Boris Kuzmin, Michael Glaß, Jürgen Teich |
J. Syst. Archit. | 1 |
| 2014 | Study on the precision evaluation index in building seismic damage information extraction from remote sensing imageabstractPrecision evaluation is an important step in remote sensing classification. The Kappa is widely used to evaluate the overall accuracy of classification result, but it cannot reflect the classification effect of a certain class. A new precision evaluation index R can apply to evaluate the accuracy of a certain category in classification image. In this paper, we extracted building damage information based on object-oriented method, and discussed the effect and robust of index R and Kappa when different categories of the classification result were involved in precision evaluation. The study shows that the standardized index R can reflect the efficiency of the classification result more stably and it has certain advantages in classification evaluation. Zhenmin Li, Aixia Dou |
IGARSS | 1 |
| 2014 | TACO: A scalable framework for timing analysis and code optimization of synchronous programsabstractStatic estimation of the Worst Case Reaction Time (WCRT) of synchronous programs is pivotal for designing hard-real time systems in these languages. The current approaches to WCRT estimation suffer from either large overestimation of the WCRT value or the state space explosion problem. In this paper, we present TACO: a framework that integrates model checking based WCRT estimation with code optimization techniques, which results in close to optimal WCRT estimates with orders of magnitude reduced worst case runtime complexity. Finally, the TACO framework also allows us to generate executables with a smaller overall memory footprint. Zhenmin Li, Avinash Malik, Zoran A. Salcic |
RTCSA | 1 |
| 2014 | Bug characteristics in open source software
Lin Tan 0001, Zhenmin Li, Xuanhui Wang, Yuanyuan Zhou 0001, ChengXiang Zhai |
Empir. Softw. Eng. | 3 |
| 2009 | Understanding Customer Problem Troubleshooting from Storage System Logs
Weihang Jiang, Chongfeng Hu, Shankar Pasupathy, Arkady Kanevsky, Zhenmin Li, Yuanyuan Zhou 0001 |
FAST | 5 |
| 2008 | CISpan: Comprehensive Incremental Mining Algorithms of Closed Sequential Patterns for Multi-Versional Software MiningabstractRecently, frequent sequential pattern mining algorithms have been widely used in software engineering field to mine various source code or specification patterns.In practice, software evolves from one version to another in its life span.The effort of mining frequent sequential patterns across multiple versions of a software can be substantially reduced by efficient incremental mining.This problem is challenging in this domain since the databases are usually updated in all kinds of manners including insertion, various modifications as well as removal of sequences.Also, different mining tools may have various mining constraints, such as low minimum support.None of the existing work can be applied effectively due to various limitations of such work.For example, our recent work, IncSpan, failed solving the problem because it could neither handle low minimum support nor removal of sequences from database.In this paper, we propose a novel, comprehensive incremental mining algorithm for frequent sequential pattern, CISpan (Comprehensive Incremental Sequential Pattern mining).CISpan supports both closed and complete incremental frequent sequence mining, with all kinds of updates to the database.Compared to IncSpan, CISpan tolerates a wide range for minimum support threshold (as low as 2).Our performance study shows that in addition to handling more test cases on which IncSpan fails, CISpan outperforms IncSpan in all test cases which IncSpan could handle, including various sequence length, number of sequences, modification ratio, etc., with an average of 3.4 times speedup.We also tested CISpan's performance on databases transformed from 20 consecutive versions of Linux Kernel source code.On average, CISpan outperforms the non-incremental CloSpan by 42 times. Ding Yuan 0004, Kyuhyung Lee, Hong Cheng 0001, Zhenmin Li, Xiao Ma 0014, Yuanyuan Zhou 0001, Jiawei Han 0001 |
SDM | 5 |
| 2007 | Address Code Optimization Exploiting Code Scheduling in DSP ApplicationsabstractExploitation of address generation units (AGUs) which are typically provided by digital signal processors (DSPs) plays an important role in DSP code generation for embedded processors. In this paper, a novel address code optimization technique for DSP code generation which integrates code scheduling has been presented. Specifically, we develop a probability based quality estimation scheme to search for the globally optimal address assignment. The proposed technique is general enough to handle the situations with commutative-input operands, loops, and conditional branches. The experimental results with DSP benchmark programs show 14%-54% improvement in terms of the number of nonzero-cost address instructions Zhenmin Li |
ISCAS | 1 |
| 2007 | MUVI: automatically inferring multi-variable access correlations and detecting related semantic and concurrency bugsabstractSoftware defects significantly reduce system dependability. Among various types of software bugs, semantic and concurrency bugs are two of the most difficult to detect. This paper proposes a novel method, called MUVI, that detects an important class of semantic and concurrency bugs. MUVI automatically infers commonly existing multi-variable access correlations through code analysis and then detects two types of related bugs: (1) inconsistent updates--correlated variables are not updated in a consistent way, and (2) multi-variable concurrency bugs--correlated accesses are not protected in the same atomic sections in concurrent programs.We evaluate MUVI on four large applications: Linux, Mozilla,MySQL, and PostgreSQL. MUVI automatically infers more than 6000 variable access correlations with high accuracy (83%).Based on the inferred correlations, MUVI detects 39 new inconsistent update semantic bugs from the latest versions of these applications, with 17 of them recently confirmed by the developers based on our reports.We also implemented MUVI multi-variable extensions to tworepresentative data race bug detection methods (lock-set and happens-before). Our evaluation on five real-world multi-variable concurrency bugs from Mozilla and MySQL shows that the MUVI-extension correctly identifies the root causes of four out of the five multi-variable concurrency bugs with 14% additional overhead on average. Interestingly, MUVI also helps detect four new multi-variable concurrency bugs in Mozilla that have never been reported before. None of the nine bugs can be identified correctly by the original race detectors without our MUVI extensions. Shan Lu 0001, Chongfeng Hu, Xiao Ma 0014, Weihang Jiang, Zhenmin Li, Raluca A. Popa, Yuanyuan Zhou 0001 |
SOSP | 6 |
| 2006 | LIFT: A Low-Overhead Practical Information Flow Tracking System for Detecting Security AttacksabstractComputer security is severely threatened by software vulnerabilities. Prior work shows that information flow tracking (also referred to as taint analysis) is a promising technique to detect a wide range of security attacks. However, current information flow tracking systems are not very practical, because they either require program annotations, source code, non-trivial hardware extensions, or incur prohibitive runtime overheads. This paper proposes a low overhead, software-only information flow tracking system, called LIFT, which minimizes run-time overhead by exploiting dynamic binary instrumentation and optimizations/or detecting various types of security attacks without requiring any hardware changes. More specifically, LIFT aggressively eliminates unnecessary dynamic information flow tracking, coalesces information checks, and efficiently switches between target programs and instrumented information flow tracking code. We have implemented LIFT on a dynamic binary instrumentation framework on Windows. Our real-system experiments with two real-world server applications, one client application and eighteen attack benchmarks show that LIFT can effectively detect various types of security attacks. LIFT also incurs very low overhead, only 6.2% for server applications, and 3.6 times on average for seven SPEC INT2000 applications. Our dynamic optimizations are very effective in reducing the overhead by a factor of 5-12 times Feng Qin 0003, Cheng Wang 0013, Zhenmin Li, Ho-Seop Kim, Yuanyuan Zhou 0001, Youfeng Wu |
MICRO | 3 |
| 2006 | CP-Miner: Finding Copy-Paste and Related Bugs in Large-Scale Software CodeabstractRecent studies have shown that large software suites contain significant amounts of replicated code. It is assumed that some of this replication is due to copy-and-paste activity and that a significant proportion of bugs in operating systems are due to copy-paste errors. Existing static code analyzers are either not scalable to large software suites or do not perform robustly where replicated code is modified with insertions and deletions. Furthermore, the existing tools do not detect copy-paste related bugs. In this paper, we propose a tool, CP-Miner, that uses data mining techniques to efficiently identify copy-pasted code in large software suites and detects copy-paste bugs. Specifically, it takes less than 20 minutes for CP-Miner to identify 190,000 copy-pasted segments in Linux and 150,000 in FreeBSD. Moreover, CP-Miner has detected many new bugs in popular operating systems, 49 in Linux and 31 in FreeBSD, most of which have since been confirmed by the corresponding developers and have been rectified in the following releases. In addition, we have found some interesting characteristics of copy-paste in operating system code. Specifically, we analyze the distribution of copy-pasted code by size (number lines of code), granularity (basic blocks and functions), and modification within copy-pasted code. We also analyze copy-paste across different modules and various software versions. Zhenmin Li, Shan Lu 0001, Suvda Myagmar, Yuanyuan Zhou 0001 |
IEEE Trans. Software Eng. | 1 |
| 2005 | PR-Miner: automatically extracting implicit programming rules and detecting violations in large software codeabstractPrograms usually follow many implicit programming rules, most of which are too tedious to be documented by programmers. When these rules are violated by programmers who are unaware of or forget about them, defects can be easily introduced. Therefore, it is highly desirable to have tools to automatically extract such rules and also to automatically detect violations. Previous work in this direction focuses on simple function-pair based programming rules and additionally requires programmers to provide rule templates.This paper proposes a general method called PR-Miner that uses a data mining technique called frequent itemset mining to efficiently extract implicit programming rules from large software code written in an industrial programming language such as C, requiring little effort from programmers and no prior knowledge of the software. Benefiting from frequent itemset mining, PR-Miner can extract programming rules in general forms (without being constrained by any fixed rule templates) that can contain multiple program elements of various types such as functions, variables and data types. In addition, we also propose an efficient algorithm to automatically detect violations to the extracted programming rules, which are strong indications of bugs.Our evaluation with large software code, including Linux, PostgreSQL Server and the Apache HTTP Server, with 84K--3M lines of code each, shows that PR-Miner can efficiently extract thousands of general programming rules and detect violations within 2 minutes. Moreover, PR-Miner has detected many violations to the extracted rules. Among the top 60 violations reported by PR-Miner, 16 have been confirmed as bugs in the latest version of Linux, 6 in PostgreSQL and 1 in Apache. Most of them violate complex programming rules that contain more than 2 elements and are thereby difficult for previous tools to detect. We reported these bugs and they are currently being fixed by developers. Zhenmin Li |
ESEC/SIGSOFT FSE | 1 |
| 2005 | Mining block correlations to improve storage performanceabstractBlock correlations are common semantic patterns in storage systems. They can be exploited for improving the effectiveness of storage caching, prefetching, data layout, and disk scheduling. Unfortunately, information about block correlations is unavailable at the storage system level. Previous approaches for discovering file correlations in file systems do not scale well enough for discovering block correlations in storage systems.In this article, we propose two algorithms, C-Miner and C-Miner *, that use a data mining technique called frequent sequence mining to discover block correlations in storage systems. Both algorithms run reasonably fast with feasible space requirement, indicating that they are practical for dynamically inferring correlations in a storage system. C-Miner is a direct application of a frequent-sequence mining algorithm with a few modifications; compared with C-Miner , C-Miner * is redesigned for mining block correlations by making concessions for the specific problem of long sequences in storage system traces. Therefore, C-Miner * can discover 7--109% more correlation rules within 2--15 times shorter time than C-Miner . Moreover, we have also evaluated the benefits of block correlation-directed prefetching and data layout through experiments. Our results using real system workloads show that correlation-directed prefetching and data layout can reduce average I/O response time by 12--30% compared to the base case, and 7--25% compared to the commonly used sequential prefetching scheme for most workloads. Zhenmin Li, Yuanyuan Zhou 0001 |
ACM Trans. Storage | 1 |
| 2005 | Performance directed energy management for main memory and disksabstractMuch research has been conducted on energy management for memory and disks. Most studies use control algorithms that dynamically transition devices to low power modes after they are idle for a certain threshold period of time. The control algorithms used in the past have two major limitations. First, they require painstaking, application-dependent manual tuning of their thresholds to achieve energy savings without significantly degrading performance. Second, they do not provide performance guarantees.This article addresses these two limitations for both memory and disks, making memory/disk energy-saving schemes practical enough to use in real systems. Specifically, we make four main contributions. (1) We propose a technique that provides a performance guarantee for control algorithms. We show that our method works well for all tested cases, even with previously proposed algorithms that are not performance-aware. (2) We propose a new control algorithm, Performance-Directed Dynamic (PD), that dynamically adjusts its thresholds periodically, based on available slack and recent workload characteristics. For memory, PD consumes the least energy when compared to previous hand-tuned algorithms combined with a performance guarantee. However, for disks, PD is too complex and its self-tuning is unable to beat previous hand-tuned algorithms. (3) To improve on PD, we propose a simpler, optimization-based, threshold-free control algorithm, Performance-Directed Static (PS). PS periodically assigns a static configuration by solving an optimization problem that incorporates information about the available slack and recent traffic variability to different chips/disks. We find that PS is the best or close to the best across all performance-guaranteed disk algorithms, including hand-tuned versions. (4) We also explore a hybrid scheme that combines PS and PD algorithms to further improve energy savings. Zhenmin Li, Yuanyuan Zhou 0001, Sarita V. Adve |
ACM Trans. Storage | 2 |
| 2004 | Performance directed energy management for main memory and disksabstractMuch research has been conducted on energy management for memory and disks. Most studies use control algorithms that dynamically transition devices to low power modes after they are idle for a certain threshold period of time. The control algorithms used in the past have two major limitations. First, they require painstaking, application-dependent manual tuning of their thresholds to achieve energy savings without significantly degrading performance. Second, they do not provide performance guarantees. In one case, they slowed down an application by 835.This paper addresses these two limitations for both memory and disks, making memory/disk energy-saving schemes practical enough to use in real systems. Specifically, we make three contributions: (1) We propose a technique that provides a performance guarantee for control algorithms. We show that our method works well for all tested cases, even with previously proposed algorithms that are not performance-aware. (2) We propose a new control algorithm, Performance-directed Dynamic (PD), that dynamically adjusts its thresholds periodically, based on available slack and recent workload characteristics. For memory, PD consumes the least energy, when compared to previous hand-tuned algorithms combined with a performance guarantee. However, for disks, PD is too complex and its self-tuning is unable to beat previous hand-tuned algorithms. (3) To improve on PD, we propose a simple, optimization-based, threshold-free control algorithm, Performance-directed Static (PS). PS periodically assigns a static configuration by solving an optimization problem that incorporates information about the available slack and recent traffic variability to different chips/disks. We find that PS is the best or close to the best across all performanceguaranteed disk algorithms, including hand-tuned versions. Zhenmin Li, Francis M. David, Pin Zhou, Yuanyuan Zhou 0001, Sarita V. Adve |
ASPLOS | 2 |
| 2004 | C-Miner: Mining Block Correlations in Storage Systems
Zhenmin Li, Sudarshan M. Srinivasan, Yuanyuan Zhou 0001 |
FAST | 1 |
| 2004 | Reducing Energy Consumption of Disk Storage Using Power-Aware Cache ManagementabstractReducing energy consumption is an important issue for data centers. Among the various components of a data center, storage is one of the biggest consumers of energy. Previous studies have shown that the average idle period for a server disk in a data center is very small compared to the time taken to spin down and spin up. This significantly limits the effectiveness of disk power management schemes. This paper proposes several power-aware storage cache management algorithms that provide more opportunities for the underlying disk power management schemes to save energy. More specifically, we present an off-line power-aware greedy algorithm that is more energy-efficient than Belady’s off-line algorithm (which minimizes cache misses only). We also propose an online power-aware cache replacement algorithm. Our trace-driven simulations show that, compared to LRU, our algorithm saves 16% more disk energy and provides 50% better average response time for OLTP I/O workloads. We have also investigated the effects of four storage cache write policies on disk energy consumption. Qingbo Zhu, Francis M. David, Christo Frank Devaraj, Zhenmin Li, Yuanyuan Zhou 0001 |
HPCA | 4 |
| 2004 | CP-Miner: A Tool for Finding Copy-paste and Related Bugs in Operating System Code
Zhenmin Li, Shan Lu 0001, Suvda Myagmar, Yuanyuan Zhou 0001 |
OSDI | 1 |
| 2003 | Fragmentation based D-MAC Protocol in Wireless Ad Hoc NetworkabstractDirectional antennas have been recently suggested to be used in wireless ad hoc networks to reduce interference outside the antenna direction and to increase spatial reuse of wireless channels. Several MAC schemes that exploit directional antennas have been proposed, among which the two D-MAC schemes proposed in [5] may have received the most attention. In this paper, we carefully analyze how the second D-MAC scheme proposed in [5] operates in the hidden/exposed terminal scenarios and show that it not only introduces new collisions, but also reduces spatial reuse to some extent. To remedy these problems, we propose an enhanced version of the second D-MAC scheme, called fragmentation based D-MAC (FD-MAC), to reduce collision, to achieve better spatial reuse (than the second D-MAC scheme), and to improve fairness. Through ns-2 simulation, we demonstrate the effectiveness of FD-MAC, and in particular, we show that in complex scenarios such as multihop forwarding and mesh topologies, FD-MAC significantly outperforms DCF with RTS/CTS and the two D-MAC schemes, both in terms of throughput and fairness. Zhenmin Li, Pin Zhou, Jennifer C. Hou |
ICDCS | 1 |