Ye Lu 0004

dblp:05/511-4 · DBLP profile ↗
← Back
28ranked-venue papers
0as first author
22since 2021 · last 2026
0000-0003-0805-6394ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 9 since 2021Computer networks · 5 · 4 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 FuXi-γ: Efficient Sequential Recommendation with Exponential-Power Temporal Encoder and Diagonal-Sparse Positional Mechanism
abstract
Sequential recommendation aims to model users' evolving preferences based on their historical interactions. Recent advances leverage Transformer-based architectures to capture global dependencies, but existing methods often suffer from high computational overhead, primarily due to discontinuous memory access in temporal encoding and dense attention over long sequences. To address these limitations, we propose FuXi-γ, a novel sequential recommendation framework that improves both effectiveness and efficiency through principled architectural design. FuXi-γ adopts a decoder-only Transformer structure and introduces two key innovations: (1) An exponential-power temporal encoder that encodes relative temporal intervals using a tunable exponential decay function inspired by the Ebbinghaus forgetting curve. This encoder enables flexible modeling of both short-term and long-term preferences while maintaining high efficiency through continuous memory access and pure matrix operations. (2) A diagonal-sparse positional mechanism that prunes low-contribution attention blocks using a diagonal-sliding strategy guided by the persymmetry of Toeplitz matrix. Extensive experiments on four real-world datasets demonstrate that FuXi-γ achieves state-of-the-art performance in recommendation quality, while accelerating training by up to 4.74× and inference by up to 6.18×, making it a practical and scalable solution for long-sequence recommendation. Code: https://github.com/Yeedzhi/FuXi-gamma.
Dezhi Yi, Wei Guo 0006, Wenyang Cui, Huifeng Guo, Yong Liu 0020, Zhenhua Dong, Ye Lu 0004
KDD (1)8
2026 ByteEye: A smart contract vulnerability detection framework at bytecode level with graph neural networks
abstract
Smart contract vulnerability detection has attracted increasing attention due to billions of economic losses caused by vulnerabilities. Existing smart contract vulnerability detection methods have high false negative and high false positive rates. To address these issues, we present ByteEye, a bytecode level smart contract vulnerability detection framework with Graph Neural Networks (GNNs). ByteEye first constructs an edge-enhanced Control Flow Graph (CFG) to maintain rich information from the low-level bytecode with low latency. ByteEye also designs and incorporates both general information and vulnerability-specific information into its detection method as bytecode level features. Furthermore, ByteEye flexibly supports machine/deep learning models, especially with graph neural networks, which can facilitate vulnerability detection precisely. The extensive experimental results highlight that ByteEye outperforms the state-of-the-art approaches on all three types of vulnerability detection. ByteEye can achieve an average of 35.29%, 43.95%, and 6.38% higher on F1 than the bytecode level best-performed baseline on reentrancy vulnerability, timestamp dependency vulnerability, and integer overflow/underflow vulnerability, respectively. Moreover, ByteEye can detect 361 new vulnerabilities in real-world smart contracts, which are reported for the first time. ByteEye enhances control flow information, designs general bytecode-level features with expert knowledge, and flexibly supports deep learning models, particularly GNNs, thus achieving high detection effectiveness.
Jinni Yang, Shuang Liu 0007, Surong Dai, Yaozheng Fang, Kunpeng Xie, Ye Lu 0004
Autom. Softw. Eng.6
2026 Towards efficient Graph-RAG via structure-aware intermediate representation: Incremental collaborative exploration on knowledge graph
Ye Lu 0004, Tao Li 0022
Knowl. Based Syst.4
2026 EENet: An Efficient and Effective Network for Large-Scale CTR Prediction
abstract
Efficient and effective modeling of feature interactions is key to large-scale Click-Through Rate (CTR) prediction. Although existing feature interaction methods have improved the model accuracy, their computational consumption still increase exponentially with the number of feature fields and become severe efficiency bottleneck in real-world industrial scenarios. To address the issues, we propose an E fficient and E ffective NET work for large-scale CTR prediction named EENet . EENet presents a new alternating stacking architecture of implicit and explicit interaction layers, and each implicit layer in EENet can reduce both local computational and parameter load remarkably. EENet also designs a unified explicit interaction operation which can only use simple matrix multiplication to capture field-wise patterns. Moreover, the order of multiplications in EENet is rearranged to further decrease the computational complexity from quadratic to linear with respect to the number of feature fields. EENet thus can support the high efficiency in real-practice industrial scenarios with hundreds of feature fields. A set of extensive experiments is performed on two public datasets and one industrial dataset for effectiveness evaluation, and five larger-scale synthetic datasets for efficiency evaluation. The results highlight that our EENet can significantly outperform the state-of-the-art models in terms of both efficiency and scalability, while also maintaining superior effectiveness. Compared with DCNv2 and FiBiNet, EENet achieves 8.06 \(\times\) and 36.72 \(\times\) efficiency improvements in training, and 2.02 \(\times\) and 48.88 \(\times\) improvements in inference, respectively. Our solution and source code are available at https://github.com/Yeedzhi/EENet .
Dezhi Yi, Bo Chen 0023, Ye Lu 0004, Suqi Shi, Yangsen Liu, Wei Guo 0006, Kenan Song, Huifeng Guo, Yong Liu 0020, Zhenhua Dong, Ruiming Tang
ACM Trans. Inf. Syst.3
2025 ORQ-ViT: Outlier resilient Post Training Quantization for vision transformers via outlier decomposition
abstract
Post-training quantization (PTQ) is critical for deploying Vision Transformers (ViTs) on resource-constrained devices. However, outliers clustered in activation channels tend to dominate the quantization range and induce significant accuracy degradation. To address the outlier challenge, this paper proposes an outlier resilient PTQ method through outlier decomposition, namely ORQ-ViT. Its core idea is to decompose outliers clustered in outlier channels into isolated outliers, so that they can be easily excluded, thus alleviating their adverse impact. Specifically, we decompose activations along the patch token dimension. Since there are very few outlier channels, decomposed rows after outlier decomposition usually cover several isolated outliers, which can be easily identified and filtered. We further design an adaptive quantization range determination strategy during quantization parameters initialization to prevent outliers from serving as boundary values of the quantization range. ORQ-ViT can improve quantization levels utilization to generate activations with higher quantization resolution, thereby achieving higher accuracy. Additionally, ORQ-ViT supports pure integer matrix multiplications to ensure the inference efficiency of quantized ViTs on edge hardware. Extensive experiments demonstrate that our method achieves state-of-the-art (SOTA) accuracy across various ViT variants under multiple low bit-width scenarios. For image classification, the top-1 accuracy of ORQ-ViT outperforms that of the SOTA methods by an average of 2.26% at W4A4. Even for object detection and instance segmentation, ORQ-ViT also delivers highly competitive results. We also evaluate the inference efficiency of pure integer matrix multiplications and the results show that our method can achieve up to 2.1× speedup.
Xinyu He 0002, Ye Lu 0004
J. Syst. Archit.2
2025 CROSC: Compilation-Runtime Joint Optimization for Fast Smart Contract Execution
abstract
State access is a critical part of smart contract execution which seriously affects the efficiency of smart contract execution in the mainstream Ethereum blockchain. To reduce state access latency, existing studies typically require manual source code modifications, which in practice may deliver limited performance gains and shift the burden to developers. In this paper, we propose CROSC to reduce state access latency and improve smart contract execution efficiency by a compilation-runtime joint optimization approach. CROSC consists of three key parts: 1) a runtime memory management mechanism named Fast State Memory (FastSM) to fully utilize the working memory and provide the context for the contract compiler; 2) a State Variable Address Relocation (SVAR) strategy to minimize costly persistent storage operations by precisely redirecting state variable access targets during compilation; 3) a one-shot unpacking design that eliminates frequent decoding overhead for low-bitwidth state variables. Preliminary experimental results highlight that, compared with the baseline compilation and runtime system of Ethereum, CROSC can achieve 2.5× and 7.5× speedups for single state load and store operations, respectively. CROSC reduces state access latency by up to 81.3%, and overall contract execution latency by 32.9% on average across 14 typical types of smart contracts. Extended evaluations on ERC20 and ERC721 token standard contracts show that CROSC delivers significant benefits in critical areas while remaining unobtrusive for less intensive state operations.
Surong Dai, Jinni Yang, Wenyang Cui, Yaozheng Fang, Ye Lu 0004
IEEE Trans. Computers5
2024 Deep Feature Surgery: Towards Accurate and Efficient Multi-exit Networks
Yao Chen 0008, Qiuyang Luo, Ye Lu 0004, Tao Li 0022
ECCV (49)4
2024 ADS-CNN: Adaptive Dataflow Scheduling for lightweight CNN accelerator on FPGAs
Xianzhong Xie, Kunpeng Xie, Dezhi Yi, Ye Lu 0004, Keke Gai
Future Gener. Comput. Syst.6
2024 AutoQNN: An End-to-End Framework for Automatically Quantizing Neural Networks
Ye Lu 0004, Surong Dai, Deng Qian, Chengkun Du, Tao Li 0022
J. Comput. Sci. Technol.2
2024 DRS: A deep reinforcement learning enhanced Kubernetes scheduler for microservice-based system
abstract
Summary Recently, Kubernetes is widely used to manage and schedule the resources of microservices in cloud‐native distributed applications, as the most famous container orchestration framework. However, Kubernetes preferentially schedules microservices to nodes with rich and balanced CPU and memory resources on a single node. The native scheduler of Kubernetes, called Kube‐scheduler, may cause resource fragmentation and decrease resource utilization. In this paper, we propose a deep reinforcement learning enhanced Kubernetes scheduler named DRS. We initially frame the Kubernetes scheduling problem as a Markov decision process with intricately designed state , action , and reward structures in an effort to increase resource usage and decrease load imbalance. Then, we design and implement DRS mointor to perceive six parameters concerning resource utilization and create a thorough picture of all available resources globally. Finally, DRS can automatically learn the scheduling policy through interaction with the Kubernetes cluster, without relying on expert knowledge about workload and cluster status. We implement a prototype of DRS in a Kubernetes cluster with five nodes and evaluate its performance. Experimental results highlight that DRS overcomes the shortcomings of Kube‐scheduler and achieves the expected scheduling target with three workloads. With only 3.27% CPU overhead and 0.648% communication delay, DRS outperforms Kube‐scheduler by 27.29% in terms of resource utilization and reduces load imbalance by 2.90 times on average.
Zhaolong Jian, Xueshuo Xie, Yaozheng Fang, Yibing Jiang, Ye Lu 0004, Ankan Dash, Tao Li 0022, Grace Guiling Wang
Softw. Pract. Exp.5
2024 Winols: A Large-Tiling Sparse Winograd CNN Accelerator on FPGAs
abstract
Convolutional Neural Networks (CNNs) can benefit from the computational reductions provided by the Winograd minimal filtering algorithm and weight pruning. However, harnessing the potential of both methods simultaneously introduces complexity in designing pruning algorithms and accelerators. Prior studies aimed to establish regular sparsity patterns in the Winograd domain, but they were primarily suited for small tiles, with domain transformation dictating the sparsity ratio. The irregularities in data access and domain transformation pose challenges in accelerator design, especially for larger Winograd tiles. This paper introduces “Winols,” an innovative algorithm-hardware co-design strategy that emphasizes the strengths of the large-tiling Winograd algorithm. Through a spatial-to-Winograd relevance degree evaluation, we extensively explore domain transformation and propose a cross-domain pruning technique that retains sparsity across both spatial and Winograd domains. To compress pruned weight matrices, we invent a relative column encoding scheme. We further design an FPGA-based accelerator for CNN models with large Winograd tiles and sparse matrix-vector operations. Evaluations indicate our pruning method achieves up to 80% weight tile sparsity in the Winograd domain without compromising accuracy. Our Winols accelerator outperforms dense accelerator by a factor of 31.7× in inference latency. When compared with prevailing sparse Winograd accelerators, Winols reduces latency by an average of 10.9×, and improves DSP and energy efficiencies by over 5.6× and 5.7×, respectively. When compared with the CPU and GPU platform, Winols accelerator with tile size 8× 8 achieves 24.6× and 2.84× energy efficiency improvements, respectively.
Kunpeng Xie, Ye Lu 0004, Xinyu He 0002, Dezhi Yi, Huijuan Dong, Yao Chen 0008
ACM Trans. Archit. Code Optim.2
2024 PaVM: A Parallel Virtual Machine for Smart Contract Execution and Validation
abstract
The performance bottleneck of blockchain has shifted from consensus to serial smart contract execution in transaction validation. Previous works predominantly focus on inter-contract parallel execution, but they fail to address the inherent limitations of each smart contract execution performance. In this paper, we propose PaVM, the first smart contract virtual machine that supports both inter-contract and intra-contract parallel execution to accelerate the validation process. PaVM consists of (1) key instructions for precisely recording entire runtime information at the instruction level, (2) a runtime system with a re-designed machine state and thread management to facilitate parallel execution, and (3) a read/write-operation-based receipt generation method to ensure both the correctness of operations and the consistency of blockchain data. We evaluate PaVM on the Ethereum testnet, demonstrating that it can outperform the mainstream blockchain client Geth. Our evaluation results reveal that PaVM speeds up overall validation performance by 33.4×, and enhances validation throughput by up to 46×.
Yaozheng Fang, Surong Dai, Jinni Yang, Hui Zhang 0002, Ye Lu 0004
IEEE Trans. Parallel Distributed Syst.6
2023 TSC-VEE: A TrustZone-Based Smart Contract Virtual Execution Environment
abstract
TrustZone as a trusted execution environment (TEE) has been proven to preserve the confidentiality of blockchain transactions supported by smart contracts. Despite some academic effort, TrustZone can only support limited languages for now. The lack of the corresponding execution environment for smart contracts seriously hinders blockchain applications from directly running on TrustZone. In this paper, we design the first virtual execution environment named TSC-VEE for performing Solidity smart contracts on TrustZone, to the best of our knowledge. TSC-VEE can be decomposed into fourfold: (1) an instruction set adapted to the isolation and world switching mechanism of TrustZone. (2) a runtime memory management mechanism that provides a pair of instructions with the corresponding processing mechanism to allocate and release the work memory. (3) a hybrid granularity resource analysis algorithm which computes and records the value of maximum stack height and static gas cost through bytecode pre-execution, avoiding runtime overflow and invalid computations. (4) a cross-isolation-environment prefetching approach that supports loading and storing the storage data from the normal world into the secure world on TrustZone before execution, thus avoiding switching the world state frequently at runtime. Extensive experimental results show that TSC-VEE can perform smart contracts correctly and efficiently on TrustZone. Compared with the most commonly used Ethereum client—Geth, TSC-VEE achieves execution performance improvements by$9.29\times$. We also implement the Ethereum virtual machine—evmoneon TrustZone. TSC-VEE can reduce the latency by 12.63% with our optimization techniques, and decrease the work memory footprint by 22.95% on average when executing various scale contracts.
Zhaolong Jian, Ye Lu 0004, Youyang Qiao, Yaozheng Fang, Xueshuo Xie, Dayi Yang, Tao Li 0022
IEEE Trans. Parallel Distributed Syst.2
2023 Differential Testing of Machine Translators Based on Compositional Semantics
abstract
Powered by the advances of deep neural networks, machine translation software has achieved rapid progresses recently. Machine translators are widely adopted in people’s daily lives, e.g., for information consumption, medical consumption and online shopping. However, machine translators are far from robust, and may produce wrong translations, which could potentially cause misunderstandings or even serious consequences. It is thus critical to detect errors in machine translators, and provide informative feedback for developers. In this work, we adopt the differential testing method to test machine translators. In particular, we use mature commercial translators as reference machine translation engines. Based on the principle of compositionality, which specifies that the meaning of a complex expression is determined by the meanings of its constituent expressions and the syntactic rules used to combine them, we design the oracle which conducts similarity comparison guided by syntactic structure and semantic encoding. In particular, we employ the constituency parsing to obtain the part-whole structure relation between a sentence and one of its component. Then we compute the semantic similarity of each sentence part with pre-trained language model and expert knowledge. We implement our approach into a tool named DCS, conduct experiments on three popular machine translators, i.e., Google translate, Baidu translate and Microsoft Bing translate, and compare DCS with two state-of-the-art approaches, i.e., CIT and CAT. The experiment results show that DCS achieves 8.6% and 35.4% higher precision, respectively. Moreover, the errors reported by DCS have the lowest redundancy in terms of the duplicated error locations in the source sentence. DCS can be used in complement with existing approaches and achieve higher detection precision. It also shows comparable efficiency with state-of-the-art approaches.
Shuang Liu 0007, Shujie Dou, Junjie Chen 0003, Zhirun Zhang, Ye Lu 0004
IEEE Trans. Software Eng.5
2022 ATOM: Architectural Support and Optimization Mechanism for Smart Contract Fast Update and Execution in Blockchain-Based IoT
abstract
Blockchain-based Internet of Things (BC-IoT) brings the advantages of blockchain into traditional IoT systems. In BC-IoT, the smart contract has been widely used for automatic, trusted, and decentralized applications. Smart contracts require frequent adjust and fast update due to various reasons, such as inevitable code bugs, changes of applications, or security requirements. However, previous smart contract architecture and updating mechanism are low speed and cause high overhead, because they are based on recompilation and redeployment in BC-IoT. Meanwhile, smart contract execution is so time consuming due to contract instruction dispatching and operand loading in the stack-based Ethereum virtual machine (EVM). To address these issues, we propose a new smart contract architecture and optimization mechanism for BC-IoTs, ATOM, which provides architectural supports to update contract economically and fast executing in instructionwise for the first time, to the best of our knowledge. We design a compact Application-oriented Instruction (AoI) set to describe application operations. We can construct the bytecode of smart contract from application by directly assembling templates prebuilt upon the AoIs rather than by compilation. We also present an optimized mechanism for AoI execution to enable access addressable storage place rather than the indirect access through stack. We perform ATOM on a BC-IoT testbed based on private Ethereum and Hyperledger Burrow. The experimental results highlight that ATOM is more efficient than state-of-the-art approaches. ATOM can reduce update latency by 62.7%, ledger size by 70%, and gas usage by 90% on average, respectively. Compared with the traditional smart contract architecture, ATOM can improve EVM Memory access efficiency significantly by up to$10\times $and achieve improvement of execution efficiency with up to$1.6\times $.
Tao Li 0022, Yaozheng Fang, Zhaolong Jian, Xueshuo Xie, Ye Lu 0004, Grace Guiling Wang
IEEE Internet Things J.5
2022 Elastic Significant Bit Quantization and Acceleration for Deep Neural Networks
abstract
Quantization has been proven to be a vital method for improving the inference efficiency of deep neural networks (DNNs). However, it is still challenging to strike a good balance between accuracy and efficiency while quantizing DNN weights or activation values from high-precision formats to their quantized counterparts. We propose a new method called elastic significant bit quantization(ESB) that controls the number of significant bits of quantized values to obtain better inference accuracy with fewer resources. We design a unified mathematical formula to constrain the quantized values of the ESB with a flexible number of significant bits. We also introduce a distribution difference aligner (DDA) to quantitatively align the distributions between the full-precision weight or activation values and quantized values. Consequently, ESB is suitable for various bell-shaped distributions of weights and activation of DNNs, thus maintaining a high inference accuracy. Benefitting from fewer significant bits of quantized values, ESB can reduce the multiplication complexity. We implement ESB as an accelerator and quantitatively evaluate its efficiency on FPGAs. Extensive experimental results illustrate that ESB quantization consistently outperforms state-of-the-art methods and achieves average accuracy improvements of 4.78%, 1.92%, and 3.56% over AlexNet, ResNet18, and MobileNetV2, respectively. Furthermore, ESB as an accelerator can achieve 10.95 GOPS peak performance of 1k LUTs without DSPs on the Xilinx ZCU102 FPGA platform. Compared with CPU, GPU, and state-of-the-art accelerators on FPGAs, the ESB accelerator can improve the energy efficiency by up to 65, 11, and 26, respectively.
Ye Lu 0004, Kunpeng Xie, Zongming Jin, Tao Li 0022, Yanzhi Wang 0001
IEEE Trans. Parallel Distributed Syst.2
2022 SmartVM: A Smart Contract Virtual Machine for Fast On-Chain DNN Computations
abstract
Blockchain-based artificial intelligence (BC-AI) has been applied for protecting deep neural network (DNN) data from being tampered with, which is expected to further boost trusted distributed AI applications in many fields. However, due to smart contract execution environment architectural defects, it is challenging for previous BC-AI systems to support computing-intensive tasks on-chain performing such as DNN convolution operations. They have to offload computations and a large amount of data from blockchain to off-chain platforms to execute smart contracts as native code. This failure to take advantage of data locality has become one of the major critical performance bottlenecks in BC-AI system. To this end, in this article, we propose SmartVM with optimization methods to support on-chain DNN inference for BC-AI system. The key idea is to design and optimize the computing mechanism and storage structure of smart contract execution environment according to the characteristics of DNN such as high computational parallelism and large data volume. We decompose SmartVM into three components: 1) a compact DNN-oriented instruction set to describe computations in a short number of instructions to reduce interpretation time. 2) a memory management mechanism to make SmartVM memory dynamic free/allocated according to the size of DNN feature maps. 3) a block-based weight prefetching and parallel computing method to organize each layer's computing and weights prefetching in a pipelined manner. We perform the typical image classification in a private Ethereum blockchain testbed to evaluate SmartVM performance. Experimental results highlight that SmartVM can support DNN inference on-chain with roughly the same efficiency against the native code execution. Compared with the traditional off-chain computing, SmartVM can speed up the overall execution by70×,16×,11×, and12×over LeNet5, AlexNet, ResNet18, and MobileNet, respectively. The memory footprint can be reduced by84%,90.8%,94.3%, and93.7%over the above four models, while offering the same level model accuracy. This article sheds light on the design space of the smart contract virtual machine for DNN computation and is promising to further boost BC-AI applications.
Tao Li 0022, Yaozheng Fang, Ye Lu 0004, Jinni Yang, Zhaolong Jian, Zhiguo Wan, Yusen Li
IEEE Trans. Parallel Distributed Syst.3
2021 WIP: Sysnif: Constructing Workflow from Interleaved Logs in Intelligent IoT System
abstract
The massive smart devices in intelligent IoT can be broken due to malicious attacks and system failures. As a nonintrusive method, workflows mined from system logs facilitate administrators to quickly locate and diagnose anomalies in time. System logs are usually interleaved since there are lots of concurrent and asynchronous operations and executions on large scale IoT devices. Consequently, it is so challenging to construct an adaptive workflow from these logs and realize the real-time anomaly detection. To meet this challenge, in this paper, we propose a two-stage workflow construction approach named Sysnif, which includes offline construction and online adjustment. First, the window-based dependence computing method is employed to obtain the context of execution paths. Second, a weight-greedy algorithm is designed to denoise the interleaved system logs effectively. Third, in order to match system mechanism variation, the online micro-iteration adjusting algorithm is presented to update the workflow model. Experiment results highlight that Sysnif can outperform state-of-the-art methods, such as Logsed, on dataset of OpenStack logs by 22.4% on recall, meanwhile maintaining the same precision roughly. Sysnif can achieve an average precision and recall of 93.8% and 94.7%, respectively.
Zongming Jin, Xueshuo Xie, Yaozheng Fang, Zhaolong Jian, Ye Lu 0004, Guangying Li
WOWMOM5
2021 OSN: Onion-ring support neighbors for correspondence selection
Ye Lu 0004, Chunying Song, Tao Li 0022, Kai Wang 0001
Inf. Sci.2
2021 A Confidence-Guided Evaluation for Log Parsers Inner Quality
Xueshuo Xie, Zhi Wang 0014, Xuhang Xiao, Ye Lu 0004, Shenwei Huang, Tao Li 0022
Mob. Networks Appl.4
2021 VecQ: Minimal Loss DNN Model Compression With Vectorized Weight Quantization
abstract
Quantization has been proven to be an effective method for reducing the computing and/or storage cost of DNNs. However, the trade-off between the quantization bitwidth and final accuracy is complex and non-convex, which makes it difficult to be optimized directly. Minimizing direct quantization loss (DQL) of the coefficient data is an effective local optimization method, but previous works often neglect the accurate control of the DQL, resulting in a higher loss of the final DNN model accuracy. In this paper, we propose a novel metric, called Vector Loss. Using this new metric, we decompose the minimization of the DQL to two independent optimization processes, which significantly outperform the traditional iterative L2 loss minimization process in terms of effectiveness, quantization loss as well as final DNN accuracy. We also develop a new DNN quantization solution called VecQ, which provides minimal direct quantization loss and achieve higher model accuracy. In order to speed up the proposed quantization process during model training, we accelerate the quantization process with a parameterized probability estimation method and template-based derivation calculation. We evaluate our proposed algorithm on MNIST, CIFAR, ImageNet, IMDB movie review and THUCNews text data sets with numerical DNN models. The results demonstrate that our proposed quantization solution is more accurate and effective than the state-of-the-art approaches yet with more flexible bitwidth support. Moreover, the evaluation of our quantized models on Salient Object Detection (SOD) tasks maintains comparable feature extraction quality with up to 16× weight size reduction.
Yao Chen 0008, Ye Lu 0004, Tao Li 0022, Cong Hao, Deming Chen
IEEE Trans. Computers3
2021 Fast Policy Interpretation and Dynamic Conflict Resolution for Blockchain-Based IoT System
abstract
Although the blockchain‐based Internet of Things (BC‐IoT) has been applied in many fields, it still faces many security attacks due to lacking policy‐based security management (PbSM). Previous PbSM is usually time‐consuming, which is difficult to integrate into BC‐IoT directly. The high‐latency policy conflict resolving in traditional PbSM cannot meet the BC‐IoT’s low‐latency requirement. Moreover, the conflict resolution rate is low as the PbSM usually neglects the runtime information. Therefore, it is challenging that achieving an efficient PbSM for BC‐IoT and overcomes both time and resource consumption. To address the problem, we propose a novel PbSM for BC‐IoT named FPICR to realize fast policy interpretation and dynamic conflict resolution efficiently. We first present policy templates based on system log to interpret policy in high speed in BC‐IoT. Benefiting from matching the characteristics of the system processing, FPICR supports interpreting a policy into the smart contract directly without complex content parsing. We then propose a weighted directed policy graph (WDPG) to evaluate the importance of the deployed policies more accurately. To improve the policy conflict resolution rate, we implement the resolution algorithm through reconstructing the WDPG. Taking the traits of these properties, FPICR thus can also remove the redundant data to compress storage space by the WDPG. Experiment results highlight that FPICR outperforms the baseline in all measure metrics. Especially, compared with the state‐of‐the‐art method, the speedup of interpretation in FPICR is about up to 2.1×. The conflict resolution rate in FPICR can be improved by 6.2% on average and achieve up to 96.1%.
Yaozheng Fang, Zhaolong Jian, Zongming Jin, Xueshuo Xie, Ye Lu 0004, Tao Li 0022
Wirel. Commun. Mob. Comput.5
2020 Confidence guided anomaly detection model for anti-concept drift in dynamic logs
Xueshuo Xie, Zongming Jin, Jiming Wang, Ye Lu 0004, Tao Li 0022
J. Netw. Comput. Appl.5
2019 An Efficient Log Parsing Algorithm Based on Heuristic Rules
Xueshuo Xie, Kunpeng Xie, Zhi Wang 0014, Ye Lu 0004, Yujun Zhang 0001
APPT5
2019 LHC: A Low-Power Heterogeneous Computing Method on Neural Network Accelerator
abstract
Accelerators can achieve high performance and low energy consumption in training or inference of neural networks. If the Non-Neural Network (Non-NN) algorithms with large amount of computation could make full use of the accelerators, it is possible to speed up its implementation, reduce energy consumption, and achieve load balancing, especially on mobile devices equipped with accelerators. However, accelerators are dedicated to neural network calculations, so that other Non-NN algorithms have difficulty in using their advantages. Furthermore, many hardware-specific restrictions have become the obstacles, such as constrained precision of operands and limited computation scale. In this paper, we propose a method named Low-power Heterogeneous Computing (LHC) to bridge the gap between Non-NN algorithms and NN accelerators. Firstly, we analyze the general principle of the accelerator and reveal the calculation model of the accelerator. To hide the details of the underlying neural network library, we extract some operators from the limited number of types of neural network computation they support. We encapsulate the low-level library, extract operators suitable for general algorithms, and implement some more advanced operators that can adapt to the constrained hardware conditions. These operators could facilitate programmers to implement some Non-NN algorithms. In the aspect of the algorithm, we extract the computationally intensive parts of the Non-NN algorithm and deploy these computational tasks on the accelerator by calling the operators. To verify our method, we implement three Non-NN algorithms by using operators and adjusting these algorithms, include Grid-based Motions Statistics, k-Nearest Neighbors, and k-Means, on a specific accelerator, Cambricon-1A. The experimental results show that the energy consumption of calculation is reduced by up to 5.4x, compared with the CPU baseline. Our method can be further applied to other similar accelerators.
Fangxin Liu, Kunpeng Xie, Shusheng Liu, Ye Lu 0004, Tao Li 0022
ICPADS5
2019 µL2Q: An Ultra-Low Loss Quantization Method for DNN Compression
abstract
Data quantization has been proved to be an effective method to compress deep neural networks (DNNs) by using less bits to represent the parameters and intermediate data. The bit width of the data directly affects the memory footprint, computing capability, and energy consumption during the computation of the DNN models. Although there have been numerous existing studies on data quantization, there is still no quantitative analysis of the existing quantization methods, which results in empirical quantization with unpredictable DNN accuracy loss. To address this problem, we propose an effective method, called ultra-low loss quantization (μL2Q), to provide DNN quantization schemes based on comprehensive quantitative data analysis. μL2Q builds the transformation of the original data to a data space with standard normal distribution, and then find the optimal parameters to minimize the loss of the quantization of a targeted bit width. In addition, we integrate the proposed μL2Q into a popular machine learning framework Caffe for convenient end-to-end DNN design and training. By comparing to the state-of-the-art DNN compression designs, μL2Q shows the greatest ability to maintain DNN accuracy after quantization. In the experiments, our proposed method can deliver 4.42%, 16.70%, 1.95%, 8.26% and 5.63% accuracy improvements on Lenet-5, Cifarnet, VGG7-64 and Resnet-18 (Top1/5), respectively, compared to the state-of-the-art solutions with the same compression ratio.
Tao Li 0022, Ye Lu 0004, Cong Hao, Xiaofan Zhang 0001, Deming Chen, Yao Chen 0008
IJCNN3
2017 How do you breathe-a non-contact monitoring method using depth data
abstract
Respiration rate is considered among the most useful biomedical signals to be observed for it has the potential of reflecting the bare biological status of human body, and by which clinic may detect and analyze symptoms in obstructive pulmonary disease. Measuring the breathing with contact methods seems to be unfriendly to patients and not as accurate as expected. There have been many techniques based on non-contact methods for respiration detecting published. Referring to those related experiment, this paper is devoted to applications of these methods using Microsoft Kinect depth sensor for non-contact monitoring of respiration. A special attention is paid to visualization of results and motion mapping over the selected chest area. The proposed methodology applies digital signal processing methods and functional transforms for acquired data filtering, frequency spectrum analysis, and feature distinction. The frequency and fluctuations of the diagram drawn from the change of abdomen areas depth data acquired by Kinect match with that of realistic breathing process, verifying the correspondence between breathing and abdomen movement and the validity of this experiment. The noninvasive and easy method for a depth-sensor to detect breathing frequency promises the possibility to analyze breathing for diagnostic and monitoring purposes at hospital, physicians office, or even home environments.
Qingcheng Li, Ye Lu 0004
Healthcom4
2017 Connecting Paper to Digitization: a Homework Data Processing System with Data Labeling and Visualization
abstract
How to monitor students' learning in daily learning behavior and guide teachers to teach more rationally are very important problems. While the digital teaching resources are popular, we can not ignore the information of traditional paper media. We design a homework processing system to solve the problems. First of all, we begin the operation from the acquisition side, by photo taking and image preprocessing, we make the homework picture with semantic meanings. Then we construct an answer reference system, which shows the answers of different kinds of incorrectbess as samples for teachers to refer to when correcting. Semantic pictures are dispatched to teachers handheld devices for intelligent and precise correct activities. We divide correct process into determining the degree of error and giving comments to evaluate and inspire students. Finally, we design a web page for homework data management and correction statistics to help teachers understand the homework completion status of students. The design of homework correction scheme is easy to deploy and can be used in most teaching environments.
Qingcheng Li, Ye Lu 0004
MobiQuitous3