Wei Li 0008

dblp:64/6025-8 · DBLP profile ↗
← Back
47ranked-venue papers
4as first author
16since 2021 · last 2026
0009-0006-0237-1034ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 29 · 2 first-author · 12 since 2021Computer networks · 8 · 1 first-authorSoftware engineering, systems software and programming languages · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Cambricon-CIM: Enabling Energy-Efficient and Error-Resilient Analog CIM Acceleration via Reformation of Coding Bases
abstract
Recently, multi-bit slicing has emerged as a promising technique to improve the energy efficiency of charge-domain Compute-In-Memory (CIM) accelerators by reducing the number of Analog-to-Digital (A/D) conversions. However, multi-bit slicing requires shift-and-add operations to reconstruct outputs, which exponentially amplify errors and cause significant accuracy degradation. Existing works mainly rely on hardware-aware retraining or noise-suppression techniques, incurring considerable design or power overhead. Thus, multi-bit CIM designs often face the dilemma of trading off energy efficiency for error resilience. In this paper, we propose Cambricon-CIM, a charge-domain multi-bit CIM accelerator that achieves both high energy efficiency and strong error resilience, without requiring retraining. The core insight is that the error amplification is proportional to digit weights; and by redefining these digit weights with smaller non-binary coding bases, it is possible to reduce the total error amplification. Leveraging this principle, CambriconCIM dynamically selects the minimal coding bases for every analog dot-product. With novel circuit and architectural support, Cambricon-CIM enables fast, low-overhead reconfiguration of coding bases at runtime. Experimental results show that Cambricon-CIM achieves 2.27× energy efficiency and 3.06× performance over RAELLA, a state-of-the-art error-resilient multi-bit slicing CIM architecture.
Hongrui Guo, Tianrui Ma, Zidong Du, Mo Zou, Yifan Hao 0001, Yongwei Zhao 0001, Rui Zhang 0040, Wei Li 0008, Xing Hu 0001, Zhiwei Xu 0002, Qi Guo 0001, Tianshi Chen 0002
HPCA8
2026 FlashAttention-T: Towards Fully Tensorized Attention by Exploiting Tensor-Vector Parallelism
abstract
The attention mechanism is central to modern deep learning, particularly in large language models (LLMs), but suffers from quadratic computational complexity. To accelerate attention computation on GPUs, fused attention techniques (e.g., FlashAttention) consolidate the matrix multiplication (GEMM) and softmax computations into a single kernel. However, these operations remain computationally decoupled: the GEMM leverages high-performance tensor units (Tensor Cores), while the softmax executes on slower vector units (CUDA cores). This imbalance induces severe vector intervals—periods where tensor units sit idle awaiting vector unit completion—significantly underutilizing tensor units. Furthermore, ongoing hardware advancements delivering faster tensor units exacerbate this bottleneck.
Jianxing Xu, Yuanbo Wen 0001, Jun Bi, Ruibai Xu, Guanglin Xu, Rui Zhang 0040, Wei Li 0008, Ling Li 0001, Tianshi Chen 0002, Qi Guo 0001, Yunji Chen
PPoPP7
2026 Cambricon-QM: A Hybrid Architecture for Microscaling Format Training
Yongwei Zhao 0001, Chang Liu 0021, Zidong Du, Xing Hu 0001, Yimin Zhuang, Yifan Hao 0001, Xinkai Song, Wei Li 0008, Xishan Zhang, Ling Li 0001, Zhiwei Xu 0002, Tianshi Chen 0002, Qi Guo 0001
IEEE Trans. Computers11
2026 UniOrch: A Unified Mixed Framework for High-Efficiency LLM Training on Heterogeneous AI Chips
abstract
Efficient coordination of heterogeneous AI chips (GPU/NPU/DCU) in data centers is crucial for Large Language Model (LLM) training, but this process is hindered by architectural mismatches, protocol fragmentation, and network partitioning. Existing solutions fail to achieve unified resource management across different chips, resulting in severe resource fragmentation and reduced allreduce efficiency. To overcome these limitations, this paper proposes the unified coordination framework UniOrch, which not only integrates three core functionalities, including hardware abstraction, software standardization, and communication coordination, but also enables the training and inference of large models across heterogeneous AI chips. UniOrch's hardware-agnostic bare-metal cloud eliminates virtualization overhead through Border Gateway Protocol Ethernet Virtual Private Network (BGP EVPN) overlay networks and gateway-based chip integration; its PyTorch-based adaptation layer masks hardware differences and reduces migration costs; the Transformer Collective Communication Library (TCCL) unifies NCCL, HCCL, and OpenMPIprotocols to support seamless hybrid parallel training. Furthermore, the framework ' score scheduling mechanism is the Heterogeneous Hybrid Estimation Model (HHEM), which employs a two stage cost model combining static analysis with dynamic runtime feedback to dynamically allocate computing power based on Transformer task loads, ensuring cross-chip task synchronization, resource pooling, and dynamic allocation. Deployment verification in real-world production environments (e.g., China Construction Bank) shows that UniOrch achieves significant improvements: resource utilization of heterogeneous AI infrastructure is increased by 35%, cross-chip latency is reduced by 42%, accuracy loss in heterogeneous environments is <0.8%.
Jia Wang 0038, Yang Zhai, Haojie Wang 0004, Wanxin Song, Wei Li 0008, Zhaofeng He 0001
IEEE Trans. Parallel Distributed Syst.10
2025 Cambricon-DG: An Accelerator for Redundant-Free Dynamic Graph Neural Networks Based on Nonlinear Isolation
Zhifei Yue, Xinkai Song, Tianbo Liu 0006, Xing Hu 0001, Rui Zhang 0040, Zidong Du, Wei Li 0008, Qi Guo 0001, Tianshi Chen 0002
HPCA7
2025 Cambricon-SR: An Accelerator for Neural Scene Representation with Sparse Encoding Table
abstract
Neural Scene Representation (NSR) is a promising technique for representing real scenes.By learning from dozens of 2D photos captured from different viewpoints, NSR computes the 3D representation of real scenes.However, the performance of NSR processing running on GPU is insufficient for applications.Cambricon-R achieves high performance of more than 60 scenes per second, but at the cost of modeling quality.
Tianbo Liu 0006, Xinkai Song, Zhifei Yue, Xing Hu 0001, Zhuoran Song, Yuanbo Wen 0001, Yifan Hao 0001, Wei Li 0008, Zidong Du, Rui Zhang 0040, Jiaming Guo, Shaohui Peng, Guangzhong Sun, Qi Guo 0001, Tianshi Chen 0002
ISCA9
2025 Morphology generalizable reinforcement learning via multi-level graph features
Yansong Pan, Rui Zhang 0040, Jiaming Guo, Shaohui Peng, Kaizhao Yuan, Yunkai Gao 0001, Siming Lan, Ruizhi Chen, Ling Li 0001, Xing Hu 0001, Zidong Du, Xin Zhang 0062, Wei Li 0008, Qi Guo 0001, Yunji Chen
Neurocomputing15
2025 SaaP: Rearchitect SoC-as-a-Processor to Orchestrate Hardware Heterogeneity
abstract
Due to the end of Moore’s Law and Dennard Scaling, Domain-Specific Accelerators (DSAs) have come to a Cambrian explosion. Especially when advancing into the intelligent era, more and more DSAs are integrated into System-on-Chips (SoCs) as intellectual property (IP) blocks to provide high performance and efficiency. Currently, IPs usually expose IP-dependent hardware interfaces, requiring SoCs to manage them as isolated devices with software running on the host CPU. However, such software-managed heterogeneity in CPU-centric SoCs leads to low IP utilization. This inefficiency arises from the dependence on software optimization, coupled with the control and data exchange overheads. To improve IP utilization of heterogeneous SoCs, in this article, we rearchitect the SoC as a processor (i.e., SaaP) to orchestrate hardware heterogeneity. SaaP features an orchestration pipeline where DSAs are integrated as execution units and managed directly by the hardware pipeline to conceal the hardware heterogeneity from software. Moreover, SaaP redesigns the register file and data paths to implement an IP-level data-forwarding mechanism, avoiding the costly control and data exchange in the CPU-centric execution model. Block data dependence among different DSAs is carefully resolved to exploit mixed-level parallelism and inter-IP data exchange. SaaP abstracts tasks as mixed-scale instructions, where each instruction can be mapped to different IPs. Experimental results show that compared against Xavier on six fully software-optimized benchmarks from different domains, SaaP-rearchitected Xavier achieves a$2.08{\times }$speedup, with an 8.21% area reduction and only 2.98% increase in power consumption.
Pengwei Jin, Zhe Fan, Yongwei Zhao 0001, Zidong Du, Hongrui Guo, Ziyuan Nan, Yifan Hao 0001, Chongxiao Li, Tianyun Ma, Xiaqing Li, Wei Li 0008, Xing Hu 0001, Qi Guo 0001, Zhiwei Xu 0002, Tianshi Chen 0002
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.12
2024 AutoOS: Make Your OS More Powerful by Exploiting Large Language Models
abstract
With the rapid development of Artificial Intelligence of Things (AIoT), customizing and optimizing operating system (OS) kernel configurations for various AIoT application scenarios is crucial for maximizing system performance. However, existing approaches falter due to the overwhelming problem complexity (i.e., over 15,000 configuration options in the Linux kernel), together with the huge evaluation costs and error-prone options that may result in OS boot-up failure, which all make it an unresolved problem to optimize the Linux kernel automatically. In this paper, we introduce AutoOS, a novel framework exploiting Large Language Models for customizing and optimizing OS kernel configurations automatically for various AIoT application scenarios.Inspired by the inherently directory-structured kernel configuration process, we first formulate our research problem as optimizing on a dynamic tree. We then propose a novel framework integrating a state machine-based traversal algorithm as the observe-prune-propose-act-correct loop, which can effectively refine the optimization space and ensure a successful OS boot-up.Experimental results show that AutoOS can automatically customize and optimize the OS kernel configurations without human effort. More importantly, AutoOS even achieves better performance by up to 25% than vendor-provided configuration.
Huilai Chen, Yuanbo Wen 0001, Limin Cheng, Shouxu Kuang, Ling Li 0001, Rui Zhang 0040, Xinkai Song, Wei Li 0008, Qi Guo 0001, Yunji Chen
ICML10
2024 Cambricon-D: Full-Network Differential Acceleration for Diffusion Models
abstract
Diffusion models have made significant progress in current image generation tasks, thus becoming a prominent area of research. Diffusion models necessitate repetitive iterations on minimally altered input data across timesteps, each timestep requiring the recalculation of the entire model, resulting in a remarkable computational redundancy and substantial hardware expenditures.Performing differential computing on input data seems to be a feasible approach for addressing such computational redundancy and improving hardware efficacy. However, non-linear operations (particularly activation functions) necessitate the merging of deltas (i.e., differential values) with raw inputs repeatedly to ensure computational correctness, leading to significant memory access for loading raw inputs, which fragmentedly blocks the forwarding of deltas throughout the network and undermines performance.To solve this problem, we propose Cambricon-D, a fullnetwork differential computing architecture with concise memory access. While maintaining the computational efficiency brought by differential computing, Cambricon-D employs a sign-mask dataflow, which requires only the loading of 1-bit signs (instead of large bitwidth raw inputs), thereby facilitating the seamless forwarding of deltas and effectively mitigating memory access overheads. Experimental results show that, compared to Diffy, Cambricon-D’s dataflow reduces 66% ~ 82% off-chip memory access. In total, Cambricon-D achieves 1.46× ~ 2.38× speedup over A100 on various diffusion models with different resolutions.
Weihao Kong, Yifan Hao 0001, Qi Guo 0001, Yongwei Zhao 0001, Xinkai Song, Xiaqing Li, Mo Zou, Zidong Du, Rui Zhang 0040, Chang Liu 0021, Yuanbo Wen 0001, Pengwei Jin, Xing Hu 0001, Wei Li 0008, Zhiwei Xu 0002, Tianshi Chen 0002
ISCA14
2024 ECFO: An Efficient Edge Classification-Based Fusion Optimizer for Deep Learning Compilers
abstract
Operation fusion is a critical technique in optimizing deep learning compilers as it enhances computational efficiency by integrating multiple operations into a single computational graph. However, finding an effective fusion strategy is challenging, requiring the definition of an optimization search space and identification of the best strategy within this space. Existing methods, such as heuristic searches and learning-based searches, have significant limitations. Heuristic searches are complex, labor-intensive, and often lack generalizability across different network architectures. On the other hand, learning-based methods demand extensive training and pro-longed search time. To address these challenges, we introduce the Edge Classification-Based Fusion Optimizer (ECFO), a novel approach that reconceptualizes operation fusion as an edge classification problem. By leveraging Graph Neural Networks (GNNs) for efficient graph feature encoding, ECFO streamline the optimization process and significantly reduces computational overhead. Comprehensive evaluations across diverse neural networks demonstrate that ECFO decrease search time by up to 23x and improves inference performance by 3.2%, representing a substantial advancement over existing strategies.
Wei Li 0008, Kangcheng Liu, Lide Xue, Zidong Du, Xishan Zhang, Xuehai Zhou
SMC1
2023 Cambricon-R: A Fully Fused Accelerator for Real-Time Learning of Neural Scene Representation
abstract
Neural scene representation (NSR) initiates a new methodology of encoding a 3D scene with neural networks by learning from dozens of photos taken from different camera positions. NSR not only achieves significant improvement in the quality of novel view synthesis and 3D reconstruction but also reduces the camera cost from the expensive laser cameras to the cheap color cameras on the shelf. However, performing 3D scene encoding using NSR is far from real-time due to the extremely low hardware utilization (only utilization of hardware peak performance), which greatly limits its applications in real-time AR/VR interactions
Xinkai Song, Yuanbo Wen 0001, Xing Hu 0001, Tianbo Liu 0006, Haoxuan Zhou, Husheng Han, Tian Zhi, Zidong Du, Wei Li 0008, Rui Zhang 0040, Chen Zhang 0001, Lin Gao 0004, Qi Guo 0001, Tianshi Chen 0002
MICRO9
2023 Local-Global Cross Fusion Network With Gaussian-Initialized Learnable Positional Prompting for Hyperspectral Image Classification
abstract
Deep learning has significantly advanced the field of hyperspectral remote sensing image classification. Among various methods, the classification method based on spectral-spatial features for hyperspectral classification has attracted wide attention because of its exceptional classification performance. However, such methods encounter challenges in handling input sample and feature extraction. Regarding the input sample, current hyperspectral image classification methods based on spectral-spatial features treat each pixel of the sample equally, resulting in inadequate attention to valuable pixels within 3D samples. Regarding feature extraction, the classification methods struggle to effectively extract both local and global information from hyperspectral images. Aiming at solving above problems, we propose the local-global cross fusion network with Gaussianinitialized positional prompting (LGGNet). LGGNet is designed with an end-to-end architecture, primarily comprising the Gaussian-initialized learnable positional prompting and the localglobal cross fusion network. The Gaussian-initialized learnable positional prompting introduces prompting technique into hyperspectral image classification, utilizing trainable parameters with prior information to learn the spatial importance of different pixels within a sample for the first time. The local-global cross fusion network combines operations such as 3D CNN feature extraction, Transformer feature extraction, and feature fusion, efficiently integrating local and global features. Extensive experiments showcase that LGGNet achieves state-of-the-art performance with limited training samples on four benchmark datasets, all within a lightweight framework. The relevant code is available at https://github.com/ibelieveican2018/LGGNet.
Xin Zhang 0062, Rui Zhang 0040, Ling Li 0001, Wei Li 0008
IEEE Trans. Geosci. Remote. Sens.4
2022 Enabling One-Size-Fits-All Compilation Optimization for Inference Across Machine Learning Computers
abstract
Machine Learning Computers (MLCs) with tensor functional units (e.g., NVIDIA's Tensor Core, Google's TPU and Habana's Tensor Processor Core) have emerged significantly over recent years. The broad diversity of MLCs makes it hard to deploy machine learning workloads with optimized performance. Though deep learning compilers (e.g., TVM) are effective to produce optimized code for different hardware back-ends, when deploying to a new MLC, it is tedious to implement platform-specific compilation optimizations by thoroughly understanding system/architectural details. To address this problem, we propose a holistic approach to achieve one-size-fits-all compilation optimization across different MLCs or inference. The key observation is that diverse MLCs share multiple key architectural characteristics for tensor processing, which can be generalized for conducting cross-platform compilation optimizations. Concretely, we propose the Tensor Abstract Machine (TAM), which features such common architectural characteristics, as the abstraction of a broad range of MLCs. To leverage architectural characteristics of the TAM, we propose the Tensor Scheduling Language (TSL) consisting of tensor computation description and tensor scheduling primitives for implementing operations with portable optimization. Experimental results demonstrate that the code generated from the same optimization schedule achieves 1.05x to 2.05x better performance than hand-tuned libraries and deep learning compilers across different platforms.
Yuanbo Wen 0001, Qi Guo 0001, Zidong Du, Jianxing Xu, Xing Hu 0001, Wei Li 0008, Rui Zhang 0040, Chao Wang 0003, Xuehai Zhou, Tianshi Chen 0002
IEEE Trans. Computers7
2022 Cambricon-G: A Polyvalent Energy-Efficient Accelerator for Dynamic Graph Neural Networks
abstract
Graph neural networks (GNNs), which extend traditional neural networks for processing graph-structured data, have been widely used in many fields. The GNN computation mainly consists of theedge processingto generate messages by combining the edge/vertex features and thevertex processingto update the vertex features with aggregated messages. In addition to nontrivial vector operations in the edge processing, huge random accesses and neural network operations in the vertex processing, the graph topology of GNNs may also vary during the computation (i.e., dynamic GNNs). The above characteristics pose significant challenges on existing architectures. In this article, we propose a novel accelerator named CAMBRICON-G for efficient processing of both dynamic and static GNNs. The key of CAMBRICON-G is to abstract the irregular computation of a broad range of GNN variants to the process of regularly tiledadjacent cuboid(which extends the traditional adjacent matrix of graph by adding the dimension of vertex features). The intuition is that the adjacent cuboid facilitates exploitation of both data locality and parallelism by offeringmultidimensional multilevel tiling(including spatial and temporal tiling) opportunities. To perform themultidimensional spatial tiling, the CAMBRICON-G architecture mainly consists of the cuboid engine (CE) and hybrid on-chip memory. The CE has multiple vertex processing units (VPUs) working in a coordinated manner to efficiently process the sparse data and dynamically update the graph topology with dedicated instructions. The hybrid on-chip memory contains the topology-aware cache and multiple scratchpad memory to reduce off-chip memory access. To perform themultidimensional temporal tiling, an easy-to-use programming model is provided to flexibly explore different tiling options for large graphs. Experimental results show that compared against Nvidia P100 GPU, the performance and energy efficiency can be improved by$7.14\times $and$20.18\times $, respectively, on various GNNs, which validates both the versatility and energy efficiency of CAMBRICON-G.
Xinkai Song, Tian Zhi, Zhe Fan, Wei Li 0008, Xing Hu 0001, Zidong Du, Qi Guo 0001, Yunji Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2021 Cambricon-Q: A Hybrid Architecture for Efficient Training
abstract
Deep neural network (DNN) training is notoriously time-consuming, and quantization is promising to improve the training efficiency with reduced bandwidth/storage requirements and computation costs. However, state-of-the-art quantized algorithms with negligible training accuracy loss, which require on-the-fly statistic-based quantization over a great amount of data (e.g., neurons and weights) and high-precision weight update, cannot be effectively deployed on existing DNN accelerators. To address this problem, we propose the first customized architecture for efficient quantized training with negligible accuracy loss, which is named as Cambricon-Q. Cambricon-Q features a hybrid architecture consisting of an ASIC acceleration core and a near-data-processing (NDP) engine. The acceleration core mainly targets at improving the efficiency of statistic-based quantization with specialized computing units for both statistical analysis (e.g., determining maximum) and data reformating, while the NDP engine avoids transferring the high-precision weights from the off-chip memory to the acceleration core. Experimental results show that on the evaluated benchmarks, Cambricon-Q improves the energy efficiency of DNN training by 6.41× and 1.62×, performance by 4.20× and 1.70× compared to GPU and TPU, respectively, with only ⩽ 0.4% accuracy degradation compared with full precision training.
Yongwei Zhao 0001, Chang Liu 0021, Zidong Du, Qi Guo 0001, Xing Hu 0001, Yimin Zhuang, Xinkai Song, Wei Li 0008, Xishan Zhang, Ling Li 0001, Zhiwei Xu 0002, Tianshi Chen 0002
ISCA9
2020 A Middleware Approach to Synchronize Transaction Data to Blockchain
abstract
Blockchain is a kind of distributed accounting system with the characteristics of decentralization, openness, transparency, non-tampering and de-trust. Recently, blockchain-based applications appear in various fields such as IoT, supply-chain, etc. A practical requirement of applying the Blockchain technology to the conventional supply-chain system is the trans-action synchronization between databases and the blockchain systems. This paper proposes a middleware that supports non-intrusive, efficient, fault-tolerant transaction data synchronization. The middleware is designed under three principles. The first is to divide the process of transaction synchronization into several independent phases. The second is to convert the representation of database transaction into an optimal form for blockchain, and transactions are packed and written into blockchain in a batch mode. The third is to guarantee the consistency between database and blockchain through failure recovery mechanism. Utilities like pipeline processing and transaction mapping built on this middleware achieve significant performance improvement. The peak TPS of Hyperledger Fabric reaches above 16000, which is more than 32 times of the direct submission method.
Taoying Liu, Wei Li 0008
ICCCN4
2019 Event Detection on Unreliable Distributed Storage Network
abstract
Event uncertainty problem happens when a distributed storage system is deployed under unreliable network environment. To address the uncertainty problem, this paper proposes an event detection scheme based on indeterminate subevent-to-event belonging relationship. For a target event, the scheme decomposes its subevent correlation into a group of total order relations, and matches them with the subevent stream to further determine its possible occurrence time range. With the occurrence range, this paper calculates event occurrence confidence based on transition probability in match components.
Zimu Yuan, Wei Li 0008
GLOBECOM2
2019 Deep Fusion: A Software Scheduling Method for Memory Access Optimization
Yimin Zhuang, Shaohui Peng, Xiaobing Chen, Shengyuan Zhou, Tian Zhi, Wei Li 0008, Shaoli Liu
NPC6
2019 A reconfigurable 4-GS/s power-efficient floating-point FFT processor design and implementation based on single-sided binary-tree decomposition
Haigang Yang, Wei Li 0008
Integr.3
2018 Error Analysis on RSS Range-Based Localization Based on General Log-Distance Path Loss Model
abstract
Received Signal Strength (RSS) is considered to be a promising measurement for indoor positioning. Many RSS range-based localization methods have been proposed due to the convenience and low cost of RSS measurements. However, a fundamental problem has not been answered, that is, how accurate are these methods? We think a key reason leading to this situation is the inappropriate assumption on RSS range models and measurement errors, which results in oversimplified analysis on those methods. In this paper, we use a more general range model and recognize the Generalized Least Square (GLS) method as an optimal estimator whose estimation error equals to the Cramer-Rao lower bound (CRLB). Through mathematical, techniques, we derive the analytic expression of the localization error for the GLS method, which reveals the key factors that affect the localization accuracy. Further studies on the minimal localization error disclose the proportional relationship between the localization accuracy and the above key factors.
Wei Li 0008, Zimu Yuan, Wei Zhao 0001
MASS1
2017 NAND-NOR: A Compact, Fast, and Delay Balanced FPGA Logic Element
Grace Zgheib, Wei Li 0008, Zhenghong Jiang, Kaihui Tu, Paolo Ienne, Haigang Yang
FPGA4
2017 A Survey of Buffer Management Strategies in Delay Tolerant Networks
abstract
When configuring a delay tolerant network (DTN), there are many aspects that need to be taken into consideration for an effective and efficient network. One aspect is a buffer management strategy. Buffer strategies are used to determine which packets need to be forwarded or dropped. This paper will focus on the variety of buffer management strategies available, providing a comprehensive survey and analysis. They have all been taken into consideration, evaluated, and then classified into different categories based on their features.
Fabie Ezife, Wei Li 0008
MASS2
2016 SparkArray: An Array-Based Scientific Data Management System Built on Apache Spark
abstract
With the highly demanded requirements for manipulating large scientific datasets, scientists are in need of flexible cluster-level software to execute fast scientific data analysis. In this paper, we discuss whether the Apache Spark framework is suitable for scientific data management. We present our system SparkArray, which extends Spark with a multidimensional array data model and a set of common used array operations (e.g., filter, subarray, smooth and join). We present analysis and performance evaluation results on different implementation methods for executing array operations on clusters. We also compared the performance of SparkArray to SciDB, a recent scientific database management system, using the workloads of the Standard Science DBMS Benchmark (SS-DB). The results show that SparkArray is feasible alternative solution for large-scale scientific data management, especially when scientists require fast data loading or one-time data analysis on large scientific datasets.
Taoying Liu, Dixin Tang, Hong Liu 0018, Wei Li 0008, Rubao Lee
NAS5
2015 A Case Study of Optimizing Big Data Analytical Stacks Using Structured Data Shuffling
abstract
Current major big data analytical stacks often consist of a general purpose, multi-staged computation framework (e.g. Hadoop) and an SQL query system (e.g. Hive) on its top. A key factor of query performance is the efficiency of data shuffling between two execution stages (e.g. Map/Reduce). In current data shuffling, various useful information about the shuffled data and the query on the data is simply wasted. In this paper, we make a strong case of cross-layer optimizations for Hive/Hadoop stack: we have designed and implemented a novel data shuffling mechanism in Hadoop, called Structured Data Shuffling (S-Shuffle), which carefully leverages the rich information in data and queries to optimize the overall query processing. Our experimental results with industry-standard TPC-H benchmark show that, by using S-Shuffle, the performance of SQL query processing on Hadoop can be improved by up to 2.4x.
Dixin Tang, Taoying Liu, Rubao Lee, Hong Liu 0018, Wei Li 0008
CLUSTER5
2015 A Parallel Algorithm for Game Tree Search Using GPGPU
abstract
Game tree search is a classical problem in the field of game theory and artificial intelligence. Fast game tree search algorithm is critical for computer games asking for real-time responses. In this paper, we focus on how to leverage massive parallelism capabilities of GPU to accelerate the speed of game tree search algorithms and propose a concise and general parallel game tree search algorithm on GPU. The performance model of our algorithm is presented and analyzed theoretically. We implement the algorithm for two real computer games called Connect6 and Chess. We also use these two games to verify the effectiveness and efficiency of our algorithm. Experiments support our theoretical results and show good performance of our approach. Compared to classical CPU-based game tree search algorithms, our algorithm can achieve speedups of 89.95x for Connect6 and 11.43x for Chess, in case of no pruning. When pruning is considered, which means the practical performance of our algorithm, the speedup can reach about 10.58x for Connect6 and 7.26x for Chess. The insight of our work is that using GPU is a feasible way to improve the performance of game tree search algorithms.
Liang Li 0003, Hong Liu 0018, Hao Wang 0002, Taoying Liu, Wei Li 0008
IEEE Trans. Parallel Distributed Syst.5
2015 A Multiphase DLL With a Novel Fast-Locking Fine-Code Time-to-Digital Converter
abstract
This brief presents a fast-locking multiphase closed-loop delay-locked loop (DLL). The proposed DLL employs a novel rapid-tracking time-to-digital converter that spends only two clock cycles to generate fine codes. This greatly reduces the fine-locking time and hence the total locking time that goes down to eight input reference clock cycles, shortened by 80%-95% compared with previously reported closed-loop architectures. Fabricated in a 130-nm technology, the proposed DLL operates at 80-450 MHz. In addition, the measured rms and peak-to-peak jitters at 180 MHz are 2.3 and 10 ps, respectively.
Haigang Yang, Wen-rui Zhu, Wei Li 0008, Tianyi Li 0007
IEEE Trans. Very Large Scale Integr. Syst.4
2014 A semi-supervised modeling approach for performance characterization of FPGA architectures
abstract
An approach to estimate the performance of FPGA architectures is proposed based on semi-supervised model tree algorithm. The proposed approach avoids synthesizing, mapping, packing, placing and routing, which are essential steps in a traditional flow to obtain the performance of FPGA. Thus it is time efficient while the performance predicted maintains quite close to the result obtained through the traditional method (a tool flow called VTR). This can be utilized effectively during the early FPGA design stage to choose an optimal architecture under a certain metric. Comparisons are made between the performance obtained by the proposed approach and by VTR on a commercial 40nm technology. Results show that the proposed approach has MRE below 7.62% compared to VTR, and improves the time cost by thousands of times when utilized in architecture design space exploration.
Liqun Yang, Haigang Yang, Wei Li 0008, Colin Yu Lin
FPL3
2014 RHJoin: A fast and space-efficient join method for log processing in MapReduce
abstract
Equi-join is heavily used in MapReduce-based log processing. With the rapid growth of dataset sizes, join methods on MapReduce are extensively studied recently. We find that existing join methods usually cannot get high query performance and affordable storage consumption at the same time when faced with a huge amount of log data. They either only optimize one aspect but significantly sacrifice the other or have limited applications. In this paper, after analyzing characteristics of the workloads and underlying MapReduce, we present a join method with specific optimizations for log processing called RHJoin (Repartition Hash Join) and its implementation on Hadoop. In RHJoin, reference tables are partitioned in the pre-processing step, the log table is partitioned on the map side and hash join is executed on the reduce side. The shuffle procedure of MapReduce is also optimized by removing the sort step and overlapping the execution of mappers and reducers. Comprehensive experiments show that RHJoin achieves high query performance with only a small extra storage cost, and has wide application circumstances for log processing.
Dixin Tang, Taoying Liu, Hong Liu 0018, Wei Li 0008
ICPADS4
2013 A performance evaluation of Hive for scientific data management
abstract
It is very important to evaluate the MapReduce-based frameworks for scientific data processing applications. Scientists need a low-cost, scalable, easy-to-use and fault-tolerance platform for large volume data processing eagerly. This paper presents an implementation of a scientific data management benchmark, SSDB, on Hive, a MapReduce-based data warehouse. A complete strategy of migrating SSDB to Hive is described in detail including query HQL implementation, data partition schema and adjustments of underlying storage facilities. We have tuned the performance using several system parameters provided by Hive, Hadoop and HDFS. This paper provides preliminary results and analysis. Evaluation results indicate that Hive achieves acceptable performance for some data analysis tasks even compared with some high efficient distributed parallel databases, but it needs subtle adjustments of underlying storage facilities and indexing mechanism.
Taoying Liu, Hong Liu 0018, Wei Li 0008
IEEE BigData4
2013 A two-tier positioning algorithm for wireless networks with diverse measurement types
abstract
A variety of measurement methods have been developed to obtain positions of nodes in wireless networks. However, most of existing positioning methods can only use limited types of measurement. These algorithms fail to exploit arbitrary measurement methods to improve the accuracy of positioning. In this paper, we propose a high-accuracy positioning algorithm which can utilize unlimited types of measurement. Our algorithm first selects a set of nodes equipped with more accurate measurement mechanisms to locate nodes whose positions are unknown. By this selection, a system of heterogenous equations is built. Then, our algorithm uses the classical genetic method to solve equations which commonly are hard to solve. The comprehensive simulation shows that our algorithm significantly improves the accuracy of positioning compared with existing methods.
Zimu Yuan, Wei Li 0008
GLOBECOM3
2012 A Node-based Parallel Game Tree Algorithm Using GPUs
abstract
Game tree search is a classical problem in the field of game theory and artificial intelligence. Fast game tree algorithm is critical for computer games asking for real-time responses. In this paper, we focus on how to leverage massive parallelism capabilities of GPUs to accelerate the speed of game tree algorithms and propose a concise and general parallel game tree algorithm on GPUs. The performance model of the algorithm is presented and analyzed theoretically. We also implement the algorithm for a real computer game called Connect6 and use it to verify the effectiveness and efficiency of our algorithm. Experiments support our theoretical results and show good performance of our approach. Compared to classical CPU-based game tree algorithms, our algorithm can achieve speedup of 70.8 in case of no pruning. When pruning is considered (which means the practical performance of our algorithm), the speedup can reach about 7.0. The insight of our work is that using GPUs is a feasible way to improve the performance of game tree algorithms.
Liang Li 0003, Hong Liu 0018, Taoying Liu, Wei Li 0008, Hao Wang 0002
CLUSTER5
2012 An efficient hybrid localization scheme for Heterogeneous Wireless Networks
abstract
The ability to track and locate physical entities is a fundamental requirement for Cyber-Physical Systems (CPSs), especially in an ad-hoc wireless environment. In Heterogeneous Wireless Networks (HWNs), hybrid localization schemes are needed due to the coexistence of both accurate and coarse measurement mechanisms. However, current localization schemes cannot fully satisfy HWNs' accuracy requirements. Therefore, we propose a universal measurement metric called Direct Proportion Distance (DPD) that can leverage most existing measurement mechanisms such as TOA/TDOA, RSS, AOA, Link Diagnosis (LD) and Signal Coverage Detection (SCD). We also prove that DPD is directly proportional to the physical distance between two wireless nodes. Based on this metric, we present three new localization algorithms and compare them with classical methods. The experiments verify that our method performs better than previous localization algorithms when both accurate and coarse measurements are fully utilized.
Zimu Yuan, Wei Li 0008, Adam C. Champion, Wei Zhao 0001
GLOBECOM2
2012 A new routing scheme based on adaptive selection of geographic directions
abstract
Geographic routing is recognized as an appealing approach to achieve efficient communications with low computational complexity and space cost. In order to apply this technology in Cyber-Physical Systems (CPSs), a comprehensive consideration must be given to performance issues such as throughput, delay, and load balance. In this paper, we provide a new routing scheme based on forwarding packets to multiple geographic directions. The proposed routing protocols are studied and analyzed theoretically. Theoretical bounds of throughput, delays and space cost are presented. Simulations show that our method performs more efficiently than traditional geographic routing schemes in terms of throughput, delay, and load balance with acceptable space cost. Our experiments also verify the tradeoff between performance metrics.
Zimu Yuan, Wei Li 0008
GLOBECOM2
2011 History-Aware Adaptive Backoff for Neighbor Discovery in Wireless Networks
abstract
The ability of discovering neighboring nodes, namely neighbor discovery, is essential for the self-organization of wireless ad hoc networks. In this paper, we propose a history-aware adaptive back off algorithm for neighbor discovery assuming collision detection and feedback mechanisms. Given successful discovery feedback, undiscovered nodes can adjust their contention window. With collision feedback and historical information, only transmission nodes enter the re-contention process, and decrease their contention window to accelerate neighbor discovery process after collision. Then, we give theoretical analysis of our algorithm on the discovery time and energy consumption, and derive the optimal size of contention windows by two rounds of optimization. Finally, we validate our theoretical analysis by simulations, and show the performance improvement over existing algorithms.
Zimu Yuan, Lizhao You, Wei Li 0008, Biao Chen 0002, Zhiwei Xu 0002
MSN3
2010 Design and Analysis of a New GPS Algorithm
abstract
In this paper, we propose and analyze a new GPS positioning algorithm. Our algorithm uses the direct linearization technique to reduce the computation time overhead. We invoke the general least squares method in order to achieve optimality in the situation when the trilateration system of equations becomes over-determined. We systematically evaluate our new algorithms and show that they indeed take much less computation time than the traditional GPS method while maintaining reasonable accuracy.
Wei Li 0008, Zhiwei Xu 0002, Wei Zhao 0001
ICDCS1
2009 Key Elements Tracing Method for Parallel XML Parsing in Multi-Core System
abstract
Though XML is applied intensively in a lot of applications, XML parsing is not practical in many fields because of its poor performance. Parallel XML parsing on multi-core is a promising choice. Previous methods all adopt data parallel approach on XML parsing. As the semi-structured nature of XML, they were obliged to divide the data into well-formed XML chunks and then parse these chunks parallel. The division process is named as preparsing. As the preparsing is serial, it becomes the bottleneck of parallel XML parsing. Related work Simultaneous Finite Transducer (SFTXP) parallelized the preparsing stage. It maintained multiple preparser results for each equal sized chunk according to enumerated all possible parsing states. In spite of finite states for each XML, the overhead by SFTXP is tremendous, including CPU time and memory for multiple results generating and storing, respectively. In this work, we address parallel XML parsing by Key Element Parse Tracing (KEPT) method which parallelizes the preparsing and parsing at element level. It remolds the preparsing as a key element extracting process and schedules the processing of key elements in the framework of KEPT. Then parsing process is parallelized as a whole. To demonstrate the effectiveness, we implement it on libxml2 and obtain good scalability on both an 8-core Linux machine and an 8-core 24 SMT Sun machine running Solaris.
Hao Wang 0002, Taoying Liu, Wei Li 0008
PDCAT4
2008 Improving the Performance of Service-based Applications by Dynamic Service Execution
abstract
Distributed applications are increasingly being built by Web services. This paper explores a dynamic service execution technique on virtual machines to exploit parallelism of service-based applications. The Dynamic service execution technique is able to determine data dependence of service invocations at run-time and executes independent service invocations concurrently with the help of stateless property of Web services. Based on this technique, we design and implement a virtual machine-based runtime environment called Abacus Virtual Machine (AVM). On AVM, service-based applications are programmed in a conventional, sequential programming model, which involves no extra parallel programming complexity. At the same time AVM also need to make balancing between the goal of exploiting parallelism and the cost brought by the dynamic service execution. Our experiments show that the cost is no more than 9% of CPU time and 10% memory cost. We believe that the penalty is acceptable for the benefits that AVM provides.
Hong Liu 0018, Wei Li 0008
PDP5
2007 An Approach to Debugging Grid or Web Services
abstract
In this paper, we first introduce some issues that are encountered in building a service debugger and briefly describe our approach to addressing them. Next, we outline some debugging modes and components of a simple composite debugger. Then, we mainly describe its message-based front-end and back-end, which are a co-existing, self-identifying, and non- intrusive. Finally, we preset some experimental results of our latest prototype.
Qiang Yue 0001, Zhiwei Xu 0002, Haiyan Yu 0002, Wei Li 0008, Li Zha
ICWS4
2006 A Service-Oriented Virtual Machine for Grid Applications
abstract
Grid computing is a new paradigm for distributed computing, and service has become building block of grid applications. However, current approaches can not free developers from low-level laborious work when building grid applications. We propose a service-oriented virtual machine called Abacus Virtual Machine to simplify the task of grid application development. As a language level virtual machine, it provides a service-oriented instruction set to abstract the operations on the services of a grid application. It also virtualizes services and creates a virtual global system image for grid applications, thus services can be transparently distributed and shared. In this way, Abacus Virtual Machine hides the cumbersome underlying details from programmers and reduces the complexity greatly in grid application development
Hong Liu 0018, Wei Li 0008, Yili Gong
PDCAT2
2005 A C/S and P2P Hybrid Resource Discovery Framework in Grid Environments
abstract
Resource discovery is crucial to efficient deployment of a grid system whose dynamic, heterogeneous characteristics make it difficult. In this paper, Vega Infrastructure for Resource Discovery (VIRD) is developed, then augmented with new features (i.e., some new algorithms) to build a C/S (client/server) and P2P (peer-to-peer) hybrid resource discovery framework. The three layered architecture of the VIRD is developed to make advantage of the physical and logical topologies of the Internet to facilitate resource discovery. With our simulations and theoretical analysis, it is proved that VIRD is of good scalability with respect to the sizes of the underlying backbone. Even when the resource density is low and the max TTL (time-to-live) is small, VIRD still achieves high search success rates in a small amount of hops. Compared with flooding and random walk algorithms via the same search success rates, VIRD outperforms them in both network traffic and response time.
Yili Gong, Wei Li 0008, Yuzhong Sun, Zhiwei Xu 0002
ICPP2
2005 System Software for China National Grid
Li Zha, Wei Li 0008, Haiyan Yu 0002, Xianghui Xie 0001, Zhiwei Xu 0002
NPC2
2004 Vega: A Computer Systems Approach to Grid Computing
Zhiwei Xu 0002, Wei Li 0008, Li Zha, Haiyan Yu 0002, Donghua Liu
J. Grid Comput.2
2004 Grid Computing in China
Guangwen Yang 0002, Hai Jin 0001, Minglu Li 0001, Wei Li 0008, Zhaohui Wu 0001, Yongwei Wu 0001, Feilong Tang 0001
J. Grid Comput.5
2003 VegaFS: A Prototype for File-Sharing Crossing Multiple Administrative Domains
abstract
Accessing remote resources is a principal challenge of grid computing. For wide-area file sharing, a most difficult problem is the inability to access files distributed in different administrative domains. In this paper, we propose a file system architecture called VegaFS, which is detached from administrative domains entirely and provides cross-domain file access abilities. The main idea is to adopt public keys in native file systems and transfer the management work of administrator to trusted certificate authorities (CAs). Compared with other systems, VegaFS provides several benefits, such as grid-wide unique user identities, cross-domain file access capability, scalabilities, fine-grained dynamic access control and better securities.
Wei Li 0008, Jianmin Liang, Zhiwei Xu 0002
CLUSTER1
2003 VEGA Infrastructure for Resource Discovery in Grids
Yili Gong, Fangpeng Dong, Wei Li 0008, Zhiwei Xu 0002
J. Comput. Sci. Technol.3
2001 Cluster and Grid Superservers: The Dawning Experiences in China
abstract
This paper summarizes recent activities at Institute of Computing Technology, Chinese Academy of Sciences, in developing superservers for cluster and grid computing. We first identify market and technical trends observed from a Chinese perspective. Then we describe the research work in developing the Dawning series high performance computers and the China computational grid. We also highlight some on-going research work in developing grid-oriented superserver systems.
Zhiwei Xu 0002, Ninghui Sun, Dan Meng 0002, Wei Li 0008
CLUSTER4