VLDB 2026 Research / reviewers in the wild / expert
Zhiwei Xu 0002
dblp:262/0620-2
· DBLP profile ↗
76ranked-venue papers
10as first author
18since 2021 · last 2026
0000-0002-1480-7265ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 43 · 3 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 13 · 4 first-author · 1 since 2021Software engineering, systems software and programming languages · 8 · 3 since 2021Computer networks · 6 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-authorTheory of computation · 3Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hardwired-Neuron Language Processing Units as General-Purpose Cognitive SubstratesabstractThe rapid advancement of Large Language Models (LLMs) has established language as a core general-purpose cognitive substrate, driving the demand for specialized Language Processing Units (LPUs) tailored for LLM inference. To overcome the growing energy consumption of LLM inference systems, this paper proposes a Hardwired-Neurons Language Processing Unit (HNLPU), which physically hardwires LLM weight parameters into the computational fabric, achieving several orders of magnitude computational efficiency improvement by extreme specialization. However, a significant challenge still lies in the scale of modern LLMs. A straightforward hardwiring of GPT-OSS-120B would require fabricating photomask sets valued at over 6 billion dollars, rendering this straightforward solution economically impractical. Yang Liu 0466, Yongwei Zhao 0001, Yifan Hao 0001, Zifu Zheng, Weihao Kong, Zhangmai Li, Dongchen Jiang, Ruiyang Xia, Zhihong Ma, Zisheng Liu, Zhaoyong Wan, Yunqi Lu, Hongrui Guo, Zhe Wang 0017, Tianrui Ma, Mo Zou, Rui Zhang 0040, Ling Li 0001, Xing Hu 0001, Zidong Du, Zhiwei Xu 0002, Qi Guo 0001, Tianshi Chen 0002, Yunji Chen |
ASPLOS (2) | 24 |
| 2026 | Cambricon-CIM: Enabling Energy-Efficient and Error-Resilient Analog CIM Acceleration via Reformation of Coding BasesabstractRecently, multi-bit slicing has emerged as a promising technique to improve the energy efficiency of charge-domain Compute-In-Memory (CIM) accelerators by reducing the number of Analog-to-Digital (A/D) conversions. However, multi-bit slicing requires shift-and-add operations to reconstruct outputs, which exponentially amplify errors and cause significant accuracy degradation. Existing works mainly rely on hardware-aware retraining or noise-suppression techniques, incurring considerable design or power overhead. Thus, multi-bit CIM designs often face the dilemma of trading off energy efficiency for error resilience. In this paper, we propose Cambricon-CIM, a charge-domain multi-bit CIM accelerator that achieves both high energy efficiency and strong error resilience, without requiring retraining. The core insight is that the error amplification is proportional to digit weights; and by redefining these digit weights with smaller non-binary coding bases, it is possible to reduce the total error amplification. Leveraging this principle, CambriconCIM dynamically selects the minimal coding bases for every analog dot-product. With novel circuit and architectural support, Cambricon-CIM enables fast, low-overhead reconfiguration of coding bases at runtime. Experimental results show that Cambricon-CIM achieves 2.27× energy efficiency and 3.06× performance over RAELLA, a state-of-the-art error-resilient multi-bit slicing CIM architecture. Hongrui Guo, Tianrui Ma, Zidong Du, Mo Zou, Yifan Hao 0001, Yongwei Zhao 0001, Rui Zhang 0040, Wei Li 0008, Xing Hu 0001, Zhiwei Xu 0002, Qi Guo 0001, Tianshi Chen 0002 |
HPCA | 10 |
| 2026 | Enabling High-Utilization and Low-Contention FaaS: A Request-Level Resource Provisioning ApproachabstractFunction-as-a-Service offers cost efficiency but often suffers from resource underutilization. This underutilization stems from the instance-level resource provisioning pattern, an issue that existing optimizations have failed to resolve fundamentally. The core problem is that static coarse-grained instance-level resource allocation cannot match the millisecond-level burstiness of dynamic requests. Consequently, it is difficult for current systems to achieve high resource utilization while maintaining high quality of service (QoS) guarantees. To address the problem, this paper advocates a shift to request-level resource provisioning, which redefines the individual request as the atomic unit for scheduling and resource management. We implement this approach in RRP, a scalable FaaS platform that enables efficient per-request resource allocation and release. RRP unifies instance placement and request routing with low-overhead, millisecond-level global visibility. Our evaluation shows that RRP significantly outperforms state-of-the-art instance-level platforms and algorithms. By matching resources to each request’s needs and isolating them from contention, RRP achieves low latency and high utilization. Specifically, on real-world Azure traces, RRP achieves speedups of 1.33 × –30.15 × for average end-to-end latency and 1.37 × –61.46 × for P99 latency, and raises CPU utilization from 44.80%–56.32% to 72.49% under bursty loads. Runfu Li, Zishu Yu, Yifan Wang 0005, Xiaohui Peng 0002, Ninghui Sun, Zhiwei Xu 0002 |
HPDC | 6 |
| 2026 | Cambricon-QM: A Hybrid Architecture for Microscaling Format Training
Yongwei Zhao 0001, Chang Liu 0021, Zidong Du, Xing Hu 0001, Yimin Zhuang, Yifan Hao 0001, Xinkai Song, Wei Li 0008, Xishan Zhang, Ling Li 0001, Zhiwei Xu 0002, Tianshi Chen 0002, Qi Guo 0001 |
IEEE Trans. Computers | 14 |
| 2026 | LASS: Reducing Cold Startup Latency in Serverless Through Loaded Library SharingabstractIn serverless scenario, function invocation runs in an individual container. Lightweight container technology has significantly reduced the startup latency of container. The library loading process now becomes a critical performance bottleneck of serverless function cold startup. The state-of-the-art approaches leverage the process fork operation to reduce the cold startup latency in serverless computing by reusing the loaded libraries. However, the fork operation can only share libraries between parent process and forked process. For security, the libraries loaded by the parent process should be a subset of those required by the forked process, which limits opportunities to eliminate library loading overhead. To address this problem, we propose theLASSsystem, which enables multiple processes to share initialized libraries in a composable and efficient manner.LASSallows a process to securely reuse libraries loaded by multiple processes, thereby reducing library loading latency to the millisecond level. Compared to the state-of-the-art approaches,LASScan improve average library loading speed by more than 10.3×, and reduce 99thpercentile end-to-end latency by 34%–57%. Zishu Yu, Runfu Li, Yifan Wang 0005, Xiaohui Peng 0002, Zhiwei Xu 0002 |
IEEE Trans. Computers | 6 |
| 2025 | Tide: A Distributed Runtime Management Framework for Things-Edge-Cloud Computing ContinuumabstractThe increasing number of connected IoT devices produces massive amounts of sensed data at the network edge. A new computing paradigm, called Things-Edge-Cloud (TEC) collaboration, has been proposed to meet real-time and high-throughput requirements. Most of the existing work focuses on workload scheduling across the computing nodes in TEC with ad-hoc implementations using runtime and management frameworks designed for the cloud. In this paper, we propose RSEP to model the entities and their relationship in the TEC computing continuum (TEC3). We then design and implement Tide—a distributed runtime management framework for TEC3based on the RSEP model, which enables elastic resource allocation and seamless computation offloading. It employs runtime environment isolation and physical resource binding to enforce strong isolation without incurring performance penalties. To decouple runtime and framework, Tide provides a set of portable application interfaces that allow the management of variety runtimes. We implement Tide from scratch and compare its latency and throughput with KubeEdge, Ray, and bare-metal implementations. Experimental results show that Tide improves throughput by 2x and reduces average latency, 95th percentile latency, and latency standard deviation by${4 5. 8 2 \%, 4 8. 3 6 \%}$, and${2 8. 9 6 \%}$, respectively. Specifically, Tide achieves${8 7. 3 \%}$of the ideal goodput, exceeding other platforms more than 10x. Xiaohui Peng 0002, Wenkai Yan, Yifan Wang 0005, Shoujian Zheng, Zhiwei Xu 0002 |
IPDPS | 5 |
| 2025 | SaaP: Rearchitect SoC-as-a-Processor to Orchestrate Hardware HeterogeneityabstractDue to the end of Moore’s Law and Dennard Scaling, Domain-Specific Accelerators (DSAs) have come to a Cambrian explosion. Especially when advancing into the intelligent era, more and more DSAs are integrated into System-on-Chips (SoCs) as intellectual property (IP) blocks to provide high performance and efficiency. Currently, IPs usually expose IP-dependent hardware interfaces, requiring SoCs to manage them as isolated devices with software running on the host CPU. However, such software-managed heterogeneity in CPU-centric SoCs leads to low IP utilization. This inefficiency arises from the dependence on software optimization, coupled with the control and data exchange overheads. To improve IP utilization of heterogeneous SoCs, in this article, we rearchitect the SoC as a processor (i.e., SaaP) to orchestrate hardware heterogeneity. SaaP features an orchestration pipeline where DSAs are integrated as execution units and managed directly by the hardware pipeline to conceal the hardware heterogeneity from software. Moreover, SaaP redesigns the register file and data paths to implement an IP-level data-forwarding mechanism, avoiding the costly control and data exchange in the CPU-centric execution model. Block data dependence among different DSAs is carefully resolved to exploit mixed-level parallelism and inter-IP data exchange. SaaP abstracts tasks as mixed-scale instructions, where each instruction can be mapped to different IPs. Experimental results show that compared against Xavier on six fully software-optimized benchmarks from different domains, SaaP-rearchitected Xavier achieves a$2.08{\times }$speedup, with an 8.21% area reduction and only 2.98% increase in power consumption. Pengwei Jin, Zhe Fan, Yongwei Zhao 0001, Zidong Du, Hongrui Guo, Ziyuan Nan, Yifan Hao 0001, Chongxiao Li, Tianyun Ma, Xiaqing Li, Wei Li 0008, Xing Hu 0001, Qi Guo 0001, Zhiwei Xu 0002, Tianshi Chen 0002 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 15 |
| 2024 | Cambricon-D: Full-Network Differential Acceleration for Diffusion ModelsabstractDiffusion models have made significant progress in current image generation tasks, thus becoming a prominent area of research. Diffusion models necessitate repetitive iterations on minimally altered input data across timesteps, each timestep requiring the recalculation of the entire model, resulting in a remarkable computational redundancy and substantial hardware expenditures.Performing differential computing on input data seems to be a feasible approach for addressing such computational redundancy and improving hardware efficacy. However, non-linear operations (particularly activation functions) necessitate the merging of deltas (i.e., differential values) with raw inputs repeatedly to ensure computational correctness, leading to significant memory access for loading raw inputs, which fragmentedly blocks the forwarding of deltas throughout the network and undermines performance.To solve this problem, we propose Cambricon-D, a fullnetwork differential computing architecture with concise memory access. While maintaining the computational efficiency brought by differential computing, Cambricon-D employs a sign-mask dataflow, which requires only the loading of 1-bit signs (instead of large bitwidth raw inputs), thereby facilitating the seamless forwarding of deltas and effectively mitigating memory access overheads. Experimental results show that, compared to Diffy, Cambricon-D’s dataflow reduces 66% ~ 82% off-chip memory access. In total, Cambricon-D achieves 1.46× ~ 2.38× speedup over A100 on various diffusion models with different resolutions. Weihao Kong, Yifan Hao 0001, Qi Guo 0001, Yongwei Zhao 0001, Xinkai Song, Xiaqing Li, Mo Zou, Zidong Du, Rui Zhang 0040, Chang Liu 0021, Yuanbo Wen 0001, Pengwei Jin, Xing Hu 0001, Wei Li 0008, Zhiwei Xu 0002, Tianshi Chen 0002 |
ISCA | 15 |
| 2024 | Cambricon-M: A Fibonacci-Coded Charge-Domain SRAM-Based CIM Accelerator for DNN InferenceabstractCharge-domain SRAM-based Computing-in-memory (CIM) proves to be a promising method for DNN inference, and benefits from avoiding data movement between computing units and memory. However, the high resolution Analog-to-Digital Converters (ADCs) dominates the energy consumption (up to 64%), limiting the energy efficiency of SRAM-CIM architectures. The main reason is the wide range of input analog values, requiring high resolution ADCs to convert the high precision averaged analog voltages into high bitwidth digital data. In this paper, to reduce the ADC overhead, we propose Cambricon-M, a novel Fibonacci-coded SRAM-based charge-domain CIM accelerator for DNN inference. Cambricon-M features the Fibonacci coding, which guarantees low density of ‘1’ in operands (i.e., the adjacent two bits of each ‘1’ are both ‘0’), narrowing the output voltage range and enabling low resolution ADCs. Further, Cambricon-M exploits the high bit-level sparsity to address the extra energy and area overhead caused by the larger bitwidth in Fibonacci coding. Specifically, Cambricon-M proposes zero-skipping methods to reduce ineffectual input/output, and the bit-slice based compression method to reduce memory capacity/bandwidth pressure. Experimental results show that Cambricon-M reduces ADC energy by 68.7%, and improves the energy efficiency 3.48× and 1.62× compared to TPUv4 and an ISAAC-based charge-domain SRAM-CIM accelerator. Hongrui Guo, Mo Zou, Yifan Hao 0001, Zidong Du, Erxiang Ren, Yang Liu 0466, Yongwei Zhao 0001, Tianrui Ma, Rui Zhang 0040, Xing Hu 0001, Fei Qiao, Zhiwei Xu 0002, Qi Guo 0001, Tianshi Chen 0002 |
MICRO | 12 |
| 2024 | Hawk: An Efficient NALM System for Accurate Low-Power Appliance RecognitionabstractNon-intrusive Appliance Load Monitoring (NALM) aims to recognize individual appliance usage from the main meter without indoor sensors. However, existing systems struggle to balance dataset construction efficiency and event/state recognition accuracy, especially for low-power appliance recognition. This paper introduces Hawk, an efficient and accurate NALM system that operates in two stages: dataset construction and event recognition. In the data construction stage, we efficiently collect a balanced and diverse dataset, HawkDATA, based on balanced Gray code and enable automatic data annotations via a sampling synchronization strategy called shared perceptible time. During the event recognition stage, our algorithm pipeline integrates steady-state differential pre-processing and voting-based post-processing for accurate event recognition from the aggregate current. Experimental results show that HawkDATA takes only 1/71.5 of the collection time to collect 6.34x more appliance state combinations than the baseline. In HawkDATA and a widely used dataset, Hawk achieves an average F1 score of 93.94% for state recognition and 97.07% for event recognition, which is a 47.98% and 11.57% increase over SOTA algorithms. Furthermore, selected appliance subsets and the model trained from HawkDATA are deployed in two real-world scenarios with many unknown background appliances. The average F1 scores of event recognition are 96.02% and 94.76%. Hawk's source code and HawkDATA are accessible at https://github.com/WZiJ/SenSys24-Hawk. Xingzhou Zhang, Yifan Wang 0005, Xiaohui Peng 0002, Zhiwei Xu 0002 |
SenSys | 5 |
| 2023 | Cambricon-U: A Systolic Random Increment Memory Architecture for Unary ComputingabstractUnary computing, whose arithmetics require only one logic gate, has enabled efficient DNN processing, especially on strictly power-constrained devices. However, unary computing still confronts the power efficiency bottleneck for buffering unary bitstreams. The buffering of unary bitstreams requires accumulating bits into large bitwidth binary numbers. The large bitwidth binary number needs to activate all bits per cycle in case of carry propagation. As a result, the accumulation process accounts for 32%-70% of the power budget. Hongrui Guo, Yongwei Zhao 0001, Zhangmai Li, Yifan Hao 0001, Chang Liu 0021, Xinkai Song, Xiaqing Li, Zidong Du, Rui Zhang 0040, Qi Guo 0001, Tianshi Chen 0002, Zhiwei Xu 0002 |
MICRO | 12 |
| 2023 | Rescue to the Curse of universality
Yongwei Zhao 0001, Zidong Du, Qi Guo 0001, Zhiwei Xu 0002, Yunji Chen |
Sci. China Inf. Sci. | 4 |
| 2023 | Self-supervised image clustering from multiple incomplete views via constrastive complementary generationabstractAbstract Incomplete Multi‐View Clustering aims to enhance clustering performance by using data from multiple modalities. Despite the fact that several approaches for studying this issue have been proposed, the following drawbacks still persist: (1) It is difficult to learn latent representations that account for complementarity yet consistency without using label information; (2) and thus fails to take full advantage of the hidden information in incomplete data results in suboptimal clustering performance when complete data is scarce. In this study, Contrastive Incomplete Multi‐View Image Clustering with Generative Adversarial Networks (CIMIC‐GAN), which uses Generative Adversarial Network (GAN) to fill in incomplete data and uses double contrastive learning to learn consistency on complete and incomplete data is proposed. More specifically, considering diversity and complementary information among multiple modalities, we incorporate autoencoding representation of complete and incomplete data into double contrastive learning to achieve learning consistency. Integrating GANs into the autoencoding process can not only take full advantage of new features of incomplete data, but also better generalise the model in the presence of high data missing rates. Experiments conducted on four extensively used data sets show that CIMIC‐GAN outperforms state‐of‐the‐art incomplete multi‐View clustering methods. Jiatai Wang, Zhiwei Xu 0002, Dongjin Guo |
IET Comput. Vis. | 2 |
| 2022 | Cambricon-P: A Bitflow Architecture for Arbitrary Precision ComputingabstractArbitrary precision computing (APC), where the digits vary from tens to millions of bits, is fundamental for scientific applications, such as mathematics, physics, chemistry, and biology. APC on existing platforms (e.g., CPUs and GPUs) is achieved by decomposing the original data into small pieces to accommodate to the low-bitwidth (e.g., 32-/64-bit) functional units. However, such fine-grained decomposition inevitably introduces large amounts of intermediates, bringing in intensive on-chip data traffic and long, complex dependency chains, so that causing low hardware utilization.To address this issue, we propose Cambricon-P, a bitflow architecture supporting monolithic large and flexible bitwidth operations for efficient APC processing, which avoids generating large amounts of intermediates from decomposition. Cambricon- P features a tightly-integrated computational architecture for processing different bitflows in parallel, where full bit-serial data paths are deployed. The bit-serial scheme still needs to eliminate the dependency chain of APC for exploiting parallelism within one monolithic large-bitwidth operation. For this purpose, Cambricon-P adopts a carry parallel computing mechanism, which enables recursively transforming the multiplication into smaller inner-products that can be performed in parallel between bit-indexed IPUs (Inner-Product Units). Furthermore, to improve the computing efficiency of APC, Cambricon- P employs a bit-indexed inner-product processing scheme, namely BIPS, to eliminate intra-IPU bit-level redundancy. Compared to Intel Xeon 6134 CPU, Cambricon-P achieves 100.98$\times$ performance on monolithic long multiplication, and 23.41$\times$/30.16$\times$ speedup and energy benefit over four real-world APC applications on average. Compared to NVidia V100 GPU, Cambricon-P also delivers the same throughput, as well as 430$\times$/60.5$\times$ lesser area and power, respectively, on batch-processing multiplications. Yifan Hao 0001, Yongwei Zhao 0001, Chenxiao Liu, Zidong Du, Shuyao Cheng, Xiaqing Li, Xing Hu 0001, Qi Guo 0001, Zhiwei Xu 0002, Tianshi Chen 0002 |
MICRO | 9 |
| 2022 | SurRF: Unsupervised Multi-View Stereopsis by Learning Surface Radiance FieldabstractThe recent success in supervised multi-view stereopsis (MVS) relies on the onerously collected real-world 3D data. While the latest differentiable rendering techniques enable unsupervised MVS, they are restricted to discretized (e.g., point cloud) or implicit geometric representation, suffering from either low integrity for a textureless region or less geometric details for complex scenes. In this paper, we propose SurRF, an unsupervised MVS pipeline by learning Surface Radiance Field, i.e., a radiance field defined on a continuous and explicit 2D surface. Our key insight is that, in a local region, the explicit surface can be gradually deformed from a continuous initialization along view-dependent camera rays by differentiable rendering. That enables us to define the radiance field only on a 2D deformable surface rather than in a dense volume of 3D space, leading to compact representation while maintaining complete shape and realistic texture for large-scale complex scenes. We experimentally demonstrate that the proposed SurRF produces competitive results over the-state-of-the-art on various real-world challenging scenes, without any 3D supervision. Moreover, SurRF shows great potential in owning the joint advantages of mesh (scene manipulation), continuous surface (high geometric resolution), and radiance field (realistic rendering). Jinzhi Zhang, Mengqi Ji, Zhiwei Xu 0002, Shengjin Wang, Lu Fang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Breaking the Interaction Wall: A DLPU-Centric Deep Learning Computing SystemabstractDue to the broad successes of deep learning, many CPU-centric artificial intelligent computing systems employ specialized devices such as GPUs, FPGAs, and ASICs, which can be named as Deep Learning Processing Units (DLPUs), for processing computation-intensive deep learning tasks. The separation between the scalar control operations mapped on CPUs and the vector computation operations mapped on DLPUs causes the frequent and costly interactions between CPUs and DLPUs, leading to theInteraction Wall. Moreover, the increasing algorithm complexity and DLPU computation speed would further aggravate the interaction wall substantially. To break the interaction wall, we propose a novel DLPU-centric deep learning computing system consisting of anexception-oriented programming (EOP) modeland the architectural support ofCPULESS DLPU. The EOP model processes scalar control operations of a deep learning task as exception handlers to maximally avoid stalling the crucial and dominated vector computation operations. Together with the CPULESS DLPU which integrates a scalar processing unit (SPU) for scalar control operations and the parallel processing unit (PPU) for vector computation operations into a fused pipeline, the proposed DLPU-centric system can cost-effectively leverage the EOP model to execute the two kinds of operations simultaneously without disturbing each other. Compared with a state-of-the-art commodity CPU-centric system with discrete V100 GPU via PCIe bus, experimental results show that our DLPU-centric system achieves 10.30× better performance and 92.99 percent energy savings, respectively. Moreover, compared with a CPU-centric version of DLPU system where the SPU serves as the host with integrated PPU, the proposed DLPU-centric system still achieves 15.60 percent better performance from avoided interactions. Zidong Du, Qi Guo 0001, Yongwei Zhao 0001, Ling Li 0001, Limin Cheng, Zhiwei Xu 0002, Ninghui Sun, Yunji Chen |
IEEE Trans. Computers | 7 |
| 2021 | Cambricon-Q: A Hybrid Architecture for Efficient TrainingabstractDeep neural network (DNN) training is notoriously time-consuming, and quantization is promising to improve the training efficiency with reduced bandwidth/storage requirements and computation costs. However, state-of-the-art quantized algorithms with negligible training accuracy loss, which require on-the-fly statistic-based quantization over a great amount of data (e.g., neurons and weights) and high-precision weight update, cannot be effectively deployed on existing DNN accelerators. To address this problem, we propose the first customized architecture for efficient quantized training with negligible accuracy loss, which is named as Cambricon-Q. Cambricon-Q features a hybrid architecture consisting of an ASIC acceleration core and a near-data-processing (NDP) engine. The acceleration core mainly targets at improving the efficiency of statistic-based quantization with specialized computing units for both statistical analysis (e.g., determining maximum) and data reformating, while the NDP engine avoids transferring the high-precision weights from the off-chip memory to the acceleration core. Experimental results show that on the evaluated benchmarks, Cambricon-Q improves the energy efficiency of DNN training by 6.41× and 1.62×, performance by 4.20× and 1.70× compared to GPU and TPU, respectively, with only ⩽ 0.4% accuracy degradation compared with full precision training. Yongwei Zhao 0001, Chang Liu 0021, Zidong Du, Qi Guo 0001, Xing Hu 0001, Yimin Zhuang, Xinkai Song, Wei Li 0008, Xishan Zhang, Ling Li 0001, Zhiwei Xu 0002, Tianshi Chen 0002 |
ISCA | 12 |
| 2021 | EdUCAS: An In-house CI/CD Platform with Cloud FPGAs for Agilely Conducting Computer Systems Course ProjectsabstractIn recent years, there has been a rapidly growing recognition of the importance of conducting hands-on HW-SW co-design labs with real hardware (e.g., programmable logic chips named FPGAs) while studying computing curricula, especially the computer systems (CSys) courses. However, using FPGA is quite atime-consuming and error-prone process for students, and manipulating FPGA development tools and boards also distracts students and instructors. To overcome these obstacles and improve agility, we introduce an in-house platform named EdUCAS, in combination with the previously designed cloud FPGA servers for students to automatically conduct computer systems course projects. EdUCAS aims at enabling students to concentrate on their logic designs using hardware description language (e.g., Verilog HDL), without wasting useless time in FPGA tools and experimental environment. Ke Zhang 0017, Yisong Chang, Mingyu Chen 0001, Yungang Bao, Zhiwei Xu 0002 |
ITiCSE (2) | 7 |
| 2020 | Self-Aware Neural Network Systems: A Survey and New PerspectiveabstractNeural network (NN) processors are specially designed to handle deep learning tasks by utilizing multilayer artificial NNs. They have been demonstrated to be useful in broad application fields such as image recognition, speech processing, machine translation, and scientific computing. Meanwhile, innovative self-aware techniques, whereby a system can dynamically react based on continuously sensed information from the execution environment, have attracted attention from both academia and industry. Actually, various self-aware techniques have been applied to NN systems to significantly improve the computational speed and energy efficiency. This article surveys state-of-the-art self-aware NN systems (SaNNSs), which can be achieved at different layers, that is, the architectural layer, the physical layer, and the circuit layer. At the architectural layer, SaNNS can be characterized from a data-centric perspective where different data properties (i.e., data value, data precision, dataflow, and data distribution) are exploited. At the physical layer, various parameters of physical implementation are considered. At the circuit layer, different logics and devices can be used for high efficiency. In fact, the self-awareness of existing SaNNS is still in a preliminary form. We propose a comprehensive SaNNS from a new perspective, that is, the model layer, to exploit more opportunities for high efficiency. The proposed system is called as MinMaxNN, which features model switching and elastic sparsity based on monitored information from the execution environment. The model switching mechanism implies that models (i.e., min and max model) dynamically switch given different inputs for both efficiency and accuracy. The elastic sparsity mechanism indicates that the sparsity of NNs can be dynamically adjusted in each layer for efficiency. The experimental results show that compared with traditional SaNNS, MinMaxNN can achieve 5.64× and 19.66% performance improvement and energy reduction, respectively, without notable loss of accuracy and negative effects on developers' productivity. Zidong Du, Qi Guo 0001, Yongwei Zhao 0001, Tian Zhi, Yunji Chen, Zhiwei Xu 0002 |
Proc. IEEE | 6 |
| 2020 | Machine Learning Computers With Fractal von Neumann ArchitectureabstractMachine learning techniques are pervasive tools for emerging commercial applications and many dedicated machine learning computers on different scales have been deployed in embedded devices, servers, and data centers. Currently, most machine learning computer architectures still focus on optimizing performance and energy efficiency instead of programming productivity. However, with the fast development in silicon technology, programming productivity, including programming itself and software stack development, becomes the vital reason instead of performance and power efficiency that hinders the application of machine learning computers. In this article, we propose Cambricon-F, which is a series of homogeneous, sequential, multi-layer, layer-similar, and machine learning computers with same ISA. A Cambricon-F machine has a fractal von Neumann architecture to iteratively manage its components: it is with von Neumann architecture and its processing components (sub-nodes) are still Cambricon-F machines with von Neumann architecture and the same ISA. Since different Cambricon-F instances with different scales can share the same software stack on their common ISA, Cambricon-Fs can significantly improve the programming productivity. Moreover, we address four major challenges in Cambricon-F architecture design, which allow Cambricon-F to achieve a high efficiency. We implement two Cambricon-F instances at different scales, i.e., Cambricon-F100 and Cambricon-F1. Compared to GPU based machines (DGX-1 and 1080Ti), Cambricon-F instances achieve 2.82x, 5.14x better performance, 8.37x, 11.39x better efficiency on average, with 74.5, 93.8 percent smaller area costs, respectively. We further propose Cambricon-FR, which enhances the Cambricon-F machine learning computers to flexibly and efficiently support all the fractal operations with a reconfigurable fractal instruction set architecture. Compared to the Cambricon-F instances, Cambricon-FR machines achieve 1.96x, 2.49x better performance on average. Most importantly, Cambricon-FR computers are able to save the code length with a factor of 5.83, thus significantly improving the programming productivity. Yongwei Zhao 0001, Zhe Fan, Zidong Du, Tian Zhi, Ling Li 0001, Qi Guo 0001, Shaoli Liu, Zhiwei Xu 0002, Tianshi Chen 0002, Yunji Chen |
IEEE Trans. Computers | 8 |
| 2019 | Engaging Heterogeneous FPGAs in the CloudabstractFPGA has become an essential infrastructural component in commercial cloud and datacenter for improving system performance and efficiency. Meanwhile, a heterogeneous FPGA chip (Hetero-FPGA) in which a multi-core System-on-Chip (SoC) is tightly integrated with an FPGA fabric has been successfully pioneered. Given its hardware-software co-programmability, Hetero-FPGA is supposed to become an independent and first-class cloud computing resource with networking capabilities in order to avoid involving brawny commodity x86 servers as carriers for FPGA fabrics which are usually the cases in current commercial FPGA clouds from several web vendors. Following this design paradigm, we present HeFA, a self-contained Hetero-FPGA Array architecture in cloud. We construct a high-level hardware template as well as a software stack for the Hetero-FPGA node, enabling the SoC as a primary engine to manage, coordinate and incorporate with the dominant FPGA fabric. We also propose a fully scripted design flow to make HeFA as an easy-to-use cloud infrastructure. Based on these techniques, we implement an academia prototype chassis of HeFA that includes 32 Hetero-FPGA nodes with Xilinx's Zynq MPSoC chips. By a customized cloud resource manager, the prototype is flexibly provisioned as either 32 individual FPGA nodes or multiple scalable sub-clusters to abstract arbitrary volume of reconfigurable fabrics as on-demand cloud services. In this manner, a versatile research and educational platform is delivered for agile hardware-software co-design in scenarios such as domain-specific accelerator development, open instruction set architecture-based chip design, computer system-related experimental project, and so on. Ke Zhang 0017, Yisong Chang, Mingyu Chen 0001, Yungang Bao, Zhiwei Xu 0002 |
FPGA | 5 |
| 2019 | Cambricon-F: machine learning computers with fractal von neumann architectureabstractMachine learning techniques are pervasive tools for emerging commercial applications and many dedicated machine learning computers on different scales have been deployed in embedded devices, servers, and data centers. Currently, most machine learning computer architectures still focus on optimizing performance and energy efficiency instead of programming productivity. However, with the fast development in silicon technology, programming productivity, including programming itself and software stack development, becomes the vital reason instead of performance and power efficiency that hinders the application of machine learning computers. Yongwei Zhao 0001, Zidong Du, Qi Guo 0001, Shaoli Liu, Ling Li 0001, Zhiwei Xu 0002, Tianshi Chen 0002, Yunji Chen |
ISCA | 6 |
| 2019 | Computer Organization and Design Course with FPGA CloudabstractComputer Organization and Design (COD) is a fundamentally required early-stage undergraduate course in most computer science and engineering curricula. During the two sessions (lecture and project part) of one COD course, educational platforms play an important role in cultivating students' computational thinking, especially the ability of viewing the hardware and software in a computer system as a whole (computer system thinking ability for short in this paper). In order to improve teaching quality, in this paper, we discuss the deployment of an inexpensive in-house Field Programmable Gate Array (FPGA) cloud platform, which can provide students with hardware-software co-design methodology and practice. The platform includes 32 FPGA nodes and the scale can be dynamically changed. Each cloud node is heterogeneously composed of an ARM processor and a tightly-coupled reconfigurable fabric to provide students with hands-on hardware and software programming experiences. We illustrate our efforts to make the FPGA cloud as an easy-to-use resource pool to elastically support a class with 92 undergrads via Internet access and to monitor students' experimental behaviors. We also present key insights in our teaching activities that indicate such appliance is feasible to provide practice of both basic principles and emerging co-design techniques for students. We believe that our cost-effective FPGA cloud is of significant interests to educators looking forward to improving computer system-related courses. Ke Zhang 0017, Yisong Chang, Mingyu Chen 0001, Yungang Bao, Zhiwei Xu 0002 |
SIGCSE | 5 |
| 2019 | T-REST: An Open-Enabled Architectural Style for the Internet of ThingsabstractComputing offloading is a key challenge of new rising computing paradigms of the Internet of Things (IoT) like edge computing, which shifts computations to data sources as near as possible to gain the benefits, such as low latency and energy efficiency. However, the fragmentation problem of IoT devices results in a heterogeneous and disordered ecosystem, hindering the interoperating demands of computing offloading. What we need is an open-enabled ecosystem which allows third-party developers to create and update functions of deployed devices dynamically. We propose Things-representational state transfer (T-REST) which is an extension of the representational state transfer (REST) architectural style to address this problem. It integrates contents and computations together to inherit the uniform interface principle from REST. Three architectural constraints are added to REST: 1) reusable remote evaluation; 2) dynamic time series representation; and 3) computational hypertext. A novel event triggering mechanism is designed to decouple the tight coupling of front-end content accesses and back-end computations for resources. A reference prototype, named T-REST engine, is implemented to verify the proposed architecture with the open-enabled style, distributed semantics, and computing offloading features. Discussions show that T-REST preserves the benefits of REST. In addition, it achieves extremely lightweight footprints and can perform computing offloading through open-enabled architectures. Zhiwei Xu 0002, Lu Chao, Xiaohui Peng 0002 |
IEEE Internet Things J. | 1 |
| 2019 | Ecosystem of Things: Hardware, Software, and ArchitectureabstractEdge computing is a continuum that includes the computing resources from cloud to things. Ecosystem of things (EoT) is a subsystem of the ecosystem of edge computing, which potentially contains trillions of devices of things and directly interacts with the physical world. This paper surveys the state of the art of EoT by focusing on the computing infrastructure aspect with a forward-looking perspective. We point out a trend of smart edge computing with four types of smartness and intelligence. We address three fundamental questions. 1) What capabilities and how much energy efficiency are the hardware providing? What is the future growth potential? 2) What abstractions are provided by the system software? Are they adequate to support smart edge computing? 3) What ecosystem architectures have been proposed for the coordination of things, the edge, and the cloud? Are they meeting the needs to encourage innovation but avoid unnecessary ecosystem fragmentation? We examine advances from both industry and academia, including research results, visions, and project concepts. We also point out future research directions. Lu Chao, Xiaohui Peng 0002, Zhiwei Xu 0002, Lei Zhang 0008 |
Proc. IEEE | 3 |
| 2018 | WebletScript: A Lightweight Distributed JavaScript Engine for Internet of ThingsabstractNowadays, there is no de facto programming language dedicated to the development of complex Internet of Things (IoT) applications containing different kinds of devices. JavaScript has great potential to become the mainstream application-development program language across IoT platforms, as JavaScript is supported by many engines which have different footprint sizes from the KB level to the MB level. However, there are still technical challenges in the coding efficiency and heterogeneous platform supports. In this paper, we propose a lightweight distributed JavaScript engine, WebletScript, to help IoT developers to build applications for a group of heterogeneous IoT devices. WebletScript can run on any device that supports normal C libraries. It enables automatic code slicing and distribution for an IoT application. Therefore, with one-time platform-independent code development, the application can be executed on multiple devices, such as a group of end devices, a smart gateway and the cloud system. Compared to the traditional solution, in which embedded application and cloud application are developed separately, our approach can reduce the lines of code effectively by 50% and reduce the cost of application update cost by 65%. Dong Li 0008, Zhiwei Xu 0002 |
GLOBECOM | 4 |
| 2018 | Post-exascale supercomputing: research opportunities aboundabstractExascale supercomputing refers to the scientific research efforts and activities to build and use supercomputers that can perform scientific computing at the speed of exaflops, or 10 18 floating-point (64-bit) operations per second.Exascale supercomputing is a major milestone in surpassing the state-of-the-art standard of petascale supercomputing, i.e., 10 15 floating-point operations per second, established a decade ago in 2008.Exascale scientific computing research is already in full bloom worldwide.The USA leads this research and development direction, with federal government funding starting in as early as 2008.Japan, Europe, India, and China soon followed suit.The most recent focus on the field is in Europe, with 1.4 billion euros budgeted for building pre-exascale supercomputers by 2020, and an additional 2.7 billion euros proposed for building an exascale supercomputer by 2023 (Feldman, 2018).It is expected that multiple exascale supercomputers will become operational in the USA, Europe, and Asia by 2020-2024, supporting cutting edge research in many scientific fields.In this context, the Chinese Academy of Engineering (CAE) organized a special issue of "Postexascale Zuoning Chen, Jack J. Dongarra, Zhiwei Xu 0002 |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2017 | DaDianNao: A Neural Network SupercomputerabstractMany companies are deploying services largely based on machine-learning algorithms for sophisticated processing of large amounts of data, either for consumers or industry. The state-of-the-art and most popular such machine-learning algorithms are Convolutional and Deep Neural Networks (CNNs and DNNs), which are known to be computationally and memory intensive. A number of neural network accelerators have been recently proposed which can offer high computational capacity/area ratio, but which remain hampered by memory accesses. However, unlike the memory wall faced by processors on general-purpose workloads, the CNNs and DNNs memory footprint, while large, is not beyond the capability of the on-chip storage of a multi-chip system. This property, combined with the CNN/DNN algorithmic characteristics, can lead to high internal bandwidth and low external communications, which can in turn enable high-degree parallelism at a reasonable area cost. In this article, we introduce a custom multi-chip machine-learning architecture along those lines, and evaluate performance by integrating electrical and optical inter-chip interconnects separately. We show that, on a subset of the largest known neural network layers, it is possible to achieve a speedup of 656.63× over a GPU, and reduce the energy by 184.05× on average for a 64-chip system. We implement the node down to the place and route at 28 nm, containing a combination of custom storage and computational units, with electrical inter-chip interconnects. Shaoli Liu, Ling Li 0001, Shijin Zhang, Tianshi Chen 0002, Zhiwei Xu 0002, Olivier Temam, Yunji Chen |
IEEE Trans. Computers | 7 |
| 2016 | sAXI: A High-Efficient Hardware Inter-Node Link in ARM Server for Remote Memory AccessabstractThe ever-growing need for fast big-data operations has made in-memory processing increasingly important in modern datacenters. To mitigate the capacity limitation of a single server node, techniques of inner-rack cross-node memory access have drawn attention recently. However, existing proposals exhibit inefficiency in remote memory access among server nodes due to inter-protocol conversions and non-transparent coarse-grained accesses. In this study, we propose the high-performance and efficient serialized AXI (sAXI) link and its associated cross-node memory access mechanism for emerging ARM-based servers. The key idea behind sAXI is directly extending the on-chip AMBA AXI-4.0 interconnection of the SoC in a local server node to the outside, and then bringing into remote server nodes via high-speed serial lanes. As a result, natively accessing remote memory in adjacent nodes in the same manner of local assets is supported by purely using existing software. Experimental results show that, using the sAXI data-path, performance of remote memory access in the user-level micro-benchmark is very promising (min. latency: 1.16µs, max. bandwidth: 1.52GB/s on our in-house FPGA prototype). In addition, through this efficient hardware inter-node link, performance of an in-memory key-value framework, Redis, can be improved up to 1.72x and large latency overhead of database query can be effectively hidden. Ke Zhang 0017, Yisong Chang, Lixin Zhang 0002, Mingyu Chen 0001, Zhiwei Xu 0002 |
CCGrid | 6 |
| 2016 | Extending On-chip Interconnects for rack-level remote resource accessabstractThe need to perform data analytics on exploding data volumes coupled with the rapidly changing workloads in cloud computing places great pressure on data-center servers. To improve hardware resource utilization across servers within a rack, we propose Direct Extension of On-chip Interconnects (DEOI), a high-performance and efficient architecture for remote resource access among server nodes. DEOI extends an SoC server node's on-chip interconnect to access resources in adjacent nodes with no protocol changes, allowing remote memory and network resources to be used as if they were local. Our results on a four-node FPGA prototype show that the latency of user-level, cross-node, random reads to DEOI-connected remote memory is as low as 1.16µs, which beats current commercial technologies. We exploit DEOI remote access to improve performance of the Redis in-memory key-value framework by 47%. When using DEOI to access remote network resources, we observe an 8.4% average performance degradation and only a 2.52µs ping-pong latency disparity compared to using local assets. These results suggest that DEOI can be a promising mechanism for increasing both performance and efficiency in next-generation data-center servers. Yisong Chang, Ke Zhang 0017, Sally A. McKee, Lixin Zhang 0002, Mingyu Chen 0001, Liqiang Ren, Zhiwei Xu 0002 |
ICCD | 7 |
| 2016 | IMR: High-Performance Low-Cost Multi-Ring NoCsabstractA ring topology is a common solution of network-on-chip (NoC) in industry, but is frequently criticized to have poor scalability. In this paper, we present a novel type of multi-ring NoC called isolated multi-ring (IMR), which can even support chip multiprocessors (CMPs) with 1,024 cores. In IMR, any pair of cores are connected via at least one isolated ring, so that each packet can reach the destination without transferring from one ring to another. Therefore, IMR no longer needs expensive routers as mesh, which not only enhances the network performance but also reduces hardware overheads. We utilize simulated evolution to design optimized IMR topologies. We compare these IMR topologies against nine representative NoCs (e.g., traditional mesh, multi mesh, low-cost mesh, Express-virtual-channels mesh (EVC), torus ring, and hierarchical ring). We observe from experiments that IMR significantly outperforms its competitors in both saturation throughput and latency across all scenarios considered. For example, in a 16 × 16 CMP, IMR improves the saturation throughput of a state-of-the-art mesh (EVC) by 265.29 percent on average, and reduces the average packet latency on SPLASH-2 application traces by 71.58 percent, while consuming 5.08 percent less area and 9.76 percent less power. In a 32 × 32 CMP, IMR averagely improves the saturation throughput of EVC by 191.58 percent, and averagely reduces the packet latency on SPLASH-2 application traces by 23.09 percent, while consuming 2.86 percent less area and 10.81 percent less power. Shaoli Liu, Tianshi Chen 0002, Ling Li 0001, Xiaoxue Feng, Zhiwei Xu 0002, Haibo Chen 0001, Fred Chong, Yunji Chen |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2015 | Accelerating Apache Hive with MPI for Data Warehouse SystemsabstractData warehouse systems, like Apache Hive, have been widely used in the distributed computing field. However, current generation data warehouse systems have not fully embraced High Performance Computing (HPC) technologies even though the trend of converging Big Data and HPC is emerging. For example, in traditional HPC field, Message Passing Interface (MPI) libraries have been optimized for HPC applications during last decades to deliver ultra-high data movement performance. Recent studies, like DataMPI, are extending MPI for Big Data applications to bridge these two fields. This trend motivates us to explore whether MPI can benefit data warehouse systems, such as Apache Hive. In this paper, we propose a novel design to accelerate Apache Hive by utilizing DataMPI. We further optimize the DataMPI engine by introducing enhanced non-blocking communication and parallelism mechanisms for typical Hive workloads based on their communication characteristics. Our design can fully and transparently support Hive workloads like Intel HiBench and TPC-H with high productivity. Performance evaluation with Intel HiBench shows that with the help of light-weight DataMPI library design, efficient job start up and data movement mechanisms, Hive on DataMPI performs 30% faster than Hive on Hadoop averagely. And the experiments on TPC-H with ORCFile show that the performance of Hive on DataMPI can improve 32% averagely and 53% at most more than that of Hive on Hadoop. To the best of our knowledge, Hive on DataMPI is the first attempt to propose a general design for fully supporting and accelerating data warehouse systems with MPI. Lu Chao, Chundian Li, Xiaoyi Lu 0001, Zhiwei Xu 0002 |
ICDCS | 5 |
| 2015 | CIUV: Collaborating information against unreliable viewsabstractIn many real world applications, the information of an object can be obtained from multiple sources. The sources may provide different point of views based on their own origin. As a consequence, conflicting pieces of information are inevitable, which gives rise to a crucial problem: how to find the truth from these conflicts. Many truth-finding methods have been proposed to resolve conflicts based on information trustworthy (i.e. more appearance means more trustworthy) as well as source reliability. However, the factor of men's involvement, i.e., information may be falsified by men with malicious intension, is more or less ignored in existing methods. Collaborating the possible relationship between information's origins and men's participation are still not studied in research. To deal with this challenge, we propose a method - Collaborating Information against Unreliable Views (CIUV) - in dealing with men's involvement for finding the truth. CIUV contains 3 stages for interactively mitigating the impact of unreliable views, and calculate the truth by weighting possible biases between sources. We theoretically analyze the error bound of CIUV, and conduct intensive experiments on real dataset for evaluation. The experimental results show that CIUV is feasible and has the smallest error compared with other methods. Zimu Yuan, Zhiwei Xu 0002, Guojie Li |
ISCC | 2 |
| 2015 | Tencent and Facebook Data Validate Metcalfe's Law
Xing-Zhou Zhang, Jingjie Liu, Zhiwei Xu 0002 |
J. Comput. Sci. Technol. | 3 |
| 2015 | Statistical Performance Comparisons of ComputersabstractAs a fundamental task in computer architecture research, performance comparison has been continuously hampered by the variability of computer performance. In traditional performance comparisons, the impact of performance variability is usually ignored (i.e., the means of performance observations are compared regardless of the variability), or in the few cases directly addressed with$t$-statistics without checking the number and normality of performance observations. In this paper, we formulate a performance comparison as a statistical task, and empirically illustrate why and how common practices can lead to incorrect comparisons. We propose a non-parametric hierarchical performance testing (HPT) framework for performance comparison, which is significantly more practical than standard$t$-statistics because it does not require to collect a large number of performance observations in order to achieve a normal distribution of sample mean. In particular, the proposed HPT can facilitate quantitative performance comparison, in which the performance speedup of one computer over another is statistically evaluated. Compared with the HPT, a common practice which uses geometric mean performance scores to estimate the performance speedup has errors of$8.0$to$56.3$percent on SPEC CPU2006 or SPEC MPI2007, which demonstrates the necessity of using appropriate statistical techniques. This HPT framework has been implemented as an open-source software, and integrated in the PARSEC 3.0 benchmark suite. Tianshi Chen 0002, Qi Guo 0001, Olivier Temam, Yungang Bao, Zhiwei Xu 0002, Yunji Chen |
IEEE Trans. Computers | 6 |
| 2014 | DataMPI: Extending MPI to Hadoop-Like Big Data ComputingabstractMPI has been widely used in High Performance Computing. In contrast, such efficient communication support is lacking in the field of Big Data Computing, where communication is realized by time consuming techniques such as HTTP/RPC. This paper takes a step in bridging these two fields by extending MPI to support Hadoop-like Big Data Computing jobs, where processing and communication of a large number of key-value pair instances are needed through distributed computation models such as MapReduce, Iteration, and Streaming. We abstract the characteristics of key-value communication patterns into a bipartite communication model, which reveals four distinctions from MPI: Dichotomic, Dynamic, Data-centric, and Diversified features. Utilizing this model, we propose the specification of a minimalistic extension to MPI. An open source communication library, DataMPI, is developed to implement this specification. Performance experiments show that DataMPI has significant advantages in performance and flexibility, while maintaining high productivity, scalability, and fault tolerance of Hadoop. Xiaoyi Lu 0001, Li Zha, Zhiwei Xu 0002 |
IPDPS | 5 |
| 2014 | ArchRanker: A ranking approach to design space explorationabstractArchitectural Design Space Exploration (DSE) is a notoriously difficult problem due to the exponentially large size of the design space and long simulation times. Previously, many studies proposed to formulate DSE as a regression problem which predicts architecture responses (e.g., time, power) of a given architectural configuration. Several of these techniques achieve high accuracy, though often at the cost of significant simulation time for training the regression models.We argue that the information the architect mostly needs during the DSEprocess is whether a given configuration will perform better than another one in the presences ofdesign constraints, or better than any other one seen so far, rather than precisely estimating the performance of that configuration. Based on this observation, we propose a novel rankingbased approach to DSE where we train a model to predict which of two architecture configurations will perform best. We show that, not only this ranking model more accurately predicts the relative merit of two architecture configurations than an ANN-based state-of-the-art regression model, but also that it requires much fewer training simulations to achieve the same accuracy, or that it can be used for and is even better at quantifying the performance gap between two configurations. We implement the framework for training and using this model, called ArchRanker, and we evaluate it on several DSE scenarios (unicore/multicore design spaces, and both time and power performance metrics). We try to emulate as closely as possible the DSE process by creating constraint-based scenarios, or an iterative DSEprocess. We find that ArchRanker makes 29.68% to 54.43% fewer incorrect predictions on pairwise relative merit of configurations (tested with 79,800 configuration pairs) than an ANN-based regression model across all DSE scenarios considered (values averaged over all benchmarks for each scenario). We also find that, to achieve the same accuracy as ArchRanker, the ANN often requires three times more training simulations. Tianshi Chen 0002, Qi Guo 0001, Ke Tang 0001, Olivier Temam, Zhiwei Xu 0002, Zhi-Hua Zhou, Yunji Chen |
ISCA | 5 |
| 2014 | DaDianNao: A Machine-Learning SupercomputerabstractMany companies are deploying services, either for consumers or industry, which are largely based on machine-learning algorithms for sophisticated processing of large amounts of data. The state-of-the-art and most popular such machine-learning algorithms are Convolutional and Deep Neural Networks (CNNs and DNNs), which are known to be both computationally and memory intensive. A number of neural network accelerators have been recently proposed which can offer high computational capacity/area ratio, but which remain hampered by memory accesses. However, unlike the memory wall faced by processors on general-purpose workloads, the CNNs and DNNs memory footprint, while large, is not beyond the capability of the on chip storage of a multi-chip system. This property, combined with the CNN/DNN algorithmic characteristics, can lead to high internal bandwidth and low external communications, which can in turn enable high-degree parallelism at a reasonable area cost. In this article, we introduce a custom multi-chip machine-learning architecture along those lines. We show that, on a subset of the largest known neural network layers, it is possible to achieve a speedup of 450.65x over a GPU, and reduce the energy by 150.31x on average for a 64-chip system. We implement the node down to the place and route at 28nm, containing a combination of custom storage and computational units, with industry-grade interconnects. Yunji Chen, Shaoli Liu, Shijin Zhang, Liqiang He, Ling Li 0001, Tianshi Chen 0002, Zhiwei Xu 0002, Ninghui Sun, Olivier Temam |
MICRO | 9 |
| 2014 | Performance Characterization of Hadoop and Data MPI Based on Amdahl's Second LawabstractAmdahl's second law has been seen as a useful guideline for designing and evaluating balanced computer systems for decades. This law has been mainly used for hardware systems and peak capacities. This paper utilizes Amdahl's second law from a new angle, i.e., Evaluating the influence on systems performance and balance of the application framework software, a key component of big data systems. We compare two big data application framework software systems, Apache Hadoop and Data MPI, with three representative application benchmarks and various data sizes. System monitors and hardware performance counters are used to record the resource utilization, characteristics of instructions execution, memory accesses, and I/O rates. These numbers are used to reveal the three runtime metrics of Amdahl's second law: CPU speed (GIPS), memory capacity (GB), and I/O rate (Gbps). The experiment and evaluation results show that a Data MPI-based big data system has better performance and is more balanced than a Hadoop-based system. Xiaoyi Lu 0001, Zhiwei Xu 0002 |
NAS | 4 |
| 2014 | On the inapproximability of minimizing cascading failures under the deterministic threshold model
Jingjie Liu, Haicang Zhang, Zhiwei Xu 0002 |
Inf. Process. Lett. | 4 |
| 2014 | Cloud-Sea Computing Systems: Towards Thousand-Fold Improvement in Performance per Watt for the Coming Zettabyte Era
Zhiwei Xu 0002 |
J. Comput. Sci. Technol. | 1 |
| 2013 | Consolidated cluster systems for data centers in the cloud age: a survey and analysis
Jian Lin 0006, Li Zha, Zhiwei Xu 0002 |
Frontiers Comput. Sci. | 3 |
| 2013 | Effective and efficient microprocessor design space exploration using unlabeled design configurationsabstractEver-increasing design complexity and advances of technology impose great challenges on the design of modern microprocessors. One such challenge is to determine promising microprocessor configurations to meet specific design constraints, which is called Design Space Exploration (DSE). In the computer architecture community, supervised learning techniques have been applied to DSE to build regression models for predicting the qualities of design configurations. For supervised learning, however, considerable simulation costs are required for attaining the labeled design configurations. Given limited resources, it is difficult to achieve high accuracy. In this article, inspired by recent advances in semisupervised learning and active learning, we propose the COAL approach which can exploit unlabeled design configurations to significantly improve the models. Empirical study demonstrates that COAL significantly outperforms a state-of-the-art DSE technique by reducing mean squared error by 35% to 95%, and thus, promising architectures can be attained more efficiently. Tianshi Chen 0002, Yunji Chen, Qi Guo 0001, Zhi-Hua Zhou, Ling Li 0001, Zhiwei Xu 0002 |
ACM Trans. Intell. Syst. Technol. | 6 |
| 2012 | An Elastic Architecture Adaptable to Millions of Application Scenarios
Yunji Chen, Tianshi Chen 0002, Qi Guo 0001, Zhiwei Xu 0002, Lei Zhang 0008 |
NPC | 4 |
| 2012 | How much power is needed for a billion-thread high-throughput server?
Zhiwei Xu 0002 |
Frontiers Comput. Sci. | 1 |
| 2011 | Effective and Efficient Microprocessor Design Space Exploration Using Unlabeled Design Configurations
Qi Guo 0001, Tianshi Chen 0002, Yunji Chen, Zhi-Hua Zhou, Weiwu Hu, Zhiwei Xu 0002 |
IJCAI | 6 |
| 2011 | Vega LingCloud: A Resource Single Leasing Point System to Support Heterogeneous Application Modes on Shared InfrastructureabstractIn large organizations or IDCs, different departments always occupy and maintain dedicated resources to satisfy their or their customers' heterogeneous application loads. This situation easily makes the infrastructure management a repeated and inefficient work. Even worse, it is difficult to share the resources owned by different departments even when they are idle, because the application modes on these resources are quite different. This paper introduces a live system, Vega Ling Cloud, which provides a Resource Single Leasing Point System for consolidated renting physical and virtual machines to support heterogeneous application modes on shared infrastructure. Furthermore, we present the asset-leasing model and the architecture of Vega Ling Cloud. According to the evaluation, Vega Ling Cloud is better than other systems like Open Nebula and Enomaly ECP in the aspects of uniformity, flexibility, security, usability, and efficiency. The experimental result of management overhead shows that the deployment speed of virtual machine in Vega Ling Cloud is 4.1 times of that in the Open Nebula and VIDA hybrid system for deploying 64 virtual machines concurrently. From a representative micro-cloud example in a research group, we show that the consolidated way of leasing physical and virtual machine in Vega Ling Cloud is approbatory. Up to now, Vega Ling Cloud has been deployed in real world environments, which include a private cloud of large organization in Beijing and a public cloud in the Dongguan IDC of China. The resource scale of the Dongguan cloud reaches about 504 cores, 625 GB memory, and 156 TB storage. The number of supported real applications in Beijing and Dongguan clouds has exceeded 35, and their modes involve high performance computing, large scale data processing, virtual machine leasing, data storage, and so on. Xiaoyi Lu 0001, Jian Lin 0006, Li Zha, Zhiwei Xu 0002 |
ISPA | 4 |
| 2011 | History-Aware Adaptive Backoff for Neighbor Discovery in Wireless NetworksabstractThe ability of discovering neighboring nodes, namely neighbor discovery, is essential for the self-organization of wireless ad hoc networks. In this paper, we propose a history-aware adaptive back off algorithm for neighbor discovery assuming collision detection and feedback mechanisms. Given successful discovery feedback, undiscovered nodes can adjust their contention window. With collision feedback and historical information, only transmission nodes enter the re-contention process, and decrease their contention window to accelerate neighbor discovery process after collision. Then, we give theoretical analysis of our algorithm on the discovery time and energy consumption, and derive the optimal size of contention windows by two rounds of optimization. Finally, we validate our theoretical analysis by simulations, and show the performance improvement over existing algorithms. Zimu Yuan, Lizhao You, Wei Li 0008, Biao Chen 0002, Zhiwei Xu 0002 |
MSN | 5 |
| 2011 | Three New Concepts of Future Computer Science
Zhiwei Xu 0002, Dandan Tu |
J. Comput. Sci. Technol. | 1 |
| 2010 | Design and Analysis of a New GPS AlgorithmabstractIn this paper, we propose and analyze a new GPS positioning algorithm. Our algorithm uses the direct linearization technique to reduce the computation time overhead. We invoke the general least squares method in order to achieve optimality in the situation when the trilateration system of equations becomes over-determined. We systematically evaluate our new algorithms and show that they indeed take much less computation time than the traditional GPS method while maintaining reasonable accuracy. Wei Li 0008, Zhiwei Xu 0002, Wei Zhao 0001 |
ICDCS | 4 |
| 2010 | CCIndex: A Complemental Clustering Index on Distributed Ordered Tables for Multi-dimensional Range Queries
Yongqiang Zou, Shicai Wang, Li Zha, Zhiwei Xu 0002 |
NPC | 5 |
| 2009 | Four styles of parallel and net programming
Zhiwei Xu 0002, Yongqiang He, Wei Lin 0004, Li Zha |
Frontiers Comput. Sci. China | 1 |
| 2009 | Classifying rendezvous tasks of arbitrary dimension
Xingwu Liu, Zhiwei Xu 0002, Jianzhong Pan |
Theor. Comput. Sci. | 2 |
| 2008 | Incentive-Based Scheduling for Market-Like Computational GridsabstractA sustainable market-like computational grid has two characteristics: it must allow resource providers and resource consumers to make autonomous scheduling decisions, and both parties of providers and consumers must have sufficient incentives to stay and play in the market. In this paper, we formulate this intuition of optimizing incentives for both parties as a dual-objective scheduling problem. The two objectives identified are to maximize the success rate of job execution and to minimize fairness deviation among resources. The challenge is to develop a grid scheduling scheme that enables individual participants to make autonomous decisions while producing a desirable emergent property in the grid system; that is, the two systemwide objectives are achieved simultaneously. We present an incentive-based scheduling scheme, which utilizes a peer-to-peer decentralized scheduling framework, a set of local heuristic algorithms, and three market instruments of job announcement, price, and competition degree. The performance of this scheme is evaluated via extensive simulation using synthetic and real workloads. The results show that our approach outperforms other scheduling schemes in optimizing incentives for both consumers and providers, leading to highly successful job execution and fair profit allocation. Lijuan Xiao, Yanmin Zhu 0006, Lionel M. Ni, Zhiwei Xu 0002 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2007 | An Approach to Debugging Grid or Web ServicesabstractIn this paper, we first introduce some issues that are encountered in building a service debugger and briefly describe our approach to addressing them. Next, we outline some debugging modes and components of a simple composite debugger. Then, we mainly describe its message-based front-end and back-end, which are a co-existing, self-identifying, and non- intrusive. Finally, we preset some experimental results of our latest prototype. Qiang Yue 0001, Zhiwei Xu 0002, Haiyan Yu 0002, Wei Li 0008, Li Zha |
ICWS | 2 |
| 2007 | Personal Grid
Zhiwei Xu 0002, Lijuan Xiao, Xingwu Liu |
NPC | 1 |
| 2007 | Revisiting the Impossibility for Boosting Service Resilience
Xingwu Liu, Zhiwei Xu 0002, Juhua Pu |
TAMC | 2 |
| 2006 | Anycast Routing in Delay Tolerant NetworksabstractAnycast routing is very useful for many applications such as resource discovery in delay tolerant networks (DTNs). In this paper, based on a new DTN model, we first analyze the any-cast semantics for DTNs. Then we present a novel metric named EMDDA (expected multi-destination delay for anycast) and a corresponding routing algorithm for anycast routing in DTNs. Extensive simulation results show that the proposed EMDDA routing scheme can effectively improve the efficiency of anycast routing in DTNs. It outperforms another algorithm, minimum expected delay (MED) algorithm, by 11.3% on average in term of routing delays and by 19.2% in term of average max queue length. Yili Gong, Yongqiang Xiong, Qian Zhang 0001, Zhensheng Zhang, Wenjie Wang 0006, Zhiwei Xu 0002 |
GLOBECOM | 6 |
| 2006 | An Approach to Exception Handling for Service-Oriented SystemsabstractThis paper sets out key issues of exception handling that relate to identifying, analyzing and dealing with an exception in service-oriented systems. Then required concepts, basic structures and algorithms for the Ml-Eh are discussed. The new concepts include the super-space of exception definitions, the meta-definition mode, and components in an exception message. The main characteristics of the Ml-Eh are a uniform definition mode, dynamical extensibility, and abilities of configurable analyses and encapsulations, and a message-level combination. Finally, implementation and evaluation of the mechanism is given Qiang Yue 0001, Hao Wang 0002, Li Zha, Zhiwei Xu 0002 |
ICWS | 5 |
| 2006 | Incentive-based scheduling in Grid computingabstractAbstract With the rapid development of high‐speed wide‐area networks and powerful yet low‐cost computational resources, Grid computing has emerged as an attractive computing paradigm. In typical Grid environments, there are two distinct parties, resource consumers and resource providers. Enabling an effective interaction between the two parties (i.e. scheduling jobs of consumers across the resources of providers) is particularly challenging due to the distributed ownership of Grid resources. In this paper, we propose an incentive‐based peer‐to‐peer (P2P) scheduling for Grid computing, with the goal of building a practical and robust computational economy. The goal is realized by building a computational market supporting fair and healthy competition among consumers and providers. Each participant in the market competes actively and behaves independently for its own benefit. A market is said to be healthy if every player in the market gets sufficient incentive for joining the market. To build the healthy computational market, we propose the P2P scheduling infrastructure, which takes the advantages of P2P networks to efficiently support the scheduling. The proposed incentive‐based algorithms are designed for consumers and providers, respectively, to ensure every participant gets sufficient incentive. Simulation results show that our approach is successful in building a healthy and scalable computational economy. Copyright © 2006 John Wiley & Sons, Ltd. Yanmin Zhu 0006, Lijuan Xiao, Zhiwei Xu 0002, Lionel M. Ni |
Concurr. Comput. Pract. Exp. | 3 |
| 2005 | Languages for the Net: From Presentation to Collaboration
Zhiwei Xu 0002, Haozhi Liu, Haiyan Yu 0002 |
APWeb | 1 |
| 2005 | Modeling and PerformanceAnalysis of the VEGA Grid SystemabstractIn this paper, we propose four general queueing models based on input and server distributions, to analyze a special grid system, VEGA grid system version 1.1 (VEGA1.1). The mean queue lengths and mean waiting times of these models are deduced. The two classic applications, the computing-oriented application (blast computing) and online transaction processing application (air booking service) are employed to analyze and evaluate the grid system and the models. Meanwhile, we validate VEGA's four prominent characteristics, versatile services, enabling intelligence, global uniformity and autonomous control. Finally, we analyze the experiment results, predict the behavior of VEGA1.1 in equilibrium and point out the bottlenecks of the system Zhiwei Xu 0002, Yuzhong Sun |
e-Science | 2 |
| 2005 | A C/S and P2P Hybrid Resource Discovery Framework in Grid EnvironmentsabstractResource discovery is crucial to efficient deployment of a grid system whose dynamic, heterogeneous characteristics make it difficult. In this paper, Vega Infrastructure for Resource Discovery (VIRD) is developed, then augmented with new features (i.e., some new algorithms) to build a C/S (client/server) and P2P (peer-to-peer) hybrid resource discovery framework. The three layered architecture of the VIRD is developed to make advantage of the physical and logical topologies of the Internet to facilitate resource discovery. With our simulations and theoretical analysis, it is proved that VIRD is of good scalability with respect to the sizes of the underlying backbone. Even when the resource density is low and the max TTL (time-to-live) is small, VIRD still achieves high search success rates in a small amount of hops. Compared with flooding and random walk algorithms via the same search success rates, VIRD outperforms them in both network traffic and response time. Yili Gong, Wei Li 0008, Yuzhong Sun, Zhiwei Xu 0002 |
ICPP | 4 |
| 2005 | Performance Analysis and Prediction on VEGA Grid
Zhiwei Xu 0002, Yuzhong Sun, Zheng Shen, Changshu Liu |
ISPA | 2 |
| 2005 | System Software for China National Grid
Li Zha, Wei Li 0008, Haiyan Yu 0002, Xianghui Xie 0001, Zhiwei Xu 0002 |
NPC | 6 |
| 2005 | An Agile Programming Model for Grid End UsersabstractGrid and service computing technologies have been explored by enterprises to promote integration, sharing, and collaboration. However, quick response to business environment changes is still a challenging issue. For end users, developing, customizing, and reengineering applications remain a difficult and timeconsuming task. Users still need to deal with excessive low-level details of platform-specific APIs. We present a high-level programming model together with a descriptive glueing language called GSML, to facilitate end-user programming. In this approach, applications could be visually composed from welldefined software components called "funnels" in an event-driven fashion. Application examples have shown that, by raising the level of abstraction as well as simplifying the programming model, GSML could empower end users to build grid applications on demand with improved productivity. Zhiwei Xu 0002, Chengchun Shu, Haiyan Yu 0002, Haozhi Liu |
PDCAT | 1 |
| 2005 | BLOSSOMS: Building Lightweight Optimized Sensor Systems on a Massive Scale
Wen Gao 0001, Lionel M. Ni, Zhiwei Xu 0002, Shing-Chi Cheung, Qiong Luo 0001 |
J. Comput. Sci. Technol. | 3 |
| 2004 | Grid Replication Coherence ProtocolabstractSummary form only given. Here we present two efficient coherence protocols for data grid applications. The data intensive grid applications often adopt the replica strategy to solve the data access bottleneck. It is important to keep coherence among all replicas of a copy of large amount of data. The first strategy for coherence is called lazy copy coherence protocol and the second aggressive copy coherence. Meanwhile, we consider in which situation we should make a replica for a copy of large amount of data. Yuzhong Sun, Zhiwei Xu 0002 |
IPDPS | 2 |
| 2004 | Design and Implementation of a 3A Accessing Paradigm Supported Grid Application and Programming Environment
Ge He, Donghua Liu, Yuzhong Sun, Zhiwei Xu 0002 |
ISPA | 4 |
| 2004 | BLOSSOMS: A CAS/HKUST Joint Project to Build Lightweight Optimized Sensor Systems on a Massive Scale
Wen Gao 0001, Lionel M. Ni, Zhiwei Xu 0002 |
NPC | 3 |
| 2004 | Vega: A Computer Systems Approach to Grid Computing
Zhiwei Xu 0002, Wei Li 0008, Li Zha, Haiyan Yu 0002, Donghua Liu |
J. Grid Comput. | 1 |
| 2003 | VegaFS: A Prototype for File-Sharing Crossing Multiple Administrative DomainsabstractAccessing remote resources is a principal challenge of grid computing. For wide-area file sharing, a most difficult problem is the inability to access files distributed in different administrative domains. In this paper, we propose a file system architecture called VegaFS, which is detached from administrative domains entirely and provides cross-domain file access abilities. The main idea is to adopt public keys in native file systems and transfer the management work of administrator to trusted certificate authorities (CAs). Compared with other systems, VegaFS provides several benefits, such as grid-wide unique user identities, cross-domain file access capability, scalabilities, fine-grained dynamic access control and better securities. Wei Li 0008, Jianmin Liang, Zhiwei Xu 0002 |
CLUSTER | 3 |
| 2003 | Community-Based Model and Access Control for Information GridabstractIt is a challenge to integrate and share resources securely across multiple autonomous administrative domains environment. We introduce the community-based model of vega enterprise information grid and its operation mechanism. The individuals and/or institutes forming different communities can not only implement diverse policies and autonomous management of resource sharing without any impact on the other communities, but also achieve a globally shared goal. Based on the model, the access control technique, including framework and uniform formalizing, is presented in detail. The static or dynamic role binding is used for entitling user's request for accessing resources within or across communities, which ensures a globally unified view for user. A prototype describes how these techniques are useful in building an enterprise information grid. The evaluation and future continuing works are presented in the conclusion. Xiaolin Li 0003, Zhiwei Xu 0002, Xingwu Liu |
Web Intelligence | 2 |
| 2003 | VEGA Infrastructure for Resource Discovery in Grids
Yili Gong, Fangpeng Dong, Wei Li 0008, Zhiwei Xu 0002 |
J. Comput. Sci. Technol. | 4 |
| 2002 | Cluster file systems: a case study
Jianyong Wang 0001, Zhiwei Xu 0002 |
Future Gener. Comput. Syst. | 2 |
| 2001 | Cluster and Grid Superservers: The Dawning Experiences in ChinaabstractThis paper summarizes recent activities at Institute of Computing Technology, Chinese Academy of Sciences, in developing superservers for cluster and grid computing. We first identify market and technical trends observed from a Chinese perspective. Then we describe the research work in developing the Dawning series high performance computers and the China computational grid. We also highlight some on-going research work in developing grid-oriented superserver systems. Zhiwei Xu 0002, Ninghui Sun, Dan Meng 0002, Wei Li 0008 |
CLUSTER | 1 |