VLDB 2026 Research / reviewers in the wild / expert
Qing Yang 0001
dblp:47/3749-1 · also Qing (Ken) Yang
· DBLP profile ↗
95ranked-venue papers
22as first author
14since 2021 · last 2026
0000-0002-3330-1542ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 77 · 18 first-author · 13 since 2021Software engineering, systems software and programming languages · 7 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-authorComputer networks · 4 · 1 first-authorSecurity and privacy · 3 · 1 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient LLM Decoding on Ryzen AI NPUsabstractWe propose an efficient and scalable LLM decoding framework optimized for AMD Ryzen AI NPUs, leveraging two novel techniques: FusedDQP and FlowKV. FusedDQP fuses dequantization with projection to minimize memory operations and latency, while FlowKV introduces a pipelined, bandwidth-optimized approach for KV cache access across compute tiles (CT). Together, these methods deliver substantial improvements in both speed and energy efficiency without altering model accuracy. Our solution achieves up to 14.2× speedup and 2.66× power efficiency gains compared to existing state-of-the-art (SOTA) NPU baselines, demonstrating linear scalability with CT count and robustness across LLaMA-3.1/3.2 model variants (1B, 3B, and 8B parameters). We also benchmark against CPU and iGPU on the same platform, our performance surpasses CPU and iGPU (up to 1.8x and 16.2x speedup), while delivering substantially improved energy efficiency (up to 3.63x and 11.38x for CPU and iGPU, respectively). Zhenyu Xu 0007, Miaoxiang Yu, Jillian Cai, Qing Yang 0001, Tao Wei 0001 |
DATE | 4 |
| 2026 | An Efficient Dataflow Framework for DiT-Based Image GenerationabstractWe present a general and efficient dataflow framework for mapping Diffusion Transformer (DiT)-based models onto tiled mesh accelerators. The framework explicitly orchestrates tiling, streaming, and Direct Memory Access (DMA)-driven data movement to efficiently execute attention and large matrix multiplication (MM). We use the AMD Ryzen AI Neural Processing Unit (NPU) [1] as a representative edge-class tiled-mesh dataflow platform. In our implementation, the DiT denoising stage runs on the NPU, while the text encoder and Variational Autoencoder (VAE) decoder remain on the CPU.The iterative denoising loop dominates end-to-end latency. We therefore focus our acceleration and optimization efforts on the denoising module, while keeping the text-embedding and VAE-decoder stages on the CPU. Each denoising step reuses the same transformer backbone, and the dominant operations within a step are MM, self-attention, and MLP submodules that combine MM with elementwise nonlinear operations. The NPU is organized as a two-dimensional array of compute tiles (CTs), also referred to as AIEs. Our design partitions input matrices into tiles and maps independent tile computations across multiple AIE cores to exploit spatial parallelism.We evaluate the framework on two representative DiT-based text-to-image models, FLUX.1-schnell and Z-Image-Turbo [2], [3], and achieve end-to-end generation of a high-quality image in just over one minute. Compared with the integrated GPU (iGPU), the NPU delivers comparable generation latency while being approximately 4.4× more power efficient (package). For FLUX.1-schnell, the iGPU achieves 10.5 s/step, while the NPU achieves 13.3 s/step. For Z-Image-Turbo, the iGPU achieves 10.1 s/step, while the NPU achieves 9.8 s/step, showing that the NPU can match and slightly surpass iGPU denoising latency. These results translate into substantially improved energy efficiency for on-device image generation. Although demonstrated on a specific NPU and two representative models, the proposed dataflow framework is neither hardware-specific nor model-specific. It generalizes naturally to other diffusion-family models and mesh-based dataflow accelerators. Our results highlight the potential of spatial NPUs as energy-efficient platforms for next-generation on-device generative AI.Code Availability: The source code and setup instructions are available at: https://github.com/jia1217/vigenflow. Yazhe Zhang, Shouyu Du, Zhenyu Xu 0007, Miaoxiang Yu, Dingjiang Yan, Zhiheng Ni, Qing Yang 0001, Tao Wei 0001 |
FCCM | 7 |
| 2025 | Tile-Level Pipeline for Linear Scalable Stencil Computation on AMD AI EnginesabstractStencil computation is an essential method, particularly useful for numerical simulations in areas like acoustics, heat transfer, and electromagnetism. Recent studies have utilized AMD AI Engines (AIEs) for stencil computations by configuring multiple AIE tiles within a Compute Unit to exploit task-level parallelism, achieving notable speedup through concurrent task execution. However, this setup suffers from suboptimal performance due to high memory bandwidth demands, resulting in underutilization of the available AIE tiles. This work introduces a Tile-Level Pipeline architecture designed for stencil computation that operates with constant memory bandwidth. This approach achieves linear scalability, where performance scales linearly with the number of AIE tiles, and ensures full utilization of all AIE tiles on the chip. We empirically demonstrate these benefits using AIEs. Zhenyu Xu 0007, Miaoxiang Yu, Yazhe Zhang, Jillian Cai, Qing Yang 0001, Tao Wei 0001 |
FPGA | 5 |
| 2024 | TwinStep Network (TwNet): a Neuron-Centric Architecture Achieving Rapid TrainingabstractRecurrent Neural Networks (RNNs) face challenges with the Back Propagation Through Time (BPTT) algorithm, leading to substantial computational and memory demands in training, especially on GPUs. Inspired by biological neural systems, we introduce the TwinStep Network (TwNet) via algorithm/architecture co-design, achieving online training via a neuron-centric design. At its core, TwinStep signifies that both the forward pass (inference) and back propagation (training) steps for each neuron happen concurrently. This approach, which more closely resembles biological neural processes, eliminates the necessity of storing the intermediate state of each neuron at each time steps, as required in BPTT. Consequently, it overcomes the limitation on the number of time steps that can be included in the BPTT training process. Uniquely, TwNet's “pipeline parallelism” facilitates serial processing and concurrent handling of multiple time steps. We implemented TwNet on FPGAs with a fully pipelined architecture. It achieves up to 885x speedup in training several popular RNN testbenches in comparison with other state-of-the-art approaches while maintaining accuracy, marking an advancement in online RNN training and potential applications. Zhenyu Xu 0007, Miaoxiang Yu, Jillian Cai, Qing Yang 0001, Tao Wei 0001 |
ASAP | 4 |
| 2024 | An FPGA-Enabled Framework for Rapid Automated Design of Photonic Integrated CircuitsabstractThis paper introduces an FPGA-enabled framework to accelerate the automated design process for Photonic Integrated Circuit (PIC) devices. PICs are foreseen as a foundation for the next-generation semiconductors. However, the complexity of PIC design presents considerable challenges. Machine Learning (ML) techniques have shown promise in the realm of PIC design. The primary hurdle, however, is the extended training duration, solely constrained by the slow electromagnetic (EM) Finite-Difference Time-Domain (FDTD) solver. We propose a fast framework with a dedicated FPGA FDTD accelerator tailor-designed to speed up the PIC simulation. Benchmarking was carried out against commercial tools, with the single-FPGA accelerator outperforming both a multicore CPU and a GPU cluster. We taped out and evaluated the PIC devices designed through the proposed framework, and the experimental outcomes aligned. This demonstrates the full design circle, showcasing that the proposed framework enabled by FPGA breaks the current bottleneck in this domain. Zhenyu Xu 0007, Miaoxiang Yu, Jillian Cai, Saddam Gafsi, Judson Douglas Ryckman, Qing Yang 0001, Tao Wei 0001 |
FPGA | 6 |
| 2023 | A Novel FPGA-Based Circuit Simulator for Accelerating Reinforcement Learning-Based Design of Power ConvertersabstractHigh-efficiency energy conversion systems have become increasingly important due to their wide use in all electronic systems such as data centers, smart mobile devices, E-vehicles, medical instruments, and so forth. Complex and interdependent parameters make optimal designs of power converters challenging to get. Recent research has shown that machine learning (ML) algorithms, such as reinforcement learning (RL), show great promise in design of such converter circuits. A trained RL agent can search for optimal design parameters for power conversion circuit topologies under targeted application requirements. Training an RL agent requires numerous circuit simulations. It requires significantly more training iterations when the tolerance of circuit components due to manufacturing inconsistency, aging, and temperature variation is considered. As a result, they may take days to complete, primarily because of the slow time-domain circuit simulation. This paper proposes a new FPGA architecture that accelerates the circuit simulation and hence substantially speeds up the RL-based design method for power converters. Our new architecture supports all power electronic circuit converters and their variations. It substantially improves the training speed of RL-based design methods. High-level synthesis (HLS) was used to build the accelerator on Amazon Web Service (AWS) F1 instance. An AWS virtual PC hosts the training algorithm. The host interacts with the FPGA accelerator by updating the circuit parameters, initiating simulation, and collecting the simulation results during training iterations. A script was created on the host side to facilitate this design method to convert a netlist containing circuit topology and parameters into core matrices in the FPGA accelerator. Experimental results showed$\mathbf{60}\times$overall speedup of our RL-based design method in comparison with using a popular commercial simulator, PowerSim. Zhenyu Xu 0007, Miaoxiang Yu, Jillian Cai, Qing Yang 0001, Yeonho Jeong, Tao Wei 0001 |
ASAP | 4 |
| 2023 | A Heterogeneous Computer Architecture Accelerating Reinforcement Learning-based Design for Silicon Photonic DevicesabstractThis paper proposes a framework to substantially accelerate Reinforcement Learning (RL)-based design method for Photonic Integrated Circuit (PIC) devices. PICs are widely anticipated to underpin the forthcoming generation of semiconductor chips. However, the complexity of PIC design, which includes hundreds of degrees of freedom (DOF), presents considerable challenges. Machine Learning (ML) techniques, inclusive of RL, have demonstrated their effectiveness in the domain of PIC design. The primary hurdle, however, is the extended training duration, primarily constrained by the sluggish electromagnetic (EM) solver, specifically, the Finite-Difference Time-Domain (FDTD) solver. We have engineered a novel computational architecture that can be deployed on cloud-based systems using a cluster of Central Processing Units (CPUs), Field-Programmable Gate Arrays (FPGAs), and Graphics Processing Units (GPUs). An FPGA-FDTD accelerator, which capitalizes on the high memory bandwidth of on-chip memory (OCM), has been specifically designed to simulate planar PIC devices. Each FPGA-FDTD accelerator, also denoted as an FPGA kernel, functions as an autonomous RL environment. The host machine, in conjunction with the FPGA kernel, is designated as a worker node within the cluster. A functional prototype has been successfully implemented on the Amazon Web Services (AWS) cloud. The framework has effectively designed numerous PIC devices, and experimental results show the architecture outperforms existing methods significantly in design speed while maintaining or exceeding their design quality. Notably, the framework's versatility requires minimal adjustments for a broad range of devices, significantly reducing design time and promising to expedite PIC innovation. Miaoxiang Yu, Zhenyu Xu 0007, Jillian Cai, Qing Yang 0001, Tao Wei 0001 |
ASAP | 4 |
| 2023 | A Finite-Difference Time-Domain (FDTD) solver with linearly scalable performance in an FPGA clusterabstractThis paper presents an FPGA cluster-based Finite-Difference Time-Domain (FDTD) accelerator that offers a linear speedup with the number of FPGAs participating in computation within the cluster. FDTD is a numeric method for simulating electromagnetic wave propagation and interactions with diverse materials and structures. Recent advancements in machine learning-based design and optimization techniques for photonic integrated circuits and microwave circuits, known as inverse design, have demonstrated remarkable success. Inverse design necessitates numerous FDTD simulations, and the high-performance FDTD accelerator enables rapid design automation, which is crucial for accelerating innovation. Our proposed accelerator comprises deeply pipelined FDTD cell update kernels that can traverse multiple FPGAs via high-speed optical links, effectively utilizing available resources across all FPGAs in a cluster. The architecture includes a head node and a flexible number of cascaded server nodes, together with custom cross-FPGA data routing kernels integrated into the "Open Cloud Testbed" (OCT) FPGA infrastructure to facilitate seamless data transfer. The proposed accelerator is developed on an existing platform, OCT FPGA. Our experiments reveal that, for a 4096×4096 2.5D FDTD simulation, each server node (Xilinx Alveo U280) can achieve 86.4 Giga-cells updates per second (GCUPS), and the head node can achieve 38.4 GCUPS. The overall speed with 4 server nodes is 38.4 + 4×86.4 = 384 GCUPS. Zhenyu Xu 0007, Miaoxiang Yu, Jillian Cai, Qing Yang 0001, Tao Wei 0001 |
CLUSTER | 4 |
| 2023 | A Novel FPGA Simulator Accelerating Reinforcement Learning-Based Design of Power ConvertersabstractHigh-efficiency energy conversion systems have become increasingly important due to their wide use in all electronic systems such as data centers, smart mobile devices, E-vehicles, medical instruments, and so forth. Complex and interdependent parameters make optimal designs of power converters challenging to get. Recent research has shown that reinforcement learning (RL) shows great promise in the design of such converter circuits. A trained RL agent can search for optimal design parameters for power conversion circuit topologies under targeted application requirements. Training an RL agent requires numerous circuit simulations. As a result, they may take days to complete, primarily because of the slow time-domain circuit simulation. Zhenyu Xu 0007, Miaoxiang Yu, Qing Yang 0001, Yeonho Jeong, Tao Wei 0001 |
FPGA | 3 |
| 2022 | Highly Scalable Runtime Countermeasure Against Microprobing Attacks on Die-to-Die Interconnections in System-in-PackageabstractThe emerging System-in-Package (SiP) technology has enabled multiple dies fabricated on a single chip for high performance and energy efficiency. Die-to-die (D2D) communication in SiP is typically unencrypted, exposing sensitive data to possible microprobing attacks. In this paper, we propose an on-chip microprobe detection circuit together with a noise canceling technique to protect D2D buses for future SiP security. The proposed method utilizes the metastable state of a flip-flop to detect the small timing variation caused by the inevitable loading effect of a microprobe. This design requires minimum digital resources with high scalability. Uniquely, the proposed design protects D2D buses at runtime without interfering with normal data transfers. In addition, it introduces zero latency to the communication channel. We built the detection circuit in a Xilinx ZYNQ Ultrascale+ SoC to prove its feasibility. Dynamic partial reconfiguration function is employed to create the test platform and emulate D2D interconnections as well as microprobing attacks on them. To demonstrate its potential to be used in standard communication protocols, we integrated the detection circuit with a fully functional Advanced eXtensible Interface (AXI) bus. Experimental results show that the proposed runtime detection method is effective, resource-efficient, and reliable under temperature-varying environments. Zhenyu Xu 0007, Thomas Mauldin, Qing Yang 0001, Tao Wei 0001 |
FPGA | 3 |
| 2022 | In-Sensor Neural Network Preprocessing for ADAS Computer SystemsabstractCurrently available automotive radars are designed to stream real-time 2D image data over high-speed links to a central ADAS (Advance Driver-Assistance System) computer for object recognition, which considerably contributes to the system's power consumption and complexity. This paper presents a preliminary work for the implementation of a new in-sensor computer architecture to extract representative features from raw sensor data to detect and identify objects with radar signals. Such new architecture makes it possible to reduce the data transferred between sensors and the central ADAS computer significantly, giving rise to significant energy savings and latency reductions, while still maintaining sufficient accuracy and preserving image details. An experimental prototype has been built using the Texas Instruments AWR1243 Frequency-Modulated Continuous Wave (FMCW) radar board. We carried out experiments using the prototype to collect radar images, to preprocess raw data, and to transfer feature vectors to the central ADAS computer for classification and object detection. Two different approaches will be presented in this paper: First, a vanilla autoencoder will demonstrate the possibility of data reduction on radar signals. Second, a convolutional neural network based cross-domain deep learning architecture is presented by using a sample dataset to show the feasibility of computing Range-Angle Heatmaps directly on the sensor board eliminating the need for the raw data preprocessing on the central ADAS computer. We show that the reconstruction of Range-Angle Heatmaps can be predicted with a very high accuracy by leveraging deep learning architectures. Implementation of such a deep learning architecture on the sensor board can reduce the amount of data transferred from sensors to the central ADAS computer implying great potential for an energy efficient deep learning architecture in such environments. Erkan Karakus, Mark Bruckner, Tao Wei 0001, Qing Yang 0001 |
NAS | 4 |
| 2022 | A Novel Interconnection Architecture for Secured Die-to-Die Communication in System-in-PackageabstractThe emerging System-in-Package (SiP) technology has enabled multiple dies fabricated on a single chip for high performance and energy efficiency. Die-to-die (D2D) communication in SiP is typically unencrypted, exposing sensitive data to possible microprobing attacks. This paper presents a new architecture design that protects D2D interconnection network from such microprobing attacks. Our new design provides instant detection of a microprobe at run time with no interference with normal data transfers and easily scalable to hundreds of D2D interconnection buses. The uniqueness of the new architecture is extremely simple and readily applicable to any system on a chip architecture. The trick is exploiting metastable state of a flip-flop (FF) to detect the small timing variation caused by the inevitable loading effect of a microprobe. To demonstrate its effectiveness and performance, a working prototype has been built on an emulation testbench in a Xilinx ZYNQ Ultrascale+ SoC to prove its feasibility. Dynamic partial reconfiguration (DPR) function is an enabler to create the test platform. DPR emulates both unprobed and probed scenarios on a D2D bus, and allow us to switch in between as required. To show our new design can be easily integrated to standard communication protocols, we implemented our prototype on a fully functional Advanced eXtensible Interface (AXI) bus. Experimental results show that the proposed runtime detection method is effective, resource-efficient, and reliable under temperature-varying environments. Zhenyu Xu 0007, Qing Yang 0001, Tao Wei 0001 |
NAS | 2 |
| 2021 | Runtime Detection of Probing/Tampering on Interconnecting BusesabstractIt has been reported that physical probing on an off-chip bus can reveal confidential information in an electronic system. An attacker can use non-invasive and inexpensive electric probes (or interposers) to measure signals from circuit traces, such as the memory bus between the memory controller and a memory module. This paper describes a method to detect any bus probing/tampering by tracking the phase shift of output digital waveforms, induced by input impedance change at the bus transmitter (Tx). A low-overhead digital circuit based on flip-flop's metastability is built around the Tx using a field-programmable logic gate array (FPGA) to precisely measure the phase shift of output signals. Uniquely, the output data launched by the Tx is used as a stimulus signal, thus, the proposed method holds the advantage of detecting probing attacks at run-time. That is, the detection action operates in parallel with the normal data transfer on a bus without any interference, imposing zero latency to the communication channel. In order to show its feasibility in a real-world communication protocol, we implemented the proposed method in the DDR memory controller on an FPGA board (Xilinx ZCU104). The working prototype is able to protect a memory bus between the FPGA board and a DDR4 DIMM with a data rate of 2400MT/s. Experimental results show that the proposed method can be used to countermeasure interposer attacks, probing attacks, and cold boot attacks. We believe that the proposed method can be implemented in a variety of communication channels. Zhenyu Xu 0007, Thomas Mauldin, Qing Yang 0001, Tao Wei 0001 |
FCCM | 3 |
| 2021 | A multidisciplinary approach to Internet of Things (IoT) cybersecurity and risk management
Kim-Kwang Raymond Choo, Keke Gai, Luca Chiaraviglio, Qing Yang 0001 |
Comput. Secur. | 4 |
| 2020 | A Bus Authentication and Anti-Probing Architecture Extending Hardware Trusted Computing Base Off CPU Chips and BeyondabstractTamper-proof hardware designs present a great challenge to computer architects. Most existing research limits hardware trusted computing base (TCB) to a CPU chip and anything off the CPU chip is vulnerable to probing and tampering. This paper introduces a new hardware design that provides strong defenses against physical attacks on interconnecting buses between chips in a computer system thereby extending the hardware TCB beyond CPU chips. The new approach is referred to as DIVOT: Detecting Impedance Variations Of Transmission-lines (Tx-lines). Every Tx-line in a computer system, such as a bus and interconnection wire has a unique, intrinsic, and fingerprint-like property: Impedance Inhomogeneity Pattern (IIP), i.e. the impedance distribution over distance. Such unpredictable, uncontrollable, and non-reproducible IIP fingerprints can be used to authenticate a Tx-line to ensure the confidentiality and integrity of data being transmitted. In addition, physical probes perturb the electromagnetic (EM) field around a Tx-line, leading to an altered IIP. As a result, runtime monitoring of IIPs can also be used to actively detect physical probing, snooping, and wire-tapping on buses. While the physics behind the IIP is known, the major technical breakthrough of DIVOT is the new integrated time domain reflectometer, iTDR, that is capable of carrying out in-situ and runtime monitoring of a Tx-line without interfering with normal data transfers. The iTDR is based on two innovations: analog-to-probability conversion (APC) and probability density modulation (PDM). The iTDR performs runtime IIP measurements noninvasively and is CMOS-compatible allowing it to be integrated with any interface logic connected to a bus. DIVOT is a generic, scalable, cost-effective, and low-overhead security solution for any computer system from servers to embedded computers in smart mobile devices and IoTs. To demonstrate the proposed architecture, a working prototype of DIVOT has been built on an FPGA as a proof of concept. Experimental results clearly showed the feasibility and performance of DIVOT for both hardware authentication and tamperproof applications. More specifically, the probability of correctly identifying a bus is close to 1 with an equal error rate (EER) of less than 0.06% at room temperature. We present an example design that incorporates DIVOT into an off-chip memory bus to protect against physical attacks including probing/snooping, tampering, and cold boot attacks. Zhenyu Xu 0007, Thomas Mauldin, Zheyi Yao, Shuyi Pei, Tao Wei 0001, Qing Yang 0001 |
ISCA | 6 |
| 2019 | WARCIP: write amplification reduction by clustering I/O pagesabstractThe storage volume of SSDs has been greatly increased recently with emerging multi-layer 3D triple-level cell and quad-level cell. However, one critical overhead of any flash memory SSD is the garbage collection (GC) process that is necessary due to the inherent physical property of flash memories. GC is a time consuming process that slows down I/O performance and decreases endurance of SSD. To minimize the negative impact of GC, we introduce Write Amplification Reduction by Clustering I/O Pages (WARCIP). The idea is to use a clustering algorithm to minimize the rewrite interval variance of pages in a flash block. As a result, pages in a flash block tend to have a similar lifetime, minimizing write amplification during a garbage collection. We have implemented WARCIP on an enterprise NVMe SSD. Both simulation and measurement experiments have been carried out. Real world I/O traces and standard I/O benchmarks are used in our experiments to assess the potential benefit of WARCIP. Experiment results show that WARCIP reduces write amplification dramatically and the number of block erasures by 4.45 times on average, implying extended lifetimes of flash SSDs. Shuyi Pei, Qing Yang 0001 |
SYSTOR | 3 |
| 2019 | REGISTOR: A Platform for Unstructured Data Processing Inside SSD StorageabstractThis article presents REGISTOR, a platform for r egular e xpression g rabbing i nside stor age. The main idea of Registor is accelerating regular expression (regex) search inside storage where large data set is stored, eliminating the I/O bottleneck problem. A special hardware engine for regex search is designed and augmented inside a flash SSD that processes data on-the-fly during data transmission from NAND flash to host. To make the speed of regex search match the internal bus speed of a modern SSD, a deep pipeline structure is designed in Registor hardware consisting of a file semantics extractor, matching candidates finder, regex matching units (REMUs), and results organizer. Furthermore, each stage of the pipeline makes the use of maximal parallelism possible. To make Registor readily usable by high-level applications, we have developed a set of APIs and libraries in Linux allowing Registor to process files in the SSD by recombining separate data blocks into files efficiently. A working prototype of Registor has been built in our newly designed NVMe-SSD. Extensive experiments and analyses have been carried out to show that Registor achieves high throughput, reduces the I/O bandwidth requirement by up to 97%, and reduces CPU utilization by as much as 82% for regex search in large datasets. Shuyi Pei, Qing Yang 0001 |
ACM Trans. Storage | 3 |
| 2019 | Editorial: IEEE Transactions on Sustainable Computing, Special Issue on Smart Data and Deep Learning in Sustainable ComputingabstractThe twelve papers in this special section focus on smart data and deep learning in sustainable computing. We are living in a data-driven era in which numerous infrastructure can be connected and the interconnected systems can perform “smart” when the large pool of the data are well utilized. Finding the way of well utilizing the large volume of data has an urgent demand in multiple realms, including academics, industries, and education. The force behind the data can be pushed out from a variety of data-driven techniques, such as machine learning and deep learning, which is a great potential for generating successful model, framework, and method for achieving sustainable computing. Therefore, gathering recent achievements in smart data and deep learning in sustainable computing is meaningful and valuable for powering the capability of datadriven domain and the various applications, implementations, and innovations in different disciplines and fields. This special issue focuses on two aspects considering the perspective of sustainable computing, which include smart data and deep learning. The smart data covers all dimensions of data usage lifecycles, such as data selections and collections, data preprocessing, data mining, and data analytics, in various application scenarios. The other aspect, deep learning, emphasizes the intelligent performance of applying data-driven techniques in practices and research explorations. Thus, this special issue aims at collecting updated outstanding papers that illustrate the latest achievements and development updates concerning the smart data and deep learning solutions, issues, applications, trends, and implementations in sustainable computing. Meikang Qiu, Sun-Yuan Kung, Qing Yang 0001 |
IEEE Trans. Sustain. Comput. | 3 |
| 2019 | Editorial: IEEE Transactions on Sustainable Computing, Special Issue on Secure Sustainable Green Smart ComputingabstractThe booming development of cloud computing has resulted in a remarkable growth of multiple industries in which dramatic demands of green computing and sustainability are addressed. The platform of smart computing has provided an efficient approach for connecting various infrastructure such that many new technologies are eventually formed, such as Internet-of-Thing and ubiquitous computing. The concept of sustainable green smart computing has become a significant issue for those enterprises or practitioners who are engaging the implementations of smart computing aiming at a longer-term strategy. Considering the achievement of the real sustainability, one of the crucial values is to ensure all operations across different computing sources are under a secure executive environment. For reaching a high performance of securing sustainable green smart computing, many problems need to be solved. For example, one of the main challenges is to balance the costs among security, energy, performance, and sustainable requirements. The distribution of the computing resources in this issue is a great challenge because the real-time executions are usually constrained by multiple elements. An efficient approach of providing an adaptive and scalable service as well as addressing sustainability is an urgent research direction for current advanced cloud computing applications. Thus, this special issue aims at collecting updated outstanding papers that illustrate the latest achievements and development updates concerning the security solutions, issues, applications, trends, and implementations in sustainable green smart computing. Meikang Qiu, Sun-Yuan Kung, Qing Yang 0001 |
IEEE Trans. Sustain. Comput. | 3 |
| 2018 | HODS: Hardware Object Deserialization Inside SSD StorageabstractThe rapid development of nonvolatile memory technologies such as flash, PCM, and Memristor has made processing in storage (PIS) a viable approach. We present an FPGA module augmented to an SSD storage controller that provides wire-speed object deserialization, referred to as HODS for hardware object deserialization in SSD. A pipelined circuit structure was designed to tailor to high-speed data conversion specifically. HODS is capable of conducting deserialization while data is being transferred on I/O bus from the storage device to host. The FPGA module has been integrated with our newly designed NVM-e SSD. The working prototype demonstrated significant performance benefits. The FPGA module can process data in line speed at 100MHz on 16 Byte data stream. For integer benchmarks, HODS showed deserialization speedup of 8~12× as compared to the traditional deserialization on a high-end host CPU. The speedup can reach 17~21× for floating-point datasets. The measured object deserialization throughput is 1GB/s on average at a clock speed of 100MHz. The overall performance improvements at the application level range from 10% to a factor of 4.3× depending on the proportion of deserialization time over total application running time. Compared to traditional SSD on the same server, HODS showed visible differences regarding application execution time while running Matlab, 3D modeling, and scientific computations. Fei Wu 0005, Yang Weng, Qing Yang 0001, Changsheng Xie 0001 |
FCCM | 4 |
| 2018 | CISC: Coordinating Intelligent SSD and CPU to Speedup Graph ProcessingabstractMinimum Spanning Tree (MST) is a fundamental problem in graph processing. The current state of the art concentrates on parallelizing its computation on multi-cores to speedup MST. Although many parallelism strategies have been explored, the actual speedup is limited, and they consume a large amount of CPU power. In this paper, we propose a new approach to the MST computation by coordinating computing power inside SSD storage with host CPU cores. A comprehensive framework of software-hardware co-design, referred to as CISC (coordinating Intelligent SSD and CPU), preprocesses MST graph edges inside storage and parallelizes the remaining computation on host CPU. Leveraging the special properties of modern SSD storage, CISC exploits a divide and conquer approach to reordering graph edges. We have implemented an FPGA circuit that reorders chunks of graph edges inside an SSD. The ordered chunks are then loaded to the system RAM and processed by the host CPU to build a B-Tree structure by repetitively picking up edges at heads of chunks. A working prototype CISC has been built using NVM-e SSD on a server. Extensive experiments have been carried out using real-world benchmarks to demonstrate the feasibility and performance of deploying CISC in NVM-e SSD storage. Our experimental results show 2.2~2.7× speedup for serial version implementation and 11.47× to 17.2× speedup for the parallel version with 96-cores. For the same number of cores, our parallel CISC outperforms the traditional software MST by up to 35%. Yafei Yang, Qing Yang 0001 |
ISPDC | 4 |
| 2018 | REGISTOR: A Platform for Unstructured Data Processing Inside SSD StorageabstractThis paper presents REGISTOR, a platform for regular expression grabbing inside storage. The main idea of Registor is accelerating regular expression (regex) search inside storage where large data set is stored, eliminating the I/O bottleneck problem. A special hardware engine for regex search is designed and augmented inside flash SSD that processes data on-the-fly during data transmission from NAND flash to host. In order to make the speed of regex search match the internal bus speed of modern SSD, a deep pipeline structure is designed in Registor hardware consisting of file semantics extractor, matching candidates finder, regex matching units (REMUs) and results organizer. Furthermore, each stage of the pipeline makes use of maximal parallelism possible. To make Registor readily usable by high level applications, we have developed a set of APIs and libraries in Linux allowing Registor to process files in SSD by recombining separate data blocks into files efficiently. A working prototype of Registor has been built in our newly designed NVMe-SSD. Extensive experiments and analyses have been carried out to show that Registor achieves high throughput, reduces I/O bandwidth requirement by up to 97% and CPU utilization by as much as 82% for regex search in large data sets. Shuyi Pei, Qing Yang 0001 |
SYSTOR | 3 |
| 2017 | DEFT-Cache: A Cost-Effective and Highly Reliable SSD Cache for RAID StorageabstractThis paper proposes a new SSD cache architecture, DEFT-cache, Delayed Erasing and Fast Taping, that maximizes I/O performance and reliability of RAID storage. First of all, DEFT-Cache exploits the inherent physical properties of flash memory SSD by making use of old data that have been overwritten but still in existence in SSD to minimize small write penalty of RAID5/6. As data pages being overwritten in SSD, old data pages are invalidated and become candidates for erasure and garbage collections. Our idea is to selectively delay the erasure of the pages and let these otherwise useless old data in SSD contribute to I/O performance for parity computations upon write I/Os. Secondly, DEFT-Cache provides inexpensive redundancy to the SSD cache by having one physical SSD and one virtual SSD as a mirror cache. The virtual SSD is implemented on HDD but using log-structured data layout, i.e. write data are quickly logged to HDD using sequential write. The dual and redundant caches provide a cost-effective and highly reliable write-back SSD cache. We have implemented DEFT-Cache on Linux system. Extensive experiments have been carried out to evaluate the potential benefits of our new techniques. Experimental results on SPC and Microsoft traces have shown that DEFT-Cache improves I/O performance by 26.81% to 56.26% in terms of average user response time. The virtual SSD mirror cache can absorb write I/Os as fast as physical SSD providing the same reliability as two physical SSD caches without noticeable performance loss. Jiguang Wan 0001, Qing Yang 0001, Xiaoyang Qu, Changsheng Xie 0001 |
IPDPS | 4 |
| 2017 | WCET-Aware Dynamic I-Cache Locking for a Single TaskabstractCaches are widely used in embedded systems to bridge the increasing speed gap between processors and off-chip memory. However, caches make it significantly harder to compute the worst-case execution time (WCET) of a task. To alleviate this problem, cache locking has been proposed. We investigate the WCET-aware I-cache locking problem and propose a novel dynamic I-cache locking heuristic approach for reducing the WCET of a task. For a nonnested loop, our approach aims at selecting a minimum set of memory blocks of the loop as locked cache contents by using the min-cut algorithm. For a loop nest, our approach not only aims at selecting a minimum set of memory blocks of the loop nest as locked cache contents but also finds a good loading point for each selected memory block. We propose two algorithms for finding a good loading point for each selected memory block, a polynomial-time heuristic algorithm and an integer linear programming (ILP)-based algorithm, further reducing the WCET of each loop nest. We have implemented our approach and compared it to two state-of-the-art I-cache locking approaches by using a set of benchmarks from the MRTC benchmark suite. The experimental results show that the polynomial-time heuristic algorithm for finding a good loading point for each selected memory block performs almost equally as well as the ILP-based algorithm. Compared to the partial locking approach proposed in Ding et al. [2012], our approach using the heuristic algorithm achieves the average improvements of 33%, 15%, 9%, 3%, 8%, and 11% for the 256B, 512B, 1KB, 4KB, 8KB, and 16KB caches, respectively. Compared to the dynamic locking approach proposed in Puaut [2006], it achieves the average improvements of 9%, 19%, 18%, 5%, 11%, and 16% for the 256B, 512B, 1KB, 4KB, 8KB, and 16KB caches, respectively. Wenguang Zheng, Hui Wu 0001, Qing Yang 0001 |
ACM Trans. Archit. Code Optim. | 3 |
| 2017 | Incorporating Intelligence in Fog Computing for Big Data Analysis in Smart CitiesabstractData intensive analysis is the major challenge in smart cities because of the ubiquitous deployment of various kinds of sensors. The natural characteristic of geodistribution requires a new computing paradigm to offer location-awareness and latency-sensitive monitoring and intelligent control. Fog Computing that extends the computing to the edge of network, fits this need. In this paper, we introduce a hierarchical distributed Fog Computing architecture to support the integration of massive number of infrastructure components and services in future smart cities. To secure future communities, it is necessary to integrate intelligence in our Fog Computing architecture, e.g., to perform data representation and feature extraction, to identify anomalous and hazardous events, and to offer optimal responses and controls. We analyze case studies using a smart pipeline monitoring system based on fiber optic sensors and sequential learning algorithms to detect events threatening pipeline safety. A working prototype was constructed to experimentally evaluate event detection performance of the recognition of 12 distinct events. These experimental results demonstrate the feasibility of the system's city-wide implementation in the future. Bo Tang 0011, Zhen Chen 0002, Gerald Hefferman, Shuyi Pei, Tao Wei 0001, Haibo He, Qing Yang 0001 |
IEEE Trans. Ind. Informatics | 7 |
| 2016 | An Efficient WCET-Aware Hybrid Global Branch Prediction ApproachabstractWe investigate the problem of reducing the number of branch mispredictions for a task such that its WCET (Worst-Case Execution Time) is minimized, and propose a novel branch correlation-based, hybrid branch prediction approach. Our approach consists of a static profile-based branch correlation analyzer and a dynamic branch predictor. The static profile-based branch correlation analyzer uses profiling and data dependency analysis to find precise correlations between branches and identifies all the branches that do not have any impact on the WCET of the task. The dynamic predictor uses the correlation information to make online predictions. We have implemented our approach and compared it with the two state-of-the-art branch prediction approaches by using a set of benchmark suite. The experimental results show that our approach outperforms the two state-of-the-art approaches. The maximum WCET improvement and the average WCET improvement of our approach over the WCET-aware static branch prediction approach are 41.35% and 11.40%, respectively. The maximum WCET improvement and the average WCET improvement of our approach over the global dynamic branch prediction approach are 32.41% and 7.67%, respectively. Xuesong Su, Hui Wu 0001, Qing Yang 0001 |
RTCSA | 3 |
| 2016 | Fit: A Fog Computing Device for Speech Tele-TreatmentsabstractThere is an increasing demand for smart fog-computing gateways as the size of cloud data is growing. This paper presents a Fog computing interface (FIT) for processing clinical speech data. FIT builds upon our previous work on EchoWear, a wearable technology that validated the use of smartwatches for collecting clinical speech data from patients with Parkinson's disease (PD). The fog interface is a low-power embedded system that acts as a smart interface between the smartwatch and the cloud. It collects, stores, and processes the speech data before sending speech features to secure cloud storage. We developed and validated a working prototype of FIT that enabled remote processing of clinical speech data to get speech clinical features such as loudness, short-time energy, zero-crossing rate, and spectral centroid. We used speech data from six patients with PD in their homes for validating FIT. Our results showed the efficacy of FIT as a Fog interface to translate the clinical speech processing chain (CLIP) from a cloud-based backend to a fog-based smart gateway. Admir Monteiro, Harishchandra Dubey, Leslie Mahler, Qing Yang 0001, Kunal Mankodiya |
SMARTCOMP | 4 |
| 2015 | F/M-CIP: Implementing Flash Memory Cache Using Conservative Insertion and PromotionabstractFlash memory SSD has emerged as a promising storage media and fits naturally as a cache between the system RAM and the disk due to its performance/cost characteristics. Managing such an SSD cache is challenging and traditional cache replacements do not work well because of SSDs asymmetric read/write performances and wearing issues. This paper presents a new cache replacement algorithm referred to as F/M-CIP that accelerates disk I/O greatly. The idea is dividing the traditional LRU list into 4 parts: candidate-list, SSD-list, RAM-list and eviction-buffer-list. Upon a cache miss, the metadata of the missed block is conservatively inserted into the candidate-list but the data itself is not cached. The block in the candidate-list is then conservatively promoted to the RAM-list upon the k-th miss. At the bottom of the RAM-list, the eviction-buffer accumulates LRU blocks to be written into the SSD cache in batches to exploit the internal parallelism of SSD. The SSD-list is managed using a combination of regency and frequency replacement policies by means of conservative promotion upon hits. To quantitatively evaluate the performance of F/M-CIP, a prototype has been built on Linux kernel at the generic block layer. Experimental results on standard benchmarks and real world traces have shown that F/M-CIP accelerates disk I/O performance up to an order of magnitude compared to the traditional hard disk storage and up to a factor of 3 compared to the traditional SSD cache algorithm in terms of application execution time. Furthermore, F/M-CIP substantially reduces write operations to the SSD implying prolonged durability. Qing Yang 0001 |
CCGRID | 2 |
| 2015 | A neural machine interface architecture for real-time artificial lower limb control
Jason Kane, Qing Yang 0001, Robert Hernandez, Willard Simoneau, Matthew Seaton |
DATE | 2 |
| 2015 | A Reconfigurable Multiclass Support Vector Machine Architecture for Real-Time Embedded Systems ClassificationabstractA great architectural challenge facing many of today's embedded systems is in combining physical sensory inputs with often power constrained computational elements to timely and accurately achieve an understanding of the current operating environment. Classification processing is one manner in which some designs accomplish this task. Support Vector Machines (SVMs) encompass one field of classification that has proven to yield high accuracy and has recently seen increased widespread use, however, the required computational intensity places a challenge on computer architects to design a hardware structure capable of performing real-time classifications while maintaining low power consumption. This paper proposes the first ever fully pipelined, floating point based, multi-use reconfigurable hardware architecture designed to act in conjunction with embedded processing as an accelerator for multiclass SVM classification. Several tasks, involving an extensive sample set of over 100,000 classifications, are evaluated for speed and accuracy against the same suite running on both an Intel Core i7-2600 system with lib SVM and an Nvidia GT 750M graphics card GPU. Our promising results show that our architecture is able to achieve a measure able speed increase in performance of up to 53x while matching 99.95% of all lib SVM predictions. Jason Kane, Robert Hernandez, Qing Yang 0001 |
FCCM | 3 |
| 2015 | A Parallel and Pipelined Architecture for Accelerating Fingerprint Computation in High Throughput Data StoragesabstractRabin fingerprints are short tags for large objects that can be used in a wide range of applications, such as data deduplication, web querying, packet routing, and caching. We present a pipelined hardware architecture for computing Rabin fingerprints on data being transferred on a high throughput bus. The design conducts real-time fingerprinting with short latencies, and can be tuned for optimized clock rate with "split fresh" technique. A pipelined sampling logic selects fingerprints based on the Minwise theory and adds only a few clock cycles of latency before returning the final results. The design can be replicated to work in parallel for higher throughput data traffic. This architecture is implemented on a Xilinx Virtex-6 FPGA, and is tested on a storage prototyping platform. The implementation shows that the design can achieve clock rates above 300 MHz with an order of magnitude improvement in latency over prior software implementations, while consuming little hardware resource. The scheme is extensible to other types of fingerprints and CRC computations, and is readily applicable to primary storages and caches in hybrid storage systems. Qing Yang 0001, Qingbo Wang, Cyril Guyot, Ashwin Narasimha, Dejan Vucinic, Zvonimir Bandic |
FCCM | 2 |
| 2015 | Reflex-Tree: A Biologically Inspired Parallel Architecture for Future Smart CitiesabstractWe introduce a new parallel computing and communication architecture, Reflex-Tree, with massive sensing, data processing, and control functions suitable for future smart cities. The central feature of the proposed Reflex-Tree architecture is inspired by a fundamental element of the human nervous system: reflex arcs, the neuromuscular reactions and instinctive motions of a part of the body in response to urgent situations. At the bottom level of the Reflex-Tree (layer 4), novel sensing devices are proposed that are controlled by low power processing elements. These "leaf" nodes are then connected to new classification engines based on machine learning techniques, including support vector machines (SVM), to form the third layer. The next layer up consists of servers that provide accurate control decisions via multi-layer adaptive learning and spatial-temporal association, before they are connected to the top level cloud where complex system behavior analysis is performed. Our multi-layered architecture mimics human neural circuits to achieve the high levels of parallelization and scalability required for efficient city-wide monitoring and feedback. To demonstrate the utility of our architecture, we present the design, implementation, and experimental evaluation of a prototype Reflex-Tree. City power supply network and gas pipeline management scenarios are used to drive our prototype as case studies. We show the effectiveness for several levels of the architecture and discuss the feasibility of implementation. Jason Kane, Bo Tang 0011, Zhen Chen 0002, Jun Yan 0007, Tao Wei 0001, Haibo He, Qing Yang 0001 |
ICPP | 7 |
| 2015 | Hardware accelerator for similarity based data dedupeabstractData deduplication has proven important in backup storage systems as large amount of identical or similar data chunks exist. Recent studies have shown the great potential of data deduplication in primary storage and storage caches. Deduplications in these environments require high speed processing not to drag down production performance. This paper presents a hardware accelerator for similarity based data deduplication. It implements three compute-intensive kernel modules to improve throughput and latency in dedupe systems: sketch computation for data blocks, index searching for reference block, and delta encoding over similar blocks. Adopting pipelined computation and parallel data lookup across multiple hardware modules, our HW design is capable of processing high throughput data traffic by working on multiple data units concurrently, thus enabling wire speed dedupe for data stream where similar blocks present. Using a PC host system connected to the FPGA-based accelerator through a PCIe Gen 2×4 interface, our experiments show that the similarity based data dedupe performs 30% better in data reduction ratio than conventional dedupe techniques that look at identical blocks only. By comparing the hardware implementation with its software counterpart, the experimental results show that our preliminary FPGA implementation with maximum clock speed of 250MHz achieves at least 6 times improvement in latency over the software implementation running on state-of-art servers. Qingbo Wang, Cyril Guyot, Ashwin Narasimha, Dejan Vucinic, Zvonimir Bandic, Qing Yang 0001 |
NAS | 7 |
| 2014 | ${\rm S}^{2}$-RAID: Parallel RAID Architecture for Fast Data RecoveryabstractAs disk volume grows rapidly with terabyte disk becoming a norm, RAID reconstruction process in case of a failure takes prohibitively long time. This paper presents a new RAID architecture, S2-RAID, allowing the disk array to reconstruct very quickly in case of a disk failure. The idea is to form skewed sub-arrays in the RAID structure so that reconstruction can be done in parallel dramatically speeding up data reconstruction process and hence minimizing the chance of data loss. We analyse the data recovery ability of this architecture and show its good scalability. A prototype S2-RAID system has been built and implemented in the Linux operating system for the purpose of evaluating its performance potential. Real world I/O traces including SPC, Microsoft, and a collection of a production environment have been used to measure the performance of S2-RAID as compared to existing baseline software RAID5, Parity Declustering, and RAID50. Experimental results show that our new S2-RAID speeds up data reconstruction time by a factor 2 to 4 compared to the traditional RAID. Meanwhile, S2-RAID keeps comparable production performance to that of the baseline RAID layouts while online RAID reconstruction is in progress. Jiguang Wan 0001, Changsheng Xie 0001, Qing Yang 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2013 | A New Metadata Update Method for Fast Recovery of SSD CacheabstractIn order to maintain data in an SSD (solid-state disk) cache durable after a crash or reboot, metadata information needs to be stored persistently in SSD. There are two typical metadata methods, update-write-update and write-update. While write-update method has one less SSD write operation than update-write-update for each write I/O, it limits the amount of cached data that can be used after a system crash. We present a design and implementation of a novel metadata update method for SSD cache, referred to as Lazy-Update Following an Update-Write (LUFUW). Our new metadata update method allows maximal amount of data in SSD cache available upon restart after a power failure or system crash with minimal additional writes to SSD. This capability makes restart run twice as fast as existing SSD caches such as Flash cache [1] that can only use dirty data in the cache after crash recovery. We present our prototype implementation on Linux kernel and performance measurements as compared with existing SSD cache solutions. Qing Yang 0001 |
NAS | 2 |
| 2012 | Implementing an FPGA system for real-time intent recognition for prosthetic legsabstractThis paper presents the design and implementation of a cyber physical system (CPS) for neural-machine interface (NMI) that continuously senses signals from a human neuromuscular control system and recognizes the user's intended locomotion modes in real-time. The CPS contains two major parts: a microcontroller unit (MCU) for sensing and buffering input signals and an FPGA device as the computing engine for fast decoding and recognition of neural signals. The real-time experiments on a human subject demonstrated its real-time, self-contained, and high accuracy in identifying three major lower limb movement tasks (level-ground walking, stair ascent, and standing), paving the way for truly neural-controlled prosthetic legs. He Huang 0002, Qing Yang 0001 |
DAC | 3 |
| 2012 | Enhancing shared RAID performance through online profilingabstractEnterprise storage systems are generally shared by multiple servers in a SAN environment. Our experiments as well as industry reports have shown that disk arrays show poor performance when multiple servers share one RAID due to resource contention as well as frequent disk head movements. We have studied IO performance characteristics of several shared storage settings of practical business operations. To avoid the IO contention, we propose a new dynamic data relocation technique on shared RAID storages, referred to as DROP, Dynamic data Relocation to Optimize Performance. DROP allocates/manages a group of cache data areas and relocates/drops the portion of hot data at a predefined sub array that is a physical partition on the top of the entire shared array. By analyzing profiling data to make each cache area owned by one server, we are able to determine optimal data relocation and partition of disks in the RAID to maximize large sequential block accesses on individual disks and at the same time maximize parallel accesses across disks in the array. As a result, DROP minimizes disk head movements in the array at run time giving rise to high IO performance. A prototype DROP has been implemented as a software module at the storage target controller. Extensive experiments have been carried out using real world IO workloads to evaluate the performance of the DROP implementation. Experimental results have shown that DROP improves shared IO performance greatly. The performance improvements in terms of average IO response time range from 20% to a factor 2.5 at no additional hardware cost. Jiguang Wan 0001, Yan Liu 0010, Qing Yang 0001, Jianzong Wang |
MSST | 4 |
| 2012 | Compression Speed Enhancements to LZO for Multi-core SystemsabstractThis paper examines several promising throughput enhancements to the Lempel-Ziv-Oberhumer (LZO) 1x-1-15 data compression algorithm. Of many algorithm variants present in the current library version, 2.06, LZO 1x-1-15 is considered to be the fastest, geared toward speed rather than compression ratio. We present several algorithm modifications tailored to modern multi-core architectures in this paper that are intended to increase compression speed while minimizing any loss in compression ratio. On average, the experimental results show that on a modern quad core system, a 3.9x speedup in compression time is achieved over the baseline algorithm with no loss to compression ratio. Allowing for a 25% loss in compression ratio, up to a 5.4x speedup in compression time was observed. Jason Kane, Qing Yang 0001 |
SBAC-PAD | 2 |
| 2012 | ST-CDP: Snapshots in TRAP for Continuous Data ProtectionabstractContinuous Data Protection (CDP) has become increasingly important as digitization continues. This paper presents a new architecture and an implementation of CDP in Linux kernel. The new architecture takes advantages of both traditional snapshot technology and recent Timely Recovery to Any Point-in-time (TRAP) architecture [CHECK END OF SENTENCE]. The idea is to periodically insert snapshots within the parity logs of changed data blocks in order to ensure fast and reliable data recovery in case of failures. A mathematical model is developed as a guide to designers to determine when and how to insert snapshots to optimize performance in terms of space usage and recovery time. Based on the mathematical model, we have designed and implemented a CDP module in the Linux system. Our implementation is at block level as a device driver that is capable of recovering data to any point-in-time in case of various failures. Extensive experiments have been carried out to show that the implementation is fairly robust and numerical results demonstrate that the implementation is efficient. Qiang Cao 0001, Changsheng Xie 0001, Qing Yang 0001 |
IEEE Trans. Computers | 5 |
| 2012 | On Design and Implementation of Neural-Machine Interface for Artificial LegsabstractThe quality of life of leg amputees can be improved dramatically by using a cyber physical system (CPS) that controls artificial legs based on neural signals representing amputees' intended movements. The key to the CPS is the neural-machine interface (NMI) that senses electromyographic (EMG) signals to make control decisions. This paper presents a design and implementation of a novel NMI using an embedded computer system to collect neural signals from a physical system - a leg amputee, provide adequate computational capability to interpret such signals, and make decisions to identify user's intent for prostheses control in real time. A new deciphering algorithm, composed of an EMG pattern classifier and a post-processing scheme, was developed to identify the user's intended lower limb movements. To deal with environmental uncertainty, a trust management mechanism was designed to handle unexpected sensor failures and signal disturbances. Integrating the neural deciphering algorithm with the trust management mechanism resulted in a highly accurate and reliable software system for neural control of artificial legs. The software was then embedded in a newly designed hardware platform based on an embedded microcontroller and a graphic processing unit (GPU) to form a complete NMI for real time testing. Real time experiments on a leg amputee subject and an able-bodied subject have been carried out to test the control accuracy of the new NMI. Our extensive experiments have shown promising results on both subjects, paving the way for clinical feasibility of neural controlled artificial legs. Yuhong Liu 0003, Fan Zhang 0015, Yan Lindsay Sun, Qing Yang 0001, He Huang 0002 |
IEEE Trans. Ind. Informatics | 6 |
| 2011 | I-CASH: Intelligently Coupled Array of SSD and HDDabstractThis paper presents a new disk I/O architecture composed of an array of a flash memory SSD (solid state disk) and a hard disk drive (HDD) that are intelligently coupled by a special algorithm. We call this architecture I-CASH: Intelligently Coupled Array of SSD and HDD. The SSD stores seldom-changed and mostly read reference data blocks whereas the HDD stores a log of deltas between currently accessed I/O blocks and their corresponding reference blocks in the SSD so that random writes are not performed in SSD during online I/O operations. High speed delta compression and similarity detection algorithms are developed to control the pair of SSD and HDD. The idea is to exploit the fast read performance of SSDs and the high speed computation of modern multi-core CPUs to replace and substitute, to a great extent, the mechanical operations of HDDs. At the same time, we avoid runtime SSD writes that are slow and wearing. An experimental prototype I-CASH has been implemented and is used to evaluate I-CASH performance as compared to existing SSD/HDD I/O architectures. Numerical results on standard benchmarks show that I-CASH reduces the average I/O response time by an order of magnitude compared to existing disk I/O architectures such as RAID and SSD/HDD storage hierarchy, and provides up to 2.8 speedup over state-of-the-art pure SSD storage. Furthermore, I-CASH reduces random writes to SSD implying reduced wearing and prolonged life time of the SSD. Qing Yang 0001 |
HPCA | 1 |
| 2010 | Design and implementation of a special purpose embedded system for neural machine interfaceabstractOur previous study has shown the potential of using a computer system to accurately decode electromyographic (EMG) signals for neural controlled artificial legs. Because of computation complexity of the training algorithm coupled with real time requirement of controlling artificial legs, traditional embedded systems generally cannot be directly applied to the system. This paper presents a new design of an FPGA-based neural-machine interface for artificial legs. Both the training algorithm and the real time controlling algorithm are implemented on an FPGA. A soft processor built on the FPGA is used to manage hardware components and direct data flows. The implementation and evaluation of this design are based on Altera Stratix II GX EP2SGX90 FPGA device on a PCI Express development board. Our performance evaluations indicate that a speedup of around 280X can be achieved over our previous software implementation with no sacrifice of computation accuracy. The results demonstrate the feasibility of a self-contained, low power, and high performance real-time neural-machine interface for artificial legs. He Huang 0002, Qing Yang 0001 |
ICCD | 3 |
| 2010 | A New Buffer Cache Design Exploiting Both Temporal and Content LocalitiesabstractThis paper presents a Least Popularly Used buffer cache algorithm to exploit both temporal locality and content locality of I/O requests. Popular data blocks are selected as reference blocks that are not only accessed frequently but also identical or similar in content to other blocks that are being accessed. Fast delta compression and decompression are used to satisfy as many I/O requests as possible using the popular reference blocks together with small deltas inside the buffer cache. The popularity of a reference block is calculated based on the statistical analysis of data contents and access frequency. A prototype LPU has been implemented as a new cache layer for Kernel Virtual Machine (KVM) on Linux system. Experimental results show LPU is effective for a variety of workloads with the maximum speed up of over 300% compared with LRU. Qing Yang 0001 |
ICDCS | 2 |
| 2010 | S2-RAID: A new RAID architecture for fast data recoveryabstractAs disk volume grows rapidly with terabyte disk becoming a norm, RAID reconstruction time in case of a failure takes prohibitively long time. This paper presents a new RAID architecture, S2-RAID, allowing the disk array to reconstruct very quickly in case of a disk failure. The idea is to form skewed sub RAIDs (S2-RAID) in the RAID structure so that reconstruction can be done in parallel dramatically speeding up data reconstruction time and hence minimizing the chance of data loss. To make such parallel reconstruction conflict-free, each sub-RAID is formed by selecting one logic partition from each disk group with size being a prime number. We have implemented a prototype S2-RAID system in Linux operating system for the purpose of evaluating its performance potential. SPC IO traces and standard benchmarks have been used to measure the performance of S2-RAID as compared to existing baseline software RAID, MD. Experimental results show that our new S2-RAID speeds up data reconstruction time by a factor of 3 to 6 compared to the traditional RAID. At the same time, S2-RAID shows similar or better production performance than baseline RAID while online RAID reconstruction is in progress. Jiguang Wan 0001, Qing Yang 0001, Changsheng Xie 0001 |
MSST | 3 |
| 2009 | Design and Analysis of Block-Level Snapshots for Data Protection and RecoveryabstractThis paper presents a comprehensive study on implementations and performance evaluations of two snapshot techniques: copy-on-write snapshot and redirect-on-write snapshot. We develop a simple Markov process model to analyze data block behavior and its impact on application performance, while the snapshot operation is underway at the block-level storage. We have implemented the two snapshots techniques on both Windows and Linux operating systems. Based on our analytical model and our implementation, we carry out quantitative performance evaluations and comparisons of the two snapshot techniques using IoMeter, PostMark, TPC-C, and TPC-W benchmarks. Our measurements reveal many interesting observations regarding the performance characteristics of the two snapshot techniques. Depending on the applications and different I/O workloads, the two snapshot techniques perform quite differently. In general, copy-on-write performs well on read-intensive applications, while redirect-on-write performs well on writeintensive applications. Weijun Xiao, Qing Yang 0001, Changsheng Xie 0001, Huaiyang Li |
IEEE Trans. Computers | 2 |
| 2009 | Securing rating aggregation systems using statistical detectors and trustabstractOnline feedback-based rating systems are gaining popularity. Dealing with unfair ratings in such systems has been recognized as an important but difficult problem. This problem is challenging especially when the number of regular ratings is relatively small and unfair ratings can contribute to a significant portion of the overall ratings. Furthermore, the lack of unfair rating data from real human users is another obstacle toward realistic evaluation of defense mechanisms. In this paper, we propose a set of statistical methods to jointly detect collaborative unfair ratings in product-rating type online rating systems. Based on detection, a framework of trust-assisted rating aggregation system is developed. Furthermore, we collect unfair rating data from real human users through a rating challenge. The proposed system is evaluated through simulations as well as experiments using real attack data. Compared with existing schemes, the proposed system can significantly reduce negative impact from unfair ratings. Yafei Yang, Yan Lindsay Sun, Steven M. Kay, Qing Yang 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2009 | A Case for Continuous Data Protection at Block Level in Disk Array StoragesabstractAbstract — This paper presents a study of data storages for continuous data protection (CDP). After analyzing the existing data protection technologies, we propose a new disk array architecture that provides Timely Recovery to Any Point-in-time, referred to as TRAP-Array. TRAP-Array stores not only the data stripe upon a write to the array, but also the time-stamped Exclusive-ORs of successive writes to each data block. By leveraging the Exclusive-OR operations that are performed upon each block write in today’s RAID4/5 controllers, TRAP does not incur noticeable performance overhead. More importantly, TRAP is able to recover data very quickly to any point-in-time upon data damage by tracing back the sequence and history of Exclusive-ORs resulting from writes. What is interesting is that TRAP architecture is very space-efficient. We have implemented a prototype TRAP architecture using software at block level and carried out extensive performance measurements using TPC-C benchmarks running on Oracle and Postgress databases, TPC-W running on MySQL database, and file system benchmarks running on Linux and Windows systems. Our experiments demonstrated that TRAP is not only able to recover data to any pointin-time very quickly upon a failure but it is also space efficient. Compared to the state-of-the-art continuous data protection technologies, TRAP saves disk storage space by one to two orders of magnitude with a simple and a fast encoding algorithm. In addition, TRAP can provide two-way data recovery with the availability of only one reference image in contrast to the one-way recovery of snapshot and incremental backup technologies. Weijun Xiao, Qing Yang 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2008 | Can We Really Recover Data if Storage Subsystem Fails?abstractThis paper presents a theoretical and experimental study on the limitations of copy-on-write snapshots and incremental backups in terms of data recoverability. We provide mathematical proofs of our new findings as well as implementation experiments to show how data recovery is done in case of various failures. Based on our study, we propose a new system architecture that will overcome the problems of existing technologies. The new architecture can provide two-way data recovery capability with the same storage overheads and can be implemented fairly easily on existing systems. We show that the new architecture has maximum data recoverability and is practically feasible. Weijun Xiao, Qing Yang 0001 |
ICDCS | 2 |
| 2006 | On Performance of Parallel iSCSI Protocol for Networked Storage SystemsabstractA newly emerging protocol for storage networking, iSCSI [1,2], was recently ratified by the Internet Engineering Task Force [3]. The iSCSI protocol is perceived as a low cost alternative to the FC protocol for networked storages [5,6,7,8]. It allows block level storage data to be transported over the popular TCP/IP network that is widely understood. This paper presents a new storage architecture allowing parallel processing of iSCSI packets. By leveraging inexpensive Ethernet ports, we are able to greatly improve the throughput of iSCSI storages through parallel processing. We have carried out a performance analysis to evaluate the performance of the new architecture as compared to the existing iSCSI storages. Our numerical results show that the new architecture has good performance potential. Qing Yang 0001 |
AINA (1) | 1 |
| 2006 | PRINS: Optimizing Performance of Reliable Internet StoragesabstractDistributed storage systems employ replicas or erasure code to ensure high reliability and availability of data. Such replicas create great amount of network traffic that negatively impacts storage performance, particularly for distributed storage systems that are geographically dispersed over a wide area network (WAN). This paper presents a performance study of our new data replication methodology that minimizes network traffic for data replications. The idea is to replicate the parity of a data block upon each write operation instead of the data block itself. The data block will be recomputed back at the replica storage site upon receiving the parity. We name the new methodology PRINS (Parity Replication in IP-Network Storages). PRINS trades off highspeed computation for communication that is costly and more likely to be the performance bottleneck for distributed storages. By leveraging the parity computation that exists in common storage systems (RAID), our PRINS does not introduce additional overhead but dramatically reduces network traffic. We have implemented PRINS using iSCSI protocol over a TCP/IP network interconnecting a cluster of PCs as storage nodes. We carried out performance measurements on Oracle database, Postgres database, MySQL database, and Ext2 file system using TPC-C, TPC-W, and Micro benchmarks. Performance measurements show up to 2 orders of magnitudes bandwidth savings of PRINS compared to traditional replicas. A queueing network model is developed to further study network performance for large networks. It is shown that PRINS reduces response time of the distributed storage systems dramatically. Qing Yang 0001, Weijun Xiao |
ICDCS | 1 |
| 2006 | TRAP-Array: A Disk Array Architecture Providing Timely Recovery to Any Point-in-timeabstractRAID architectures have been used for more than two decades to recover data upon disk failures. Disk failure is just one of the many causes of damaged data. Data can be damaged by virus attacks, user errors, defective software/firmware, hardware faults, and site failures. The risk of these types of data damage is far greater than disk failure with today's mature disk technology and networked information services. It has therefore become increasingly important for today's disk array to be able to recover data to any point in time when such a failure occurs. This paper presents a new disk array architecture that provides timely recovery to any point-in-time, referred to as TRAP-array. TRAP-array stores not only the data stripe upon a write to the array, but also the time-stamped exclusive-ORs of successive writes to each data block. By leveraging the exclusive-OR operations that are performed upon each block write in today's RAID4/5 controllers, TRAP does not incur noticeable performance overhead. More importantly, TRAP is able to recover data very quickly to any point-in-time upon data damage by tracing back the sequence and history of exclusive-ORs resulting from writes. What is interesting is that TRAP architecture is amazingly space-efficient. We have implemented a prototype TRAP architecture using software at block device level and carried out extensive performance measurements using TPC-C benchmark running on Oracle and Postgress databases, TPC-W running on MySQL database, and file system benchmarks running on Linux and Windows systems. Our experiments demonstrated that TRAP is not only able to recover data to any point-in-time very quickly upon a failure but it also uses less storage space than traditional daily differential backup/snapshot. Compared to the state-of-the-art continuous data protection technologies, TRAP saves disk storage space by one to two orders of magnitude with a simple and a fast encoding algorithm. From an architecture point of view, TRAP-array opens up another dimension for storage arrays. It is orthogonal and complementary to RAID in the sense that RAID protects data in the dimension along an array of physical disks while TRAP protects data in the dimension along the time sequence Qing Yang 0001, Weijun Xiao |
ISCA | 1 |
| 2005 | SPEK: A Storage Performance Evaluation Kernel Module for Block-Level Storage Systems under Faulty ConditionsabstractThis paper introduces a new benchmark tool, SPEK (storage performance evaluation kernel module), for evaluating the performance of block-level storage systems in the presence of faults as well as under normal operations. SPEK can work on both direct attached storage (DAS) and block level networked storage systems such as storage area networks (SAN). Each SPEK consists of a controller, several workers, one or more probers, and several fault injection modules. Since it runs at kernel level and eliminates skews and overheads caused by file systems, SPEK is highly accurate and efficient. It allows a storage architect to generate configurable workloads to a system under test and to inject different faults into various system components such as network devices, storage devices, and controllers. Available performance measurements under different workloads and faulty conditions are dynamically collected and recorded in SPEK over a spectrum of time. To demonstrate its functionality, we apply SPEK to evaluate the performance of two direct attached storage systems and two typical SANs under Linux with different fault injections. Our experiments show that SPEK is highly efficient and accurate to measure performance for block-level storage systems. Xubin He, Ming Zhang 0026, Qing Yang 0001 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2004 | BUCS - A Bottom-Up Cache Structure for Networked Storage ServersabstractThis paper introduces a new caching structure to improve server performance by minimizing data traffic over the system bus. The idea is to form a bottom-up caching hierarchy in a networked storage server. The bottom level cache is located on an embedded controller that is a combination of a network interface card (NIC) and a storage host bus adapter (HBA). Storage data coming from or going to a network are cached at this bottom level cache and meta-data related to these data are passed to the host for processing. When cached data exceed the capacity of the bottom level cache, some data are moved to the host RAM that is usually larger than the bottom level cache. This new cache hierarchy is referred to as bottom-up cache structure (BUGS) in contrast to a traditional CPU-centric top-down cache where the top-level cache is the smallest and fastest, and the lower in the hierarchy the larger and slower the cache. Such data caching at the controller level dramatically reduces bus traffic and leads to great performance improvement for networked storages. We have implemented a proof-of-concept prototype using Intel's IQ80310 reference board and Linux network block device. Through performance measurements on the prototype implementation, we observed up to 3 times performance improvement of BUCS over traditional systems in terms of response time and system throughput. Ming Zhang 0026, Qing Yang 0001 |
ICPP | 2 |
| 2004 | Cost-Effective Remote Mirroring Using the iSCSI Protocol
Ming Zhang 0026, Qing Yang 0001 |
MSST | 3 |
| 2004 | STICS: SCSI-to-IP cache for storage area networks
Xubin He, Ming Zhang 0026, Qing Yang 0001 |
J. Parallel Distributed Comput. | 3 |
| 2003 | Performability Evaluation of Networked Storage Systems Using N-SPEKabstractThis paper introduces a new benchmark tool for evaluating performance and availability (performability) of networked storage systems, specifically storage area network (SAN) that is intended for providing block-level data storage with high performance and availability. The new benchmark tool, named N-SPEK (Networked-Storage Performability Evaluation Kernel module), consists of a controller, several workers, one or more probers, and several fault injection modules. N-SPEK is highly accurate and efficient since it runs at kernel level and eliminates skews and overheads caused by file systems. It allows a SAN architect to generate configurable storage workloads to the SAN under test and to inject different faults into various SAN components such as network devices, storage devices, and controllers. Available performances under different workloads and failure conditions are dynamically collected and recorded in the N-SPEK over a spectrum of time. To demonstrate its functionality, we apply N-SPEK to evaluate the performability of a specific iSCSI-based SAN under Linux environment. Our experiments show that N-SPEK not only efficiently generates quantitative performability results but also reveals a few optimization opportunities for future iSCSI implementations. Ming Zhang 0026, Qing Yang 0001, Xubin He |
CCGRID | 2 |
| 2003 | A unified, low-overhead framework to support continuous profiling and optimizationabstractWe propose a unified, low-overhead framework (ULF) to support continuous system profiling and optimization based on a specifically designed embedded board. Instead of building a new profiling tool from scratch, ULF provides a unified interface to integrate various existing profiling tools and optimizers, and helps to build future tools easily. ULF uses an embedded processor to off-load tasks of post-processing profiling data, which reduces system overhead caused by profiling tools and makes ULF especially suitable for continuous profiling on production systems. By processing the profiling data in parallel and providing feedback promptly, ULF supports on-line optimization. Our case study on I/O profiling demonstrated that ULF-enhanced profiling tool dramatically reduces overhead, making continuous profiling on production systems feasible. Ming Zhang 0026, Xubin He, Qing Yang 0001 |
IPCCC | 3 |
| 2003 | RORIB: An Economic and Efficient Solution for Real-Time Online Remote Info BackupabstractData plays an essential role in business today. Most, if not all, E-business applications are database driven, and data backup is a necessary element of managing data. Backup and recovery techniques have always been critical to any database, and as real-time databases are used more often, real-time online backup strategies become critical to optimize performance. In this paper, current backup methods are discussed and evaluated for response time and cost. A prototype device driver, RORIB (Real-time Online Remote Information Backup) is presented and discussed. An experiment is conducted comparing the performance, in terms of response time, of the prototype and several current backup strategies. RORIB provides an economic and efficient solution for real-time online remote backup. Significant improvement in response time is demonstrated using this prototype device driver when compared to other types of software-driven backup protocols. Another advantage of RORIB is that the cost is negligible when compared to other hardware solutions for backup, such as Storage Area Networks (SANs) and Private Backup Networks (PBNs). Additionally, this multi-layered device-driver uses TCP/IP (Telecommunications Protocol/Internet Protocol) which allows the driver to be a “drop in” filter between existing hardware layers and thus reduces the implementation overhead and improves portability. Linux is used as the operating system in this experiment because of its open source nature and its similarity to UNIX. This also increases the portability of this approach. The driver is transparent to both the user and the database management system. Other potential applications and future research directions for this technology are presented. Scott J. Lloyd, Joan Peckham, Jian Li 0059, Qing Yang 0001 |
J. Database Manag. | 4 |
| 2002 | Introducing SCSI-to-IP Cache for Storage Area NetworksabstractData storage plays an essential role in today's fast-growing data-intensive network services. iSCSI is one of the most recent standards that allow SCSI protocols to be carried out over IP networks. However, the disparities between SCSI and IP prevent fast and efficient deployment of SAN (storage area network) over IP. This paper introduces STICS (SCSI-To-IP cache storage), a novel storage architecture that couples reliable and high-speed data caching with low-overhead conversion between SCSI and IP protocols. Through the efficient caching algorithm and localization of certain unnecessary protocol overheads, STICS significantly improves performance over current iSCSI system. Furthermore, STICS can be used as a basic plug-and-play building block for data storage over IP. We have implemented software STICS prototype on Linux operating system. Numerical results using popular PostMark benchmark program and EMC's trace have shown dramatic performance gain over the current iSCSI implementation. Xubin He, Qing Yang 0001, Ming Zhang 0026 |
ICPP | 2 |
| 2002 | A Caching Strategy to Improve iSCSI PerformanceabstractiSCSI is one of the most recent standards that allows SCSI protocols to be carried out over IP networks. However, to encapsulate the SCSI protocol over IP requires a significant amount of overhead traffic for SCSI commands transfers and handshaking over the Internet. In this paper, we propose a caching scheme, called iCache, to improve the iSCSI performance. iCache uses a log disk along with a piece of non-volatile RAM to cache the iSCSI traffic. Through an efficient caching algorithm, iCache can significantly improve performance over current iSCSI systems. Numerical results using popular benchmark program and real world trace have shown dramatic performance gain. Xubin He, Qing Yang 0001, Ming Zhang 0026 |
LCN | 2 |
| 2002 | RAPID-Cache-A Reliable and Inexpensive Write Cache for High Performance Storage SystemsabstractModern high performance disk systems make extensive use of nonvolatile RAM (NVRAM) write caches. A single-copy NVRAM cache creates a single point of failure while a dual-copy NVRAM cache is very expensive because of the high cost of NVRAM. This paper presents a new cache architecture called RAPID-Cache for Redundant, Asymmetrically Parallel, and Inexpensive Disk Cache. A typical RAPID-Cache consists of two redundant write buffers on top of a disk system. One of the buffers is a primary cache made of RAM or NVRAM and the other is a backup cache containing a two-level hierarchy: a small NVRAM buffer on top of a log disk. The small NVRAM buffer combines small write data and writes them into the log disk in large sizes. By exploiting the locality property of I/O accesses and taking advantage of well-known Log-structured File Systems, the backup cache has nearly equivalent write performance as the primary RAM cache. The read performance of the backup cache is not as critical because normal read operations are performed through the primary RAM cache and reads from the backup cache happen only during error recovery periods. The RAPID-Cache presents an asymmetric architecture with a fast-write-fast-read RAM being a primary cache and a fast-write-slow-read NVRAM-disk hierarchy being a backup cache. The asymmetrically parallel architecture and an algorithm that separates actively accessed data from inactive data in the cache virtually eliminate the garbage collection overhead, which are the major problems associated with previous solutions such as Log-structured File Systems and Disk Caching Disk. The asymmetric cache allows cost-effective designs for very large write caches for high-end parallel disk systems that would otherwise have to use dual-copy, costly NVRAM caches. It also makes it possible to implement reliable write caching for low-end disk I/O systems since the RAPID-Cache makes use of inexpensive disks to perform reliable caching. Our analysis and trace-driven simulation results show that the RAPID-Cache has significant reliability/cost advantages over conventional single NVRAM write caches and has great cost advantages over dual-copy NVRAM caches. The RAPID-Cache architecture opens a new dimension for disk system designers to exercise trade-offs among performance, reliability, and cost. Yiming Hu, Tycho Nightingale, Qing Yang 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 1999 | RAPID-Cache - A Reliable and Inexpensive Write Cache for Disk I/O SystemsabstractThis paper presents a new cache architecture called RAPID-Cache for Redundant, Asymmetrically Parallel, and Inexpensive Disk Cache. A typical RAPID-Cache consists of two redundant write buffers on top of a disk system. One of the buffers is a primary cache made of RAM or NVRAM and the other is a backup cache containing a two level hierarchy: a small NVRAM buffer on top of a log disk. The backup cache has nearly equivalent write performance as the primary RAM cache, while the read performance of the backup cache is not as critical because normal read operations are performed through the primary RAM cache and reads from the backup cache happen only during error recovery periods. The RAPID-Cache presents an asymmetric architecture with a fast-write-fast-read RAM being a primary cache and a fast-write-slow-read NVRAM-disk hierarchy being a backup cache. The asymmetric cache architecture allows cost-effective designs for very large write caches for high-end disk I/O systems that would otherwise have to use dual-copy, costly NVRAM caches. It also makes it possible to implement reliable write caching for low-end disk I/O systems since the RAPID-Cache makes use of inexpensive disks to perform reliable caching. Our analysis and trace-driven simulation results show that the RAPID-Cache has significant reliability/cost advantages over conventional single NVRAM write caches and has great cost advantages over dual-copy NVRAM caches. The RAPID-Cache architecture opens a new dimension for disk system designers to exercise trade-offs among performance, reliability and cost. Yiming Hu, Qing Yang 0001, Tycho Nightingale |
HPCA | 2 |
| 1999 | Measurement, analysis and performance improvement of the Apache Web serverabstractPerformance of Web servers is critical to the success of many corporations and organizations. However, very few results have been published that quantitatively study the server behavior and identify the performance bottlenecks. In this paper we measured and analyzed the behavior of the popular Apache Web server on a uniprocessor system and a 4-CPU SMP (Symmetric Multi-Processor) system running the IBM AIX operating system. Using the AIX built-in tracing facility, we obtained detailed information on kernel events and system activities while running Apache driven by the SPECweb96 and the WebStone benchmarks. After quantitatively identifying the performance bottlenecks, we proposed and implemented 6 techniques that improve the throughput of Apache by 61%. These techniques are general purpose and can be applied to other Web servers as well. Yiming Hu, Ashwini K. Nanda, Qing Yang 0001 |
IPCCC | 3 |
| 1999 | The Design and Implementation of a DCD Device Driver for Unix
Tycho Nightingale, Yiming Hu, Qing Yang 0001 |
USENIX ATC, General Track | 3 |
| 1999 | A Comparative Analysis of Cache Designs for Vector ProcessingabstractThis paper presents an experimental study on cache memory designs for vector computers. We use an execution-driven simulator to evaluate vector cache performance of a set of application programs from Perfect Club and SPEC92 benchmark suites. Our simulation results uncover a few important facts which were unknown before: First of all, the prime-mapped cache that we newly proposed shows great performance potential in vector processing environment. Because of its conflict-free property, the prime-mapped cache performs significantly better than conventional cache designs for all applications considered. Second, performance results on the benchmarks indicate that data locality in vector processing does exist, although the effects of line size, associativity, replacement algorithm, and prefetching scheme on cache performance are very different from what has been commonly believed. A medium size vector cache (e.g., 128 Kbytes) eliminates the necessity for a large number of interleaved memory banks in vector computers. Our experiments show that the vector computer that has a medium size prime-mapped cache with small cache line size and limited amount of prefetching provides significant speedup over conventional vector computers without cache. Performance results reported in this paper can also provide guidance to general-purpose computer designers to enhance cache performance for numerical applications. Tong Sun 0001, Qing Yang 0001 |
IEEE Trans. Computers | 2 |
| 1998 | Performance of One's Complement Caches
Qing Yang 0001, Sridhar Adina, Tong Sun 0001 |
J. Parallel Distributed Comput. | 1 |
| 1997 | Minimizing Area Cost of On-Chip Cache Memories by Caching Address TagsabstractThis paper presents a technique for minimizing chip-area cost of implementing an on-chip cache memory of microprocessors. The main idea of the technique is Caching Address Tags, or CAT cache, for short. The CAT cache exploits locality property that exists among addresses of memory references. By keeping only a limited number of distinct tags of cached data, rather than having as many tags as cache lines, the CAT cache can reduce the cost of implementing tag memory by an order of magnitude without noticeable performance difference from ordinary caches. Therefore, CAT represents another level of caching for cache memories. Simulation experiments are carried out to evaluate performance of CAT cache as compared to existing caches. Performance results of SPEC92 programs show that the CAT cache, with only a few tag entries, performs as well as ordinary caches, while chip-area saving is significant. Such area saving will increase as the address space of a processor increases. By allocating the saved chip-area for larger cache capacity, or more powerful functional units, CAT is expected to have a great impact on overall system performance. Hong Wang 0003, Tong Sun 0001, Qing Yang 0001 |
IEEE Trans. Computers | 3 |
| 1996 | DCD - Disk Caching Disk: A New Approach for Boosting I/O PerformanceabstractThis paper presents a novel disk storage architecture called DCD, Disk Caching Disk, for the purpose of optimizing I/O performance. The main idea of the DCD is to use a small log disk, referred to as cache-disk, as a secondary disk cache to optimize write performance. While the cache-disk and the normal data disk have the same physical properties, the access speed of the former differs dramatically from the latter because of different data units and different ways in which data are accessed. Our objective is to exploit this speed difference by using the log disk as a cache to build a reliable and smooth disk hierarchy. A small RAM buffer is used to collect small write requests to form a log which is transferred onto the cache-disk whenever the cache-disk is idle. Because of the temporal locality that exists in office/engineering work-load environments, the DCD system shows write performance close to the same size RAM (i.e. solid-state disk) for the cost of a disk. Moreover, the cache-disk can also be implemented as a logical disk in which case a small portion of the normal data disk is used as the log disk. Trace-driven simulation experiments are carried out to evaluate the performance of the proposed disk architecture. Under the office/engineering work-load environment, the DCD shows superb disk performance for writes as compared to existing disk systems. Performance improvements of up to two orders of magnitude are observed in terms of average response time for write operations. Furthermore, DCD is very reliable and works at the device or device driver level. As a result, it can be applied directly to current file systems without the need of changing the operating system. Yiming Hu, Qing Yang 0001 |
ISCA | 2 |
| 1996 | A comparative analysis of different arbitration protocols for multiple-bus multiprocessors
Chi-Ming Chung, Ding-An Chiang, Qing Yang 0001 |
J. Comput. Sci. Technol. | 3 |
| 1996 | Guest editors' introduction
Qing Yang 0001 |
J. Comput. Sci. Technol. | 1 |
| 1996 | A Compiler-Directed Approach to Network Latency Reduction for Distributed Shared Memory Multiprocessors
Sibabrata Ray, Qing Yang 0001 |
J. Parallel Distributed Comput. | 3 |
| 1995 | CAT - Caching Address Tags: A Technique for Reducing Area Cost of On-Chip CachesabstractThis paper presents a technique for minimizing chip-area cost of implementing an on-chip cache memory of microprocessors. The main idea of the technique Caching Address Tags, or CAT cache for short. The CAT cache exploits locality property that exists among addresses of memory references for the purpose of minimizing chip area-cost of address tags. By keeping only a limited number of distinct tags of cached data rather than having as many tags as cache lines, the CAT cache can reduce the cost of implementing tag memory by an order of magnitude without noticeable performance difference from ordinary caches. Therefore, CAT represents another level of caching for cache memories. Simulation experiments are carried out to evaluate performance of CAT cache as compared to existing caches. Performance results of SPEC92 programs show that the CAT cache with only a few tag entries performs as well as ordinary caches while chip-area saving is significant. Such area saving will increase as the address space of a processor increases. By allocating the saved chip area for larger cache capacity, or more powerful functional units, CAT is expected to have a great impact on overall system performance. Hong Wang 0003, Tong Sun 0001, Qing Yang 0001 |
ISCA | 3 |
| 1995 | A Memory Interference Model for Regularly Patterned Multiple Stream Vector AccessesabstractMost existing analytical models for memory interference generally assume random bank selection for each memory access. In vector computers, however, memory accesses are typically regularly patterned with a number of data items being accessed concurrently from different banks. Very little is known about the queueing behavior of memory interferences in multiple stream vector accesses. This paper presents an analytical model for memory interferences due to vector accesses in multiple vector processor systems. The model captures the effects of both bank conflicts among elements within one vector access stream and conflicts among multiple vector access streams on system performance. The model is based on a closed queueing network assuming an ideal interconnection network. An approximation technique is proposed to solve the memory queueing system that serves customers in a complicated way (non-FIFO). We also carry out extensive simulation experiments to study memory interference and validate our analytical model. Simulation results and analytical results are in a very good agreement, indicating that the model is very accurate. We further validate our analysis by comparing the numerical results obtained from our analytical model with those measurement results that were published by other researchers. Based on our analytical model and simulations, we carry out performance evaluation of the multiple vector processor systems. Our numerical results show that memory access conflicts pose a severe limitation on the number of useful processors in the system, implying that memory system design is essential to high-performance computing.> Qing Yang 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1994 | A New and Efficient FFT Algorithm for Distributed Memory SystemsabstractThis paper presents a new and optimal parallel implementation of multidimensional fast Fourier transform algorithm on distributed memory multiprocessors. Its optimality is obtained by minimizing the number of message passings necessary, at the cost of increase in message length. This distinctive feature of the new algorithm effectively utilizes the important architectural property of most of today's distributed memory multiprocessors-wormhole routing for interprocessor communications. By using the algebra of stride permutations and tenser products as a mathematical tool, we are able to derive and formulate an efficient data partition and communication scheme that reduces communication cost from O(N/sup 2/) required for the best known FFT to O(N) on an N/sup 2/-processor machine. Our data partition scheme is natural and efficient for solving discretized boundary value problems such as partial differential equations and finite element analysis. To evaluate the actual performance of our new algorithm in comparison with other existing parallel FFT algorithms, we have carried out implementation experiments on the Intel's Touchstone Delta machine. Nagesh Anupindi, Myoung An, James W. Cooley, Qing Yang 0001 |
ICPADS | 4 |
| 1994 | A One's Complement Cache MemoryabstractMost of today's microprocessors have an on-chip cache to reduce average memory access latency. These on-chip caches generally have low associativity and small sizes. Cache line conflicts are the main source of cache misses which are essential to overall system performance. This paper introduces an innovative, conflict-free cache design, called one's complement cache. By means of parallel computation of cache addresses and memory addresses of data, the new design does not increase critical hit time of cache accesses. Cache misses caused by line interferences are minimized by means of evenly distributing data items referenced by program loops across all sets in a cache. Evenly distribution of data in the cache is achieved by making the number of sets in the cache a prime or an odd number thereby the chance of related data being mapped to a same set is small. Trace-driven simulations are used to evaluate the performance of the new design. Performance results on a set of programs from SPEC92 benchmarks show that the new design improves cache performance over the conventional set-associative cache by about 100% with negligibly additional hardware cost. Qing Yang 0001, Sridhar Adina |
ICPP (1) | 1 |
| 1994 | A Closed-Form Formula for Queueing Delays in Disk ArraysabstractDisk arrays have become increasingly popular as a means of improving performance of secondary storage systems. In making design decisions, it is essential to understand the performances of different configurations as well as effects of system parameters on their performance. Analytical performance models are desirable in both configuring and designing disk arrays since they provide designers with a quick and efficient tool to evaluate and understand a system with a wide range of system and workload parameters. Shengbin Hu, Qing Yang 0001 |
ICPP (2) | 3 |
| 1994 | An Analytical Model for Load Balancing on Symmetric Multiprocessor Systems
Xiaoshu Qian, Qing Yang 0001 |
J. Parallel Distributed Comput. | 2 |
| 1994 | Parallel All-Row Preconditioned Interval Linear Solver for Nonlinear Equations on Multiprocessors
Qi Gan, Qing Yang 0001, Chenyi Hu |
Parallel Comput. | 2 |
| 1993 | Performance of Cache Memories for Vector Computers
Qing Yang 0001 |
J. Parallel Distributed Comput. | 1 |
| 1993 | Introducing a New Cache Design into Vector ComputersabstractIntroduces an innovative cache design for vector computers, called prime-mapped cache. By utilizing the special properties of a Mersenne prime, the new design does not increase the critical path length of a processor, nor does it increase the cache access time as compared to existing cache organizations. The prime-mapped cache minimizes cache miss ratio caused by line interferences that have been shown to be critical for numerical applications by previous investigators. With negligibly additional hardware cost, significant performance gains are obtained by adding the proposed cache memory to an existing vector computer. The performance of the design is studied analytically, using a generic vector computation model. The analytical model is validated through extensive simulation experiments. A performance analysis for various vector access patterns shows that the prime-mapped cache performs significantly better than conventional cache organizations in the vector processing environment. The performance gain will increase with the increase of the speed gap between processors and memories.> Qing Yang 0001 |
IEEE Trans. Computers | 1 |
| 1993 | A New Graph Approach to Minimizing Processor Fragmentation in Hypercube MultiprocessorsabstractThe authors propose a new approach for subcube and noncubic processor allocations for hypercube multiprocessors. The main idea is to represent available processors in the system by means of a prime cube graph (PC-graph). The PC-graph maintains the inter-relationships between free subcubes and hence reduces both internal and external processor fragmentations. Their simulation results show that the PC-graph approach outperforms the existing allocation strategies by 25% to 50% under certain load conditions.> Qing Yang 0001, Hong Wang 0003 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1992 | On Fault-Tolerant Computation of Orthogonal Transforms on Hypercube Computers
Qing Yang 0001, Hong Wang 0003 |
ICPP (1) | 1 |
| 1992 | A Novel Cache Design for Vector ProcessingabstractThis paper introduces an innovative cache design for vector computers, called prime-mapped cache. By utilizing the special properties of a Mersenne prime, the new design does not increase the critical path length of a processor, nor does it increase the cache access time as compared to a direct-mapped cache. The prime-mapped cache minimizes cache miss ratio caused by line interferences that have been shown to be critical for numerical applications by previous investigators. We show that significant performance gains are possible by adding the proposed cache memory into an existing vector computer provided that application programs can be blocked. The performance gain will increase with the increase of the speed gap between processors and memories. We develop an analytical performance model based on a generic vector computation model to study the performance of the design. Our preliminary performance analysis on various vector access patterns shows that the prime-mapped cache can provide as much as a factor of 2 to 3 performance improvement over the conventional direct-mapped cache in the vector processing environment. Moreover, the additional hardware cost introduced by the new mapping scheme is negligible. Qing Yang 0001, Liping Wu Yang |
ISCA | 1 |
| 1992 | Performance study of two protocols for voice/data integration on ring networks
Qing Yang 0001, Dipak Ghosal, Satish K. Tripathi |
Comput. Networks ISDN Syst. | 1 |
| 1992 | Design of an Adaptive Cache Coherence Protocol for Large Scale MultiprocessorsabstractA large scale, cache-based multiprocessor that is interconnected by a hierarchical network such as hierarchical buses or a multistage interconnection network (MIN) is considered. An adaptive cache coherence scheme for the system is proposed based on a hardware approach that handles multiple shared reads efficiently. The new protocol allows multiple copies of a shared data block in the hierarchical network, but minimizes the cache coherence overhead by dynamically partitioning the network into sharing and nonsharing regions based on program behavior. The new cache coherence scheme effectively utilizes the bandwidth of the hierarchical networks and exploits the locality properties of parallel algorithms. Simulation experiments have been carried out to analyze the performance of the new protocol. The simulation results show that the new protocol gives 15% to 30% performance improvement over some existing cache coherence schemes on similar systems for a wide range of workload parameters.> Qing Yang 0001, George Thangadurai, Laxmi N. Bhuyan |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 1991 | Load balancing on generalized hypercube and mesh multiprocessors with LALabstractA typical nearest neighbor balancing strategy, called LAL (local average load), in which the workload of a processor is averaged among its nearest neighbors at discrete time steps is investigated. The underlying systems considered are multiprocessor systems interconnected by generalized hypercube (GHC), mesh and loop structures. It is assumed that the amount of computation tasks arriving at or finished by a processor at each time step can be described by a random variable with some general distribution. With some general assumptions about these random variables, it is shown that the expected difference between the actual load of a processor and the average load of the system is zero and the variance of this difference is bounded by a constant independent of time.> Xiaoshu Qian, Qing Yang 0001 |
ICDCS | 2 |
| 1991 | Prime Cube Graph Approach for Processor Allocation in Hypercube Multiprocessors
Hong Wang 0003, Qing Yang 0001 |
ICPP (1) | 2 |
| 1991 | Effects of Arbitration Protocols on the Performance of Multiple-Bus Multiprocessors
Qing Yang 0001 |
ICPP (1) | 1 |
| 1991 | Analysis of Packet-Switched Multiple-Bus Multiprocessor SystemsabstractPerformance analyses of packet-switched multiple-bus multiprocessor systems are presented. Approximate queuing network models are developed for both synchronous and asynchronous control schemes, and the results are shown to be in good agreement with simulation results. The analysis of the synchronous system is based on a decomposition technique, with each of the shared resources in the system being represented as a single-server queue. For asynchronous systems, the analysis is based on the flow equivalence technique. Numerical results obtained from the analyses indicate that packet-switched multiple-bus multiprocessors with only a few buses perform almost as well as crossbar-based multiprocessors.> Qing Yang 0001, Laxmi N. Bhuyan |
IEEE Trans. Computers | 1 |
| 1990 | Performance of Multiple-Bus Interconnections for Multiprocessors
Qing Yang 0001, Laxmi N. Bhuyan |
J. Parallel Distributed Comput. | 1 |
| 1989 | Approximate Analysis of Single and Multiple Ring NetworksabstractAsynchronous packet-switched interconnection networks with decentralized control are very appropriate for multiprocessing and data-flow architectures. The authors present performance models of single- and multiple-ring networks based on token-ring, slotted-ring, and register-insertion-ring protocols. The multiple ring networks have the advantage of being reliable, expandable, and cost effective. An approximate and uniform analysis, based on the gate M/G/1 queuing model, has been developed to evaluate the performance of both existing single-ring networks and the proposed multiple-ring networks. Approximations are good for low and medium load. The analyses are based on symmetric ring structure with nonexhaustive service policy and infinite queue length at each station. They essentially involve modeling of queues with single- and multiple-walking servers. The results obtained from the analytical models are compared to those obtained from simulation.> Laxmi N. Bhuyan, Dipak Ghosal, Qing Yang 0001 |
IEEE Trans. Computers | 3 |
| 1989 | Analysis and Comparison of Cache Coherence Protocols for a Packet-Switched MultiprocessorabstractAnalytical models are developed for seven existing cache protocols, namely, Write-Once, Write-Through, Synapse, Berkeley, Illinois, Firefly, and Dragon. The protocols are implemented on a multiprocessor with a packet-switched shared bus. The models are based on queuing networks that consist of both open and closed classes of customers. The models incorporate the requests for invalidation signals, write-through, and write-back operations, and the solution is based on the mean value analysis (MVA) algorithm. The performance of these protocols under various system parameters is compared on the basis of the models. It is found that Firefly and Dragon perform better than the others.> Qing Yang 0001, Laxmi N. Bhuyan, Bao-Chyn Liu |
IEEE Trans. Computers | 1 |
| 1988 | A Queueing Network Model for a Cache Coherence Protocol on Multiple-bus Multiprocessors
Qing Yang 0001, Laxmi N. Bhuyan |
ICPP (1) | 1 |
| 1987 | Design and Analysis of a Decentralized Multiple-Bus Multiprocessor
Qing Yang 0001, Laxmi N. Bhuyan |
ICPP | 1 |
| 1987 | Performance Analysis of Packet-Switched Multiple-Bus Multiprocessor Systems
Qing Yang 0001, Laxmi N. Bhuyan, R. Pavaskar |
RTSS | 1 |