EDBT 2026 Demo / reviewers in the wild / expert
Youhui Zhang
dblp:00/5995
· DBLP profile ↗
57ranked-venue papers
14as first author
24since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 44 · 8 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CoGrad: Profiling Competence Boundaries for Efficient Mathematical Alignment in LLMs
Youhui Zhang, WenLiang Wu, Xifan Liu, Shouqiang Liu |
ICIC (22) | 1 |
| 2026 | Look Before You Leap : Precision Instruction Supply via SmartScoutabstractModern high-performance processors extensively employ Fetch-Directed Instruction Prefetching (FDIP) to mitigate instruction supply bottlenecks. However, the efficacy of FDIP is fundamentally constrained by the accuracy of the Branch Prediction Unit (BPU). As the critical component within the BPU, the Branch Target Buffer (BTB) faces severe capacity bottlenecks. While pre-decoding-based prefetching offers a remedy, existing approaches suffer from two critical impediments: (1) The Noise Dilemma: Suboptimal trade-off between coverage and accuracy. (2) Inefficient Miss Resolution: Current designs rely on reactive recovery or stalls, failing to leverage available front-end slack for proactive correction. Peng Qu 0001, Tingji Zhang, Fang Su, Zhe Pan 0001, Youhui Zhang |
ICS | 6 |
| 2026 | PANA: A Fine-Grained Runtime-Adaptive Load Balancing for Parallel SpMV on Multicore CPUsabstractSpMV has been widely utilized and is regarded as a significant kernel in various scientific and engineering computing applications, where its parallel performance is heavily influenced by matrix sparsity and hardware architecture. Despite extensive prior research, static partitioning strategies that narrowly target computation or memory access remain a key performance bottleneck, severely stifling performance advancement of SpMV on modern multicore CPUs. Haodong Bian, Youhui Zhang, Jianqiang Huang 0002, Xiaoying Wang 0002 |
PPoPP | 2 |
| 2026 | VDHA: Vector-Driven Hash Aggregation for Sparse Matrix-Sparse Vector Multiplication on GPUsabstractSparse matrix-sparse vector multiplication (SpMSpV) is a core primitive in graph analytics and scientific computing, also arising in spiking neural networks for event-driven spike propagation. On GPUs, the performance of the prevalent and efficient SpMSpV paradigm is often bottlenecked by the write-back phase of accumulating non-zero multiply–accumulate results; its many-to-one index scatter pattern causes severe conflicts and poor bandwidth utilization on GPUs. We present VDHA, a GPU-based weighted SpMSpV kernel that leverages block-private hash tables for local aggregation, substantially reducing write conflicts and improving memory coalescing. To further amplify this benefit, we incorporate column splitting with lightweight reordering to expose more locality, and employ a fetch–compute–writeback pipeline to overlap hash computation with memory accesses. Extensive evaluation on over 300 matrices with more than 5 million nonzeros, including web-scale graphs (Konect/LAW) and scientific workloads (SuiteSparse), shows that VDHA consistently outperforms state-of-the-art baselines. On web graphs, it achieves a 1.41× geometric-mean speedup (up to 3.42×), while on SuiteSparse it delivers 1.13× (up to 2.55×). We also provide a lightweight predictive model that identifies matrices favorable to VDHA with 91.3% accuracy. Zhe Pan 0001, Peng Qu 0001, Youhui Zhang |
PPoPP | 4 |
| 2026 | Root-Down Exposure for Maximal Clique Enumeration on GPUsabstractMaximal clique enumeration (MCE) in large-scale graphs is critical across various application domains, including social network analysis, bioinformatics, and computer vision. However, existing GPU-based MCE solutions suffer from inefficient load-balancing mechanisms. These mechanisms force busy workers to pause and hand over workloads to idle workers, which introduces significant synchronization overhead and increases memory usage. To address these limitations, we introduce a root-down exposure mechanism where busy workers dynamically expose their current root, enabling idle workers to pull workloads from the exposed node directly without synchronization. We then propose a bitmap-centric MCE scheme and an aggressive node generation rule to further simplify the memory layout and accelerate enumeration. We combine them into RDMCE, a Root-Down MCE solution on GPUs. Across large real-world graphs with up to 146 billion maximal cliques, RDMCE is 1.25-5.38× faster than any next-best state-of-the-art GPU-based solutions and is the only one that completes enumeration on every test dataset, offering a more efficient and scalable MCE solution. Zhe Pan 0001, Peng Qu 0001, Youhui Zhang |
PPoPP | 3 |
| 2025 | Hierarchical Prefetching: A Software-Hardware Instruction Prefetcher for Server ApplicationsabstractThe large working set of instructions in server-side applications causes a significant bottleneck in the front-end, even for high-performance processors equipped with fetch-directed instruction prefetching (FDIP). Prefetchers specifically designed for server scenarios typically rely on a record-and-replay mechanism that exploits the repetitiveness of instruction sequences. However, the efficacy of these techniques is compromised by discrepancies between actual and predicted control flows, resulting in loss of coverage and timeliness. This paper proposes Hierarchical Prefetching, a novel approach that tackles the limitations of existing prefetchers. It identifies common coarse-grained functionality blocks (called Bundles) within the server code and prefetches them as a whole. Bundles are significantly larger than typical prefetch targets, encompassing tens to hundreds of kilobytes of code. The approach combines simple software analysis of code for bundle formation and light-weight hardware for record-and-replay prefetching. The prefetcher requires under 2KB of on-chip storage by keeping most of the metadata in main memory. Experiments with 11 popular server workloads reveal that Hierarchical Prefetching significantly improves miss coverage and timeliness over prior techniques, achieving a 6.6% average performance gain over FDIP. Tingji Zhang, Boris Grot, Wenjian He, Yashuai Lv, Peng Qu 0001, Fang Su, Guowei Zhang 0002, Youhui Zhang |
ASPLOS (2) | 10 |
| 2025 | SoftGuide: A Hardware-Software Co-design Predictor for Data-Dependent BranchesabstractBranch predictors based on historical information exhibit exemplary performance. However, a subset of data-dependent branches still pose a significant challenge, often resulting in severe mispredictions. Such branches are commonly encountered during the processing of various data structures, and increasing the capacity of predictors has shown limited effective-ness. Thus, a dedicated branch predictor is necessary to enhance performance for varying data structures while maintaining lower hardware complexity. This paper proposes SoftGuide, a novel hardware-software cooperative branch predictor designed to tackle two main chal-lenges inherent in data-dependent branch prediction: (1) the detection and identification of branch dependencies(software-friendly) and (2) data prefetching and dependency chain pre-execution(hardware-friendly). SoftGuide leverages software to convey the memory access patterns and the dependency chains associated with the branch, thereby circumventing the overhead of hardware-based detection. Utilizing the information provided by the software, the enhanced hardware prefetches data and triggers pre-execution in advance. SoftGuide can perfectly unify prefetching and prediction tasks. For SPEC2006 and GAP benchmarks with the method of SimPoint, SoftGuide realizes a decrease in branch mispredictions per 1K instructions (MPKI) by 46.4% and an increase in Instructions Per Cycle (IPC) by 1.25x average on the processor equipped with the state-of-the-art branch predictor. Moreover, the storage overhead is just 3.88KB. Peng Qu 0001, Tingji Zhang, Youhui Zhang |
CCGrid | 5 |
| 2025 | AlphaSparseTensor: Discovering Faster Sparse Matrix Multiplication Algorithms on GPUs for LLM Inference
Xuanzheng Wang, Shuo Miao, Zihan Zhu, Youhui Zhang |
Euro-Par (2) | 5 |
| 2025 | Improving crowdsourced label quality by peer-to-peer federated learning
Xiangming Lu, Jiangbo Qian, Chong Wang 0001, Diqun Yan, Youhui Zhang |
Appl. Intell. | 5 |
| 2024 | DTrans: A Dataflow-transformation FPGA Accelerator with Nonlinear-operators fusion aiming for the Generative ModelabstractFPGA-based accelerators have emerged as an effective solution for GPT inference, given their inherent flexibility and capacity for domain-specific customization. However, the effective utilization of GPT has been hindered by two main obstacles: the unequal ratios between compute and memory access during the prefilling and generation stages, and the growing hardware resource requirements for nonlinear operations caused by longer text lengths and larger embedding dimensions. To address these challenges, we introduce DTrans, an FPGA accelerator specifically designed for GPT, which leverages nonlinear-operator fusion and a customized pipeline to efficiently handle long-input and long-generation tasks. We introduce a sequence-length-decoupled nonlinear operator design that enables in-place execution while preserving computational accuracy. Additionally, we employ a customized pipeline design that utilizes a two-level alternating input pipeline mapping for long-input tasks to alleviate the overhead associated with residual computation and on-chip buffers. Furthermore, for long-token generation tasks, our pipeline overlaps computational delays in operations such as Softmax and layer normalization with matrix operations. We also propose a novel two-stage dataflow transformation strategy that adopts different reuse strategies for the prefilling and generating stages with different computational/memory access characteristics. Our comparative analyses reveal that DTrans outperforms the GPU(V100) in terms of throughput and energy efficiency, achieving improvements of $11.99 \times$ and $11.7 \times$, respectively. When compared with state-of-the-art GPT inference accelerators, DTrans demonstrates more than $5.64 \times$ and $5.22 \times$ enhancements in these metrics. Xuanzheng Wang, Shuo Miao, Youhui Zhang |
FPL | 4 |
| 2024 | ChameSC: Virtualizing Superscalar Core of a SIMD Architecture for Vector Memory AccessabstractIn modern computing, the persistent issue of the memory wall significantly inhibits performance efficiency, especially in Single Instruction, Multiple Data (SIMD) architectures processing data-parallel workloads. Contemporary applications, including machine learning, data analytics, and computer vision, often feature complex chains of dependent, indirect, stride, and indexed vector memory accesses. These intricate access patterns place substantial pressure on vector memory units, leading to underutilization of SIMD resources and overall performance degradation. This study introduces ChameSC11A combination of Chameleon and Superscalar Core., a novel, adaptive SIMD architecture addressing the escalating memory wall issue. ChameSC exploits the typically idle memory units of the superscalar core during data-parallel workload execution by dynamically virtualizing the core to function as a precise data prefetcher, fetching data exactly as needed by the SIMD architecture in advance. The L1 cache of the superscalar core serves as a proactive cache to store prefetched data, which can be directly accessed by the vector memory unit. This strategic utilization of idle resources significantly enhances memory-level parallelism and effective memory bandwidth. Evaluations of ChameSC demonstrate substantial performance enhancements, achieving an average speedup of 34.3% compared to conventional SIMD architectures across a range of representative data-parallel applications with various memory access behaviors. Compared to state-of-the-art data prefetching techniques such as Berti and IPCP, ChameSC offers better performance without introducing extra storage overhead. Zhongzhu Pu, Guangda Zhang, Tiejian Zhang, Youhui Zhang |
ICCD | 5 |
| 2024 | ActiveN: A Scalable and Flexibly-Programmable Event-Driven Neuromorphic ProcessorabstractAt present, most neuromorphic chips utilize custom circuits and/or on-chip memory to achieve neural computations and parameter storage. However, with the diversified development of brain-inspired applications, this approach faces significant challenges regarding programming flexibility and scalability. To address these issues, we propose a RISC-V-based many-core neuromorphic architecture, ActiveN. Each neuro-core is equipped with an active-message-enabled micro-architecture to support the event-driven programming model of Spiking Neural Networks (SNNs), and the memory subsystem is enhanced to identify sparse data and forward them directly. These mechanisms significantly mitigate the impact of memory latency on performance to enable the storage of synapse data in off-chip (or off-die) bulk storages (e.g., DRAMs, HBMs), which not only improve storage scalability but also increase computing density. Furthermore, the core's instruction set is customized to include a compact and complete set of fixed-and floating-point computing instructions, supporting various neural models and SNN computation algorithms flexibly. End-to-end prototype testing demonstrates that, compared to a state-of-the-art chip based on custom circuitry and on-chip SRAM that also supports event-driven operations, ActiveN can integrate over 10x more processing units (512) with strong scalability to fully utilize memory bandwidth. Additionally, it achieves 7.9 times the performance of this counterpart and 96.6 times the performance of an NVIDIA A100 GPU, while maintaining flexible programmability. Zhongzhu Pu, Youhui Zhang |
MICRO | 5 |
| 2024 | A Row Decomposition-based Approach for Sparse Matrix Multiplication on GPUsabstractSparse-Matrix Dense-Matrix Multiplication (SpMM) and Sampled Dense Dense Matrix Multiplication (SDDMM) are important sparse kernels in various computation domains. The uneven distribution of nonzeros in the sparse matrix and the tight data dependence between sparse and dense matrices make it a challenge to run sparse matrix multiplication efficiently on GPUs. By analyzing the aforementioned problems, we propose a row decomposition (RoDe)-based approach to optimize the two kernels on GPUs, using the standard Compressed Sparse Row (CSR) format. Specifically, RoDe divides the sparse matrix rows into regular parts and residual parts, to fully optimize their computations separately. We also devise the corresponding load balancing and finegrained pipelining technologies. Profiling results show that RoDe can achieve more efficient memory access and reduce warp stall cycles significantly. Compared to the state-of-the-art (SOTA) alternatives, RoDe achieves a speedup of up to 7.86× with a geometric mean of 1.45× for SpMM, and a speedup of up to 8.99× with a geometric mean of 1.49× for SDDMM; the dataset is SuiteSparse. RoDe also outperforms its counterpart in the deep learning dataset. Furthermore, its preprocessing overhead is significantly smaller, averaging only 16% of the SOTA. Peng Qu 0001, Youhui Zhang, Zhaolin Li |
PPoPP | 4 |
| 2023 | FABLE: A Development and Computing Framework for Brain-inspired Learning AlgorithmsabstractSpiking neural networks (SNNs) have received extensive attention in multi-disciplinary fields, due to their rich spatiotemporal dynamics and the potential for low processing delay and high energy efficiency on neuromorphic hardware. The research on SNN learning algorithms is active and diverse, and many algorithms differ significantly from those of DNN in terms of computation model/features and weight adjustment mechanisms. This paper proposes FABLE, a multi-level framework for building and running SNN learning algorithms efficiently. Its kernel is an adaptable computation model based on synchronous data flow, which can well express the spatiotemporal parallelism of SNN and then can organize and schedule the underlying SNN-custom tensor operators (OPs) to construct optimized computing procedures. It also provides a flexible programming interface for users to design or customize their learning algorithms. In addition, the implementation of FABLE has high compatibility: It extends PyTorch's OP library, scheduler, and APIs to take advantage of the ecology and usability of the latter. To show the flexibility of the framework, we have ported five different learning algorithms, each with less programming than its original implementation. Further experiments demonstrate that FABLE outperforms all of them (up to 2.61 x) in terms of computing performance, while the original implementations are either based on PyTorch, based on some third-party tool using PyTorch, or based on GPGPU's runtime directly. Zhaolin Li, Youhui Zhang |
IJCNN | 4 |
| 2023 | Multi-Objective Optimization for Floating Point Mix-Precision TuningabstractThis paper proposes a multi-objective optimization method for mixed-precision computation. Unlike previous studies that often take mantissa length reduction as the only optimization target, our work models the actual performance and power consumption of mixed precision programs on the corresponding hardware platforms, and based on this model searches for the pareto optimal set of all precision configurations. Experiments show that this tool can obtain performance improvements of 15% - 71 % on floating-point benchmarks while satisfying accuracy requirements. Compared to some typical counterpart-work, an average 21 % improvement can be obtained in SIMD scenarios. Zeqing Li, Yongwei Wu 0001, Youhui Zhang |
ISLPED | 3 |
| 2023 | MAICC : A Lightweight Many-core Architecture with In-Cache Computing for Multi-DNN Parallel InferenceabstractThe growing complexity and diversity of neural networks in the fields of autonomous driving and intelligent robots have facilitated the research of many-core architectures, which can offer sufficient programming flexibility to simultaneously support multi-DNN parallel inference with different network structures and sizes compared to domain-specific architectures. However, due to the tight constraints of area and power consumption, many-core architectures typically use lightweight scalar cores without vector units and are almost unable to meet the high-performance computing needs of multi-DNN parallel inference. To solve the above problem, we design an area- and energy-efficient many-core architecture by integrating large amounts of lightweight processor cores with RV32IMA ISA. The architecture leverages the emerging SRAM-based computing-in-memory technology to implement vector instruction extensions by reusing memory cells in the data cache instead of conventional logic circuits. Thus, the data cache in each core can be reconfigured as the memory part and the computing part with the latter tightly coupled with the core pipeline, enabling parallel execution of the basic RISC-V instructions and the extended multi-cycle vector instructions. Furthermore, a corresponding execution framework is proposed to effectively map DNN models onto the many-core architecture by using intra-layer and inter-layer pipelining, which potentially supports multi-DNN parallel inference. Experimental results show that the proposed MAICC architecture obtains a 4.3 × throughput and 31.6 × energy efficiency over CPU (Intel i9-13900k). MAICC also achieves a 1.8 × energy efficiency over GPU (RTX 4090) with only 4MB on-chip memory and 28 mm2 area. Renhao Fan, Yikai Cui, Qilin Chen, Mingyu Wang 0003, Youhui Zhang, Zhaolin Li |
MICRO | 5 |
| 2023 | ENLARGE: An Efficient SNN Simulation Framework on GPU ClustersabstractSpiking Neural Networks (SNNs) are currently the most widely used computing model for neuroscience communities. There is also an increasing research interest in exploring the potential of SNN in brain-inspired computing, artificial intelligence, and other areas. As SNNs possess distinguished characteristics that originate from biological authenticity, they require dedicated simulation frameworks to achieve usability and efficiency. However, there is no widely-used, easily accessible, high performance SNN simulation framework for GPU clusters. In this paper, we propose ENLARGE, an efficient SNN simulation framework on GPU clusters. ENLARGE provides a multi-level architecture that deals with computation, communication, and synchronization hierarchically. We also propose an efficient communication method with an all-to-all communication pattern. To deal with the delay of spike delivery, which is the most distinguished SNN characteristic, several delay-aware optimization methods are also proposed. We further propose a multilevel workload management method. Various experiments are carried out to demonstrate the performance and scalability of the framework, as well as the effects of the optimization methods. Test results show that ENLARGE can achieve$3.17\times \sim 28.12\times$speedup compared with the most widely used NEST simulator and$3.26\times \sim 13.57\times$speedup compared with the widely used NEST GPU simulator for GPU clusters. Peng Qu 0001, Youhui Zhang |
IEEE Trans. Parallel Distributed Syst. | 6 |
| 2022 | Accelerating Neural Network Training with Processing-in-Memory GPUabstractProcessing-in-memory (PIM) architecture is promising for accelerating deep neural network (DNN) training due to its low-latency and energy-efficient data movement between computation units and the memory. This paper explores a novel GPU-PIM architecture for DNN training, where streaming multiprocessors of GPU are integrated into the logic layer of 3D memory stack, and multiple such stacks are connected to form a PIM-network. Two corresponding optimization strategies are proposed. The first is to increase the computational parallelism of the data-parallel training mode with the large memory, high bandwidth/high network transmission speed of GPU-PIM. The second is further utilizing the optimized model-parallel training to significantly reduce the communication overhead: We propose a mapping scheme to decide the proper parallelization for different DNN layers on the proposed architecture. Experiments show that the proposed architecture outperforms the baseline GPU by 35.5% and 59.9% and reduces energy consumption by 28.2% and 27.8% for the two benchmarks we evaluated. Jianhui Han, Youhui Zhang |
CCGRID | 5 |
| 2022 | GaBAN: a generic and flexibly programmable vector neuro-processor on FPGAabstractSpiking neural network (SNN) is the main computational model of brain-inspired computing and neuroscience, which also acts as the bridge between them. With the rapid development of neuroscience, accurate and flexible SNN simulation with high performance is becoming important. This paper proposes GaBAN, a generic and flexibly programmable neuro-processor on FPGA. Different from the majority of current designs that realize neural components by custom hardware directly, it is centered on a compact, versatile vector instruction set, which supports multiple-precision vector calculation, indexed-/strided-memory access, and conditional execution to accommodate computational characteristics. By software and hardware co-design, the compiler extracts memory-accesses from SNN programs to generate micro-ops executed by an independent hardware unit; the latter interacts with the computing pipeline through an asynchronous buffering mechanism. Thus memory access delay can fully cover the calculation. Tests show that GaBAN can not only outperform the SOTA ISA-based FPGA solution remarkably but also be comparable with counterparts of the hardware-fixed model on some tasks. Moreover, in end-to-end testing, its simulation performance exceeds that of high-performance X86 processor (1.44--3.0x). Youhui Zhang |
DAC | 3 |
| 2022 | A review of basic software for brain-inspired computing
Peng Qu 0001, Youhui Zhang |
CCF Trans. High Perform. Comput. | 4 |
| 2022 | EcoForecast: An interpretable data-driven approach for short-term macroeconomic forecasting using N-BEATS neural network
Xuanzheng Wang, Changwang Li, Chengqi Yi, Xinan Xu, Youhui Zhang |
Eng. Appl. Artif. Intell. | 6 |
| 2022 | Polyhedral-Based Compilation Framework for In-Memory Neural Network AcceleratorsabstractMemristor-based processing-in-memory architecture is a promising solution to the memory bottleneck in the neural network ( NN ) processing. A major challenge for the programmability of such architectures is the automatic compilation of high-level NN workloads, from various operators to the memristor-based hardware that may provide programming interfaces with different granularities. This article proposes a source-to-source compilation framework for such memristor-based NN accelerators, which can conduct automatic detection and mapping of multiple NN operators based on the flexible and rich representation capability of the polyhedral model. In contrast to previous studies, it implements support for pipeline generation to exploit the parallelism in the NN loads to leverage hardware resources for higher efficiency. The evaluation based on synthetic kernels and NN benchmarks demonstrates that the proposed framework can reliably detect and map the target operators. Case studies on typical memristor-based architectures also show its generality over various architectural designs. The evaluation further demonstrates that compared with existing polyhedral-based compilation frameworks that do not support the pipelined execution, the performance can upgrade by an order of magnitude with the pipelined execution, which emphasizes the necessity of our improvement. Jianhui Han, Zhaolin Li, Youhui Zhang |
ACM J. Emerg. Technol. Comput. Syst. | 4 |
| 2021 | Regu2D: Accelerating Vectorization of SpMV on Intel Processors through 2D-partitioning and Regular ArrangementabstractSparse matrix-vector multiplication (SpMV) is an elementary kernel of many high-performance computing (HPC) applications, and it is often one of the performance bottlenecks of them. Accelerating SpMV on vector processors usually faces several issues including irregular data accesses, memory bandwidth limitation, and the short vector problem. Based on a detailed analysis of the effects and interactions of various technologies introduced by state-of-the-art studies (ALBUS, CVR, CSR5, SELL-C-σ etc.), we propose Regu2D, a comprehensive solution to accelerate vectorization of SpMV through three methods: adaptive 2D-partitioning, the regular arrangement of matrix elements, and indices compression. Dynamic programming algorithms are used to optimize the first two methods. We conduct experiments on Intel Xeon processors (Skylake architecture) which support AVX-512 SIMD instructions and use sparse matrices from the University of Florida Sparse Matrix Collection. Experiments show that Regu2D achieves an average speedup of 1.69X, 1.93X, 1.40X, and 1.20X over ALBUS, CVR, CSR5, and SELL-C-σ for 30 scale-free sparse matrices, respectively. For 16 HPC sparse matrices, Regu2D achieves an average speedup of 1.34X, 1.89X, 1.34X, and 1.50X over them, respectively. Youhui Zhang |
ICPP | 2 |
| 2021 | A Reduced Architecture for ReRAM-Based Neural Network Accelerator and Its Software StackabstractNeural network (NN) accelerators based on resistive random access memory (ReRAM) have been widely investigated as a promising solution to address the memory wall challenge, due to its capability of processing-in-memory with extremely high density. However, the performance of these accelerators is bounded by the peripheral circuits and the interconnection. And they also suffer from accuracy issue and flexibility issue caused by the device-variation and the in-situ computation mode respectively. Solving these issues with hardware will further offset the performance. Enlightened by the design principle of conventional reduced instruction set computer (RISC), this article proposes a new software/hardware system for ReRAM-based NN acceleration, which achieves complex functions with sophisticated software tools while making the hardware much more compact and efficient, in order to fully utilize ReRAM potential. The hardware architecture, Field Programmable Synapse Array (FPSA), provides high-density reconfigurable logic, wires, and computational resources based on ReRAM-crossbar. Accordingly, the software can convert various types of NNs, including dynamic NNs, into equivalent networks that meet hardware constraints with negligible accuracy loss and optimize the scheduling and mapping of the latter onto FPSA. Further evaluations show that compared to one state-of-the-art ReRAM-based NN accelerator, PRIME, our approach achieves up to 1000× speedup. Yu Ji 0002, Youhui Zhang |
IEEE Trans. Computers | 3 |
| 2020 | SuSy: A Programming Model for Productive Construction of High-Performance Systolic Arrays on FPGAsabstractSystolic algorithms are one of the killer applications on spatial architectures such as FPGAs and CGRAs. However, it requires a tremendous amount of human effort to design and implement a high-performance systolic array for a given algorithm using the traditional RTL-based methodology. On the other hand, existing high-level synthesis (HLS) tools either (1) force the programmers to do "micro-coding" where too many optimizations must be carried out through tedious code restructuring and insertion of vendor-specific pragmas, or (2) give them too little control to influence a push-button compilation flow to achieve high quality of results. Yi-Hsiang Lai, Hongbo Rong, Size Zheng 0001, Xiuping Cui, Yunshan Jia, Jie Wang 0022, Brendan Sullivan, Zhiru Zhang, Yun Liang 0001, Youhui Zhang, Jason Cong, Nithin George, Christopher J. Hughes, Pradeep Dubey |
ICCAD | 11 |
| 2020 | ERA-LSTM: An Efficient ReRAM-Based Architecture for Long Short-Term MemoryabstractProcessing-in-memory (PIM) architecture based on resistive random access memory (ReRAM) crossbars is a promising solution to the memory bottleneck that long short-term memory (LSTM) faces. Based on the dataflow analysis of the LSTM computing paradigm, this article proposes to adopt the ReRAM-based analog approximate computing to conduct the LSTM-specific element-wise computation. Combined with the dot-product computation implemented with ReRAM crossbars, a new LSTM processing tile is designed to significantly reduce the demand for analog-to-digital converters (ADCs), which is the major part of power consumption of existing designs. Next, we elaborate on a mapping scheme to efficiently deploy large-scale LSTM onto multiple processing tiles. Finally, an architecture enhancement is proposed to support crossbar-friendly LSTM pruning to further improve efficiency. This overall design, named ERA-LSTM, is presented. Our evaluation shows that it can outperform two state-of-the-art FPGA-based LSTM accelerators by 103.6 and 35.9 times, respectively; compared with a state-of-the-art ReRAM-based LSTM accelerator with digital element-wise computation, it is 6.1 times more efficient. Moreover, our experiments demonstrate that the impact of hardware constraints and approximation errors on the inference accuracy can be effectively reduced by the proposed fine-tuning scheme and by optimizing the design of the approximator. Jianhui Han, Mingyu Wang 0003, Zhaolin Li, Youhui Zhang |
IEEE Trans. Parallel Distributed Syst. | 5 |
| 2020 | High Performance Simulation of Spiking Neural Network on GPGPUsabstractSpiking neural network (SNN) is the most commonly used computational model for neuroscience and neuromorphic computing communities. It provides more biological reality and possesses the potential to achieve high computational power and energy efficiency. Because existing SNN simulation frameworks on general-purpose graphics processing units (GPGPUs) do not fully consider the biological oriented properties of SNNs, like spike-driven, activity sparsity, etc., they suffer from insufficient parallelism exploration, irregular memory access, and load imbalance. In this article, we propose specific optimization methods to speed up the SNN simulation on GPGPU. First, we propose a fine-grained network representation as a flexible and compact intermediate representation (IR) for SNNs. Second, we propose the cross-population/-projection parallelism exploration to make full use of GPGPU resources. Third, sparsity aware load balance is proposed to deal with the activity sparsity. Finally, we further provide dedicated optimization to support multiple GPGPUs. Accordingly, BSim, a code generation framework for high-performance simulation of SNN on GPGPUs is also proposed. Tests show that, compared to a state-of-the-art GPU-based SNN simulator GeNN, BSim achieves 1.41× - 9.33× speedup for SNNs with different configurations; it outperforms other simulators much more. Peng Qu 0001, Youhui Zhang |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2019 | FPSA: A Full System Stack Solution for Reconfigurable ReRAM-based NN Accelerator ArchitectureabstractNeural Network (NN) accelerators with emerging ReRAM (resistive random access memory) technologies have been investigated as one of the promising solutions to address the memory wall challenge, due to the unique capability of processing-in-memory within ReRAM-crossbar-based processing elements (PEs). However, the high efficiency and high density advantages of ReRAM have not been fully utilized due to the huge communication demands among PEs and the overhead of peripheral circuits. In this paper, we propose a full system stack solution, composed of a reconfigurable architecture design, Field Programmable Synapse Array (FPSA) and its software system including neural synthesizer, temporal-to-spatial mapper, and placement & routing. We highly leverage the software system to make the hardware design compact and efficient. To satisfy the high-performance communication demand, we optimize it with a reconfigurable routing architecture and the placement & routing tool. To improve the computational density, we greatly simplify the PE circuit with the spiking schema and then adopt neural synthesizer to enable the high density computation-resources to support different kinds of NN operations. In addition, we provide spiking memory blocks (SMBs) and configurable logic blocks (CLBs) in hardware and leverage the temporal-to-spatial mapper to utilize them to balance the storage and computation requirements of NN. Owing to the end-to-end software system, we can efficiently deploy existing deep neural networks to FPSA. Evaluations show that, compared to one of state-of-the-art ReRAM-based NN accelerators, PRIME, the computational density of FPSA improves by 31x; for representative NNs, its inference performance can achieve up to 1000x speedup. Yu Ji 0002, Youyang Zhang, Xinfeng Xie, Shuangchen Li, Peiqi Wang 0001, Xing Hu 0001, Youhui Zhang, Yuan Xie 0001 |
ASPLOS | 7 |
| 2019 | Design Guidelines of RRAM based Neural-Processing-Unit: A Joint Device-Circuit-Algorithm AnalysisabstractRRAM based neural-processing-unit (NPU) is emerging for processing general purpose machine intelligence algorithms with ultra-high energy efficiency, while the imperfections of the analog devices and cross-point arrays make the practical application more complicated. In order to improve accuracy and robustness of the NPU, device-circuit-algorithm codesign with consideration of underlying device and array characteristics should outperform the optimization of individual device or algorithm. In this work, we provide a joint device-circuit-algorithm analysis and propose the corresponding design guidelines. Key innovations include: 1) An end-to-end simulator for RRAM NPU is developed with an integrated framework from device to algorithm. 2) The complete design of circuit and architecture for RRAM NPU is provided to make the analysis much close to the real prototype. 3) A large-scale neural network as well as other general-purpose networks are processed for the study of device-circuit interaction. 4) Accuracy loss from non-idealities of RRAM, such as I-V nonlinearity, noises of analog resistance levels, voltage-drop for interconnect, ADC/DAC precision, are evaluated for the NPU design. Xiaochen Peng, Huaqiang Wu, Bin Gao 0006, Hu He 0001, Youhui Zhang, Shimeng Yu, He Qian |
DAC | 6 |
| 2018 | Bridge the Gap between Neural Networks and Neuromorphic Hardware with a Neural Network CompilerabstractDifferent from developing neural networks (NNs) for general-purpose processors, the development for NN chips usually faces with some hardware-specific restrictions, such as limited precision of network signals and parameters, constrained computation scale, and limited types of non-linear functions. This paper proposes a general methodology to address the challenges. We decouple the NN applications from the target hardware by introducing a compiler that can transform an existing trained, unrestricted NN into an equivalent network that meets the given hardware's constraints. We propose multiple techniques to make the transformation adaptable to different kinds of NN chips, and reliable for restrict hardware constraints. We have built such a software tool that supports both spiking neural networks (SNNs) and traditional artificial neural networks (ANNs). We have demonstrated its effectiveness with a fabricated neuromorphic chip and a processing-in-memory (PIM) design. Tests show that the inference error caused by this solution is insignificant and the transformation time is much shorter than the retraining time. Also, we have studied the parameter-sensitivity evaluations to explore the tradeoffs between network error and resource utilization for different transformation strategies, which could provide insights for co-design optimization of neuromorphic hardware and software. Yu Ji 0002, Youhui Zhang, Yuan Xie 0001 |
ASPLOS | 2 |
| 2018 | TETRIS: TilE-matching the TRemendous Irregular SparsityabstractCompressing neural networks by pruning weights with small magnitudes can significantly reduce the computation and storage cost. Although pruning makes the model smaller, it is difficult to get practical speedup in modern computing platforms such as CPU and GPU due to the irregularity. Structural pruning has attract a lot of research interest to make sparsity hardware-friendly. Increasing the sparsity granularity can lead to better hardware utilization, but it will compromise the sparsity for maintaining accuracy. In this work, we propose a novel method, TETRIS, to achieve both better hardware utilization and higher sparsity. Just like a tile-matching game, we cluster the irregularly distributed weights with small value into structured groups by reordering the input/output dimension and structurally prune them. Results show that it can achieve comparable sparsity with the irregular element-wise pruning and demonstrate negligible accuracy loss. The experiments also shows ideal speedup, which is proportional to the sparsity, on GPU platforms. Our proposed method provides a new solution toward algorithm and architecture co-optimization for accuracy-efficiency trade-off. Yu Ji 0002, Ling Liang 0003, Lei Deng 0003, Youyang Zhang, Youhui Zhang, Yuan Xie 0001 |
NeurIPS | 5 |
| 2017 | POSTER: Bridge the Gap Between Neural Networks and Neuromorphic HardwareabstractDifferent from training common neural networks (NNs) for inference on general-purpose processors, the development of NNs for neuromorphic chips is usually faced with a number of hardware-specific restrictions. This paper proposes a systematic methodology to address the challenge. It can transform an existing trained, unrestricted NN (usually for software execution substrate) into an equivalent network that meets the given hardware constraints, which decouples NN applications from target hardware. We have built such a software tool that supports both spiking neural networks (SNNs) and traditional artificial neural networks (ANNs). Its effectiveness has been demonstrated with a real neuromorphic chip and a processor-in-memory(PIM) design. Tests show that the extra inference error caused by this solution is very limited and the transformation time is much less than the retraining time. Yu Ji 0002, Youhui Zhang, Yuan Xie 0001 |
PACT | 2 |
| 2017 | Parallel Turing Machine, a Proposal
Youhui Zhang, Guang R. Gao |
J. Comput. Sci. Technol. | 3 |
| 2017 | Service-Oriented Architecture on FPGA-Based MPSoCabstractThe integration of software services-oriented architecture (SOA) and hardware multiprocessor system-on-chip (MPSoC) has been pursued for several years. However, designing and implementing a service-oriented system for diverse applications on a single chip has posed significant challenges due to the heterogeneous architectures, programming interfaces, and software tool chains. To solve the problem, this paper proposes SoSoC, a service-oriented system-on-chip framework that integrates both embedded processors and software defined hardware accelerators s as computing services on a single chip. Modeling and realizing the SOA design principles, SoSoC provides well-defined programming interfaces for programmers to utilize diverse computing resources efficiently. Furthermore, SoSoC can provide task level parallelization and significant speedup to MPSoC chip design paradigms by providing out-of-order execution scheme with hardware accelerators. To evaluate the performance of SoSoC, we implemented a hardware prototype on Xilinx Virtex5 FPGA board with EEMBC benchmarks. Experimental results demonstrate that the service componentization over original version is less than 3 percent, while the speedup for typical software Benchmarks is up to 372x. To show the portability of SoSoC, we implement the convolutional neural network as a case study on both Xilinx Zynq and Altera DE5 FPGA boards. Results show the SoSoC outperforms state-of-the-art literature with great flexibility. Chao Wang 0003, Xi Li 0003, Yunji Chen, Youhui Zhang, Oliver Diessel, Xuehai Zhou |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2016 | Neural network transformation under hardware constraintsabstractThere are a number of mature ways to train various kinds of ANNs (artificial neural networks), including the BP (back propagation) based algorithm and so on. These training procedures are usually carried out on some GPU-enabled machine(s); 16-/32-bit-width floating point numbers are used as the NN parameters, without any limitation on the maximum fan-in/fan-out of a single neuron or on the type of activation functions. In contrast, for neuromorphic chips [1][2][3], quite a few hardware-specific constraints (the limited fan-in/fan-out of a single neuron, the limited range of synaptic weights, and the hardware types of neurons or activation functions are usually simpler than the software counterparts) do exist, which makes programming such chips difficult. Youhui Zhang, Yu Ji 0002, Yuan Xie 0001 |
CASES | 1 |
| 2016 | Optimized Mapping Spiking Neural Networks onto Network-on-Chip
Yu Ji 0002, Youhui Zhang |
ICA3PP | 2 |
| 2016 | NEUTRAMS: Neural network transformation and co-design under neuromorphic hardware constraintsabstractWith the recent reincarnations of neuromorphic computing comes the promise of a new computing paradigm, with a focus on the design and fabrication of neuromorphic chips. A key challenge in design, however, is that programming such chips is difficult. This paper proposes a systematic methodology with a set of tools to address this challenge. The proposed toolset is called NEUTRAMS (Neural network Transformation, Mapping and Simulation), and includes three key components: a neural network (NN) transformation algorithm, a configurable clock-driven simulator of neuromorphic chips and an optimized runtime tool that maps NNs onto the target hardware for better resource utilization. To address the challenges of hardware constraints on implementing NN models (such as the maximum fan-in/fan-out of a single neuron, limited precision, and various neuron models), the transformation algorithm divides an existing NN into a set of simple network units and retrains each unit iteratively, to transform the original one into its counterpart under such constraints. It can support both spiking neural networks (SNNs) and traditional artificial neural networks (ANNs), including convolutional neural networks (CNNs) and multilayer perceptrons (MLPs) and recurrent neural networks (RNNs). With the combination of these tools, we have explored the hardware/software co-design space of the correlation between network error-rates and hardware constraints and consumptions. Doing so provides insights which can support the design of future neuromorphic architectures. The usefulness of such a toolset has been demonstrated with two different designs: a real Complementary Metal-Oxide-Semiconductor (CMOS) neuromorphic chip for both SNNs and ANNs and a processing-in-memory architecture design for ANNs. Yu Ji 0002, Youhui Zhang, Shuangchen Li, Ping Chi, Cihang Jiang, Yuan Xie 0001 |
MICRO | 2 |
| 2016 | Modelling Spiking Neural Network from the Architecture Evaluation Perspective
Yu Ji 0002, Youhui Zhang |
J. Comput. Sci. Technol. | 2 |
| 2016 | A Cloud Gaming System Based on User-Level Virtualization and Its Resource SchedulingabstractMany believe the future of gaming lies in the cloud, namely Cloud Gaming, which renders an interactive gaming application in the cloud and streams the scenes as a video sequence to the player over Internet. This paper proposes GCloud, a GPU/CPU hybrid cluster for cloud gaming based on the user-level virtualization technology. Specially, we present a performance model to analyze the server-capacity and games' resource-consumptions, which categorizes games into two types: CPU-critical and memory-of-critical. Consequently, several scheduling strategies have been proposed to improve the resource-utilization and compared with others. Simulation tests show that both of the First-Fit-like and the Best-Fit-like strategies outperform the other(s); especially they are near optimal in the batch processing mode. Other test results indicate that GCloud is efficient: An off-the-shelf PC can support five high-end video-games run at the same time. In addition, the average per-frame processing delay is 8~19 ms under different image-resolutions, which outperforms other similar solutions. Youhui Zhang, Cihang Jiang |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2015 | Software-Based Lightweight Multithreading to Overlap Memory-Access Latencies of Commodity ProcessorsabstractEmerging services applications operate on vast datasets that are kept in DRAM to minimize latency and to improve throughput. A considerable part of them have irregular memory references and then caused the serious locality issue. This paper presents a Software-based LIght weight Multithreading framework, SLIM, to conquer this problem for commodity hardware, which still keeps the simple style of multithreading programming. The principle is fairly straight: as issuing an irregular memory reference, the current fine-granularity thread uses some primitive of asynchronous memory-accesses and then switches itself out for others' execution to overlap long memory-latencies. Meanwhile, SLIM tries to maintain most contents of thread-contexts in the on-chip cache to reduce cache-misses. Therefore, the main challenge lies in how to improve the cache behavior at the expense of more instructions involved for context-switches and smaller cache-space left for applications. Consequently, we have proposed a corresponding performance model to guide the design, which is also verified by tests. Moreover, an optimized synchronization mechanism has been designed. For some classic irregular application, excessive tests have been carried out to explore the effects on performance of system configurations, including the aggressiveness of data-pre-fetch, the distribution of tasks among cores / CPUs, etc. Results show that it can achieve higher performance than the counterpart using traditional threads, under different data scales. Even compared to some tricky codes with manual optimizations, its performance is comparable and it has still reserved the simple programing manner of high-concurrency applications. Cihang Jiang, Youhui Zhang |
ICPP | 2 |
| 2015 | Solving the Global Atmospheric Equations through Heterogeneous Reconfigurable PlatformsabstractOne of the most essential and challenging components in climate modeling is the atmospheric model. To solve multiphysical atmospheric equations, developers have to face extremely complex stencil kernels that are costly in terms of both computing and memory resources. This article aims to accelerate the solution of global shallow water equations (SWEs), which is one of the most essential equation sets describing atmospheric dynamics. We first design a hybrid methodology that employs both the host CPU cores and the field-programmable gate array (FPGA) accelerators to work in parallel. Through a careful adjustment of the computational domains, we achieve a balanced resource utilization and a further improvement of the overall performance. By decomposing the resource-demanding SWE kernel, we manage to map the double-precision algorithm into three FPGAs. Moreover, by using fixed-point and reduced-precision floating point arithmetic, we manage to build a fully pipelined mixed-precision design on a single FPGA, which can perform 428 floating-point and 235 fixed-point operations per cycle. The mixed-precision design with four FPGAs running together can achieve a speedup of 20 over a fully optimized design on a CPU rack with two eight-core processorsand is 8 times faster than the fully optimized Kepler GPU design. As for power efficiency, the mixed-precision design with four FPGAs is 10 times more power efficient than a Tianhe-1A supercomputer node. Lin Gan 0001, Haohuan Fu, Wayne Luk, Chao Yang 0002, Wei Xue 0003, Xiaomeng Huang, Youhui Zhang, Guangwen Yang 0002 |
ACM Trans. Reconfigurable Technol. Syst. | 7 |
| 2014 | An approach of processor core customization for stencil computationabstractArchitecture customization is believed as one of the most promising methods to meet ever-increasing computing needs and power density limitations. This paper presents an approach to enhance a preliminary customizable core with some common architecture features, to adapt to the specific applications while keeping the programming flexibility. Those features include several effective software/hardware co-optimizing strategies, such as loop tiling, pre-fetching, cache customization, customized Single Instruction Multiple Data (SIMD) and Direct Memory Access (DMA), as well as the necessary ISA extensions. Currently we select stencil computation as the research target. Detailed tests of power-efficiency to evaluate the effect of all these optimizations comprehensively shows impressive performance speedup and power efficiency, even compared to X86, GPU and FPGA platforms. All these proposed customizations here could be applied to other computing applications. Youhui Zhang, Wayne Luk, Guangwen Yang 0002 |
ASAP | 2 |
| 2014 | Customized Network-on-Chip for Message Reduction
Youhui Zhang |
ICA3PP (1) | 3 |
| 2013 | Accelerating solvers for global atmospheric equations through mixed-precision data flow engineabstractOne of the most essential and challenging components in a climate system model is the atmospheric model. To solve the multi-physical atmospheric equations, developers have to face extremely complex stencil kernels. In this paper, we propose a hybrid CPU-FPGA algorithm that applies single and multiple FPGAs to compute the upwind stencil for the global shallow water equations. Through mixed-precision arithmetic, we manage to build a fully pipelined upwind stencil design on a single FPGA, which can perform 428 floating-point and 235 fixed-point operations per cycle. The CPU-FPGA algorithm using one Virtex-6 FPGA provides 100 times speedup over a 6-core CPU and 4 times speedup over a hybrid node with 12 CPU cores and a Fermi GPU card. The algorithm using four FPGAs provides 330 times speedup over a 6-core CPU; it is also 14 times faster and 9 times more power efficient than the hybrid CPU-GPU node. Lin Gan 0001, Haohuan Fu, Wayne Luk, Chao Yang 0002, Wei Xue 0003, Xiaomeng Huang, Youhui Zhang, Guangwen Yang 0002 |
FPL | 7 |
| 2013 | Cache Optimizations of Distributed Storage for Software Streaming Services
Youhui Zhang |
ICA3PP (1) | 1 |
| 2013 | Aegis: partitioning data block for efficient recovery of stuck-at-faults in phase change memoryabstractWhile Phase Change Memory (PCM) holds a great promise as a complement or even replacement of DRAM-based memory and flash-based storage, it must effectively overcome its limit on write endurance to be a reliable device for an extended period of intensive use. The limited write endurance can lead to permanent stuck-at faults after a certain number of writes, which causes some memory cells permanently stuck at either '0' or '1'. State-of-the-art solutions apply a bit inversion technique on selected bit groups of a data block after its partitioning. The effectiveness of this approach hinges on how a data block is partitioned into bit groups. While all existing solutions can separate faults into different groups for error correction, they are inadequate on three fundamental capabilities desired for any partition scheme. First, it can maximize probability of successfully re-partitioning a block so that two faults currently in the same group are placed into two new groups. Second, it can partition a block into a small number of groups for space efficiency. Third, it should spread out faults across the groups as uniformly as possible, so that more faults can be accommodated within the same number of groups. A recovery solution with these capabilities can provide strong fault tolerance with minimal overhead. Jie Fan 0004, Song Jiang 0001, Jiwu Shu, Youhui Zhang, Weimin Zhen |
MICRO | 4 |
| 2013 | Software/Hardware Hybrid Network-on-Chip Simulation on FPGA
Youhui Zhang, Ziqiang Qian |
NPC | 1 |
| 2013 | Employing intelligence in object-based storage devices to provide attribute-based file access
Youhui Zhang, Ziqiang Qian |
Sci. China Inf. Sci. | 1 |
| 2013 | Automatic software deployment using user-level virtualization for cloud-computing
Youhui Zhang |
Future Gener. Comput. Syst. | 1 |
| 2011 | A user-space file system for on-demand legacy desktop software
Youhui Zhang, Gelin Su |
Sci. China Inf. Sci. | 1 |
| 2010 | Efficient Monte Carlo-based options pricing on graphics processors and its optimizations
Li Liu 0020, Youhui Zhang, Li Liu 0021 |
Sci. China Inf. Sci. | 2 |
| 2010 | Converting Legacy Desktop Applications into On-Demand Personalized SoftwareabstractTo access personalized software applications on demand is an attractive usage mode of software. This paper presents such a solution based on lightweight virtualization technologies, which can convert the enormous existing desktop software into on-demand software across the Internet without any modification of source code. First, a construction and runtime model of software is proposed. It regards software as an entity containing three parts: Part1 includes all resources provided by the OS; Part2 contains what is created by the installation process; and Part3 is the data produced/modified during the runtime. Moreover, the software is executed in a lightweight virtualization environment where the APIs accessing Parts 2 and 3 are intercepted and redirected to the real storage positions (like the network and the portable storage device) as needed. In addition, a network-resource access protocol is developed for the software on demand, which implements content-addressable storage, p2p (peer-to-peer) transfer acceleration, content integrity check, and the prevention of illegal copy. From the viewpoint of the user, he/she can run his/her personalized software on any compatible computer although it does not exist on the host. Finally, its performance analysis and tests show that this proposed solution can be efficient in the performance. Youhui Zhang, Gelin Su |
IEEE Trans. Serv. Comput. | 1 |
| 2008 | IDRS: Combining File-level Intrusion Detection with Block-level Data Recovery based on iSCSIabstractOver the past years, researches on the intrusion detection have been parallelized with those on data recovery. Most of them scarcely try to combine the two issues together to propose an integrated solution, which is employed to defend the pivotal data and to recover the data when the intrusion has taken place. In this paper, we propose a framework of intrusion detection/recovery system (IDRS). This system is capable of detecting the intrusion on the file-level and recovering data on the block-level. Its advantages include that the file-level detection simplifies the implementation and the recovery based on the block-level can decrease the recovery time and raise the utilization ratio of storage devices. Again, considering that iSCSI has increasingly played an important role in network storage systems, we implement the IDRS prototype based on this promising protocol. The result of tests shows the extra storage overheads, arising from IDRS, are small (less than 10%). Therefore, we believe it is feasible to deploy IRDS on iSCSI systems to protect key files from damage. Youhui Zhang, Yu Gu 0005, Dongsheng Wang 0002 |
ARES | 1 |
| 2008 | Portable Desktop Applications Based on P2P Transportation and Virtualization
Youhui Zhang |
LISA | 1 |
| 2006 | Virtual-Machine-based Intrusion Detection on File-aware Block Level StorageabstractIn this paper we present a storage-based intrusion detection system (IDS) that makes use of advantages of virtual machine (VM) and smart disk technologies. The virtual machine monitor (VMM) can prevent the IDS itself from potential attacks while the smart disk technology provides IDS with a whole view of the file system of the monitored VM. We show how to use a tool and some file system knowledge to enable the virtual disk to maintain a sector-to-file mapping table (called file-aware block level storage) as well as how to detect the changes to file content on-line. Based on these features, normal file-level intrusion detection (ID) rules can be converted to sector-level ones in order to integrate ID functions to the virtual storage. We implement such a prototype based on QEMU VMM and the OS of VM is Windows XP. Moreover the time overhead introduced by this solution is tested Youhui Zhang, Yu Gu 0005, Dongsheng Wang 0002 |
SBAC-PAD | 1 |
| 2004 | Parallel Checkpoint/Recovery on Cluster of IA-64 Computers
Youhui Zhang, Dongsheng Wang 0002 |
ISPA | 1 |
| 2004 | The Flexible Replication Method in an Object-Oriented Data Storage System
Youhui Zhang, Jinfeng Hu |
NPC | 1 |