EDBT 2026 Demo / reviewers in the wild / expert
Yongwen Wang
dblp:64/3980
· DBLP profile ↗
37ranked-venue papers
0as first author
27since 2021 · last 2026
0009-0008-2514-2052ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 31 · 24 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Vector Value Prediction with Element-wise Stride CompressionabstractThe increasing emphasis on vectorization and Single Instruction, Multiple Data (SIMD) processing reflects their central role in modern processors. However, as workloads in data processing, multimedia, and algorithmic operations grow in complexity, they introduce more pronounced data dependencies, leading to longer execution times compared to scalar instructions. To address these evolving challenges, we present the Vector Value TAGE predictor (VVTAGE), a novel value predictor specifically designed for vector instructions. Although value prediction has been proposed as a fundamental strategy to enhance processor performance, it has traditionally focused on predicting 64-bit scalar values to mitigate data dependencies and improve pipeline throughput. VVTAGE extends the prediction capabilities of existing scalar predictors to accommodate the wide vector registers used in contemporary processors. Our research demonstrates that VVTAGE can significantly improve processor performance, achieving performance gains of up to 20.1% and an average increase of 4.53% in the evaluated SIMD benchmarks. This innovative approach and surprising results represent a significant advancement in optimizing the performance of SIMD processors. Furthermore, to enhance the scalability of VVTAGE, we propose an element-wise stride compression method to reduce its storage overhead. Experimental results show that VVTAGE still retains 64% performance gain while reducing 15.4KB overhead. Yanmeng Huang, Ling Yang 0008, Yuanhu Cheng, Quan Deng 0003, Junbo Tie, Yongwen Wang, Hai Zhong, Libo Huang 0002 |
ACM Great Lakes Symposium on VLSI | 7 |
| 2026 | A Deadlock-Free Bridge Module for Inter-Chiplet Cache-Coherent Communication in an Open Chiplet EcosystemabstractThe envisioned open chiplet ecosystem promises significant reductions in chip design complexity and cost by enabling designers to rapidly assemble standard chiplets from diverse vendors. Constructing such an open chiplet ecosystem requires support from Network-on-Chip (NoC) routing algorithms, as integrating multiple chiplets onto an interposer can potentially lead to inter-chiplet deadlock. Prior work avoids deadlock through methods like turn restrictions, virtual channel isolation, or packet injection control, or recovers from deadlock using mechanisms such as escape channels or bubble flow control. These approaches achieve a favorable balance regarding modularity, performance, and cost. However, they still rely on designers possessing detailed knowledge of the internal NoC architecture within each chiplet. This requirement impedes the development of a truly open ecosystem, as it constrains chiplet interoperability and vendor independence. Addressing this limitation, we propose a Deadlock-Free Bridge Module (DFBM) designed to resolve interchiplet deadlock without relying on the specifics of individual chiplet NoC implementations. The DFBM infers the transmission behavior of inter-chiplet packets by analyzing the dependency relationships among coherence protocol transaction flows. It then employs a packet injection control mechanism to isolate inter- and intra-chiplet packets, thereby preventing deadlock. DFBMs can be seamlessly interconnected between arbitrary chiplets to achieve deadlock-freedom, eliminating the need for modifications to their internal NoC architectures. Experimental results demonstrate that DFBM incurs only 2.5% area overhead, while achieving a performance improvement ranging from 1% to 7%. Zhiqiang Chen 0006, Wenwen Fu, Yongwen Wang |
HPCA | 3 |
| 2026 | Revisiting Global Value Prediction: A Resurgent Complement to Local Predictors
Ling Yang 0008, Libo Huang 0002, Bingcai Sui, Sheng Ma, Yongwen Wang, Li Shen 0007, Qianming Yang, Songwen Pei |
ISCA | 6 |
| 2026 | Reducing Off-Chip Prefetch Request Latency of LLC Hardware Prefetchers via Neural PredictionabstractLLC hardware prefetchers are a pivotal technique in modern high-performance processors for hiding memory access long latency. However, even accurate prefetch requests generated by LLC prefetchers experience long latency when they must traverse the LLC before accessing off-chip main memory. Zhengwei Huang, Yongwen Wang |
SPAA | 3 |
| 2026 | Apollo: Accelerating Load Requests via Multi-Level Cache Miss Load PredictionabstractAccelerating load requests is an effective way to improve processor performance through reducing load request latency. Currently, state-of-the-art methods, including Hermes and TLP, are limited to accelerating off-chip load requests that are correctly predicted and cannot accelerate mispredicted off-chip load requests and any on-chip load requests. To overcome this limitation, we propose a new technique called Apollo that integrates TLP. The key innovation of Apollo lies in its transformation of the perceptron-based off-chip prediction paradigm, pioneered by Hermes, into a comprehensive multi-level cache miss prediction technique. The workflow of Apollo is as follows: (1) predicting whether a load request will miss the L1D or L2, and (2) performing arbitration to decide whether to issue a speculative load request and to which cache level the request is issued, and (3) issuing a speculative load request to the lower-level cache (either L2 or LLC) after arbitration for those predicted to miss the L1D or L2, while allowing the regular load request to concurrently access the cache hierarchy. If the prediction is correct, the regular load request eventually misses the L1D or L2 and waits for the speculative load request to finish. Therefore, Apollo can hide the L1D access latency for correctly predicted L1D miss load requests, and both L1D and L2 access latency for correctly predicted L2 miss load requests. To enable Apollo, we propose a lightweight L1D miss load predictor (L1MP), a lightweight L2 miss load predictor (L2MP), and an arbiter called Athena. L1MP and L2MP predict whether load requests will miss in the L1D and L2, respectively, while Athena performs arbitration to control the issuance of speculative load requests. Our evaluation using a diverse set of workloads shows that Apollo provides a geometric mean (geomean) performance improvement of 15.2% for the single-core processor, outperforming Hermes by 9.6% and TLP by 6.4%. For the multi-core processor, Apollo provides a geomean performance improvement of 18.5%, which is 16.8% higher than Hermes and 5.8% higher than TLP. Zhengwei Huang, Yongwen Wang |
ACM Trans. Archit. Code Optim. | 3 |
| 2025 | Late Breaking Results: AFS: Improving Accuracy of Quantized Mamba via Aggressive Forgetting StrategyabstractMamba overcomes the quadratic complexity problem inherent in Transformer models while maintaining comparable contextual modeling capabilities. However, Mamba-based foundation models encounter challenges in achieving efficient inference on resource-constrained devices, primarily due to their considerable size. Model compression techniques, such as linear quantization, offer a viable solution to this problem. Nevertheless, the introduction of significant outliers during Mamba's past-state forgetting process can lead to a notable decrease in accuracy when employing linear quantization. To overcome these challenges, this paper introduces the Aggressive Forgetting Strategy (AFS), an innovative and efficient algorithm designed to mitigate the quantization issues caused by outliers in the state forgetting mechanism. AFS incorporates a computation-free approach for handling outliers, facilitating both efficient and accurate linear quantization for Mamba, which is essential for applications in resource-constrained scenarios. By leveraging the AFS strategy, Mamba can perform more efficient inference, while significantly improving the accuracy by up to 21.5× compared to conventional methods. Zhouquan Liu, Libo Huang 0002, Ling Yang 0008, Gang Chen 0023, Yongwen Wang |
DATE | 7 |
| 2025 | SONet: Towards Practical Online Neural Network for Enhancing Hard-to-Predict Branches
Zhenxuan Xiong, Libo Huang 0002, Ling Yang 0008, Hui Guo 0004, Songwen Pei, Gang Chen 0023, Yongwen Wang |
Euro-Par (2) | 9 |
| 2025 | Low-Cost Approximate Floating-Point Multiplier Design Based on SSA and Sparse Processing
Gang Chen 0023, Qianmin Yang, Yongwen Wang, Libo Huang 0002 |
ICA3PP (6) | 7 |
| 2025 | Facess: A Fast Access Method for Shared Data Among Multi-cores in Parallel Programs
Jierong Tang, Yijing Peng, Quanyou Feng, Libo Huang 0002, Yongwen Wang |
ICA3PP (1) | 8 |
| 2025 | PolyPE: An Efficient Multi-Precision Multi-Mode Floating-Point Processing Element for HPC and AIabstractIn this paper, an efficient multi-precision multimode floating-point Processing Element is designed for HPCenabled AI workloads, called PolyPE, in which Poly means multiprecision multi-mode. It supports both conventional and mixedprecision FMA operations, including single-FMA, dual-FMA, and quad-FMA modes, as well as quad-FMA-add for enhanced throughput. The supported precisions include double precision, single precision, half precision, TF32, and BF16. At each clock cycle, the processing element can perform one double-precision, two single-precision, or four half-precision operations. Compared to existing designs, it offers broader precision support, including TF32 and BF16, with higher throughput and lower hardware overhead, achieving up to 5× improvement over standard FMA. We integrated the design into an open-source GPGPU and extended its instruction set. Experimental results show up to 2.17× performance gain, with 27.2% and 41.2% reductions in LUT and FF usage, respectively, while preserving functional equivalence. Zhenzhen Jia, Hongbing Tan, Ling Yang 0008, Hui Guo 0004, Junsheng Chang, Yongwen Wang, Libo Huang 0002 |
ICCD | 7 |
| 2025 | Initial-Key Cache: An Efficient KV Cache Strategy Focusing on Initial and Key Tokens for LLMsabstractLarge language models (LLMs) have developed rapidly in recent years and have demonstrated excellent performance in various application fields. Despite their outstanding performance, LLMs also introduce significant challenges in the practical inference process, mainly because of their computational and memory-intensive characteristics. Due to the autoregressive nature of the attention mechanism, KV caching can effectively accelerate the inference of LLMs by substituting quadratic-complexity computation with linear-complexity memory accesses. However, during the calculation, it is necessary to transfer the KV cache value to the computing unit, which not only requires a larger storage capacity but also places higher demands on storage bandwidth. In the process of inference and text generation for long texts, storage has become a bottleneck that limits the performance of the inference. In this paper, we first noticed the attention sink phenomenon that high attention scores are allocated towards initial tokens as "attention sink" even if they are not semantically significant in short input sentences and further observed that the attention sink tends to diminish as the length of the input sentence increases. Based on the above preliminary empirical results, we proposed Initial-key Cache, an efficient KV cache selection algorithm that not only effectively addresses the disappearance of the sink phenomenon in long inputs but also comprehensively considers the importance of middle-sentence tokens. We conducted a series of experiments on four baseline models, namely LLaMA, Qwen, Pythia, and OPT, to assess the performance of the Initial-Key Cache. The experiment results indicate that the Initial-Key Cache saves KV cache memory usage, while almost not losing the model’s accuracy. Zhongyi Tang, Zejiang He, Junzhong Shen, Yiyue Hu, Luchen Zhou, Yongzhang Nie, Yongwen Wang |
IJCNN | 7 |
| 2025 | Brief Announcement: LCTree: A Fast Hardware BVH Constructor for Real-Time Ray TracingabstractUnlike traditional rasterization rendering, ray tracing is a groundbreaking technology that has revolutionized the realistic rendering of images, marking a significant leap forward. However, achieving real-time ray tracing in dynamic scene applications remains a challenging task. This difficulty arises primarily from the substantial technical bottlenecks related to the frequent need for reconstructing or incrementally updating acceleration structures essential for efficient ray calculations. Run Yan, Su Yin, Hui Guo 0004, Yongwen Wang, Gang Chen 0023, Nong Xiao 0001, Libo Huang 0002 |
SPAA | 4 |
| 2025 | RVAM16: a low-cost multiple-ISA processor based on RISC-V and ARM Thumb
Libo Huang 0002, Ling Yang 0008, Sheng Ma, Yongwen Wang, Yuanhu Cheng |
Frontiers Comput. Sci. | 5 |
| 2025 | Optimizing value prediction for ILP processors: A design space exploration approach
Ling Yang 0008, Libo Huang 0002, Run Yan, Sheng Ma, Yongwen Wang, Weixia Xu 0001 |
Integr. | 6 |
| 2025 | Steered Bubble: An Interposer-based Deadlock Recovery Algorithm for Multi-chiplet SystemsabstractDividing a single System-on-Chip (SoC) into multiple chiplets and integrating them via an interposer can achieve an optimal balance between continuous transistor integration and monetary cost. However, potential deadlock may arise between the chiplets and the interposer. This deadlock can be avoided by applying turn restriction or injection control on the boundary routers, at the cost of additional latency and suboptimal performance. Compared to deadlock avoidance, deadlock recovery exerts less impact on network performance. Nevertheless, accurate and timely deadlock detection, along with efficient deadlock recovery, continues to pose significant challenges. Additionally, modularity is a specific concern, which involves integrating chiplets of various functions, sizes, manufacturing processes, and so on. Minimizing the negative impact of deadlock resolution while maximizing modularity is crucial for achieving the benefit of chiplets. This article proposes a modular deadlock detection strategy, Up-Down, which monitors both the upward and downward directions of vertical channels, facilitating information exchange through the congestion-sense network. When a pair of blocked upward and downward vertical channels is detected simultaneously, it is considered that an inter-chiplet deadlock has occurred. This significantly enhances the accuracy of deadlock detection by two orders of magnitude compared to time-out deadlock detection. Furthermore, this article introduces Steered Bubble, a low-cost deadlock recovery algorithm. It does so by injecting bubbles into potential deadlock cycles identified by Up-Down. These bubbles follow preset paths, ensuring efficient deadlock recovery. Experimental results indicate that the Steered Bubble results in an average performance enhancement of 1% to 10% during full-system simulations, with an area overhead of less than 2%. Zhiqiang Chen 0006, Yongwen Wang, Jian Zhang 0022 |
ACM Trans. Archit. Code Optim. | 2 |
| 2024 | Out-of-Order and Recursive RAS: A Return Address Stack Design on High Performance ProcessorabstractIn high-performance processor design, maintaining Return-Address Stack (RAS) integrity is crucial for efficient instruction flow. Yet, separating multi-level branch predictors from L1I-caches brings substantial hurdles, especially when speculative execution corrupts the RAS through out-of-order branches. Past remedies struggle with storage demands and recursive call inefficiencies. Hence, we introduce the Out-of-Order and Recursive RAS (OR-RAS), an innovative architecture. It employs a advanced RAS for out-of-order control and a compressed LUT, enhancing efficiency. OR-RAS targets IPC boost and RAS error reduction (RAS-MPKI). Evaluations on powerful processors reveal a 0.5% IPC uplift, with effectively maintaining the MPKI below 0.02%. In essence, OR-RAS constitutes a holistic strategy against speculative execution issues and RAS corruption, auguring well for peak performance and dependability in contemporary microarchitectures. Yude Fang, Libo Huang 0002, Yongwen Wang, Weixia Xu 0001 |
ASAP | 4 |
| 2024 | QuickTree: A Fast Hardware BVH Construction EngineabstractRay tracing has emerged as a powerful technique for generating visually stunning and realistic images compared to rasterization. With the continuous advancements in computer hardware, modern GPUs have integrated specialized ray tracing acceleration units to enhance rendering capabilities further. However, achieving realtime ray tracing presents a challenge in dynamic scenes, where spatial data structures used for accelerated rendering must be reconstructed or updated when there are changes in the scene primitives. This paper introduces QuickTree, a novel Bounding Volume Hierarchy (BVH) construction engine based on the linear BVH (LBVH) optimization algorithm. QuickTree addresses the challenge of dynamic scenes support by employing a highly parallel and pipelined system design. This innovative approach ensures fast construction speed. QuickTree demonstrates significant performance improvements. Compared to the currently fastest MergeTree, it has increased construction speed by 10% and reduced area by 45% compared to RayCore, which has the smallest chip area. Yin Su, Hui Guo 0004, Run Yan, Yongwen Wang, Nong Xiao 0001, Gang Chen 0023, Libo Huang 0002 |
CF | 5 |
| 2024 | ImSPU: Implicit Sharing of Computation Resources Between Vector and Scalar Processing Units
Hongbing Tan, Guichu Sun, Liquan Xiao, Yuanhu Cheng, Quan Deng 0003, Bingcai Sui, Yongwen Wang, Libo Huang 0002 |
Euro-Par (2) | 10 |
| 2024 | Cost-Effective Value Predictor for ILP processors through Design Space ExplorationabstractValue prediction is a microarchitectural technique that enhances processor performance by speculatively breaking true data dependencies. It has demonstrated improved performance in both single-threaded and multi-threaded workloads, rendering it an appealing microarchitectural approach. While high-performance value predictors can achieve impressive accuracy, they may also incur significant costs in terms of area, power consumption, and complexity. Therefore, there is a demand for lightweight value prediction techniques capable of striking a favorable balance between performance and overhead. However, designing value predictors with superior performance using limited resources presents an urgent challenge. Consequently, this work proposes a design space exploration framework for the state-of-the-art EVES value predictor, aiming to efficiently configure the design parameters of the value predictor within constrained RAM resources. Additionally, the article evaluates the performance of the explored value predictor across a wide range of workloads. The explored value predictors exhibit high efficiency across RAM sizes ranging from 2KB to 16KB while maintaining acceptable computational complexity. Furthermore, the results indicate that the explored value predictor achieves optimal efficiency under the 2KB constraint, with the highest acceleration-to-cost ratio reaching 4.02%/KB. Ling Yang 0008, Libo Huang 0002, Run Yan, Sheng Ma, Yongwen Wang, Weixia Xu 0001 |
ACM Great Lakes Symposium on VLSI | 6 |
| 2024 | SSC: An SRAM-Based Silence Computing Design for On-chip Memory
Quan Deng 0003, Yiyue Hu, Libo Huang 0002, Yongwen Wang |
ICA3PP (4) | 6 |
| 2024 | Priority-aware deadlock recovery algorithm for multi-chiplet systemsabstractIn order to tackle complex tasks, multi-chiplet systems integrate various types of chiplets such as CPUs, GPUs, and other accelerators on the active interposer. Interposer Network-on-Chip (NoC) serves as the foundational infrastructure for interconnecting chiplets. The rich and varied traffic characteristics of these chiplets lead to diverse network requirements and pose additional challenges for the design of NoC. Firstly, different chiplets exhibit unique traffic patterns and network requirements. The competition among these different types of chiplets can prevent the system from achieving optimal performance. Secondly, adaptive routing is more adept at managing the dynamic traffic patterns within heterogeneous systems than deterministic routing. However, managing deadlock is critical for network correctness. Prior strategies for deadlock resolution often fail to adequately consider packets of different priorities, potentially hindering overall performance enhancement in multi-chiplet systems. This paper introduces the Priority-Aware Deadlock Recovery Algorithm (PARA) for multi-chiplet systems. By analyzing packet distribution characteristics within a deadlock cycle, it is observed that breaking the deadlock cycle can be accomplished by permitting head-of-line (HoL) blocking packets to depart. Additionally, the Priority-aware FIFO (PFIFO) has been introduced, supporting three modes: normal, priority, and reserve. The proposed PARA achieves a balance between deadlock recovery, HoL blocking alleviation, and priority scheduling. The experimental results indicate that PARA achieves comparable performance with reduced overhead. Yongwen Wang, Mengjin Li |
ICPADS | 2 |
| 2024 | Hybrid Deadlock Recovery Algorithm for Irregular NoC in Multi-Chiplet SystemsabstractDividing a single System-on-Chip (SoC) into multiple chiplets and connecting them using 2.5D packaging technology is becoming a widely adopted approach to enhance chip scale and performance. However, when multiple chiplets are integrated to form a chiplet-based system, the Network-on-Chip (NoC) that interconnects these chiplets can be susceptible to deadlock. Additionally, modularity is a specific concern, as it involves integrating chiplets of different functions, sizes, manufacturing processes, and so on. However, physical layout constraints and potential vertical link failures may result in irregular topologies, which further complicates the design of both fault-tolerant and load-balanced routing algorithms. To tackle these challenges, a Hybrid Deadlock Recovery algorithm for Irregular NoC (HDRI) in multi-chiplet systems is proposed. This algorithm shows adaptability in irregular topologies, additionally exhibiting fault-tolerance capabilities. HDRI features two modes that can dynamically adapt to the network status. Upon detecting a deadlock, it recovers from the deadlock through the coordination of inter-chiplet packets. To be specific, HDRI seeks to break the deadlock by either forwarding or recycling the blocked packets, which respectively correspond to low and high load modes. Experimental results demonstrate that HDRI offers a 7.5% reduction in latency and an area overhead of less than 1.1%. Yongwen Wang |
ISPA | 2 |
| 2024 | A survey of compute nodes with 100 TFLOPS and beyond for supercomputers
Junsheng Chang, Kai Lu 0001, Yang Guo 0003, Yongwen Wang, Libo Huang 0002, Yao Wang 0002, Biwei Zhang |
CCF Trans. High Perform. Comput. | 4 |
| 2024 | MPRTA: An Efficient Multilevel Parallel Mobile Accelerator for High-Performance Ray TracingabstractRay tracing has been regarded as the future of graphics rendering technology for a long time. However, interactive ray tracing still faces challenges, especially in mobile devices, such as high computational intensity and multiple branches. In this brief, we aim to maximize overall efficiency by leveraging all forms of potential parallelism, including task, basic block, loop, and pipeline levels. We present multilevel parallel ray tracing accelerator (MPRTA), an innovative mobile accelerator that offers high performance and optimal efficiency for ray tracing. Experimental results indicate that MPRTA is$1.67\times $more efficient than the currently best-reported mobile accelerator. Run Yan, Yin Su, Hui Guo 0004, Yashuai Lü, Nong Xiao 0001, Li Shen 0007, Yongwen Wang, Libo Huang 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2023 | A Multi-level Parallel Integer/Floating-Point Arithmetic Architecture for Deep Learning Instructions
Hongbing Tan, Libo Huang 0002, Dezun Dong, Yongwen Wang, Liquan Xiao |
Euro-Par | 6 |
| 2023 | Fast Approximate LUT-based Vector Multiplication in DRAMabstractVector multiplication is widely used in real-world applications. To accelerate vector multiplication, processing-in-memory-based domain-specific architectures leverage lookup tables (LUTs) to decrease the computational complexity of multi-plication. However, the overhead of LUTs increases exponentially with the result space, which throttles the system performance.To decouple the LUT size and the performance improvement, we make a trade-off between accuracy and performance. We propose fast approximate LUT-based vector multiplication in DRAM, which builds the partial LUT of higher-bit locations and reuses the LUT to calculate the lower-bit data. We propose an operand reorganizing optimization to group more zero, which does not affect the results. To reduce the LUT and pre-calculation cost further, we propose a value encoding. Our experiment shows that the performance of the proposed design can be improved by 4x and 1.25x compared with LAcc in AlexNet and MobileNetV2 without any accuracy loss, respectively. Besides, the performance of the proposed design improves by up to 3.74x compared with pLUTo in multi-vector multiplication. The area and power of the proposed design are 54.8mm2and 5.35W, respectively. Quan Deng 0003, Yongwen Wang |
ICPADS | 3 |
| 2022 | RV16: An Ultra-Low-Cost Embedded RISC-V Processor Core
Yuanhu Cheng, Libo Huang 0002, Yi-Jun Cui, Sheng Ma, Yongwen Wang, Bingcai Sui |
J. Comput. Sci. Technol. | 5 |
| 2020 | CSMO-DSE: Fast and Precise Application-driven DSE Guided by Criticality and Sensitivity AnalysisabstractDetermining the optimal microarchitecture configuration of a processor at the early stages of design is undeniably a challenge. Due to many parameters at the microarchitecture level, finding the proper combination of these parameters to arrive at a balanced design is difficult. Application-specific Design Space Exploration (DSE) is even more difficult, since the property of application needs to be considered during the DSE process. Improving the speed and accuracy of the DSE process remains a particular challenge in microprocessor design. In this article, we propose a novel processor DSE methodology based on criticality and sensitivity analysis, named Criticality and Sensitivity-based Multi-Objective DSE (CSMO-DSE). In our methodology, a dependence-graph is derived from the profile generated by running a program on an instrumented cycle-accurate microprocessor simulator. Then, the criticality of the processor’s performance events is obtained through critical path analysis. The sensitivity of microarchitecture parameters to various performance events is also analyzed. Then, this information is used to optimize performance, power/area, and energy efficiency of the design. Experiments with SPEC 2006 show that CSMO-DSE methodology is 4.73× faster than the baseline DSE methodology and that the quality of result (QoR) is better than the baseline methodology for all the benchmark programs. Lei Wang 0011, Yu Deng 0001, Yongwen Wang |
ACM J. Emerg. Technol. Comput. Syst. | 6 |
| 2017 | Effective Optimization of Branch Predictors through Lightweight SimulationabstractBranch predictors are important components that affect the performance of modern microprocessors. Many of the existing prediction simulation platforms either only consider accuracy computed with coarse-grained updating model, or just are the low speed full system simulators. In this paper, We present SimpleBP, a lightweight prediction simulator based on trace driven. It leverages the SystemC language to simulate branch predictor at clock cycle granularity. Also, the CACTI tools is integrated to evaluate area and power consumption. Using SimpleBP, we first make hardware design analysis in a 64Kits storage budget with new added features(feedback delay, fetch width and RAM port number). The experimental results show that the prediction accuracy loss is less than 2% when the feedback delay increased. Under the same RAM area, when fetch width modified, various traces show different accuracy changes. When the area used for dual-port is saved to construct more entries of the subpredictors, the accuracy improvement is about 2.45%. Then, considering these new features, this paper does example optimization exploration of RAM size and subpredictors number through SimpleBP, and the results show many meaningful differences with the results simulated through other simulation platforms. Chaobing Zhou, Libo Huang 0002, Tan Zhang, Yongwen Wang, Qiang Dou |
ICCD | 4 |
| 2015 | Fast FPGA system for microarchitecture optimization on synthesizable modern processor designabstractMicroarchitecture optimization for processor design is a must to achieve target system performance. Provided the register transfer level (RTL) model in real chip design, this paper proposes MOFPGA system, which uses field programmable gate array (FPGA) prototyping as an effective method for fine-grain microarchitecture optimization. It is a fast, reconfigurable, and visible platform with zero impact on the performance of the monitored processor. MOFPGA implements a complete computing platform equipped with a modern out-of-order processor and is able to achieve 60 MHz processor frequency. Besides general FPGA implementation techniques such as multi-port SRAM design and gate-clock conversion, extensive optimization efforts are done to improve the FPGA performance of mapping such a large core. To our knowledge, MOFPGA is the first published FPGA system that implements a modern out-of-order processor running at such high frequency and can report the real SPEC CPU2000 evaluation results. Libo Huang 0002, Yongwen Wang, Qiang Dou, Caixia Sun |
FPL | 2 |
| 2014 | A High-Dynamic Invocation Load Balancing Algorithm for Distributed Servers in the Cloud
Zhao Yang Qu, Jiannan Zang, Huiyu Sun, Yongwen Wang |
ICIC (1) | 5 |
| 2014 | Integrated Coherence Prediction: Towards Efficient Cache Coherence on NoC-Based Multicore ArchitecturesabstractMulticore architectures with Network-on-Chips (NoCs) have been widely recognized as the de facto design for the efficient utilization of the continuously increasing density of transistors on a chip. A key challenge in designing such an NoC-based multicore processor is maintaining cache coherence in an efficient manner. Directory-based protocols avoid the bandwidth overhead of snoop-based protocols, therefore scaling to a large number of cores. However, conventional directory structures add significant indirection delay to cache-to-cache accesses in larger multicore processor. In this article we propose a novel hardware coherence technique, called integrated coherence prediction (ICP). This approach adopts a prediction technique for managing shared data to reduce or eliminate the cache-to-cache delay in coherence accesses. ICP has two unique features that differ from previous coherence prediction techniques. First, ICP introduces a new integrated prediction scheme that combines two kinds of predictors: owner predictor, which predicts the data writers and avoids the indirection through directory, and data predictor, which predicts the access address and prefetches data from remote nodes directly. Second, ICP uses a request replication method to reduce the negative effect of wrong owner prediction operations, thus facilitating overall performance improvement. We present the design and implementation details of the ICP approach. Using detailed full-system simulations, we conclude that the ICP provides a cost-effective solution for designing high-performance multicore processors. Libo Huang 0002, Zhiying Wang 0003, Nong Xiao 0001, Yongwen Wang, Qiang Dou |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2013 | Efficient multimedia coprocessor with enhanced SIMD engines for exploiting ILP and DLP
Libo Huang 0002, Nong Xiao 0001, Zhiying Wang 0003, Yongwen Wang |
Parallel Comput. | 4 |
| 2013 | Adaptive communication mechanism for accelerating MPI functions in NoC-based multicore processorsabstractMulticore designs have emerged as the dominant organization for future high-performance microprocessors. Communication in such designs is often enabled by Networks-on-Chip (NoCs). A new trend in such architectures is to fit a Message Passing Interface (MPI) programming model on NoCs to achieve optimal parallel application performance. A key issue in designing MPI over NoCs is communication protocol, which has not been explored in previous research. This article advocates a hardware-supported communication mechanism using a protocol-adaptive approach to adjust to varying NoC configurations (e.g., number of buffers) and workload behavior (e.g., number of messages). We propose the ADaptive Communication Mechanism (ADCM), a hybrid protocol that involves behavior similar to buffered communication when sufficient buffer is available in the receiver to that similar to a synchronous protocol when buffers in the receiver are limited. ADCM adapts dynamically by deciding communication protocol on a per-request basis using a local estimate of recent buffer utilization. ADCM attempts to combine both the advantages of buffered and synchronous communication modes to achieve enhanced throughput and performance. Simulations of various workloads show that the proposed communication mechanism can be effectively used in future NoC designs. Libo Huang 0002, Zhiying Wang 0003, Nong Xiao 0001, Yongwen Wang, Qiang Dou |
ACM Trans. Archit. Code Optim. | 4 |
| 2013 | Dynamic Streamization Model Execution for SIMD Engines on Multicore ArchitecturesabstractThis paper proposes dynamic streamization model execution (DSME), a dynamic vectorization technique for single instruction multiple data (SIMD) engines on multicore architectures. The technique uses stream model as intermediate representation for programs to optimize the combination of computation and memory accesses of SIMD engines in general-purpose (GP) designs. DSME allows the dynamic placement of computations on different cores when they are not in use to utilize multiple SIMD engines. This study also discusses hardware extensions to existing GP processor designs as well as related compiler extensions that use the special hardware components. Our extensive experiments demonstrate that performance gains of DSME can be achieved. Libo Huang 0002, Zhiying Wang 0003, Nong Xiao 0001, Yongwen Wang, Qiang Dou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2011 | A Novel Chaining Approach to Indirect Control Transfer Instructions
Wei Chen 0009, Zhiying Wang 0003, Qiang Dou, Yongwen Wang |
ARES | 4 |
| 2004 | An Efficient Broadcast Algorithm Based on Connected Dominating Set in Unstructured Peer-to-Peer Network
Qianbing Zheng, Wei Peng 0005, Yongwen Wang, Xicheng Lu |
WISE | 3 |