Xi Wang 0009

dblp:08/5760-9 · DBLP profile ↗
← Back
24ranked-venue papers
10as first author
16since 2021 · last 2026
0000-0002-1998-6733ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 8 first-author · 12 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TokenPowerBench: Benchmarking the Power Consumption of LLM Inference
abstract
Large language model (LLM) services now answer billions of queries per day, and industry reports show that inference, not training, accounts for more than 90% of total power consumption. However, existing benchmarks focus on either training/fine-tuning or performance of inference and provide little support for power consumption measurement and analysis of inference. We introduce TokenPowerBench, the first lightweight and extensible benchmark designed for LLM-inference power consumption studies. The benchmark combines a declarative configuration interface covering model choice, prompt set, and inference engine, a measurement layer that captures GPU-, node-, and system-level power without specialized power meters, and a phase-aligned metrics pipeline that attributes energy to the prefill and decode stages of every request. These elements make it straightforward to explore the power consumed by an LLM inference run; furthermore, by varying batch size, context length, parallelism strategy and quantization, users can quickly assess how each setting affects joules per token and other energy-efficiency metrics. We evaluate TokenPowerBench on four of the most widely used model series (Llama, Falcon, Qwen, and Mistral). Our experiments cover from 1 billion parameters up to the frontier-scale Llama3-405B model. Furthermore, we release TokenPowerBench as open source to help users to measure power consumption, forecast operating expenses, and meet sustainability targets when deploying LLM services.
Chenxu Niu 0001, Wei Zhang 0097, Jie Li 0057, Tongyang Wang, Xi Wang 0009, Yong Chen 0001
AAAI6
2026 FIXME: Towards End-to-End Benchmarking of LLM-Aided Design Verification
abstract
Despite the transformative potential of Large Language Models (LLMs) in hardware design, a comprehensive evaluation of their capabilities in design verification remains underexplored. Current efforts predominantly focus on RTL generation and basic debugging, overlooking the critical domain of functional verification, which is the primary bottleneck in modern design methodologies due to the rapid escalation of hardware complexity. We present FIXME, the first end-to-end, multi-model, and open-source evaluation framework for assessing LLM performance in hardware functional verification (FV) to address this crucial gap. FIXME introduces a structured three-level difficulty hierarchy spanning six verification sub-domains and 180 diverse tasks, enabling in-depth analysis across the design lifecycle. Leveraging a collaborative AI-human approach, we construct a high-quality dataset using 100% silicon-proven designs, ensuring comprehensive coverage of real-world challenges. Furthermore, we enhance the functional coverage by 45.57% through expert-guided optimization. By rigorously evaluating state-of-the-art LLMs such as GPT-4, Claude3, and LlaMA3, we identify key areas for improvement and outline promising research directions to unlock the full potential of LLM-driven automation in hardware design verification. The benchmark is available at https://github.com/ChatDesignVerification/FIXME.
Gwok-Waa Wan, Sam-Zaak Wong, Shengchu Su, Chenxu Niu 0001, Ning Wang 0071, Xinlai Wan, Qixiang Chen, Mengnv Xing, Jianmin Ye, Rongchang Song, Qiang Xu 0001, Nan Guan, Zhe Jiang 0004, Xi Wang 0009, Yong Chen 0001, Jun Yang 0006
AAAI17
2026 ChipMind: Retrieval-Augmented Reasoning for Long-Context Circuit Design Specifications
abstract
While Large Language Models (LLMs) demonstrate immense potential for automating integrated circuit (IC) development, their practical deployment is fundamentally limited by restricted context windows. Existing context-extension methods struggle to achieve effective semantic modeling and thorough multi-hop reasoning over extensive, intricate circuit specifications. To address this, we introduce ChipMind, a novel knowledge graph-augmented reasoning framework specifically designed for lengthy IC specifications. ChipMind first transforms circuit specifications into a domain-specific knowledge graph (ChipKG) through the Circuit Semantic-Aware Knowledge Graph Construction methodology. It then leverages the ChipKG-Augmented Reasoning mechanism, combining information-theoretic adaptive retrieval to dynamically trace logical dependencies with intent-aware semantic filtering to prune irrelevant noise, effectively balancing retrieval completeness and precision. Evaluated on an industrial-scale specification reasoning benchmark, ChipMind significantly outperforms state-of-the-art baselines, achieving an average improvement of 34.59% (up to 72.73%). Our framework bridges a critical gap between academic research and practical industrial deployment of LLM-aided Hardware Design (LAD).
Changwen Xing, Sam-Zaak Wong, Xinlai Wan, Mengli Zhang, Zebin Ma, Lei Qi 0001, Zhengxiong Li, Nan Guan, Zhe Jiang 0004, Xi Wang 0009, Jun Yang 0006
AAAI11
2026 ChatHLS: Towards Systematic Design Automation and Optimization for High-Level Synthesis
abstract
Runkai Li, Jia Xiong, Xiuyuan He, Jieru Zhao, Jiaqi Lv, Haowen Fang, Lei Qi, Xi Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Runkai Li, Jia Xiong, Xiuyuan He, Jieru Zhao, Haowen Fang, Lei Qi 0001, Xi Wang 0009
ACL (1)8
2026 ChatTest: Coverage-Enhanced Testbench Generation for Agile Hardware Verification with LLMs
abstract
The growing complexity of modern hardware designs has rendered traditional functional verification increasingly time-consuming, with verification costs now dominating the design cycle. While large language models (LLMs) show promise in automating testbench generation, existing approaches struggle with real-world scalability, suffering from poor comprehension of long specifications and complex designs. To address these challenges, we propose ChatTest, a novel, end-to-end, multi-agent LLM framework for coverage-aware, agile hardware verification. Our key innovation lies in a function-mapped, divide-and-conquer architecture that integrates a Verification Description Language (VDL)—a structured, LLM-friendly DSL for precise specification encoding—with Constraint-Aware Segmental Adaptation (CASA) to enable coherent processing of long, heterogeneous design documents. By leveraging retrieval-augmented generation and supervised fine-tuning using multi-hierarchical specification-code alignment, ChatTest ensures accurate translation of functional points into targeted test stimuli. Furthermore, we introduce a coverage-driven feedback loop for automated test augmentation. Evaluated on a new benchmark of 20 complex RTL designs (up to 31K tokens of specification and 4K line-of-code), ChatTest achieves 1.46× higher toggle coverage and 2.28× higher line coverage than SOTA, with a 24.23% improvement in functional coverage, demonstrating its effectiveness in accelerating verification convergence.
Gwok-Waa Wan, Shengchu Su, Sam-Zaak Wong, Mengnv Xing, Zhe Jiang 0004, Xi Wang 0009, Jun Yang 0006
DATE8
2025 UVLLM: An Automated Universal RTL Verification Framework using LLMs
abstract
Verifying hardware designs in embedded systems is crucial but often labor-intensive and time-consuming. While existing solutions have improved automation, they frequently rely on unrealistic assumptions. To address these challenges, we introduce a novel framework, UVLLM, which combines Large Language Models (LLMs) with the Universal Verification Methodology (UVM) to relax these assumptions. UVLLM significantly enhances the automation of testing and repairing error-prone Register Transfer Level (RTL) codes, a critical aspect of verification development. Unlike existing methods, UVLLM ensures that all errors are triggered during verification, achieving a syntax error fix rate of 86.99% and a functional error fix rate of 71.92% on our proposed benchmark. These results demonstrate a substantial improvement in verification efficiency. Additionally, our study highlights the current limitations of LLM applications, particularly their reliance on extensive training data. We emphasize the transformative potential of LLMs in hardware design verification and suggest promising directions for future research in AI-driven hardware design methodologies. The Repo. of dataset and code: https://github.com/SEU-ACAL/reproduce-UVLLM-DAC-25/.
Junhao Ye, Xinyao Jiao, Dingrong Pan, Jie Zhou 0001, Ning Wang 0071, Weiwei Shan, Xinwei Fang, Xi Wang 0009, Nan Guan, Zhe Jiang 0004
DAC12
2025 ReChisel: Effective Automatic Chisel Code Generation by LLM with Reflection
abstract
Coding with hardware description languages (HDLs) such as Verilog is a time-intensive and laborious task. With the rapid advancement of large language models (LLMs), there is increasing interest in applying LLMs to assist with HDL coding. Recent efforts have demonstrated the potential of LLMs in translating natural language to traditional HDL Verilog. Chisel, a next-generation HDL based on Scala, introduces higher-level abstractions, facilitating more concise, maintainable, and scalable hardware designs. However, the potential of using LLMs for Chisel code generation remains largely unexplored. This work proposes ReChisel, an LLM-based agentic system designed to enhance the effectiveness of Chisel code generation. ReChisel incorporates a reflection mechanism to iteratively refine the quality of generated code using feedback from compilation and simulation processes, and introduces an escape mechanism to break free from non-progress loops. Experiments demonstrate that ReChisel significantly improves the success rate of Chisel code generation, achieving performance comparable to state-of-the-art LLM-based agentic systems for Verilog code generation.
Juxin Niu, Xiangfeng Liu, Dan Niu, Xi Wang 0009, Zhe Jiang 0004, Nan Guan
DAC4
2025 Location is Key: Leveraging LLM for Functional Bug Localization in Verilog Design
abstract
In Verilog code design, identifying and locating functional bugs is an important yet challenging task. Existing automatic bug localization methods have limited capabilities; they only suggest a set of potential buggy lines rather than precisely identifying the bug. Moreover, they depend on verification tools like testbenches and reference models, which require expert input and are time-consuming to develop. This paper introduces LiK (Location is Key), an open-source Large Language Model (LLM) to precisely locate functional bugs in Verilog code without the need for expert-written verification tools. LiK is developed from the open-source coding LLM Deepseek-Coder-Lite-Base-16B through a threestep training process: continuous pre-training to enhance foundational knowledge, supervised fine-tuning to learn how to output localization results, and reinforcement learning to reduce output errors. Experiment results demonstrate that LiK achieves superior functional bug localization accuracy, outperforming both the SOTA traditional method Strider, and SOTA closed-source LLMs like GPT-o1-preview and Claude-3.5-Sonnet. Moreover, integrating LiK into the SOTA LLM-based Verilog debugging tool significantly boosts its functional bug fixing success rate from $76.47 \%$ to $90.54 \%$. This underscores LiK’s potential to enhance the performance of end-to-end automatic Verilog debugging tools.
Bingkun Yao, Ning Wang 0071, Jie Zhou 0001, Xi Wang 0009, Hong Gao 0001, Zhe Jiang 0004, Nan Guan
DAC4
2025 Revolution or Hype? Seeking the Limits of Large Models in Hardware Design
abstract
Recent breakthroughs in Large Language Models (LLMs) and Large Circuit Models (LCMs) have sparked excitement across the electronic design automation (EDA) community, promising a revolution in circuit design and optimization. Yet, this excitement is met with significant skepticism: Are these AI models a genuine revolution in circuit design, or a temporary wave of inflated expectations? This paper serves as a foundational text for the corresponding ICCAD 2025 panel, bringing together perspectives from leading experts in academia and industry. It critically examines the practical capabilities, fundamental limitations, and future prospects of large AI models in hardware design. The paper synthesizes the core arguments surrounding reliability, scalability, and interpretability, framing the debate on whether these models can meaningfully outperform or complement traditional EDA methods. The result is an authoritative overview offering fresh insights into one of today’s most contentious and impactful technology trends.
Qiang Xu 0001, Leon Stok, Rolf Drechsler, Xi Wang 0009, Grace Li Zhang, Igor L. Markov
ICCAD4
2025 From Concept to Practice: an Automated LLM-aided UVM Machine for RTL Verification
abstract
Verification presents a major bottleneck in Integrated Circuit (IC) development, consuming nearly 70% of the total development effort. While the Universal Verification Methodology (UVM) is widely used in industry to improve verification efficiency through structured and reusable testbenches, constructing these testbenches and generating sufficient stimuli remain challenging. These challenges arise from the considerable manual coding effort required, repetitive manual execution of multiple EDA tools, and the need for in-depth domain expertise to navigate complex designs. Here, we present UVM2, an automated verification framework that leverages Large Language Models (LLMs) to generate UVM testbenches and iteratively refine them using coverage feedback, significantly reducing manual effort while maintaining rigorous verification standards. To evaluate UVM2, we introduce a benchmark suite comprising Register Transfer Level (RTL) designs of up to 1.6K lines of code. The results show that UVM2reduces testbench setup time by up to 38.82× compared to experienced engineers, and achieve average code and function coverage of 87.44% and 89.58%, outperforming state- of-the-art solutions by 20.96% and 23.51%, respectively.
Junhao Ye, Dingrong Pan, Qichun Chen, Jie Zhou 0001, Shuai Zhao 0004, Xinwei Fang, Xi Wang 0009, Nan Guan, Zhe Jiang 0004
ICCAD9
2025 MMCircuitEval: A Comprehensive Multimodal Circuit-Focused Benchmark for Evaluating LLMs
abstract
The emergence of multimodal large language models (MLLMs) presents promising opportunities for automation and enhancement in Electronic Design Automation (EDA). However, comprehensively evaluating these models in circuit design remains challenging due to the narrow scope of existing benchmarks. To bridge this gap, we introduce MMCircuitEval, the first multimodal benchmark specifically designed to assess MLLM performance comprehensively across diverse EDA tasks. MMCircuitEval comprises 3614 meticulously curated question-answer (QA) pairs spanning digital and analog circuits across critical EDA stages—ranging from general knowledge and specifications to front-end and back-end design. Derived from textbooks, technical question banks, datasheets, and real-world documentation, each QA pair undergoes rigorous expert review for accuracy and relevance. Our benchmark uniquely categorizes questions by design stage, circuit type, tested abilities (knowledge, comprehension, reasoning, computation), and difficulty level, enabling detailed analysis of model capabilities and limitations. Extensive evaluations reveal significant performance gaps among existing LLMs, particularly in back-end design and complex computations, highlighting the critical need for targeted training datasets and modeling approaches. MMCircuitEval provides a foundational resource for advancing MLLMs in EDA, facilitating their integration into real-world circuit design workflows. Our benchmark is available at https://github.com/cure-lab/MMCircuitEval.
Chenchen Zhao 0001, Zhengyuan Shi, Xiangyu Wen 0001, Yi Liu 0081, Yunhao Zhou, Hefei Feng, Yinan Zhu, Gwok-Waa Wan, Yongqi Fu, Chujie Chen, Chenhao Xue, Ying Wang 0001, Yibo Lin, Jun Yang 0006, Ning Xu 0009, Xi Wang 0009, Qiang Xu 0001
ICCAD20
2024 ChatCPU: An Agile CPU Design and Verification Platform with LLM
abstract
The increasing complexity of semiconductor designs necessitates agile hardware development methodologies to keep pace with rapid technological advancements. Following this trend, the Large Language Models (LLMs) emerge as a potential solution, providing new opportunities in hardware design automation. However, existing LLMs exhibit challenges in HDL design and verification, especially for complicated hardware systems. Addressing this need, we introduce ChatCPU, the first end-to-end agile hardware design and verification platform with LLM. ChatCPU streamlines the ASIC design and verification process, guiding it from initial specifications to the final RTL implementations with enhanced design agility. Incorporating the LLM fine-tuning and the processor description language design for CPU design automation, ChatCPU significantly enhances the hardware design capability using LLM. Utilizing ChatCPU, we developed a 6-stage in-order RISC-V CPU prototype, achieving successful tape-out using SkyWater 130nm MPW project with Efabless, which is currently the largest CPU design generated by LLM. Our results demonstrate a remarkable improvement in CPU design efficiency, accelerating the design iteration process by an average of 3.81X, and peaking at 12X and 9.33X in HDL implementations and verification stages, respectively. The ChatCPU also enhances the design capability of LLM by 2.63X as compared to base LLama2. These advancements in ChatCPU represent a significant milestone in LLM-driven ASIC design and optimization.
Xi Wang 0009, Gwok-Waa Wan, Sam-Zaak Wong, Layton Zhang, Jianmin Ye
DAC1
2024 Jailbreaking Pre-trained Large Language Models Towards Hardware Vulnerability Insertion Ability
abstract
We introduce RTLAttack, the first prompt-based jailbreak model designed to activate the hardware attack capabilities of LLM. Unlike conventional approaches, RTLAttack combines LLM and hardware security traits, enabling models to execute sensitive tasks like hardware Trojans insertion persistently. Extensive experiments across 10 prominent LLMs, including Claude and ChatGPT, demonstrated RTLAttack efficacy, achieving an 88.90% hardware vulnerability insertion success rate. Moreover, we analyzed the integrity and usability of LLM-injected vulnerabilities, unveiling inherent attack capabilities and harmful applications in current LLMs. This study aims to foster a comprehensive understanding of LLM capabilities in LLM-aided hardware design and drive finer-grained alignment.
Gwok-Waa Wan, Sam-Zaak Wong, Xi Wang 0009
ACM Great Lakes Symposium on VLSI3
2024 MEIC: Re-thinking RTL Debug Automation using LLMs
abstract
The deployment of Large Language Models (LLMs) for code debugging (e.g., C and Python) is widespread, benefiting from their ability to understand and interpret intricate concepts. However, in the semiconductor industry, utilising LLMs to debug Register Transfer Level (RTL) code is still insufficient, largely due to the underrepre-sentation of RTL-specific data in training sets. This work introduces a novel framework, Make Each Iteration Count (MEIC), which contrasts with traditional one-shot LLM-based debugging methods that heavily rely on prompt engineering, model tuning, and model training. MEIC utilises LLMs in an iterative process to overcome the limitation of LLMs in RTL code debugging, which is suitable for identifying and correcting both syntax and function errors, while effectively managing the uncertainties inherent in LLM operations. To evaluate our framework, we provide an open-source dataset comprising 178 common RTL programming errors. The experimental results demonstrate that the proposed debugging framework achieves fix rate of 93% for syntax errors and 78% for function errors, with up to 48x speedup in debugging processes when compared with experienced engineers. The Repo. of dataset and code: https://github.com/SEU-ACAL/reproduce-MEIC-ICCAD.
Xinwei Fang, Weiwei Shan, Xi Wang 0009, Zhe Jiang 0004
ICCAD6
2021 xBGAS: A Global Address Space Extension on RISC-V for High Performance Computing
abstract
The tremendous expansion of data volume has driven the transition from monolithic architectures towards systems integrated with discrete and distributed subcomponents in modern scalable high performance computing (HPC) systems. As such, multi-layered software infrastructures have become essential to bridge the gap between heterogeneous commodity devices. However, operations across synthesized components with divergent interfaces inevitably lead to redundant software footprints and undesired latency. Therefore, a scalable and unified computing platform, capable of supporting efficient interactions between individual components, is desirable for largescale data-intensive applications. In this work, we introduce the Extended Base Global Address Space, or xBGAS, microarchitecture extension to the RISC-V instruction set architecture (ISA) for scalable high performance computing. The xBGAS extension provides native ISA-level support for direct accesses to remote shared memory by mapping remote data objects into a system's extended address space. We perform both software and hardware evaluations of the xBGAS design. The results show that xBGAS reduces instruction count generated by interprocess communication by 69.26% on average. Overall, xBGAS achieves an average performance gain of 21.96% (up to 37.29%) across the tested workloads.
Xi Wang 0009, John D. Leidel, Brody Williams, Alan Ehret, Miguel Mark, Michel A. Kinsy, Yong Chen 0001
IPDPS1
2021 HAM: Hotspot-Aware Manager for Improving Communications With 3D-Stacked Memory
abstract
Emerging High-Performance Computing (HPC) workloads, such as graph analytics, machine learning, and big data science, are data-intensive. Data-intensive workloads usually present fine-grained memory accesses with limited or no data locality, and thus incur frequent cache misses and low utilization of memory bandwidth. 3D-stacked memory devices such as Hybrid Memory Cube (HMC) and High Bandwidth Memory (HBM) can provide significantly higher bandwidth than conventional memory modules. However, the traditional interfaces and optimization methods for JEDEC DDR devices do not allow to fully exploit the potential performance of 3D-stacked memory with the massive amount of irregular memory accesses of data-intensive applications. In this article, we propose a novel Hotspot-Aware Manager (HAM) infrastructure for 3D-stacked memory devices capable of optimizing memory access streams via request aggregation, hotspot detection, and in-memory prefetching. We present the HAM design and implementation, and simulate it on a system using RISC-V embedded cores with attached HMC devices. We extensively evaluate HAM with over 12 benchmarks and applications representing diverse irregular memory access patterns. The results show that, on average, HAM reduces redundant requests by 37.51 percent and increases the prefetch buffer hit rate by 4.2 times, compared to a baseline streaming prefetcher. On the selected benchmark set, HAM provides performance gains of 21.81 percent in average (up to 34.28 percent), and power savings of 35.07 percent over a standard 3D-stacked memory.
Xi Wang 0009, Antonino Tumeo, John D. Leidel, Jie Li 0057, Yong Chen 0001
IEEE Trans. Computers1
2020 Remote Atomic Extension (RAE) for Scalable High Performance Computing
abstract
Emerging data-intensive applications such as graph analytics, machine learning, and data-driven scientific computing are driving the evolution of high-performance computing (HPC) systems from monolithic to scaled-out, heterogeneous, and complex architectures. In these systems, enormous data sets are mapped to discrete nodes to improve the performance of the system by using distributed storage and computing resources. As such, these data distributions induce frequent cross-node data transactions which challenge the performance of large-scale systems. Global atomic operations are one emerging class of the remote data operations that enable lock-free remote shared data operations. However, the cross-node read-modify-write operations consist of multiple distinct data operations and specific atomicity management, which induces a large amount of overhead. As such, these global atomic operations require an efficient communication methodology Existing advanced compo-nents, such as network interface controllers, network fabrics, network-on-chip (NoC) interconnects, are architected together to improve the system performance. However, complex software infrastructures are needed to provide integration between each discrete component. As a result, the redundant software routines across distinct devices induce a large amount of overhead that causes performance degradationIn this paper, we propose a remote atomic extension (RAE) design that provides inherent ISA-level instructions and micro-architecture support for remote atomic operations based on the RISC-V instruction set architecture (ISA). We design a toolchain and evaluate the RAE infrastructure via simulation. Our experiment results show that RAE eliminates 89.71% of the redundant software instructions used for remote atomic accesses and improves the performance by 17.61% on average (up to 23.35%), compared with the OpenSHMEM.
Xi Wang 0009, Brody Williams, John D. Leidel, Alan Ehret, Michel A. Kinsy, Yong Chen 0001
DAC1
2020 PAC: Paged Adaptive Coalescer for 3D-Stacked Memory
abstract
Many contemporary data-intensive applications exhibit irregular and highly concurrent memory access patterns and thus challenge the performance of conventional memory systems. Driven by an expanding need for high-bandwidth memory featuring low access latency, 3D-stacked memory devices, such as the Hybrid Memory Cube (HMC) and High Bandwidth Memory (HBM), were designed to provide significantly higher throughput as compared to standard JEDEC DDR devices. However, existing memory interfaces and coalescing models, designed for conventional DDR devices, are unable to fully exploit the bandwidth potential inherent in these new 3D-stacked memory devices. In order to remedy this disparity, we introduce in this work a novel paged adaptive coalescer (PAC) infrastructure with a scalable coalescing network for 3D-stacked memory. We present the design and simulated implementation of this approach on RISC-V embedded cores with attached HMC devices. We have carried out extensive evaluations and the results show that the proposed PAC methodology yields an average coalescing efficiency of 56.01%. Further, our evaluation results also show that the PAC reduces bank conflicts and the power consumption by 85.16% and 59.21%, respectively. Overall, PAC achieves an average performance gain of 14.35% (and up to 26.06%) across 14 test suites. These results showcase the potential of the PAC methodology as applied to architecture design for increasingly critical data-intensive algorithms and applications.
Xi Wang 0009, John D. Leidel, Brody Williams, Yong Chen 0001
HPDC1
2019 POSTER: Memory Hotspot Optimization for Data-Intensive Applications
abstract
Emerging High-Performance Computing (HPC) workloads, such as graph analytics, machine learning, big data science, are data-intensive. The data-intensive workloads usually present irregular memory footprints with limited data locality, and thus incur frequent cache misses and a growing desire for memory bandwidth. Driven by this need, 3D-stacked memory devices such as Hybrid Memory Cube (HMC) and High Bandwidth Memory (HBM) are introduced to yield significantly higher throughput. However, the traditional interfaces and optimization methods for JEDEC DDR devices cannot fully exploit the potential performance of 3D-stacked memory to handle massive irregular memory accesses accompanied with data-intensive applications. In this paper, we propose a novel Hotspot-Aware Manager (HAM) infrastructure for 3D-stacked memory devices that is capable of optimizing memory access streams via request aggregation, hotspot detection, prefetching, and an associated hotspot-aware page policy. We present the HAM design and simulation implementation on RISC-V embedded cores with attached HMC devices. We have conducted extensive evaluations with over 12 benchmarks and applications representing diverse irregular memory access patterns. The results reveal that HAM reduces redundant memory accesses by 37.51% and achieves a 4.19X enhancement on the prefetch buffer hit rate on average. Overall, HAM exhibits an average of 21.81% performance gain (up to 34.28%) and 35.07% power saving over the standard 3D-stacked memory.
Xi Wang 0009, Jie Li 0057, Antonino Tumeo, John D. Leidel, Yong Chen 0001
PACT1
2019 MAC: Memory Access Coalescer for 3D-Stacked Memory
abstract
Emerging data-intensive applications, such as graph analytics and data mining, exhibit irregular memory access patterns. Research has shown that with these memory-bound applications, traditional cache-based processor architectures, which exploit locality and regular patterns to mitigate the memory-wall issue, are inefficient. Meantime, novel 3D-stacked memory devices, such as Hybrid Memory Cube (HMC) and High Bandwidth Memory (HBM), promise significant increases in bandwidth that appear extremely appealing for memory-bound applications. However, conventional memory interfaces designed for cache-based architectures and JEDEC DDR devices fit poorly with the 3D-stacked memory, which leads to significant under-utilization of the promised high bandwidth.
Xi Wang 0009, Antonino Tumeo, John D. Leidel, Jie Li 0057, Yong Chen 0001
ICPP1
2018 Memory Coalescing for Hybrid Memory Cube
abstract
Arguably, many data-intensive applications pose significant challenges to conventional architectures and memory systems, especially when applications exhibit non-contiguous, irregular, and small memory access patterns. The long memory access latency can dramatically slow down the overall performance of applications. The growing desire of high memory bandwidth and low latency access stimulate the advent of novel 3D-staked memory devices such as the Hybrid Memory Cube (HMC), which provides significantly higher bandwidth compared with the conventional JEDEC DDR devices. Even though many existing studies have been devoted to achieving high bandwidth throughput of HMC, the bandwidth potential cannot be fully exploited due to the lack of highly efficient memory coalescing and interfacing methodology for HMC devices. In this research, we introduce a novel memory coalescer methodology that facilitates memory bandwidth efficiency and the overall performance through an efficient and scalable memory request coalescing interface for HMC. We present the design and implementation of this approach on RISC-V embedded cores with attached HMC devices. Our evaluation results show that the new memory coalescer eliminates 47.47% memory accesses to HMC and improves the overall performance by 13.14% on average.
Xi Wang 0009, John D. Leidel, Yong Chen 0001
ICPP1
2017 OpenSoC system architect: An open toolkit for building soft-cores on FPGAs
abstract
Given the recent difficulty in continuing the classic CMOS manufacturing density and power scaling curves, also known as Moore's Law and Dennard Scaling, respectively, we find that modern complex system architectures are increasingly relying upon accelerators in order to optimize the placement of specific computational workloads. In addition, large-scale computing infrastructures utilized in HPC, data intensive computing, and cloud computing must rely almost exclusively upon commodity device architectures provided by third-party manufacturers. The end result being a final system architecture that lacks specificity for the target software workload. At the same time, there is a trend in the FPGA space of much larger FPGAs with a lot more resources and hardened IP blocks, making this type of architecture design space exploration much easier. The OpenSoC System Architect infrastructure combines several open source design tools and methodologies into a central infrastructure for designing, developing, and verifying the necessary hardware and software modules required to implement application-specific processors for use in FPGAs. The end result is an infrastructure that permits rapid development and deployment of application-specific accelerators and softcores, including a fully functional software development tool chain.
Farzad Fatollahi-Fard, David Donofrio, John Shalf, John D. Leidel, Xi Wang 0009, Yong Chen 0001
FPL5
2009 Clustering aggregation by probability accumulation
Xi Wang 0009, Chunyu Yang 0005, Jie Zhou 0001
Pattern Recognit.1
2008 Spectral aggregation for clustering ensemble
abstract
Since a large number of clustering algorithms exist, aggregating different clustered partitions into a single consolidated one to obtain better results has become an important problem. We propose a new algorithm for clustering ensemble based on spectral clustering. We also propose a criteria along with this algorithm, for the detection of cluster numbers. Our algorithm can determine the number of clusters more accurately with less volatility, and therefore can deduce a better combined clustering result. Experimental results on both synthesis and real data-sets show the capability and robustness of our approach.
Xi Wang 0009, Chunyu Yang 0005, Jie Zhou 0001
ICPR1