EDBT 2026 Demo / reviewers in the wild / expert
Chongxi Wang
dblp:254/8747
· DBLP profile ↗
5ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CUTE-XS: Bottleneck Analysis and Collaborative Optimization for Integrating CUTE Into XiangShan
Junyu Yue, Chongxi Wang, Jinpeng Ye, Jianan Xie, Longbing Zhang, Fuxin Zhang |
APPT | 2 |
| 2025 | ACRS: Adjacent Computation Resource Sharing among Partitioned GPU Sub-CoresabstractModern GPUs typically segment Streaming Multiprocessors (SMs) into sub-cores (e.g. 4 sub-cores) to reduce power consumption and chip area. However, this partitioned design prevents potential task distributions across sub-cores, impairing overall execution efficiency. In this paper, we explore the performance benefit of sharing hardware resources among sub-cores and identify functional units (FUs) as critical components for compute-intensive applications. Moreover, our observations reveal that instructions residing in operand collectors can be obstructed by back-end FUs, but there is a high probability that unoccupied FUs are available in adjacent sub-cores during such blockages. In response, we introduce the adjacent computation resource sharing (ACRS) framework to efficiently utilize these unoccupied units among sub-cores. ACRS has two key modules: Shared FU Issue (SF_ISSUE) and Shared FU Write Back (SF_WriteBack). SF_ISSUE monitors the status of operand collectors and functional units, and offloads instructions from blocked sub-cores to unoccupied resources. Meanwhile, SF_WriteBack routes results back to the original sub-core.To minimize wiring overhead, each sub-core is assigned a fixed target core for sharing. We design a series of matching policies and finally filter out the most effective sequential method. Evaluation results show that ACRS improves performance by up to $46.4 \%$, with an average of $14.1 \%$ over the traditional partitioned architecture, while reducing energy consumption by $8.3 \%$. Besides, ACRS achieves an additional 12.3% performance improvement compared with the SOTA method. Penghao Song, Chongxi Wang, Chenji Han |
DAC | 2 |
| 2025 | The Expressibility of Polynomial based Attention SchemeabstractLarge language models (LLMs) have significantly improved various aspects of our daily lives. These models have impacted numerous domains, from healthcare to education, enhancing productivity, decision-making processes, and accessibility. However, the quadratic complexity of attention in transformer architectures makes it impractical to train very large models on lengthy texts or use them efficiently during inference. While a recent study by [Kacham, Mirrokni and Zhong 2023] introduced a technique that replaces the softmax with a polynomial function and polynomial sketching to speed up attention mechanisms, the theoretical understandings of this new approach are not yet well understood. In this paper, we offer a theoretical analysis of the expressive capabilities of polynomial attention. Our study reveals a disparity in the ability of high-degree and low-degree polynomial attention. Specifically, we construct two carefully designed datasets, namely D 0 and D 1, where D 1 includes a feature with a significantly larger value compared to D 0. We demonstrate that with a sufficiently high degree β, a single-layer polynomial attention network can distinguish between D 0 and D 1. However, with a low degree β, the network cannot effectively separate the two datasets. This analysis underscores the greater effectiveness of high-degree polynomials in amplifying large values and distinguishing between datasets. Our analysis offers insight into the representational capacity of polynomial attention and provides a rationale for incorporating higher-degree polynomials in attention mechanisms to capture intricate linguistic correlations. Zhao Song 0002, Chongxi Wang, Guangyi Xu 0002, Junze Yin |
KDD (2) | 2 |
| 2024 | High-Utilization GPGPU Design for Accelerating GEMM Workloads: An Incremental ApproachabstractGeneral Purpose Graphics Processing Units (GPGPUs) have been employed primarily in domains such as graphics acceleration and high-performance computing in the past. However, the rise of artificial intelligence (AI), particularly the computational demands associated with matrix multiplications in AI models, has presented formidable demands on the computational power of GPGPUs. Consequently, the design of matrix multiplication units within GPGPUs and ensuring their utilization have become key issues in optimizing AI workloads. This paper explores an incremental design approach, building upon a Single Instruction Multiple Threads (SIMT) GPGPU architecture, to facilitate General Matrix Multiply (GEMM) acceleration. This approach encompasses not only the design of matrix units within the stream processors but, more crucially, the optimization of the data path within the GPGPU to maximize the utilization of the matrix units. We present a practical demonstration of our approach through the fabrication of a GPGPU on a 12 nm CMOS process node, achieving a core clock speed of 1 GHz and INT8 peak performance of 8 TOPS with memory bandwidth limited to 32 GB/s LPDDR4-4000. Notably, this design results in only a 6.57% increase in chip area compared to the original GPGPU design. In a series of fair GEMM workload tests, the GPGPU implemented in this work outperforms the recent three generations of NVIDIA GPGPUs—V100, T4, and A100—in terms of matrix unit utilization. Chongxi Wang, Penghao Song, Fuxin Zhang, Longbing Zhang |
ISCAS | 1 |
| 2023 | Randomized Testing Framework for Dissecting NVIDIA GPGPU Thread Block-To-SM Scheduling MechanismsabstractNVIDIA General Purpose Graphics Processing Units (GPGPUs) have been extensively utilized in scientific computing, artificial intelligence acceleration, and numerous other fields. The compute unified device architecture (CUDA) programming interface facilitates access to the enormous computing power of NVIDIA GPGPUs. However, the thread block-to-stream multiprocessor (SM) scheduler of NVIDIA GPGPUs has remained a black box, presenting challenges for performance modeling of CUDA programs, especially those featuring concurrent kernel execution, as well as hindering further exploration of the full potential of NVIDIA GPGPUs. To address these issues, we have devised a randomized testing framework targeting thread block-to-SM scheduling. In this framework, concurrent kernel sequences are randomly generated and assessed by a self-implemented scheduling predictor along with an actual NVIDIA GPGPU. Through iteratively scrutinizing the discrepancies between predictions and actual results, we dissect the key thread block-to-SM scheduling mechanisms intrinsic to NVIDIA GPGPUs, revealing their underlying logic and considerations. This process continuously refines the scheduling predictor until it accurately aligns with real-world NVIDIA GPGPU scheduling. Furthermore, we investigate the monitoring and allocation mechanisms for distinct resources in NVIDIA GPGPUs and SMs, elucidating the essential connection between resource management and thread block scheduling. Our comprehensive and precise analysis of Nvidia GPGPU thread block-to-SM scheduling contributes significantly to constructing more accurate simulators for GPGPU performance modeling. Moreover, it facilitates optimizing CUDA programs based on the scheduling mechanisms to exploit more parallelism among thread blocks of concurrent kernels. Chongxi Wang, Penghao Song, Fuxin Zhang, Longbing Zhang |
ICPADS | 1 |