EDBT 2026 Demo / reviewers in the wild / expert
Dong Chen 0015
dblp:44/3371-15
· DBLP profile ↗
13ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0003-3751-7997ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 5 since 2021Software engineering, systems software and programming languages · 3 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DyGen: A Constant-Time Kernel Generator for Dynamic-Shape Neural NetworksabstractIn recent years, dynamic-shape neural networks have been widely adopted in intelligent applications, such as Mixture-of-Experts based large language models and computer vision tasks. However, in dynamic scenarios, operator shapes are determined at runtime. This leads to prohibitively expensive compilation times for existing static compilers, as they must search across a vast optimization space to identify the best configuration. To address the need for efficient optimization of dynamic-shape neural networks, we present DyGen (Dynamic-shape Kernel Generator)—a lightweight, two-stage compiler plug-in on GPU platforms. In the offline stage, DyGen employs deliberately crafted pruning rules to construct a compact candidate configuration set for the target hardware, then select the configuration of the high-performance kernel to train a configuration generation model. During the online stage, dynamic operator information is directly fed into the generator, which can quickly produce efficient kernel configurations without the need for costly search. Compared to state-of-the-art tensor compilers, DyGen improves inference performance by an average of 36%, while significantly reducing generation overhead from 9 seconds to 0.3 seconds. Yuhan Kang, Dong Chen 0015, Yang Shi 0008, Jianchao Yang, Zeyu Xue, Mei Wen |
DATE | 3 |
| 2026 | Precision boundary modeling for area-efficient Block Floating Point accumulation
Xin Ju 0005, Yasong Cao, Zhongdi Luo, Jianchao Yang, Jingkui Yang, Dong Chen 0015, Mei Wen |
J. Syst. Archit. | 10 |
| 2024 | HyFiSS: A Hybrid Fidelity Stall-Aware Simulator for GPGPUsabstractThe widespread adoption of GPUs has driven the development of GPU simulators, which, in turn, lead advancements in both GPU architectures and software optimization. Trace-driven cycle-accurate Cycle-accurate simulators, which provide detailed microarchitectural models and clock-level precision, come at the cost of extended simulation times and require high computational resources. Their scalability has become a bottleneck. A growing trend is the adoption of cycle-approximate simulators, which introduce mathematical modeling of partial hardware units and utilize sampling to accelerate simulation. However, this approach faces challenges regarding the accuracy of performance predictions. To address these limitations, we introduce HyFiSS, a hybrid fidelity stall-aware GPU simulator. HyFiSS features fine-grained stall events tracking and attribution by constructing a detailed execution pipeline model for various stall events on Streaming Multiprocessors (SMs). It accurately emulates the thread block scheduler behavior using real-time scheduling logs and utilizes sampling based on thread block sets to minimize the precision loss due to fine-grained sampling points on the microarchitectural state. We achieve a balance between reliability, speed, and the level of simulation detail, especially regarding bottlenecks. By evaluating a diverse set of benchmarks, HyFiSS achieves a mean absolute percentage error in predicting active cycles that is comparable to the state-of-the-art cycle-accurate simulator Accel-Sim. Moreover, HyFiSS achieves a substantial 12.8 × speedup in the simulation efficiency compared to Accel-Sim. HyFiSS also requires at least 3.2 × less disk storage than both Accel-Sim and another state-of-the-art cycle-approximate simulator PPT-GPU due to its efficient SASS (Streaming Assembler) traces compression. With precise, per-cycle stall events statistics, HyFiSS can provide accurate GPU performance metrics and stall cause reporting. This significantly simplifies performance analysis, bottleneck identification, and performance optimization tasks for researchers, making it easier to enhance GPU performance effectively. Jianchao Yang, Mei Wen, Dong Chen 0015, Zhaoyun Chen, Zeyu Xue, Junzhong Shen, Yang Shi 0008 |
MICRO | 3 |
| 2023 | VIDGCN: Embracing input data diversity with a configurable graph convolutional network accelerator
Tingting Pan, Dong Chen 0015, Chencheng Ye 0001, Haikun Liu, Liting Tang, Xiaofei Liao, Hai Jin 0001 |
J. Syst. Archit. | 3 |
| 2022 | CARL: Compiler Assigned Reference LeasingabstractData movement is a common performance bottleneck, and its chief remedy is caching. Traditional cache management is transparent to the workload: data that should be kept in cache are determined by the recency information only, while the program information, i.e., future data reuses, is not communicated to the cache. This has changed in a new cache design named Lease Cache . The program control is passed to the lease cache by a compiler technique called Compiler Assigned Reference Lease (CARL). This technique collects the reuse interval distribution for each reference and uses it to compute and assign the lease value to each reference. In this article, we prove that CARL is optimal under certain statistical assumptions. Based on this optimality, we prove miss curve convexity, which is useful for optimizing shared cache, and sub-partitioning monotonicity, which simplifies lease compilation. We evaluate the potential using scientific kernels from PolyBench and show that compiler insertions of up to 34 leases in program code achieve similar or better cache utilization (in variable size cache) than the optimal fixed-size caching policy, which has been unattainable with automatic caching but now within the potential of cache programming for all tested programs and most cache sizes. Chen Ding 0001, Dong Chen 0015, Fangzhou Liu 0004, Benjamin Reber, Wesley Smith |
ACM Trans. Archit. Code Optim. | 2 |
| 2021 | Uniform lease vs. LRU cache: analysis and evaluationabstractLease caching is a new technique that provides greater control of the cache than what is allowed in conventional caches. The simplest control is uniform lease (UL), which means that all leases are identical in length. The UL cache is prescriptive and based on allocation. In comparison, a conventional cache is reactive and based on replacement. They represent two fundamentally different approaches to cache management. Dong Chen 0015, Chen Ding 0001, Fangzhou Liu 0004, Benjamin Reber, Wesley Smith, Pengcheng Li 0001 |
ISMM | 1 |
| 2020 | PLUM: static parallel program locality analysis under uniform multiplexingabstractData movement has a significant impact on program performance. For multithread programs, this impact is amplified, since different threads often interfere with each other by competing for shared cache space. However, recent de facto locality metrics consider either sequential execution only, or derive locality for multithread programs in an inefficient way, i.e. exhaustive simulation. Fangzhou Liu 0004, Dong Chen 0015, Wesley Smith, Chen Ding 0001 |
PPoPP | 2 |
| 2018 | Locality analysis through static parallel samplingabstractLocality analysis is important since accessing memory is much slower than computing. Compile-time locality analysis can provide detailed program-level feedback for compilers or runtime systems faster than trace-based locality analysis. Dong Chen 0015, Fangzhou Liu 0004, Chen Ding 0001, Sreepathi Pai |
PLDI | 1 |
| 2017 | LD: Low-Overhead GPU Race Detection Without Access MonitoringabstractData race detection has become an important problem in GPU programming. Previous designs of CPU race-checking tools are mainly task parallel and incur high overhead on GPUs due to access instrumentation, especially when monitoring many thousands of threads routinely used by GPU programs. This article presents a novel data-parallel solution designed and optimized for the GPU architecture. It includes compiler support and a set of runtime techniques. It uses value-based checking, which detects the races reported in previous work, finds new races, and supports race-free deterministic GPU execution. More important, race checking is massively data parallel and does not introduce divergent branching or atomic synchronization. Its slowdown is less than 5 × for over half of the tests and 10 × on average, which is orders of magnitude more efficient than the cuda-memcheck tool by Nvidia and the methods that use fine-grained access instrumentation. Pengcheng Li 0001, Dong Chen 0015, Jacob Brock, Hao Luo 0007, Eddy Z. Zhang, Chen Ding 0001 |
ACM Trans. Archit. Code Optim. | 3 |
| 2017 | Cache Exclusivity and Sharing: Theory and OptimizationabstractA problem on multicore systems is cache sharing, where the cache occupancy of a program depends on the cache usage of peer programs. Exclusive cache hierarchy as used on AMD processors is an effective solution to allow processor cores to have a large private cache while still benefitting from shared cache. The shared cache stores the “victims” (i.e., data evicted from private caches). The performance depends on how victims of co-run programs interact in shared cache. This article presents a new metric called the victim footprint (VFP). It is measured once per program in its solo execution and can then be combined to compute the performance of any exclusive cache hierarchy, replacing parallel testing with theoretical analysis. The work evaluates the VFP by using it to analyze cache sharing by parallel mixes of sequential programs, comparing the accuracy of the theory to hardware counter results, and measuring the benefit of exclusivity-aware analysis and optimization. Chencheng Ye 0001, Chen Ding 0001, Hao Luo 0007, Jacob Brock, Dong Chen 0015, Hai Jin 0001 |
ACM Trans. Archit. Code Optim. | 5 |
| 2015 | Improving performance portability for GPU-specific OpenCL kernels on multi-core/many-core CPUs by analysis-based transformationsabstractOpenCL is an open heterogeneous programming framework. Although OpenCL programs are functionally portable, they do not provide performance portability, so code transformation often plays an irreplaceable role. When adapting GPU-specific OpenCL kernels to run on multi-core/many-core CPUs, coarsening the thread granularity is necessary and thus has been extensively used. However, locality concerns exposed in GPU-specific OpenCL code are usually inherited without analysis, which may give side-effects on the CPU performance. Typically, the use of OpenCL’s local memory on multi-core/many-core CPUs may lead to an opposite performance effect, because local-memory arrays no longer match well with the hardware and the associated synchronizations are costly. To solve this dilemma, we actively analyze the memory access patterns using array-access descriptors derived from GPU-specific kernels, which can thus be adapted for CPUs by (1) removing all the unwanted local-memory arrays together with the obsolete barrier statements and (2) optimizing the coalesced kernel code with vectorization and locality re-exploitation. Moreover, we have developed an automated tool chain that makes this transformation of GPU-specific OpenCL kernels into a CPU-friendly form, which is accompanied with a scheduler that forms a new OpenCL runtime. Experiments show that the automated transformation can improve OpenCL kernel performance on a multi-core CPU by an average factor of 3.24. Satisfactory performance improvements are also achieved on Intel’s many-integrated-core coprocessor. The resultant performance on both architectures is better than or comparable with the corresponding OpenMP performance. Mei Wen, Dafei Huang, Changqing Xun, Dong Chen 0015 |
Frontiers Inf. Technol. Electron. Eng. | 4 |
| 2014 | Automated Transformation of GPU-Specific OpenCL Kernels Targeting Performance Portability on Multi-Core/Many-Core CPUs
Dafei Huang, Mei Wen, Changqing Xun, Dong Chen 0015, Xing Cai, Yuran Qiao, Nan Wu 0003, Chunyuan Zhang |
Euro-Par | 4 |
| 2013 | Efficient fine-grained shared buffer management for multiple OpenCL devicesabstractOpenCL programming provides full code portability between different hardware platforms, and can serve as a good programming candidate for heterogeneous systems, which typically consist of a host processor and several accelerators. However, to make full use of the computing capacity of such a system, programmers are requested to manage diverse OpenCL-enabled devices explicitly, including distributing the workload between different devices and managing data transfer between multiple devices. All these tedious jobs pose a huge challenge for programmers. In this paper, a distributed shared OpenCL memory (DSOM) is presented, which relieves users of having to manage data transfer explicitly, by supporting shared buffers across devices. DSOM allocates shared buffers in the system memory and treats the on-device memory as a software managed virtual cache buffer. To support fine-grained shared buffer management, we designed a kernel parser in DSOM for buffer access range analysis. A basic modified, shared, invalid cache coherency is implemented for DSOM to maintain coherency for cache buffers. In addition, we propose a novel strategy to minimize communication cost between devices by launching each necessary data transfer as early as possible. This strategy enables overlap of data transfer with kernel execution. Our experimental results show that the applicability of our method for buffer access range analysis is good, and the efficiency of DSOM is high. Changqing Xun, Dong Chen 0015, Qiang Lan, Chunyuan Zhang |
J. Zhejiang Univ. Sci. C | 2 |