VLDB 2026 Research / reviewers in the wild / expert
Xiuping Cui
dblp:279/8367
· DBLP profile ↗
5ranked-venue papers
2as first author
4since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Dolphium: Co-Optimizing Quantization Dataflow and Paradigms on Poly-Hierarchical NPUsabstractPoly-hierarchical NPUs integrate distributed memory modules with heterogeneous computation units, posing significant challenges for mapping quantized operators. The difficulty arises from the need to coordinate data transfers across memory hierarchies and to assign diverse operations to suitable computation units. In this work, we systematically construct the mapping space from quantization to dataflows by addressing three key aspects: generation of NPU-friendly computation flows, integrated operation–data co-mapping, and determination of transfer granularity and frequency. Building on this foundation, we further exploit quantization dataflows to guide the selection of quantization paradigms. Compared with the state-of-the-art quantization compiler, our mapping achieves a 1.67–2.03× speedup. Moreover, the selected quantization paradigms deliver an average 2.18× efficiency improvement on NPUs without accuracy loss. Xiuping Cui, Chengrui Zhang, Yun Liang 0001 |
DATE | 1 |
| 2026 | LATIAS: A General Architecture-Operator Model for Spatial Accelerators with Complex Topology and Memory HierarchyabstractSpatial accelerators are widely deployed for deep neural networks, but their architectural diversity—from hierarchical to dataflow designs—makes accurate architecture–operator modeling difficult, limiting operator optimization and hardware utilization. Existing models abstract hardware as hierarchical chains and operators as loop trees, which cannot capture essential features of modern dataflow accelerators, including heterogeneous processing elements (PEs), uni-directional interconnects, and cross-PE memory hierarchies, leading to inaccurate latency prediction. We propose LATIAS, a unified framework that introduces (1) an architecture graph with uni-directional edges to represent arbitrary topologies, and (2) a dataflow-aware tile-centric notation that augments loop trees with transfer nodes to model diverse dataflows. Building on these, LATIAS further provides a graph-guided tree analysis that accurately resolves tensor residency and latency under hardware constraints. Experiments on representative operators (GEMM, vector, fused vector) and operator shapes extracted from DNNs (BERT, ViT, T5) on Huawei Ascend 910B3 show that LATIAS achieves over 0.99 correlation with runtime measurements—substantially outperforming prior models—and provides actionable insights for architectural design. Chengrui Zhang, Liancheng Jia, Renze Chen, Xiuping Cui, Size Zheng 0001, Shengen Yan, Yu Wang 0002, Yun Liang 0001 |
DATE | 6 |
| 2026 | DynaMo: Runtime Switchable Quantization for MoE with Cross-Dataset AdaptationabstractAs the Mix-of-Experts (MoE) architecture increases the number of parameters in large models, there is an even greater need for model quantization. However, existing quantization methods overlook the expert dynamics of MoE across multiple datasets. Moreover, the existing static quantization cannot adapt MoE to various data change scenarios. In this paper, we perform a multi-level analysis to reveal MoE dynamics and define the significance of each channel/each expert. Based on the analysis results, we propose DynaMo, an end-to-end MoE quantization framework. DynaMo adopts an expert-level mixed-precision baseline quantization strategy, which ensures the quantized MoEs are compatible with multiple existing datasets. Furthermore, DynaMo incorporates a channel-level dynamic switching mechanism to adapt these quantized MoE models to novel datasets. Experiments show that DynaMo achieves a 2.78~4.54 PPL decrease and a 1.85%~3.77% accuracy improvement in various datasets, with ~3× inference speedup and negligible overhead. Xiuping Cui, Size Zheng 0001, Maoliang Li, Yun Liang 0001 |
DATE | 2 |
| 2023 | ARES: A Mapping Framework of DNNs Towards Diverse PIMs with General AbstractionsabstractNumerous architectures based on processing-in-memory (PIM) have recently emerged, exhibiting diversity in memory types, compute functions, memory mapping constraints, etc. To effectively utilize PIM hardware for deploying deep neural networks (DNNs), programmers face the challenge of mapping computations and data across multiple memory arrays, scheduling computation and data transfers, while adhering to various hardware constraints. Existing mapping approaches, however, are tailored to specific architectures and lack a general formulation for mapping optimization, limiting their applicability and performance. In this paper, we present ARES, a comprehensive mapping framework designed for diverse PIM architectures. The core of the framework is hardware abstractions for PIMs, which is inspired by the fact that DNNs on PIM hardware can be represented by a tensorized compute function and data layout constraints in the memory array. This abstraction forms the basis for constructing a mapping space that encompasses both compute and memory constraints. Through exploration of this mapping space, we derive efficient mapping strategies tailored to different PIM hardware configurations. Experimental evaluation conducted on four distinct hardware architectures demonstrates that compared to state-of-the-art mapping methods, ARES yields up to a 70% speed improvement for single operator mapping and a 50% speedup for overall network mapping. Xiuping Cui, Size Zheng 0001, Le Ye, Yun Liang 0001 |
ICCAD | 1 |
| 2020 | SuSy: A Programming Model for Productive Construction of High-Performance Systolic Arrays on FPGAsabstractSystolic algorithms are one of the killer applications on spatial architectures such as FPGAs and CGRAs. However, it requires a tremendous amount of human effort to design and implement a high-performance systolic array for a given algorithm using the traditional RTL-based methodology. On the other hand, existing high-level synthesis (HLS) tools either (1) force the programmers to do "micro-coding" where too many optimizations must be carried out through tedious code restructuring and insertion of vendor-specific pragmas, or (2) give them too little control to influence a push-button compilation flow to achieve high quality of results. Yi-Hsiang Lai, Hongbo Rong, Size Zheng 0001, Xiuping Cui, Yunshan Jia, Jie Wang 0022, Brendan Sullivan, Zhiru Zhang, Yun Liang 0001, Youhui Zhang, Jason Cong, Nithin George, Christopher J. Hughes, Pradeep Dubey |
ICCAD | 5 |