VLDB 2026 Research / reviewers in the wild / expert
Chongqing Zhao
dblp:48/7993
· DBLP profile ↗
3ranked-venue papers
0as first author
2since 2021 · last 2025
0009-0000-1345-9484ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Cloud and datacenter computing · 33% Hardware accelerators and domain-specific architectures · 29% Processor architecture and microarchitecture · 29% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% | |
| Computer networks
1 paper |
Datacenter networks · 100% |
Topics — the 6 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization
deep learning compiler |
0.9 | 1 | 2025 | Mosaic: Exploiting Instruction-Level Parallelism on Deep Learning Accelerators with iTex Tessellation · ASPLOS (2) 2025 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
DNN accelerator |
0.9 | 1 | 2025 | Mosaic: Exploiting Instruction-Level Parallelism on Deep Learning Accelerators with iTex Tessellation · ASPLOS (2) 2025 |
Processor architecture and microarchitecture
instruction-level parallelism |
0.9 | 1 | 2025 | Mosaic: Exploiting Instruction-Level Parallelism on Deep Learning Accelerators with iTex Tessellation · ASPLOS (2) 2025 |
Datacenter networks
RDMA |
0.8 | 1 | 2024 | Turbo: Efficient Communication Framework for Large-scale Data Processing Cluster · SIGCOMM 2024 |
Cloud and datacenter computing
cluster data processing |
0.8 | 1 | 2024 | Turbo: Efficient Communication Framework for Large-scale Data Processing Cluster · SIGCOMM 2024 |
Cloud and datacenter computing › big data platform
shuffle service |
0.2 | 1 | 2024 | Turbo: Efficient Communication Framework for Large-scale Data Processing Cluster · SIGCOMM 2024 |
Methods — techniques the papers use, named apart from their topics
itex tessellation · 1.7instruction mapping · 1.7non-blocking communication middleware · 1.5flowlet transmission · 1.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Mosaic: Exploiting Instruction-Level Parallelism on Deep Learning Accelerators with iTex TessellationabstractDeep learning has achieved great success in numerous application areas at the cost of high computational complexity. To meet the ever-increasing computational demand, commodity hardware platforms (e.g., CPUs and GPUs) offer abundant computing resources including scalar, vector, and tensor units for deep learning that could execute in parallel. However, existing top-down tiling-based deep learning compilers often generate a homogeneous mapping from the given tensor computation task to hardware arithmetic instructions, failing to utilize different computing units simultaneously to achieve higher performance. Jianxing Xu, Yuanbo Wen 0001, Ruibai Xu, Tingfeng Ruan, Jun Bi, Rui Zhang 0040, Xinkai Song, Yifan Hao 0001, Xing Hu 0001, Zidong Du, Chongqing Zhao, Jiang Jie, Qi Guo 0001 |
ASPLOS (2) | 13 |
| 2024 | Turbo: Efficient Communication Framework for Large-scale Data Processing ClusterabstractBig data processing clusters are suffering from a long job completion time due to the inefficient utilization of the RDMA capability. Our production measurement results in a large-scale cluster with hundreds of server nodes to process large-scale jobs have shown that the existing deployment of RDMA technique results in a long-tail job completion time, with some jobs even taking up more than twice the average time to complete. In this paper, we present the design and implementation of Turbo, an efficient communication framework for the large-scale data processing cluster to achieve high performance and scalability. The core of Turbo's approach is to leverage a dynamic block-level flowlet transmission mechanism and a non-blocking communication middleware to improve the network throughput and enhance system's scalability. Furthermore, Turbo ensures high system reliability by utilizing an external shuffle service as well as TCP serving as a backup. We integrate Turbo into Apache Spark and evaluate Turbo in a small-scale testbed and a large-scale cluster consisting of hundreds of server nodes. The small-scale testbed evaluation results show that Turbo improves the network throughput by 15.1% while maintaining high system reliability. The large-scale production results have shown Turbo can reduce the job completion time by 23.9% and increase the job completion rate by 2.03× over the existing RDMA solutions. Xuya Jia, Zhiyi Yao, Edison Liu, Xiang Li 0223, Zekun He, Yachen Wang, Xianneng Zou, Chongqing Zhao, Jinhui Chu, Jilong Wang 0001, Congcong Miao |
SIGCOMM | 11 |
| 2009 | Synthesis Constraints Optimized Genetic Algorithm for Autonomous Task Planning and Allocating in MASabstractNow, autonomous tasks planning and allocating (TPA) in Multi Agent System (MAS) has been one key and fundamental problem to promote the intelligent level of such system. Autonomous TPA means that, all tasks should be (re)planned and (re)allocated automatically according to the synthesis constraints and the dynamic environment aspects, such as the changing mission, status of each member, and topology, etc. In this article, the formal descriptions of hierarchical tasks and models of logic constraints are studied firstly. And then, some new methods are proposed to evaluate the efficiency of synthesis constraints. Moreover, the key elements, e.g. task allocation vector (TAV), are designed with the theory of genetic algorithm (GA), and a TPA problem can be mapped to the solving model of GA. Based on above, the crossover and mutation operators of GA are optimized with the domain knowledge to perfect the solving efficiency and quality while ensuring the randomicity of evolution. The simulation results show that the solving quality and velocity are improved with studied methods. Kailong Zhang, Xingshe Zhou 0001, Chongqing Zhao, Yuan Yao 0004 |
SERA | 3 |