VLDB 2026 Research / reviewers in the wild / expert
Zicong Wang
dblp:190/5197
· DBLP profile ↗
15ranked-venue papers
2as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 1 first-author · 10 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MoProteus: LLM-Driven Multi-Version Operator Generation for Energy-Aware Scheduling in Heterogeneous Cloud-Edge Environments
Haotian Wang 0006, Junshuang Ma, Zicong Wang, Wangdong Yang, Kenli Li 0001 |
Euro-Par (2) | 4 |
| 2026 | Cohet: A CXL-Driven Coherent Heterogeneous Computing Framework with Hardware-Calibrated Full-System SimulationabstractConventional heterogeneous computing systems built on PCIe interconnects suffer from inefficient fine-grained host-device interactions and complex programming models. In recent years, many proprietary and open cache-coherent interconnect standards have emerged, among which compute express link (CXL) prevails in the open-standard domain after acquiring several competing solutions. Although CXL-based coherent heterogeneous computing holds the potential to fundamentally transform the collaborative computing mode of CPUs and XPUs, research in this direction remains hampered by the scarcity of available CXL-supported platforms, immature software/hardware ecosystems, and unclear application prospects. This paper presents Cohet, the first CXL-driven coherent heterogeneous computing framework. Cohet decouples the compute and memory resources to form unbiased CPU and XPU pools which share a single unified and coherent memory pool. It exposes a standard malloc/mmap interface to both CPU and XPU compute threads, which share a single per-process page table for user applications, leaving the OS dealing with smart memory allocation, page auto-migration, and management of heterogeneous resources. This design significantly simplifies heterogeneous parallel programming to a level comparable to homogeneous programming. To facilitate Cohet research, we also present a fullsystem cycle-level simulator named SimCXL, which is capable of modeling all CXL sub-protocols and device types. SimCXL has been rigorously calibrated against a real CXL testbed with various CXL memory and accelerators, showing an average simulation error of 3 %. Our evaluation reveals that CXL.cache reduces latency by 68 % and increases bandwidth by$14.4 \times$compared to DMA transfers at cacheline granularity. Building upon these insights, we demonstrate the benefits of Cohet with two killer apps, which are remote atomic operation (RAO) and remote procedure call (RPC). Compared to PCIe-NIC design, CXL-NIC achieves a 5.5 to$40.2 \times$speedup for RAO offloading and an average speedup of$\mathbf{1. 8 6} \times$for$\mathbf{R P C}$(de)serialization offloading. Yanjing Wang 0007, Lizhou Wu, Sunfeng Gao, Yibo Tang, Junhui Luo, Zicong Wang, Dezun Dong, Nong Xiao 0001 |
HPCA | 6 |
| 2026 | From Memorization to Generalization: A Practical Neural Network Prefetching Framework
Zicong Wang, Shuiyi He, Dezun Dong, Xiangke Liao |
ISCA | 2 |
| 2026 | CXL-DMSim: A Full-System CXL Disaggregated Memory Simulator With Comprehensive Silicon ValidationabstractCompute eXpress Link (CXL) has emerged as a key enabler of memory disaggregation for future heterogeneous computing systems to expand memory on-demand and improve resource utilization. However, CXL is still in its infancy stage and lacks commodity products on the market, thus necessitating a reliable system-level simulation tool for research and development. In this paper, we propose CXL-DMSim1, an open-source full-system simulator to simulate CXL disaggregated memory systems with high fidelity at a gem5-comparable simulation speed. CXL-DMSim incorporates a flexible CXL memory expander model along with its associated device driver, and CXL protocol support with CXL.io and CXL.mem. It can operate in both app-managed mode and kernel-managed mode, with the latter using a dedicated NUMA-compatible mechanism. The simulator has been rigorously verified against a real hardware testbed with both FPGA- and ASIC-based CXL memory devices, which demonstrates the qualification of CXL-DMSim in simulating the characteristics of various CXL memory devices at an average simulation error of 3.4%. The experimental results using LMbench and STREAM benchmarks suggest that the CXL-FPGA memory exhibits a ~2.88× higher latency than local DDR while the CXL-ASIC latency is ~2.18×; CXL-FPGA achieves 45-69% of local DDR memory bandwidth, whereas the number for CXL-ASIC is 82-83%. The study also reveals that CXL memory can significantly enhance the performance of memory-intensive applications, improved by 23× at most with limited local memory for Viper key–value database and approximately 60% in memory-bandwidth-sensitive scenarios such as MERCI. Moreover, the simulator’s observability and expandability are showcased with detailed case-studies, highlighting its great potential for research on future CXL-interconnected hybrid memory pool. Yanjing Wang 0007, Lizhou Wu, Wentao Hong, Zicong Wang, Sunfeng Gao, Jie Zhang 0048, Sheng Ma, Dezun Dong, Xingyun Qi, Nong Xiao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2025 | Amphi: Practical and Intelligent Data Prefetching for the First-Level CacheabstractData prefetchers play a crucial role in alleviating the memory wall by predicting future memory accesses. First-level cache prefetchers can observe all memory instructions but often rely on simpler strategies due to limited resources. While emerging machine learning-based approaches cover more memory access patterns, they typically require higher computational and storage resources and are usually deployed in the last-level cache. Other intelligent solutions for the first-level cache show only modest performance gains. To address this, we propose Amphi, the first practical and intelligent data prefetcher specifically designed for the first-level cache. Applying a binarized temporal convolutional network, Amphi significantly reduces storage overhead while maintaining performance comparable to the SOTA intelligent prefetcher. With a storage overhead of only 3.4 KB, Amphi requires only one-eighth of Pythia's storage needs. Amphi paves the way for the broader adoption of intelligence-driven prefetching solutions. Zicong Wang, Shuiyi He, Dezun Dong, Xiangke Liao |
DATE | 2 |
| 2025 | Elevating Temporal Prefetching Through Instruction Correlation
Shuiyi He, Zicong Wang, Dezun Dong, Liquan Xiao |
MICRO | 2 |
| 2025 | vtism: Efficient Tiered Memory Management for Virtual Machines with CXLabstractVirtual machines (VMs) impose increasing memory demands, exposing the capacity and cost limitations of traditional DRAM only memory architectures. To address this problem, heterogeneous DRAM+CXL tiered memory management systems have emerged as a promising solution. However, in virtualization environments, the semantic gap between guest and host abstraction layers, coupled with dynamic workload behaviors, hinders precise page tracking, classification, and efficient page migration across memory tiers. Zhixing Lu, Lizhou Wu, Zicong Wang, Xuran Ge, Zhenlong Song |
SYSTOR | 4 |
| 2024 | Chimera: Leveraging Hybrid Offsets for Efficient Data PrefetchingabstractData prefetching is an essential technique in contemporary high-performance processors for mitigating the effects of long-latency memory accesses. With the increasing demand for prefetcher to learn complex memory access patterns, many state-of-the-art prefetchers adopt methods such as using the program counter or access delta to separate memory access streams. This allows them to learn detailed memory access features and thereby improve memory system performance. However, the separation-based approach is prone to missing global correlations, leading to miss prefetching opportunities. In this paper, we propose Chimera, a hybrid offsets prefetcher that captures multiple offset features from the overall stream of memory access instructions, thus overcoming the drawbacks of traditional prefetchers that tend to lose memory access information when learning from one-sided features. We evaluated Chimera using SPEC CPU 2006 and 2017 through simulation, and the results show that Chimera improves system performance by 39.5% over a baseline with no data prefetcher and by 6.6% over the state-of-the-art data prefetcher. Shuiyi He, Zicong Wang, Qiyao Sun, Dezun Dong |
PACT | 2 |
| 2024 | EIGP: document-level event argument extraction with information enhancement generated based on prompts
Zicong Wang, Qianxi Hou |
Knowl. Inf. Syst. | 3 |
| 2024 | Pimo: memory-efficient privacy protection in video streaming and analyticsabstractAbstract Video streaming from cameras to backend cloud or edge servers for neural-based analytics has gained significant popularity. However, the transmission of data from cameras to a backend raises substantial privacy concerns, particularly regarding sensitive information like facial data. To offer privacy protection, visual processing techniques, such as Generative Adversarial Networks (GANs), have been employed on cameras to blur and safeguard such data intelligently. However, these techniques frequently face memory challenges, particularly when dealing with high-resolution videos. In this paper, we propose PIMO, a memory-efficient visual privacy protection scheme designed to effectively blur video content leveraging adaptive slicing of frames and resolution degradation. Our extensive experimental evaluations validate that PIMO’s adaptive mechanism proficiently navigates fluctuating memory constraints. Furthermore, utilizing a content-based blur scheme, our approach can maintain an impressive mean precision of 95.2%, as compared to the original, non-blurred images. Jie Yuan 0001, Zicong Wang, Tingting Yuan 0001 |
Multim. Syst. | 2 |
| 2023 | LARE: A Linear Approximate Reinforcement Learning Based Adaptive Routing for Network-on-ChipsabstractThe routing algorithm is crucial for network performance in network-on-chips (NoCs). With emerging applications bring new features to NoCs with more complex and time-varying traffic, which turns the routing computation process into multi-objective optimization. However, We found that the existing routing algorithms cannot effectively achieve load-balanced between different traffic due to the static method of routing design. The routing algorithm design space will increase if all factors affecting routing are considered. Reinforcement learning (RL) methods have demonstrated promising opportunities applied to architecture design exploration. In this paper, we proposed a novel RL framework for adaptive routing design in NoCs. This method uses network information to select the best path to achieve load balance and lower communication latency at the same time. Unfortunately, with this method, the implementation overhead of RL increases rapidly as the network scale increases. Therefore, we introduce a linear function of approximate RL-based adaptive routing (LARE) to reduce implementation over-head. We conduct extensive experiments against state-of-the-art routing algorithms to evaluate our design. Simulation results demonstrate the benefits of our design under synthetic traffic workloads and real applications. In addition, LARE can achieve similar network performance with traditional RL implementation with a much lower hardware overhead. Dezun Dong, Cunlu Li, Zicong Wang, Zongmao Zhang |
ISCAS | 5 |
| 2022 | MUSH: Multi-scale Hierarchical Feature Extraction for Semantic Image Synthesis
Zicong Wang, Junli Wang 0001, ChunGang Yan, Changjun Jiang 0002 |
ACCV (7) | 1 |
| 2021 | RELAR: A Reinforcement Learning Framework for Adaptive Routing in Network-on-ChipsabstractAdaptive routing is crucial to the overall performance of network-on-chips (NoCs), and still faces great challenges, especially when emerging applications on many-core architecture exhibit complicated and time-varying traffic patterns. When witnessing most existing heuristic adaptive routing algorithms fail to address multi-objective optimization for complex traffic well, we make the first attempt to propose a novel and comprehensive reinforcement learning framework for adaptive routing on NoCs, called RELAR. RELAR is suitable for diversified traffic patterns and resolve multi-objective optimization simultaneously. We conduct experiments against state-of-the-art routing algorithms to evaluate our design. The results show that RELAR achieves 14.82% reduction in packet latency on average, and reduces packet latency by up to 34.24% under heavy synthetic traffic workload. Dezun Dong, Zicong Wang |
CLUSTER | 3 |
| 2017 | Fairness-Oriented and Location-Aware NUCA for Many-Core SoCabstractNon-uniform cache architecture (NUCA) is often employed to organize the last level cache (LLC) by Networks-on-Chip (NoC). However, along with the scaling up for network size of Systems-on-Chip (SoC), two trends gradually begin to emerge. First, the network latency is becoming the major source of the cache access latency. Second, the communication distance and latency gap between different cores is increasing. Such gap can seriously cause the network latency imbalance problem, aggravate the degree of non-uniform for cache access latencies, and then worsen the system performance. Zicong Wang, Chen Li 0015, Yang Guo 0003 |
NOCS | 1 |
| 2016 | DLL: A dynamic latency-aware load-balancing strategy in 2.5D NoC architectureabstractAs the 3D stacking technology still faces several challenges, the 2.5D stacking technology gains better application prospects nowadays. With the silicon interposer, the 2.5D stacking can improve the bandwidth and capacity of the memory system. To satisfy the communication requirements of the integrated memory system, the free routing resources in the interposer should be explored to implement an additional network. Yet, the performance is strongly limited by the unbalanced loads between the CPU-layer network and the interposer-layer network. In this paper, to address this issue, we propose a dynamic latency-aware load-balancing (DLL) strategy. Our key innovations are detecting congestion of the network layer via the average latency of recent packets and making the network layer selection at each source node. We leverage the free routing resources in the interposer to implement a latency propagation ring. With the ring, the latency information tracked at destination nodes is propagated back to source nodes. We achieve load-balance by using these information. Experimental results show that compared with the baseline design, a destination-detection strategy and a buffer-aware strategy, our DLL strategy achieves 45%, 14.9% and 6.5% of average throughput improvements with minor overheads. Chen Li 0015, Sheng Ma, Lu Wang 0019, Zicong Wang, Xia Zhao 0004, Yang Guo 0003 |
ICCD | 4 |