VLDB 2026 Research / reviewers in the wild / expert
Zikai Zhou
dblp:197/4350
· DBLP profile ↗
17ranked-venue papers
6as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 5 first-author · 10 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improved and Accelerated Text-to-Image Generation With Collect, Reflect, and RefineabstractRecently, enhancing the generative capability of text-to-image (T2I) models has become a promising direction in both academia and industry. Prior studies often focused on either improving generative quality or reducing inference latency, but typically failed to improve both quality and speed simultaneously. Moreover, existing inference-enhancement methods do not achieve significant improvements simultaneously across both diffusion models (DMs) and autoregressive models (ARMs). In this paper, we introduce a general tuning-based inference-enhancement framework, named CoRe$^{2}$2, which is the first to simultaneously achieve significant generative quality and reduced inference overhead across DMs and ARMs, to the best of our knowledge. CoRe$^{2}$2 comprises three stages: Collect, Reflect, and Refine. During the Collect stage, classifier-free guidance (CFG) trajectories are collected and subsequently used in the Reflect stage to train a weak model capable of reflecting the "easy-to-learn" content. Finally, during the Refine stage, CoRe$^{2}$2 can utilize the trained weak model to achieve speedup and performance gain in inference. Specifically, in the early sampling steps, CoRe$^{2}$2 employs weak-to-strong guidance to refine the "difficult-to-learn" and realistic content, thereby improving generative quality. In the later sampling steps, CoRe$^{2}$2 can use the weak model to generate "easy-to-learn" content instead of CFG, dramatically reducing inference time. Experimental outcomes substantiates CoRe$^{2}$2 achieve significant performance improvements on HPD v2, Pick-of-Pic, Drawbench, GenEval, and T2I-Compbench across SDXL, SD3.5, FLUX and LlamaGen. Notably, for SD3.5, CoRe$^{2}$2 can be seamlessly integrated with the state-of-the-art inference-enhancement algorithm Z-Sampling, outperforming it even with less time. Shitong Shao, Zikai Zhou, Dian Xie, Yuetong Fang, Tian Ye 0001, Lichen Bai, Bo Han 0003, Zeke Xie |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Blend the Separated: Mixture of Synergistic Experts for Data-Scarcity Drug-Target Interaction PredictionabstractDrug-target interaction prediction (DTI) is essential in various applications including drug discovery and clinical application. There are two perspectives of input data widely used in DTI prediction: Intrinsic data represents how drugs or targets are constructed, and extrinsic data represents how drugs or targets are related to other biological entities. However, any of the two perspectives of input data can be scarce for some drugs or targets, especially for those unpopular or newly discovered. Furthermore, ground-truth labels for specific interaction types can also be scarce. Therefore, we propose the first method to tackle DTI prediction under input data and/or label scarcity. To make our model functional when only one perspective of input data is available, we design two separate experts to process intrinsic and extrinsic data respectively and fuse them adaptively according to different samples. Furthermore, to make the two perspectives complement each other and remedy label scarcity, two experts synergize with each other in a mutually supervised way to exploit the enormous unlabeled data. Extensive experiments on 3 real-world datasets under different extents of input data scarcity and/or label scarcity demonstrate our model outperforms states of the art significantly and steadily, with a maximum improvement of 53.53%. We also test our model without any data scarcity and it still outperforms current methods. Xinlong Zhai, Chunchen Wang, Jiazheng Kang, Shujie Li 0003, Zikai Zhou, Cheng Yang 0002, Chuan Shi 0001 |
AAAI | 8 |
| 2025 | Golden Noise for Diffusion Models: A Learning FrameworkabstractText-to-image diffusion model is a popular paradigm that synthesizes personalized images by providing a text prompt and a random Gaussian noise. While people observe that some noises are ``golden noises'' that can achieve better text-image alignment and higher human preference than others, we still lack a machine learning framework to obtain those golden noises. To learn golden noises for diffusion sampling, we mainly make three contributions in this paper. First, we identify a new concept termed the \textit{noise prompt}, which aims at turning a random Gaussian noise into a golden noise by adding a small desirable perturbation derived from the text prompt. Following the concept, we first formulate the \textit{noise prompt learning} framework that systematically learns ``prompted'' golden noise associated with a text prompt for diffusion models. Second, we design a noise prompt data collection pipeline and collect a large-scale \textit{noise prompt dataset}~(NPD) that contains 100k pairs of random noises and golden noises with the associated text prompts. With the prepared NPD as the training dataset, we trained a small \textit{noise prompt network}~(NPNet) that can directly learn to transform a random noise into a golden noise. The learned golden noise perturbation can be considered as a kind of prompt for noise, as it is rich in semantic information and tailored to the given text prompt. Third, our extensive experiments demonstrate the impressive effectiveness and generalization of NPNet on improving the quality of synthesized images across various diffusion models, including SDXL, DreamShaper-xl-v2-turbo, and Hunyuan-DiT. Moreover, NPNet is a small and efficient controller that acts as a plug-and-play module with very limited additional inference and computational costs, as it just provides a golden noise instead of a random noise without accessing the original pipeline. Zikai Zhou, Shitong Shao, Lichen Bai, Shufei Zhang, Zhiqiang Xu 0003, Bo Han 0003, Zeke Xie |
ICCV | 1 |
| 2025 | Zigzag Diffusion Sampling: Diffusion Models Can Self-Improve via Self-ReflectionabstractDiffusion models, the most popular generative paradigm so far, can inject conditional information into the generation path to guide the latent towards desired directions. However, existing text-to-image diffusion models often fail to maintain high image quality and high prompt-image alignment for those challenging prompts. To mitigate this issue and enhance existing pretrained diffusion models, we mainly made three contributions in this paper. First, we propose **diffusion self-reflection** that alternately performs denoising and inversion and demonstrate that such diffusion self-reflection can leverage the guidance gap between denoising and inversion to capture prompt-related semantic information with theoretical and empirical evidence. Second, motivated by theoretical analysis, we derive Zigzag Diffusion Sampling (Z-Sampling), a novel self-reflection-based diffusion sampling method that leverages the guidance gap between denosing and inversion to accumulate semantic information step by step along the sampling path, leading to improved sampling results. Moreover, as a plug-and-play method, Z-Sampling can be generally applied to various diffusion models (e.g., accelerated ones and Transformer-based ones) with very limited coding and computational costs. Third, our extensive experiments demonstrate that Z-Sampling can generally and significantly enhance generation quality across various benchmark datasets, diffusion models, and performance evaluation metrics. For example, DreamShaper with Z-Sampling can self-improve with the HPSv2 winning rate up to **94%** over the original results. Moreover, Z-Sampling can further enhance existing diffusion models combined with other orthogonal methods, including Diffusion-DPO. The code is publicly available at
[github.com/xie-lab-ml/Zigzag-Diffusion-Sampling](https://github.com/xie-lab-ml/Zigzag-Diffusion-Sampling). Lichen Bai, Shitong Shao, Zikai Zhou, Zipeng Qi, Zhiqiang Xu 0003, Haoyi Xiong, Zeke Xie |
ICLR | 3 |
| 2025 | IV-mixed Sampler: Leveraging Image Diffusion Models for Enhanced Video SynthesisabstractExploring suitable solutions to improve performance by increasing the computational cost of inference in visual diffusion models is a highly promising direction. Sufficient prior studies have demonstrated that correctly scaling up computation in the sampling process can successfully lead to improved generation quality, enhanced image editing, and compositional generalization. While there have been rapid advancements in developing inference-heavy algorithms for improved image generation, relatively little work has explored inference scaling laws in video diffusion models (VDMs). Furthermore, existing research shows only minimal performance gains that are perceptible to the naked eye. To address this, we design a novel training-free algorithm IV-Mixed Sampler that leverages the strengths of image diffusion models (IDMs) to assist VDMs surpass their current capabilities. The core of IV-Mixed Sampler is to use IDMs to significantly enhance the quality of each video frame and VDMs ensure the temporal coherence of the video during the sampling process. Our experiments have demonstrated that IV-Mixed Sampler achieves state-of-the-art performance on 4 benchmarks including UCF-101-FVD, MSR-VTT-FVD, Chronomagic-Bench-150/1649, and VBench. For example, the open-source Animatediff with IV-Mixed Sampler reduces the UMT-FVD score from 275.2 to 228.6, closing to 223.1 from the closed-source Pika-2.0. Shitong Shao, Zikai Zhou, Bai Lichen, Haoyi Xiong, Zeke Xie |
ICLR | 2 |
| 2025 | LowRA: Accurate and Efficient LoRA Fine-Tuning of LLMs under 2 BitsabstractFine-tuning large language models (LLMs) is increasingly costly as models scale to hundreds of billions of parameters, and even parameter-efficient fine-tuning (PEFT) methods like LoRA remain resource-intensive. We introduce LowRA, the first framework to enable LoRA fine-tuning below 2 bits per parameter with minimal performance loss. LowRA optimizes fine-grained quantization—mapping, threshold selection, and precision assignment—while leveraging efficient CUDA kernels for scalable deployment. Extensive evaluations across 4 LLMs and 4 datasets show that LowRA achieves a superior performance–precision trade-off above 2 bits and remains accurate down to 1.15 bits, reducing memory usage by up to 50%. Our results highlight the potential of ultra-low-bit LoRA fine-tuning for resource-constrained environments. Zikai Zhou, Qizheng Zhang, Hermann Kumbong, Kunle Olukotun |
ICML | 1 |
| 2025 | Hierarchical Adaptive Learning-Based Congestion Control With Low Training Overhead for Datacenter NetworksabstractMost congestion control mechanisms perform well in specific datacenter networks, but none can consistently deliver good performance across varying scenarios. Recently proposed frameworks based on reinforcement learning can flexibly select congestion control algorithms to adapt to dynamic network. However, frequently altering the congestion control mechanisms during relatively stable periods of the network actually leads to instability and unnecessary computational overhead. In this paper, we propose a lightweight and hierarchical adaptive congestion control algorithm (LACC) to be resilient to the varying network. LACC dynamically selects the appropriate congestion control mechanism only when the current congestion control algorithm is not suitable for the current network state, rather than changing the congestion control scheme every training cycle to ensure network stability. The simulation results show that LACC significantly reduces the average overhead by 31% and improves throughput by up to 47%, 35%, 23% and 15% compared to Cubic, Reno, BBR and Antelope, respectively. Jinbin Hu 0001, Zikai Zhou, Jing Wang 0209 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2024 | Rethinking Centered Kernel Alignment in Knowledge Distillation
Zikai Zhou, Yunhang Shen, Shitong Shao, Linrui Gong, Shaohui Lin |
IJCAI | 1 |
| 2024 | Enhancing Adversarial Attacks: The Similar Target MethodabstractAdversarial examples are notably characterized by their strong transferability, allowing attackers to craft these examples on their models and subsequently deploy them against other models. This poses significant threats to existing deep learning systems and raises substantial security concerns. While several methods have been developed to enhance transferability, ensemble attacks stand out for their effectiveness. However, previous approaches of ensemble attacks typically rely on simple averaging of logits, probabilities, or losses for ensembling, without delving into the underlying reasons for the improved transferability. In this work, we propose a new approach that makes full use of the information of each surrogate model by regularizing the optimization direction to concurrently attack all surrogate models. This is achieved by promoting cosine similarity between their gradients. Extensive experiments conducted on the ImageNet dataset demonstrate the superior efficacy of our method in enhancing adversarial transferability. Notably, our approach outperforms leading state-of-the-art attackers across 18 discriminative classifiers and adversarially trained models, underscoring its potential in this domain. Ziruo Wang, Zikai Zhou, Jiyao Liu, Huanran Chen |
IJCNN | 3 |
| 2024 | Learning-Based Hierarchical Adaptive Congestion Control with Low Training OverheadabstractMost congestion control mechanisms perform well in specific network environments, but none can consistently deliver good performance across all scenarios. Recently proposed frameworks based on reinforcement learning can flexibly select congestion control algorithms to adapt to dynamic changes in network conditions. However, frequently altering the congestion control mechanisms during relatively stable periods of the network actually leads to instability and unnecessary computational overhead. In this paper, we propose a hierarchical adaptive congestion control algorithm (HACC) to be resilient to the varying network. HACC dynamically selects the appropriate congestion control mechanism only when the current congestion control algorithm is not suitable for the current network state, rather than changing the congestion control scheme every training cycle to ensure network stability. The simulation results show that under different realistic workloads, HACC significantly reduces the computational overhead and improves throughput. Specifically, HACC reduces average overhead by 31% and improves throughput by up to 47%, 35%, 23%, and 15% compared to Cubic, Reno, BBR, and Antelope, respectively. Jinbin Hu 0001, Zikai Zhou, Shuying Rao, Yujie Peng, Bowen Bao, Chang Ruan |
ISPA | 2 |
| 2024 | Elucidating the Design Space of Dataset CondensationabstractDataset condensation, a concept within $\textit{data-centric learning}$, aims to efficiently transfer critical attributes from an original dataset to a synthetic version, meanwhile maintaining both diversity and realism of syntheses. This approach can significantly improve model training efficiency and is also adaptable for multiple application areas. Previous methods in dataset condensation have faced several challenges: some incur high computational costs which limit scalability to larger datasets ($\textit{e.g.,}$ MTT, DREAM, and TESLA), while others are restricted to less optimal design spaces, which could hinder potential improvements, especially in smaller datasets ($\textit{e.g.,}$ SRe$^2$L, G-VBSM, and RDED). To address these limitations, we propose a comprehensive designing-centric framework that includes specific, effective strategies like implementing soft category-aware matching, adjusting the learning rate schedule and applying small batch-size. These strategies are grounded in both empirical evidence and theoretical backing. Our resulting approach, $\textbf{E}$lucidate $\textbf{D}$ataset $\textbf{C}$ondensation ($\textbf{EDC}$), establishes a benchmark for both small and large-scale dataset condensation. In our testing, EDC achieves state-of-the-art accuracy, reaching 48.6% on ImageNet-1k with a ResNet-18 model at an IPC of 10, which corresponds to a compression ratio of 0.78\%. This performance surpasses those of SRe$^2$L, G-VBSM, and RDED by margins of 27.3%, 17.2%, and 6.6%, respectively. Code is available at: https://github.com/shaoshitong/EDC. Shitong Shao, Zikai Zhou, Huanran Chen |
NeurIPS | 2 |
| 2024 | Attention-optimized vision-enhanced prompt learning for few-shot multi-modal sentiment analysis
Zikai Zhou, Baiyou Qiao, Haisong Feng, Donghong Han, Gang Wu 0007 |
Neural Comput. Appl. | 1 |
| 2024 | Lightweight Automatic ECN Tuning Based on Deep Reinforcement Learning With Ultra-Low Overhead in Datacenter NetworksabstractIn modern datacenter networks (DCNs), mainstream congestion control (CC) mechanisms essentially rely on Explicit Congestion Notification (ECN) to reflect congestion. The traditional static ECN threshold performs poorly under dynamic scenarios, and setting a proper ECN threshold under various traffic patterns is challenging and time-consuming. The recently proposed reinforcement learning (RL) based ECN Tuning algorithm (ACC) consumes a large number of computational resources, making it difficult to deploy on switches. In this paper, we present a lightweight and hierarchical automated ECN tuning algorithm called LAECN, which can fully exploit the performance benefits of deep reinforcement learning with ultra-low overhead. The simulation results show that LAECN improves performance significantly by reducing latency and increasing throughput in stable network conditions, and also shows consistent high performance in small flows network environments. For example, LAECN effectively improves throughput by up to 47%, 34%, 32% and 24% over DCQCN, TIMELY, HPCC and ACC, respectively. Jinbin Hu 0001, Zikai Zhou, Jin Zhang 0018 |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2023 | HAECN: Hierarchical Automatic ECN Tuning with Ultra-Low Overhead in Datacenter Networks
Jinbin Hu 0001, Youyang Wang, Zikai Zhou, Shuying Rao, Rundong Xin, Jing Wang 0209, Shiming He |
ICA3PP (3) | 3 |
| 2021 | A Framework for Reproducible Data Plane Performance ModelingabstractLanguages for programming data planes like P4 sparked a plethora of new applications in the data plane. The dynamic, evolving environment makes it challenging to understand what performance can be expected when running a program in a specific data plane target. However, knowing this is crucial for network operators when upgrading their networks. Dominik Scholz, Hasanin Harkous, Sebastian Gallenmüller, Henning Stubbe, Max Helm, Benedikt Jaeger, Nemanja Deric, Endri Goshi, Zikai Zhou, Wolfgang Kellerer, Georg Carle |
ANCS | 9 |
| 2021 | P4Update: fast and locally verifiable consistent network updates in the P4 data planeabstractProgrammable networks come with the promise of logically centralized control, in order to optimize the network's routing behavior. However, until now, controllers are heavily involved in network operations to prevent inconsistencies such as blackholes, loops, and congestion. In this paper, we propose the P4Update framework, based on the network programming language P4, to shift the consistency control and most of the routing update logic out of the overloaded and slow control plane. As such P4Update avoids high and unnecessary control plane delays by mainly scheduling and offloading the update process to the data plane. Zikai Zhou, Wolfgang Kellerer, Andreas Blenk, Klaus-Tycho Förster |
CoNEXT | 1 |
| 2019 | A Distant Supervised Relation Extraction Model with Two Denoising StrategiesabstractDistant supervised relation extraction has been an effective way to find relational facts from text. However, distant supervised method inevitably accompanies with wrongly labeled sentences. Noisy sentences lead to poor performance of relation extraction models. Though existing piecewise convolutional neural network model with sentence-level attention (PCNN+ATT) is an effective way to reduce the effect of noisy sentences, it still has two limitations. On one hand, it adopts a PCNN module as sentence encoder, which only captures local contextual features of words and might lose important information. On the other hand, it neglects the fact that not all words contribute equally to the semantics of sentences. To address these two issues, we propose a hierarchical attention-based bidirectional GRU (HA-BiGRU) model. For the first limitation, our model utilizes a BiGRU module in place of PCNN, so as to extract global contextual information. For the second limitation, our model combines word-level and sentence-level attention mechanisms, which help get accurate sentence representations. To further alleviate the wrongly labeling problem, we first calculate the co-occurrence probabilities (CP) between the shortest dependency path (SDP) and the relation labels. Based on these co-occurrence probabilities, two denoising strategies are proposed to reduce noise interference respectively from aspect of filtering labeled data and integrating CP information into model. Experimental results on the corpus of Freebase and New York Times (Freebase+NYT) show that the HA-BiGRU model outperforms baseline models, and the two co-occurrence probabilities based denoising strategies can improve robustness of HA-BiGRU model. Zikai Zhou, Yi Cai 0001, Jiayuan Xie, Qing Li 0001, Haoran Xie 0001 |
IJCNN | 1 |