VLDB 2026 Research / reviewers in the wild / expert
Ziwei Zheng
dblp:188/2801
· DBLP profile ↗
20ranked-venue papers
7as first author
19since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Computer networks · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Reciprocal encoding and aligning of semantic and topological information for causality graph event prediction
Ziwei Zheng, Chuanhong Zhan, Bang Wang 0001 |
Expert Syst. Appl. | 2 |
| 2026 | HeadHunt-VAD: Hunting Robust Anomaly-Sensitive Heads in MLLM for Tuning-Free Video Anomaly DetectionabstractVideo Anomaly Detection (VAD) aims to locate events that deviate from normal patterns in videos. Traditional approaches often rely on extensive labeled data and incur high computational costs. Recent tuning-free methods based on Multimodal Large Language Models (MLLMs) offer a promising alternative by leveraging their rich world knowledge. However, these methods typically rely on textual outputs, which introduces information loss, exhibits normalcy bias, and suffers from prompt sensitivity, making them insufficient for capturing subtle anomalous cues. To address these constraints, we propose HeadHunt-VAD, a novel tuning-free VAD paradigm that bypasses textual generation by directly hunting robust anomaly-sensitive internal attention heads within the frozen MLLM. Central to our method is a Robust Head Identification module that systematically evaluates all attention heads using a multi-criteria analysis of saliency and stability, identifying a sparse subset of heads that are consistently discriminative across diverse prompts. Features from these expert heads are then fed into a lightweight anomaly scorer and a temporal locator, enabling efficient and accurate anomaly detection with interpretable outputs. Extensive experiments show that HeadHunt-VAD achieves state-of-the-art performance among tuning-free methods on two major VAD benchmarks while maintaining high efficiency, validating head-level probing in MLLMs as a powerful and practical solution for real-world anomaly detection. Zhaolin Cai, Fan Li 0003, Ziwei Zheng, Haixia Bi, Lijun He 0001 |
AAAI | 3 |
| 2026 | Experience-driven Multi-turn Reinforcement Learning for GUI AgentsabstractZhengxi Lu, Jiabo Ye, Fei Tang, Yongliang Shen, Haiyang Xu, Ziwei Zheng, Weiming Lu, Ming Yan, Fei Huang, Jun Xiao, Yueting Zhuang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhengxi Lu, Jiabo Ye, Fei Tang 0005, Yongliang Shen 0001, Haiyang Xu 0001, Ziwei Zheng, Weiming Lu 0001, Ming Yan 0008, Fei Huang 0002, Jun Xiao 0001, Yueting Zhuang |
ACL (1) | 6 |
| 2025 | Nullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace ProjectionabstractRecent studies have shown that large vision-language models (LVLMs) often suffer from the issue of object hallucinations (OH). To mitigate this issue, we introduce an efficient method that edits the model weights based on an unsafe subspace, which we call HalluSpace in this paper. With truthful and hallucinated text prompts accompanying the visual content as inputs, the HalluSpace can be identified by extracting the hallucinated embedding features and removing the truthful representations in LVLMs. By orthog-onalizing the model weights, input features will be projected into the Null space of the HalluSpace to reduce OH, based on which we name our method Nullu. We reveal that Hal-luSpaces generally contain prior information in the large language models (LLMs) applied to build LVLMs, which have been shown as essential causes of OH in previous studies. Therefore, null space projection suppresses the LLMs’ priors to filter out the hallucinated features, resulting in contextually accurate outputs. Experiments show that our method can effectively mitigate OH across different LVLM families without extra inference costs and also show strong performance in general LVLM benchmarks. Code is released at https://github.com/Ziwei-Zheng/Nullu. Le Yang 0007, Ziwei Zheng, Boxu Chen, Zhengyu Zhao 0001, Chenhao Lin, Chao Shen 0001 |
CVPR | 2 |
| 2025 | DeepCatl: A Combination of Channel Attention Mechanism and Transformer Encoding to Predict Transcription Factor Binding Sites
Ziwei Zheng, Guangsheng Wu, Xianfang Wang |
ICIC (26) | 2 |
| 2025 | LVLM-FDA: Protecting Large Vision-Language Models via Fast Detection of Malicious Attempts
Boxu Chen, Le Yang 0007, Ziwei Zheng, Cong Wang 0001, Qian Wang 0002, Chao Shen 0001 |
KSEM (1) | 4 |
| 2025 | HiProbe-VAD: Video Anomaly Detection via Hidden States Probing in Tuning-Free Multimodal LLMsabstractVideo Anomaly Detection (VAD) aims to identify and locate deviations from normal patterns in video sequences. Traditional methods often struggle with substantial computational demands and a reliance on extensive labeled datasets, thereby restricting their practical applicability. To address these constraints, we propose HiProbe-VAD, a novel framework that leverages pre-trained Multimodal Large Language Models (MLLMs) for VAD without requiring fine-tuning. In this paper, we discover that the intermediate hidden states of MLLMs contain information-rich representations, exhibiting higher sensitivity and linear separability for anomalies compared to the output layer. To capitalize on this, we propose a Dynamic Layer Saliency Probing (DLSP) mechanism that intelligently identifies and extracts the most informative hidden states from the optimal intermediate layer during the MLLMs reasoning. Then a lightweight anomaly scorer and temporal localization module efficiently detects anomalies using these extracted hidden states and finally generate explanations. Experiments on the UCF-Crime and XD-Violence datasets demonstrate that HiProbe-VAD outperforms existing training-free and most traditional approaches. Furthermore, our framework exhibits remarkable cross-model generalization capabilities in different MLLMs without any tuning, unlocking the potential of pre-trained MLLMs for video anomaly detection and paving the way for more practical and scalable solutions. Zhaolin Cai, Fan Li 0003, Ziwei Zheng, Yanjun Qin |
ACM Multimedia | 3 |
| 2025 | When Average Delay Optimization Meets Deterministic Delay Constraint: A Renewal Framework for Resource AllocationabstractWhile the average delay is traditionally an importance metric for system performance, the emerging technologies have given birth to a variety of critical applications, for which the deterministic delay guarantee is highly desired. When optimizing the average delay objective is embraced with satisfying the deterministic delay constraint, it leads to complicated coupling and brings new challenges to resource allocation. In this paper, we propose a renewal framework for multi-user power control and subband allocation to improve the comprehensive delay performance. The average delay objective is optimized under the Markov decision process (MDP) problem, while the deterministic delay constraint is satisfied through Lyapunov optimization with virtual queues. Due to the conflict of the inter-slot influence in MDP and the i.i.d. state requirement in Lyapunov approach, we exploit the recurrent property of queue states and construct a renewal system. By solving the equivalent infinite-horizon MDP in the renewal framework, we propose a resource allocation algorithm, which is proved to be asymptotically optimal. Finally, the simulation results demonstrate that the proposed scheme meets the deterministic delay constraint and achieves better average delay performance than existing baselines. Yuze Jin, Wei Wang 0021, Ziwei Zheng, Yitu Wang, Rui Yin 0001, Zhaoyang Zhang 0001 |
IEEE Trans. Commun. | 3 |
| 2024 | DyFADet: Dynamic Feature Aggregation for Temporal Action Detection
Le Yang 0007, Ziwei Zheng, Yizeng Han, Hao Cheng 0015, Shiji Song, Gao Huang 0001, Fan Li 0003 |
ECCV (46) | 2 |
| 2024 | Fine-Grained Dynamic Network for Generic Event Boundary Detection
Ziwei Zheng, Lijun He 0001, Le Yang 0007, Fan Li 0003 |
ECCV (43) | 1 |
| 2024 | On Learnable Parameters of Optimal and Suboptimal Deep Learning Models
Ziwei Zheng, Huizhi Liang 0001, Václav Snásel, Vito Latora, Panos M. Pardalos, Guiseppe Nicosia, Varun Ojha 0001 |
ICONIP (3) | 1 |
| 2024 | Rethinking the Architecture Design for Efficient Generic Event Boundary Detection
Ziwei Zheng, Zechuan Zhang, Yulin Wang 0002, Shiji Song, Gao Huang 0001, Le Yang 0007 |
ACM Multimedia | 1 |
| 2024 | Bias-Rectified Multi-way Learning with Data Augmentation for Implicit Discourse Relation Recognition
Ziwei Zheng, Wei Xiang 0005, Bang Wang 0001 |
NLPCC (1) | 1 |
| 2024 | Adaptive Modulation and Coding for URLLC RetransmissionabstractFor ultra-reliable low-latency communication (URLLC), retransmission data should be scheduled with different priorities due to the urgency, which brings new challenges in delay-oriented optimization, especially with deterministic delay constraint. In this paper, we propose an adaptive modulation and coding (AMC) scheme for both initial transmissions and retransmissions in separate data queues. Different from most of the existing works focusing on average delay, we jointly consider delay performance and deterministic delay requirement by combining the Markov decision process (MDP) with Lyapunov optimization technique. To overcome the coupling in the objective and the deterministic constraint, we transform this problem into an infinite horizon MDP by constructing a renewal system with sampling. Based on this, we propose a delay-optimal modulation and coding scheme (MCS) selection policy using reinforcement learning. Simulation results show that the proposed scheme achieves better delay performance than the conventional AMC schemes. Yuze Jin, Wei Wang 0021, Ziwei Zheng, Yitu Wang, Zhaoyang Zhang 0001 |
WCNC | 3 |
| 2024 | A granularity-level information fusion strategy on hypergraph transformer for predicting synergistic effects of anticancer drugsabstractCombination therapy has exhibited substantial potential compared to monotherapy. However, due to the explosive growth in the number of cancer drugs, the screening of synergistic drug combinations has become both expensive and time-consuming. Synergistic drug combinations refer to the concurrent use of two or more drugs to enhance treatment efficacy. Currently, numerous computational methods have been developed to predict the synergistic effects of anticancer drugs. However, there has been insufficient exploration of how to mine drug and cell line data at different granularity levels for predicting synergistic anticancer drug combinations. Therefore, this study proposes a granularity-level information fusion strategy based on the hypergraph transformer, named HypertranSynergy, to predict synergistic effects of anticancer drugs. HypertranSynergy introduces synergistic connections between cancer cell lines and drug combinations using hypergraph. Then, the Coarse-grained Information Extraction (CIE) module merges the hypergraph with a transformer for node embeddings. In the CIE module, Contranorm is a normalization layer that mitigates over-smoothing, while Gaussian noise addresses local information gaps. Additionally, the Fine-grained Information Extraction (FIE) module assesses fine-grained information's impact on predictions by employing similarity-aware matrices from drug/cell line features. Both CIE and FIE modules are integrated into HypertranSynergy. In addition, HypertranSynergy achieved the AUC of 0.93${\pm }$0.01 and the AUPR of 0.69${\pm }$0.02 in 5-fold cross-validation of classification task, and the RMSE of 13.77${\pm }$0.07 and the PCC of 0.81${\pm }$0.02 in 5-fold cross-validation of regression task. These results are better than most of the state-of-the-art models. Wei Wang 0166, Gaolin Yuan, Shitong Wan, Ziwei Zheng, Dong Liu 0008, Juntao Li 0001, Xianfang Wang |
Briefings Bioinform. | 4 |
| 2024 | Dynamic Spatial Focus for Efficient Compressed Video Action RecognitionabstractRecent years have witnessed a growing interest in compressed video action recognition due to the rapid growth of online videos. It remarkably reduces the storage by replacing raw videos with sparsely sampled RGB frames and other compressed motion cues (motion vectors and residuals). However, existing compressed video action recognition methods face two main issues: First, the inefficiency caused by the usage of coarse-level information under full resolution, and second, the disturbing due to the noisy dynamics in motion vectors. To address the two issues, this paper proposes a dynamic spatial focus method for efficient compressed video action recognition (CoViFocus). Specifically, we first use a light-weighted two-stream architecture to localize the task-relevant patches for both the RGB frames and motion vectors. Then the selected patch pair will be processed by a high-capacity two-stream deep model for the final prediction. Such a patch selection strategy crops out the irrelevant motion noise in motion vectors, as well as reduces the spatial redundancy of the inputs, leading to the high efficiency of our method in the compressed domain. Moreover, we found that the motion vectors can help our method to address the possibly happened static-issue, which means that the focus patches get stuck at some regions related to static objects rather than target actions, which further improves our method. Extensive results on both the HMDB-51 and UCF-101 datasets demonstrate the effectiveness and efficiency of our method in compressed video action recognition tasks. Ziwei Zheng, Le Yang 0007, Yulin Wang 0002, Miao Zhang 0041, Lijun He 0001, Gao Huang 0001, Fan Li 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | OStr-DARTS: Differentiable Neural Architecture Search Based on Operation StrengthabstractDifferentiable architecture search (DARTS) has emerged as a promising technique for effective neural architecture search, and it mainly contains two steps to find the high-performance architecture. First, the DARTS supernet that consists of mixed operations will be optimized via gradient descent. Second, the final architecture will be built by the selected operations that contribute the most to the supernet. Although DARTS improves the efficiency of neural architecture search (NAS), it suffers from the well-known degeneration issue which can lead to deteriorating architectures. Existing works mainly attribute the degeneration issue to the failure of its supernet optimization, while little attention has been paid to the selection method. In this article, we cease to apply the widely-used magnitude-based selection method and propose a novel criterion based on operation strength that estimates the importance of an operation by its effect on the final loss. We show that the degeneration issue can be effectively addressed by using the proposed criterion without any modification of supernet optimization, indicating that the magnitude-based selection method can be a critical reason for the instability of DARTS. The experiments on NAS-Bench-201 and DARTS search spaces show the effectiveness of our method. Le Yang 0007, Ziwei Zheng, Yizeng Han, Shiji Song, Gao Huang 0001, Fan Li 0003 |
IEEE Trans. Cybern. | 2 |
| 2024 | Retransmission Aware Adaptive Modulation and Coding Toward Deterministic Delay PerformanceabstractUltra-reliable low-latency communication (URLLC) is an indispensable element towards supporting various latency-sensitive and reliability-critical applications. To optimize the average delay while satisfying the deterministic delay constraint, initial transmission and retransmission should be handled with different priorities due to the differentiated urgency, which creates complex interdependency and brings new technical challenges to delay-oriented optimization. In this paper, we propose a retransmission-aware adaptive modulation and coding (RAMC) scheme to improve the delay performance in URLLC scenarios. Specifically, we first establish a cascaded queue system, including an initial transmission queue and a retransmission queue. The deterministic delay constraint is satisfied through Lyapunov optimization, where we transform the Lyapunov drift-plus-penalty problem into an infinite horizon Markov decision process (MDP) by constructing a renewal system with sampling to overcome the challenge brought by queue coupling. Next, we propose the delay-optimal RAMC scheme by solving the associated Bellman equation by improved reinforcement learning, which is proved to be asymptotically optimal. Finally, the superiority of the proposed RAMC scheme is verified through simulations. Yuze Jin, Wei Wang 0021, Yitu Wang, Rui Yin 0001, Ziwei Zheng, Zhaoyang Zhang 0001 |
IEEE Trans. Wirel. Commun. | 5 |
| 2021 | Frame-Level Video Caching and Transmission Scheduling via Stochastic LearningabstractTo meet the ever-increasing demand for mobile video services, one of the effective solutions is caching some popular videos in edge nodes. In this paper, we propose an online stochastic learning algorithm with two time scales for joint caching and transmission optimization in the video frame level. To overcome the drift distortion caused by the dependency among video frames, the transmission process is formulated as an infinite horizon Markov decision process (MDP). We derive the equivalent Bellman equation and design the online value iteration algorithm via stochastic approximation for transmission. Due to the lack of the expression between the system performance and the caching policy, we design a gradient-free stochastic optimization algorithm to update the caching policy. Finally, simulation results show that our proposed algorithm achieves better performance than conventional caching algorithms. Ziwei Zheng, Wei Wang 0021, Hangguan Shan, Zhaoyang Zhang 0001 |
GLOBECOM | 1 |
| 2016 | Graph-Based Multi-Modality Learning for Clinical Decision SupportabstractThe task of clinical decision support (CDS) involves retrieval and ranking of medical journal articles for medical records of diagnosis, test or treatment. Previous studies on this task are based on bag-of-words representations of document texts and general retrieval models. In this paper, we propose to use the paragraph vector technique to learn the latent semantic representation of texts and treat the latent semantic representations and the original bag-of-words representations as two different modalities. We then propose to use the graph-based multi-modality learning algorithm for document re-ranking. Experimental results on two TREC-CDS benchmark datasets demonstrate the excellent performance of our proposed approach. Ziwei Zheng, Xiaojun Wan 0001 |
CIKM | 1 |