VLDB 2026 Research / reviewers in the wild / expert
Zenan Zhou
dblp:41/9582
· DBLP profile ↗
18ranked-venue papers
2as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A frequency mixing single-stream framework with LoRA prompt tuning for RGBD tracking
Dawei Zhang 0002, Kaiwei Jiang, Zhou Ou, Yufan Zhu, Zenan Zhou, Xiaowei He 0003, Zhonglong Zheng, Jun Zhang 0003 |
Neurocomputing | 5 |
| 2026 | Toward Graph Data Collaboration in a Data-Sharing-Free Manner: A Novel Privacy-Preserving Graph Pretraining ModelabstractGraph data, prevalent in various domains such as telecommunication, supply chain, and social networks, holds significant potential for business, operations, and social administration. Collaborating on graph data across institutions or users can further unleash its value, making it a highly sought-after practice. However, such collaboration poses risks to information privacy and commercial confidentiality. In response, we introduce an innovative new model-sharing strategy for graph data collaboration. Here, a data owner pretrains a graph neural network (GNN) model on their private graph data and then provides model users with query access to this model. The pretrained GNN acts as an intermediary, encapsulating knowledge from the private data without exposing it directly. Two fundamental principles are essential for such a pretrained GNN model: model generalizability and privacy preservation. However, current efforts often fail to achieve both concurrently. To tackle this challenge and promote an open yet secure graph data collaboration framework, we propose a novel privacy-preserving operator. This operator integrates smoothly with graph data augmentation and graph contrastive learning, allowing the pretraining of a GNN that effectively eliminates private links at high risk of exposure while maintaining generalizability. Additionally, to improve model generalizability, we introduce a new method called generalizability learning to enhance the model’s adaptability when deployed on unseen data of model user. This approach is designed to simulate diverse environments and develop representations that remain invariant across these varied environments. Extensive experiments suggest that our model surpasses existing state-of-the-art approaches in striking an effective balance between privacy preservation and generalizability. History: Accepted by Ram Ramesh, Area Editor for Data Science & Machine Learning. Funding: This work was supported (to J. Xu) by the National Natural Science Foundation of China [Grants 62206056, 72271059, and 72442011] and the CIPSC-SMP-Zhipu Large Model Cross-Disciplinary Fund. Supplemental Material: The software that supports the findings of this study is available within the paper and its Supplemental Information ( https://pubsonline.informs.org/doi/suppl/10.1287/ijoc.2023.0115 ) as well as from the IJOC GitHub software repository ( https://github.com/INFORMSJoC/2023.0115 ). The complete IJOC Software and Data Repository is available at https://informsjoc.github.io/ . Jiarong Xu, Jiaan Wang, Zenan Zhou, Tian Lu 0002 |
INFORMS J. Comput. | 3 |
| 2025 | MM-Verify: Enhancing Multimodal Reasoning with Chain-of-Thought VerificationabstractLinzhuang Sun, Hao Liang, Jingxuan Wei, Bihui Yu, Tianpeng Li, Fan Yang, Zenan Zhou, Wentao Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Linzhuang Sun, Hao Liang 0017, Jingxuan Wei, Bihui Yu, Tianpeng Li, Fan Yang 0132, Zenan Zhou, Wentao Zhang 0001 |
ACL (1) | 7 |
| 2025 | CFBench: A Comprehensive Constraints-Following Benchmark for LLMsabstractTao Zhang, ChengLIn Zhu, Yanjun Shen, Wenjing Luo, Yan Zhang, Hao Liang, Tao Zhang, Fan Yang, Mingan Lin, Yujing Qiao, Weipeng Chen, Bin Cui, Wentao Zhang, Zenan Zhou. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Tao Zhang 0194, Chenglin Zhu, Wenjing Luo, Yan Zhang 0109, Hao Liang 0017, Fan Yang 0132, Yujing Qiao, Weipeng Chen, Bin Cui 0001, Wentao Zhang 0001, Zenan Zhou |
ACL (1) | 13 |
| 2025 | Beyond Sight: Towards Cognitive Alignment in LVLM via Enriched Visual KnowledgeabstractDoes seeing always mean knowing? Large Vision-Language Models (LVLMs) integrate separately pre-trained vision and language components, often using CLIP-ViT as vision backbone. However, these models frequently encounter a core issue of "cognitive misalignment" between the vision encoder (VE) and the large language model (LLM). Specifically, the VE’s representation of visual information may not fully align with LLM’s cognitive framework, leading to a mismatch where visual features exceed the language model’s interpretive range. To address this, we investigate how variations in VE representations influence LVLM comprehension, especially when the LLM faces VE-Unknown data—images whose ambiguous visual representations challenge the VE’s interpretive precision. Accordingly, we construct a multi-granularity landmark dataset and systematically examine the impact of VE-Known and VE-Unknown data on interpretive abilities. Our results show that VE-Unknown data limits LVLM’s capacity for accurate understanding, while VE-Known data, rich in distinctive features, helps reduce cognitive misalignment. Building on these insights, we propose Entity-Enhanced Cognitive Alignment (EECA), a method that employs multi-granularity supervision to generate visually enriched, well-aligned tokens that not only integrate within the embedding space but also align with the LLM’s cognitive framework. This alignment markedly enhances LVLM performance in landmark recognition. Our findings underscore the challenges posed by VE-Unknown data and highlight the essential role of cognitive alignment in advancing multimodal systems. Yuanyang Yin, Victor Shea-Jay Huang, Weipeng Chen, Baoqun Yin, Zenan Zhou |
CVPR | 9 |
| 2025 | FB-Bench: A Fine-Grained Multi-Task Benchmark for Evaluating LLMs' Responsiveness to Human FeedbackabstractHuman feedback is crucial in the interactions between humans and Large Language Models (LLMs). However, existing research primarily focuses on benchmarking LLMs in single-turn dialogues. Even in benchmarks designed for multi-turn dialogues, the user utterances are often independent, neglecting the nuanced and complex nature of human feedback within real-world usage scenarios. To fill this research gap, we introduce FB-Bench, a fine-grained, multi-task benchmark designed to evaluate LLMs’ responsiveness to human feedback under real-world usage scenarios in Chinese. Drawing from the two main interaction scenarios, FB-Bench comprises 591 meticulously curated samples, encompassing eight task types, five deficiency types of response, and nine feedback types. We extensively evaluate a broad array of popular LLMs, revealing significant variations in their performance across different interaction scenarios. Further analysis indicates that task, human feedback, and deficiencies of previous responses can also significantly impact LLMs’ responsiveness. Our findings underscore both the strengths and limitations of current models, providing valuable insights and directions for future research. Youquan Li, Miao Zheng, Fan Yang 0132, Guosheng Dong, Bin Cui 0001, Weipeng Chen, Zenan Zhou, Wentao Zhang 0001 |
EMNLP | 7 |
| 2025 | Training Data Distribution Estimation for Optimized Pre-training Data ManagementabstractLarge language models (LLMs) have demonstrated exceptional performance across a wide range of tasks and domains, with data preparation playing a critical role in achieving these results. Pretraining data typically combines information from multiple domains. To maximize performance when integrating data from various domains, determining the optimal data distribution is essential. However, state-of-the-art (SOTA) LLMs rarely disclose details about their pretraining data, making it difficult for researchers to identify ideal data distributions. In this paper, we introduce a new approach, data distribution estimation, which enables the automatic estimation of pretraining data distributions by analyzing the generated outputs of LLMs. We provide rigorous theoretical proofs, practical algorithms, and preliminary experimental results for data distribution estimation. Based on these findings, we offer valuable insights into the challenges and future directions for effective data distribution estimation and data management. The source code, data, and other artifacts are available at https://github.com/yangyajie0625/data_detection Hao Liang 0017, Keshi Zhao, Yajie Yang, Bin Cui 0001, Zenan Zhou, Wentao Zhang 0001 |
ICDE | 5 |
| 2025 | DataSculpt: A Holistic Data Management Framework for Long-Context LLMs TrainingabstractIn recent years, foundation models, particularly large language models (LLMs), have demonstrated significant improvements across a variety of tasks. One of their most important features is long-context capability, which enables them to generate extended text with high semantic coherence, retrieving relevant information, and handling tasks with substantial amounts of text efficiently. The key to improving long-context performance lies in effective data organization and management strategies that integrate data from multiple domains and optimize the context window during training. Through extensive experimental analysis, we identified three key challenges in designing effective data management strategies that enable the model to achieve long-context capability without sacrificing performance in other tasks: (1) a shortage of long documents across multiple domains, (2) effective construction of context windows, and (3) efficient organization of large-scale datasets. To address these challenges, we introduce DataSculpt, a novel data management framework designed for long-context training. We first formulate the organization of training data as a multi-objective optimization problem, focusing on attributes including the relevance among documents within the same training sequence, the quantity of concatenated instances, individual document integrity, and computational cost. Specifically, our approach utilizes a coarse-to-fine method to optimize training data organization effectively. We begin by clustering the data based on semantic similarity (coarse), followed by a multi-objective greedy search within each cluster to score and concatenate documents into various context windows (fine). We have deployed DataSculpt as the data management backend for long-context training in Baichuan Inc. Extensive experiments with diverse downstream tasks show that DataSculpt enhances the model's long-context performance by an average of 15.73%, while maintaining the general capabilities with a 4.63% improvement. Keer Lu, Xiaonan Nie, Da Pan 0003, Shusen Zhang, Keshi Zhao, Weipeng Chen, Zenan Zhou, Guosheng Dong, Bin Cui 0001, Wentao Zhang 0001 |
ICDE | 8 |
| 2025 | PAS: Plug-and-Play Prompt Augmentation SystemabstractIn recent years, the rise of Large Language Models (LLMs) has spurred a growing demand for plug-and-play AI systems. Among the various AI techniques, prompt engineering stands out as particularly significant. However, users often face challenges in writing prompts due to the steep learning curve and significant time investment, and existing automatic prompt engineering (APE) models can be difficult to use. To address this issue, we propose PAS, an LLM-based plug-and-play APE system. PAS utilizes LLMs trained on high-quality, automatically generated prompt augmentation datasets, resulting in exceptional performance. In comprehensive benchmarks, PAS achieves state-of-the-art (SOTA) results compared to previous APE models, with an average improvement of 6.09 points. Moreover, PAS is highly efficient, achieving SOTA performance with only 9000 data points. Additionally, PAS can autonomously generate prompt augmentation data without requiring additional human labor. Its flexibility also allows it to be compatible with all existing LLMs and applicable to a wide range of tasks. Moreover, we deployed PAS for Baichuan online model, and then tested PAS using internal human evaluations in Baichuan underscoring its strong performance. This combination of high performance, efficiency, and flexibility makes PAS a valuable system for enhancing the usability and effectiveness of LLMs through automatic prompt engineering. The codebase is available at https://github.com/PKU-Baichuan-MLSystemLab/PAS. Miao Zheng, Hao Liang 0017, Fan Yang 0132, Bin Cui 0001, Zenan Zhou, Wentao Zhang 0001 |
ICDE | 5 |
| 2025 | Facilitating Multi-turn Function Calling for LLMs via Compositional Instruction TuningabstractLarge Language Models (LLMs) have exhibited significant potential in performing diverse tasks, including the ability to call functions or use external tools to enhance their performance. While current research on function calling by LLMs primarily focuses on single-turn interactions, this paper addresses the overlooked necessity for LLMs to engage in multi-turn function calling—critical for handling compositional, real-world queries that require planning with functions but not only use functions. To facilitate this, we introduce an approach, BUTTON, which generates synthetic compositional instruction tuning data via bottom-up instruction construction and top-down trajectory generation. In the bottom-up phase, we generate simple atomic tasks based on real-world scenarios and build compositional tasks using heuristic strategies based on atomic tasks. Corresponding function definitions are then synthesized for these compositional tasks. The top-down phase features a multi-agent environment where interactions among simulated humans, assistants, and tools are utilized to gather multi-turn function calling trajectories. This approach ensures task compositionality and allows for effective function and trajectory generation by examining atomic tasks within compositional tasks. We produce a dataset BUTTONInstruct comprising 8k data points and demonstrate its effectiveness through extensive experiments across various LLMs. Mingyang Chen 0002, Haoze Sun, Tianpeng Li, Fan Yang 0132, Hao Liang 0017, Keer Lu, Bin Cui 0001, Wentao Zhang 0001, Zenan Zhou, Weipeng Chen |
ICLR | 9 |
| 2025 | SysBench: Can LLMs Follow System Message?abstractLarge Language Models (LLMs) have become instrumental across various applications, with the customization of these models to specific scenarios becoming increasingly critical. System message, a fundamental component of LLMs, is consist of carefully crafted instructions that guide the behavior of model to meet intended goals. Despite the recognized potential of system messages to optimize AI-driven solutions, there is a notable absence of a comprehensive benchmark for evaluating how well LLMs follow system messages. To fill this gap, we introduce SysBench, a benchmark that systematically analyzes system message following ability in terms of three limitations of existing LLMs: constraint violation, instruction misjudgement and multi-turn instability. Specifically, we manually construct evaluation dataset based on six prevalent types of constraints, including 500 tailor-designed system messages and multi-turn user conversations covering various interaction relationships. Additionally, we develop a comprehensive evaluation protocol to measure model performance. Finally, we conduct extensive evaluation across various existing LLMs, measuring their ability to follow specified constraints given in system messages. The results highlight both the strengths and weaknesses of existing models, offering key insights and directions for future research. Yanzhao Qin, Tao Zhang 0194, Wenjing Luo, Haoze Sun, Yan Zhang 0109, Yujing Qiao, Weipeng Chen, Zenan Zhou, Wentao Zhang 0001, Bin Cui 0001 |
ICLR | 10 |
| 2025 | MathScape: Benchmarking Multimodal Large Language Models in Real-World Mathematical ContextsabstractWith the rapid progress of Multimodal LLMs, evaluating their mathematical reasoning capabilities has become an increasingly important research direction. In particular, visual-textual mathematical reasoning serves as a key indicator of an MLLM's ability to comprehend and solve complex, multi-step quantitative problems. While existing benchmarks such as MathVista and MathVerse have advanced the evaluation of multimodal math proficiency, they primarily rely on digitally rendered content and fall short in capturing the complexity of real-world scenarios. To bridge this gap, we introduce MathScape, a novel benchmark focused on assessing MLLMs' reasoning ability in realistic mathematical contexts. MathScape comprises 1,369 high-quality math problems paired with human-captured real-world images, closely reflecting the challenges encountered in practical educational settings. We conduct a thorough multi-dimensional evaluation across nine leading closed-source MLLMs, three open-source MLLMs with over 20 billion parameters, and seven smaller-scale MLLMs. Our results show that even SOTA models struggle with real-world math tasks, lagging behind human performance-highlighting critical limitations in current model capabilities. Moreover, we find that strong performance on synthetic or digitally rendered images does not guarantee similar effectiveness on real-world tasks. This underscores the necessity of MathScape in the next stage of multimodal mathematical reasoning. Hao Liang 0017, Linzhuang Sun, zhouminxuan zhouminxuan, Zirong Chen, Meiyi Qiang, Tianpeng Li, Fan Yang 0132, Zenan Zhou, Wentao Zhang 0001 |
ACM Multimedia | 9 |
| 2025 | ReSearch: Learning to Reason with Search for LLMs via Reinforcement LearningabstractLarge Language Models (LLMs) have shown remarkable capabilities in reasoning, exemplified by the success of OpenAI-o1 and DeepSeek-R1. However, integrating reasoning with external search processes remains challenging, especially for complex multi-hop questions requiring multiple retrieval steps. We propose ReSearch, a novel framework that trains LLMs to Reason with Search via reinforcement learning without using any supervised data on reasoning steps. Our approach treats search operations as integral components of the reasoning chain, where when and how to perform searches is guided by text-based thinking, and search results subsequently influence further reasoning. We train ReSearch on Qwen2.5-7B(-Instruct) and Qwen2.5-32B(-Instruct) models and conduct extensive experiments. Despite being trained on only one dataset, our models demonstrate strong generalizability across various benchmarks. Analysis reveals that ReSearch naturally elicits advanced reasoning capabilities such as reflection and self-correction during the reinforcement learning process. Mingyang Chen 0002, Linzhuang Sun, Tianpeng Li, Haoze Sun, Chenzheng Zhu, Haofen Wang, Jeff Z. Pan, Wen Zhang 0015, Huajun Chen, Fan Yang 0132, Zenan Zhou, Weipeng Chen |
NeurIPS | 12 |
| 2025 | Social media meets FinTech platforms: How do online emotions support credit risk decision-making?
Zenan Zhou, Zhichen Chen, Yingjie Zhang 0003, Tian Lu 0002, Xianghua Lu |
Decis. Support Syst. | 1 |
| 2024 | EARNet: Error-Aware Reconstruction Network for no-reference image quality assessment
Zhiheng Zhou 0001, Zenan Zhou, Xiyuan Tao, Zerui Yu, Yinglie Cao |
Expert Syst. Appl. | 2 |
| 2022 | SAR Ground Moving Target Refocusing by Combining mRe³ Network and TVβ-LSTMabstractThis article proposes a novel framework by combining a modified real-time recurrent regression (mRe³) network and a newly designed trajectory smoothing long short-term memory (LSTM) network for refocusing the ground moving target (GMT) in the synthetic aperture radar (SAR) image. The mRe^3 network that consists of a convolutional neural network (CNN) backbone and two LSTM modules is designed to track the GMT's shadow in an SAR video. Furthermore, we find that the complex trajectory obtained by the tracking network cannot directly be used for refocusing the GMT because of the estimation error. To address the abovementioned problem, a β-order total variation loss-based smoothing LSTM (TVβ-LSTM) is proposed to recover the GMT's trajectory to meet the requirement of refocusing. Besides, the effect of TVβ on the performance of smoothing LSTM is analyzed. By the experiments on simulated and real SAR videos, we find that the mRe^3 has stronger robustness and a better trajectory reconstruction precision compared with the existing tracking methods, especially for the strong interference cases. In addition, the smoothing LSTM can recover the trajectory of the GMT with higher precision and better smoothness. When β is set to 3, with the TVβ-LSTM, the center distance error of a recovered complex trajectory can be reduced from 0.82 to 0.782, while its fluctuation can be suppressed from 6 to 1 mm. By using our framework, the focused GMT with bountiful geometrical features can be obtained even for the K_a-band SAR. Yuanyuan Zhou 0007, Jun Shi 0002, Chen Wang 0041, Yao Hu 0006, Zenan Zhou, Xiaqing Yang, Xiaoling Zhang 0002, Shunjun Wei |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2021 | Semisupervised Learning-Based SAR ATR via Self-Consistent AugmentationabstractIn synthetic aperture radar (SAR) automatic target recognition, it is expensive and time-consuming to annotate the targets. Thus, training a network with a few labeled data and plenty of unlabeled data attracts attention of many researchers. In this article, we design a semisupervised learning framework including self-consistent augmentation rule, mixup-based mixture, and weighted loss, which allows a classification network to utilize unlabeled data during training and ultimately alleviates the demand of labeled data. The proposed self-consistent augmentation rule forces the samples before and after augmentation to share the same labels to utilize the unlabeled data, which can ensure the prominent effect of supervised learning part of the framework for training by balancing amounts of labeled and unlabeled samples in a minibatch, and makes the network achieve better performance. Then, a mixture method is introduced to mix the labeled, unlabeled, and augmented samples for the better involvement of label information in the mixed samples. By using cross-entropy loss for the mixed-labeled mixtures and mean-squared error loss for the mixed-unlabeled mixtures, the total loss is defined as the weighted sum of them. The experiments on the MSTAR data set and OpenSARShip data set show that the performance of the method is not only far better than the state of the art among current semisupervised-based classifiers but also near to the state of the art among the supervised learning-based networks. Chen Wang 0041, Jun Shi 0002, Yuanyuan Zhou 0007, Xiaqing Yang, Zenan Zhou, Shunjun Wei, Xiaoling Zhang 0002 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2011 | MaxFirst for MaxBRkNNabstractThe MaxBRNN problem finds a region such that setting up a new service site within this region would guarantee the maximum number of customers by proximity. This problem assumes that each customer only uses the service provided by his/her nearest service site. However, in reality, a customer tends to go to his/her k nearest service sites. To handle this, MaxBRNN can be extended to the MaxBRkNN problem which finds an optimal region such that setting up a service site in this region guarantees the maximum number of customers who would consider the site as one of their k nearest service locations. We further generalize the MaxBRkNN problem to reflect the real world scenario where customers may have different preferences for different service sites, and at the same time, service sites may have preferred targeted customers. In this paper, we present an efficient solution called MaxFirst to solve this generalized MaxBRkNN problem. The algorithm works by partitioning the space into quadrants and searches only in those quadrants that potentially contain an optimal region. During the space partitioning, we compute the upper and lower bounds of the size of a quadrant's BRkNN, and use these bounds to prune the unpromising quadrants. Experiment results show that MaxFirst can be two to three orders of magnitude faster than the state-of-the-art algorithm. Zenan Zhou, Wei Wu 0020, Xiaohui Li 0002, Mong-Li Lee, Wynne Hsu |
ICDE | 1 |