VLDB 2026 Research / reviewers in the wild / expert
Ziyu Ye
dblp:191/8005
· DBLP profile ↗
11ranked-venue papers
5as first author
10since 2021 · last 2025
0000-0003-1256-7576ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Reinforcement learning · 34% Language models and text generation · 26% Multi-agent systems · 21% | |
| Databases, data mining, and information retrieval
2 papers |
Data mining · 72% Machine learning and data management · 28% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-AI interaction · 100% |
Topics — the 21 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
LLM agents |
1.0 | 2 | 2025 | Can Large Language Model Agents Simulate Human Trust Behavior? · NeurIPS 2024 OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation · NeurIPS 2025 |
Knowledge, reasoning and agents › Multi-agent systems › agent architecture
hierarchical multi-agent framework |
0.9 | 1 | 2025 | OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation · NeurIPS 2025 |
Knowledge, reasoning and agents › Multi-agent systems
LLM-based multi-agent systems |
0.9 | 1 | 2025 | OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation · NeurIPS 2025 |
Machine learning › Reinforcement learning
multi-agent reinforcement learning |
0.9 | 1 | 2025 | OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation · NeurIPS 2025 |
Natural language and speech › Language models and text generation › prompting
prompt evolution |
0.9 | 1 | 2025 | Reward-Guided Prompt Evolving in Reinforcement Learning for LLMs · ICML 2025 |
Natural language and speech › Language models and text generation › prompting › prompt engineering
prompt optimization |
0.9 | 1 | 2025 | Reward-Guided Prompt Evolving in Reinforcement Learning for LLMs · ICML 2025 |
Machine learning › Reinforcement learning › reinforcement learning for NLP
reinforcement learning for language models |
0.9 | 1 | 2025 | Reward-Guided Prompt Evolving in Reinforcement Learning for LLMs · ICML 2025 |
Machine learning › Reinforcement learning
reinforcement learning from human feedback |
0.9 | 1 | 2025 | Reward-Guided Prompt Evolving in Reinforcement Learning for LLMs · ICML 2025 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.8 | 1 | 2024 | Understanding the Role of Equivariance in Self-supervised Learning · NeurIPS 2024 |
Machine learning › Representation and self-supervised learning › contrastive learning › self-supervised contrastive learning
equivariant contrastive learning |
0.8 | 1 | 2024 | Understanding the Role of Equivariance in Self-supervised Learning · NeurIPS 2024 |
Knowledge, reasoning and agents › Multi-agent systems › agent-based simulation
social simulation |
0.8 | 1 | 2024 | Can Large Language Model Agents Simulate Human Trust Behavior? · NeurIPS 2024 |
Machine learning › Reinforcement learning › bandit
contextual bandit |
0.7 | 1 | 2023 | Follow-ups Also Matter: Improving Contextual Bandits via Post-serving Contexts · NeurIPS 2023 |
Machine learning › Reinforcement learning
regret minimization |
0.7 | 1 | 2023 | Follow-ups Also Matter: Improving Contextual Bandits via Post-serving Contexts · NeurIPS 2023 |
Data mining › predictive modeling › classification
decision tree learning |
0.7 | 1 | 2023 | Efficient Online Decision Tree Learning with Active Feature Acquisition · IJCAI 2023 |
Machine learning and data management
online learning |
0.7 | 1 | 2023 | Efficient Online Decision Tree Learning with Active Feature Acquisition · IJCAI 2023 |
Machine learning › Trustworthy machine learning
robustness |
0.5 | 1 | 2021 | Understanding the Effect of Bias in Deep Anomaly Detection · IJCAI 2021 |
Data mining
anomaly detection |
0.5 | 1 | 2021 | Understanding the Effect of Bias in Deep Anomaly Detection · IJCAI 2021 |
Data mining › anomaly detection
deep anomaly detection |
0.5 | 1 | 2021 | Understanding the Effect of Bias in Deep Anomaly Detection · IJCAI 2021 |
Natural language and speech › Language models and text generation › LLM agents › tool use
tool-calling agents |
0.3 | 1 | 2025 | OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation · NeurIPS 2025 |
Computational finance and economics
behavioral economics |
0.2 | 1 | 2024 | Can Large Language Model Agents Simulate Human Trust Behavior? · NeurIPS 2024 |
Machine learning › Probabilistic and Bayesian machine learning › experimental design
active feature acquisition |
0.2 | 1 | 2023 | Efficient Online Decision Tree Learning with Active Feature Acquisition · IJCAI 2023 |
Methods — techniques the papers use, named apart from their topics
behavioral alignment analysis · 2.3trust game · 1.5active learning · 1.3self-play · 0.9reinforcement learning · 0.9modular architecture · 0.9hierarchical planning · 0.9direct preference optimization · 0.9RLOO · 0.9trust games · 0.8information-theoretic analysis · 0.8posterior sampling · 0.7adaptive submodularity · 0.7finite sample rate analysis · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Explain Implicit Semantic of Cyberbullying Text by Knowledge EnhancementabstractCyberbullying detection remains a challenging due to the prevalence of implicit aggression and context-dependent language in online speech. Existing methods often rely on surfacelevel lexical meaning, which limits their ability to identify hostile intent conveyed through sarcasm, indirect references, or cultural nuance. This work proposes ISKE (Implicit Semantic Knowledge Enhancement), a novel framework that integrates external commonsense knowledge into pre-trained language models (PLMs) to improve understanding implicit meaning of cyberbullying content without context. ISKE firstly extracts sentiment-bearing entity pairs using aspect-based sentiment triplet extraction (ASTE), then retrievals a knowledge graph to collect semantically coherent relational paths centered on negative sentiment. These paths are transformed into structured triplets and injected into PLMs via adapter training on a relation prediction task. Experimental results on three benchmark datasets demonstrate that ISKE consistently improves$F_{1}$-scores over PLM baselines, with gains up to$\mathbf{1 1. 1 6 \%}$, especially in cases involving implicit meaning. This work highlights the value of structured knowledge in enhancing semantic reasoning for abusive language detection. Ziyu Ye, Huakang Li |
CW | 1 |
| 2025 | Reward-Guided Prompt Evolving in Reinforcement Learning for LLMsabstractExisting reinforcement learning (RL) methods for large language models (LLMs) rely on static prompt sets, where prompts are curated a priori, and sampled in a fixed schedule for training, regardless of their usefulness to the RL process. We design eva, the first method that allows LLMs to prioritize and adaptively create useful prompts during RL training by reward signals. In principle, eva (Evolving via A symmetric Self-Play) casts language model training as a game between: (1) a creator, who samples and generates training prompts, and (2) a solver, who generates responses to the prompts. eva is simple, suits both offline and online RL for LLMs, and sets a new state-of-the-art on challenging benchmarks without extra human prompts: it improves gemma-2-9b-it’s win-rate on Arena-Hard from 51.6% to 60.1% by DPO and 52.6% to 62.4% by RLOO, surpassing claude-3-opus and nearing gemini-1.5-pro, both are orders of magnitude larger. Further ablation studies show eva can induce meaningful learning curriculum, and effectively scale RL for LLMs beyond static human prompts. Ziyu Ye, Rishabh Agarwal, Tianqi Liu 0002, Rishabh Joshi, Sarmishta Velury, Quoc V. Le, Qijun Tan |
ICML | 1 |
| 2025 | OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task AutomationabstractLarge Language Model (LLM)-based multi-agent systems show promise for automating real-world tasks but struggle to transfer across domains due to their domain-specific nature.
Current approaches face two critical shortcomings: they require complete architectural redesign and full retraining of all components when applied to new domains.
We introduce **Workforce**, a hierarchical multi-agent framework that decouples strategic planning from specialized execution through a modular architecture comprising:
*(i)* a *domain-agnostic* **Planner** for task decomposition,
*(ii)* a **Coordinator** for subtask management, and
*(iii)* specialized **Workers** with *domain-specific* tool-calling capabilities.
This decoupling enables cross-domain transferability during both inference and training phases:
During inference, Workforce seamlessly adapts to new domains by adding or modifying worker agents;
For training, we introduce **Optimized Workforce Learning (OWL)**, which improves generalization across domains by optimizing a domain-agnostic planner with reinforcement learning from real-world feedback.
To validate our approach, we evaluate Workforce on the GAIA benchmark, covering various realistic, multi-domain agentic tasks.
Experimental results demonstrate Workforce achieves open-source state-of-the-art performance (**69.70%**), outperforming commercial systems like OpenAI's Deep Research by **2.34%**.
More notably, our OWL-trained 32B model achieves **52.73%** accuracy (**+16.37%**) and demonstrates performance comparable to GPT-4o on challenging tasks.
To summarize, by enabling scalable generalization and modular domain transfer, our work establishes a foundation for the next generation of general-purpose AI assistants.
*Our code is available at [Anonymous URL](https://anonymous.4open.science/r/annonymous-owl/), and our data is available at [Anonymous URL](https://huggingface.co/anonymous21016).* Mengkang Hu, Wendong Fan, Yuzhou Nie, Ziyu Ye, Bowei Xia, Zhaoxuan Jin, Yingru Li, Qianshuo Ye, Bernard Ghanem, Ping Luo 0002, Guohao Li 0001 |
NeurIPS | 5 |
| 2025 | A Data Augmentation Approach Using Sentiment Analysis and WordNet-Based Comments Transformation
Ziyu Ye, Zhuowen Zhang, Huakang Li |
PRICAI (4) | 1 |
| 2024 | Don't Be Pessimistic Too Early: Look K Steps Ahead!
Chaoqi Wang, Ziyu Ye, Kevin Murphy 0002, Yuxin Chen 0001 |
AISTATS | 2 |
| 2024 | Understanding the Role of Equivariance in Self-supervised LearningabstractContrastive learning has been a leading paradigm for self-supervised learning, but it is widely observed that it comes at the price of sacrificing useful features (\eg colors) by being invariant to data augmentations. Given this limitation, there has been a surge of interest in equivariant self-supervised learning (E-SSL) that learns features to be augmentation-aware. However, even for the simplest rotation prediction method, there is a lack of rigorous understanding of why, when, and how E-SSL learns useful features for downstream tasks. To bridge this gap between practice and theory, we establish an information-theoretic perspective to understand the generalization ability of E-SSL. In particular, we identify a critical explaining-away effect in E-SSL that creates a synergy between the equivariant and classification tasks. This synergy effect encourages models to extract class-relevant features to improve its equivariant prediction, which, in turn, benefits downstream tasks requiring semantic features. Based on this perspective, we theoretically analyze the influence of data transformations and reveal several principles for practical designs of E-SSL. Our theory not only aligns well with existing E-SSL methods but also sheds light on new directions by exploring the benefits of model equivariance. We believe that a theoretically grounded understanding on the role of equivariance would inspire more principled and advanced designs in this field. Code is available at
https://github.com/kaotty/Understanding-ESSL. Yifei Wang 0001, Kaiwen Hu, Sharut Gupta, Ziyu Ye, Yisen Wang 0001, Stefanie Jegelka |
NeurIPS | 4 |
| 2024 | Can Large Language Model Agents Simulate Human Trust Behavior?abstractLarge Language Model (LLM) agents have been increasingly adopted as simulation tools to model humans in social science and role-playing applications. However, one fundamental question remains: can LLM agents really simulate human behavior? In this paper, we focus on one critical and elemental behavior in human interactions, trust, and investigate whether LLM agents can simulate human trust behavior. We first find that LLM agents generally exhibit trust behavior, referred to as agent trust, under the framework of Trust Games, which are widely recognized in behavioral economics. Then, we discover that GPT-4 agents manifest high behavioral alignment with humans in terms of trust behavior, indicating the feasibility of simulating human trust behavior with LLM agents. In addition, we probe the biases of agent trust and differences in agent trust towards other LLM agents and humans. We also explore the intrinsic properties of agent trust under conditions including external manipulations and advanced reasoning strategies. Our study provides new insights into the behaviors of LLM agents and the fundamental analogy between LLMs and humans beyond value alignment. We further illustrate broader implications of our discoveries for applications where trust is paramount. Chengxing Xie, Canyu Chen, Feiran Jia, Ziyu Ye, Shiyang Lai, Kai Shu, Jindong Gu, Adel Bibi, Ziniu Hu, David Jurgens, Philip Torr 0001, Bernard Ghanem, Guohao Li 0001 |
NeurIPS | 4 |
| 2023 | Efficient Online Decision Tree Learning with Active Feature AcquisitionabstractConstructing decision trees online is a classical machine learning problem. Existing works often assume that features are readily available for each incoming data point. However, in many real world applications, both feature values and the labels are unknown a priori and can only be obtained at a cost. For example, in medical diagnosis, doctors have to choose which tests to perform (i.e., making costly feature queries) on a patient in order to make a diagnosis decision (i.e., predicting labels). We provide a fresh perspective to tackle this practical challenge. Our framework consists of an active planning oracle embedded in an online learning scheme for which we investigate several information acquisition functions. Specifically, we employ a surrogate information acquisition function based on adaptive submodularity to actively query feature values with a minimal cost, while using a posterior sampling scheme to maintain a low regret for online prediction. We demonstrate the efficiency and effectiveness of our framework via extensive experiments on various real-world datasets. Our framework also naturally adapts to the challenging setting of online learning with concept drift and is shown to be competitive with baseline models while being more flexible. Arman Rahbar, Ziyu Ye, Yuxin Chen 0001, Morteza Haghir Chehreghani |
IJCAI | 2 |
| 2023 | Follow-ups Also Matter: Improving Contextual Bandits via Post-serving ContextsabstractStandard contextual bandit problem assumes that all the relevant contexts are observed before the algorithm chooses an arm. This modeling paradigm, while useful, often falls short when dealing with problems in which additional valuable contexts can be observed after arm selection. For example, content recommendation platforms like Youtube, Instagram, Tiktok receive much additional features about a user's reward after the user clicks a content (e.g., how long the user stayed, what is the user's watch speed, etc.). To improve online learning efficiency in these applications, we study a novel contextual bandit problem with post-serving contexts and design a new algorithm, poLinUCB, that achieves tight regret under standard assumptions. Core to our technical proof is a robustified and generalized version of the well-known Elliptical Potential Lemma (EPL), which can accommodate noise in data. Such robustification is necessary for tackling our problem, though we believe it could also be of general interest.
Extensive empirical tests on both synthetic and real-world datasets demonstrate the significant benefit of utilitzing post-serving contexts as well as the superior performance of our algorithm over the state-of-the-art approaches. Chaoqi Wang, Ziyu Ye, Zhe Feng 0004, Ashwinkumar Badanidiyuru |
NeurIPS | 2 |
| 2021 | Understanding the Effect of Bias in Deep Anomaly DetectionabstractAnomaly detection presents a unique challenge in machine learning, due to the scarcity of labeled anomaly data. Recent work attempts to mitigate such problems by augmenting training of deep anomaly detection models with additional labeled anomaly samples. However, the labeled data often does not align with the target distribution and introduces harmful bias to the trained model. In this paper, we aim to understand the effect of a biased anomaly set on anomaly detection. Concretely, we view anomaly detection as a supervised learning task where the objective is to optimize the recall at a given false positive rate. We formally study the relative scoring bias of an anomaly detector, defined as the difference in performance with respect to a baseline anomaly detector. We establish the first finite sample rates for estimating the relative scoring bias for deep anomaly detection, and empirically validate our theoretical results on both synthetic and real-world datasets. We also provide an extensive empirical study on how a biased training anomaly set affects the anomaly score function and therefore the detection performance on different anomaly classes. Our study demonstrates scenarios in which the biased anomaly set can be useful or problematic, and provides a solid benchmark for future research. Ziyu Ye, Yuxin Chen 0001, Haitao Zheng 0001 |
IJCAI | 1 |
| 2018 | Joint Energy Optimization of Video Encoding and TransmissionabstractDisposable wireless video sensors have many potential applications but are subject to stringent energy constraints. We studied the minimization of end-to-end distortion under an total energy constraint, by means of optimizing FEC code rate, number of source bits, and energy allocation between video encoding and wireless transmission. A two-step approach is employed. First, the FEC rate is optimized by exhaustive search. Then a binary-search-based algorithm is proposed to optimize the energy allocation and number of source bits. Experiments show that the algorithm achieves a PSNR gain up to 1dB over some reasonable baselines. A simpler suboptimal algorithm is also tested and exhibits similar performance. Ziyu Ye, Rana D. Hegazy, Pamela C. Cosman, Laurence B. Milstein |
PCS | 1 |