Ziyu Ye

dblp:191/8005 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
10since 2021 · last 2025
0000-0003-1256-7576ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Reinforcement learning · 34% Language models and text generation · 26% Multi-agent systems · 21%
Databases, data mining, and information retrieval
2 papers
Data mining · 72% Machine learning and data management · 28%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 100%

Topics — the 21 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
LLM agents
1.022025
Can Large Language Model Agents Simulate Human Trust Behavior? · NeurIPS 2024
OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation · NeurIPS 2025
Knowledge, reasoning and agents › Multi-agent systems › agent architecture
hierarchical multi-agent framework
0.912025
OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation · NeurIPS 2025
Knowledge, reasoning and agents › Multi-agent systems
LLM-based multi-agent systems
0.912025
OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation · NeurIPS 2025
Machine learning › Reinforcement learning
multi-agent reinforcement learning
0.912025
OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation · NeurIPS 2025
Natural language and speech › Language models and text generation › prompting
prompt evolution
0.912025
Reward-Guided Prompt Evolving in Reinforcement Learning for LLMs · ICML 2025
Natural language and speech › Language models and text generation › prompting › prompt engineering
prompt optimization
0.912025
Reward-Guided Prompt Evolving in Reinforcement Learning for LLMs · ICML 2025
Machine learning › Reinforcement learning › reinforcement learning for NLP
reinforcement learning for language models
0.912025
Reward-Guided Prompt Evolving in Reinforcement Learning for LLMs · ICML 2025
Machine learning › Reinforcement learning
reinforcement learning from human feedback
0.912025
Reward-Guided Prompt Evolving in Reinforcement Learning for LLMs · ICML 2025
Machine learning › Representation and self-supervised learning
contrastive learning
0.812024
Understanding the Role of Equivariance in Self-supervised Learning · NeurIPS 2024
Machine learning › Representation and self-supervised learning › contrastive learning › self-supervised contrastive learning
equivariant contrastive learning
0.812024
Understanding the Role of Equivariance in Self-supervised Learning · NeurIPS 2024
Knowledge, reasoning and agents › Multi-agent systems › agent-based simulation
social simulation
0.812024
Can Large Language Model Agents Simulate Human Trust Behavior? · NeurIPS 2024
Machine learning › Reinforcement learning › bandit
contextual bandit
0.712023
Follow-ups Also Matter: Improving Contextual Bandits via Post-serving Contexts · NeurIPS 2023
Machine learning › Reinforcement learning
regret minimization
0.712023
Follow-ups Also Matter: Improving Contextual Bandits via Post-serving Contexts · NeurIPS 2023
Data mining › predictive modeling › classification
decision tree learning
0.712023
Efficient Online Decision Tree Learning with Active Feature Acquisition · IJCAI 2023
Machine learning and data management
online learning
0.712023
Efficient Online Decision Tree Learning with Active Feature Acquisition · IJCAI 2023
Machine learning › Trustworthy machine learning
robustness
0.512021
Understanding the Effect of Bias in Deep Anomaly Detection · IJCAI 2021
Data mining
anomaly detection
0.512021
Understanding the Effect of Bias in Deep Anomaly Detection · IJCAI 2021
Data mining › anomaly detection
deep anomaly detection
0.512021
Understanding the Effect of Bias in Deep Anomaly Detection · IJCAI 2021
Natural language and speech › Language models and text generation › LLM agents › tool use
tool-calling agents
0.312025
OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation · NeurIPS 2025
Computational finance and economics
behavioral economics
0.212024
Can Large Language Model Agents Simulate Human Trust Behavior? · NeurIPS 2024
Machine learning › Probabilistic and Bayesian machine learning › experimental design
active feature acquisition
0.212023
Efficient Online Decision Tree Learning with Active Feature Acquisition · IJCAI 2023

Methods — techniques the papers use, named apart from their topics

behavioral alignment analysis · 2.3trust game · 1.5active learning · 1.3self-play · 0.9reinforcement learning · 0.9modular architecture · 0.9hierarchical planning · 0.9direct preference optimization · 0.9RLOO · 0.9trust games · 0.8information-theoretic analysis · 0.8posterior sampling · 0.7adaptive submodularity · 0.7finite sample rate analysis · 0.5
YearPublicationVenuePosition
2025 Explain Implicit Semantic of Cyberbullying Text by Knowledge Enhancement
abstract
Cyberbullying detection remains a challenging due to the prevalence of implicit aggression and context-dependent language in online speech. Existing methods often rely on surfacelevel lexical meaning, which limits their ability to identify hostile intent conveyed through sarcasm, indirect references, or cultural nuance. This work proposes ISKE (Implicit Semantic Knowledge Enhancement), a novel framework that integrates external commonsense knowledge into pre-trained language models (PLMs) to improve understanding implicit meaning of cyberbullying content without context. ISKE firstly extracts sentiment-bearing entity pairs using aspect-based sentiment triplet extraction (ASTE), then retrievals a knowledge graph to collect semantically coherent relational paths centered on negative sentiment. These paths are transformed into structured triplets and injected into PLMs via adapter training on a relation prediction task. Experimental results on three benchmark datasets demonstrate that ISKE consistently improves$F_{1}$-scores over PLM baselines, with gains up to$\mathbf{1 1. 1 6 \%}$, especially in cases involving implicit meaning. This work highlights the value of structured knowledge in enhancing semantic reasoning for abusive language detection.
Ziyu Ye, Huakang Li
CW1
2025 Reward-Guided Prompt Evolving in Reinforcement Learning for LLMs
abstract
Existing reinforcement learning (RL) methods for large language models (LLMs) rely on static prompt sets, where prompts are curated a priori, and sampled in a fixed schedule for training, regardless of their usefulness to the RL process. We design eva, the first method that allows LLMs to prioritize and adaptively create useful prompts during RL training by reward signals. In principle, eva (Evolving via A symmetric Self-Play) casts language model training as a game between: (1) a creator, who samples and generates training prompts, and (2) a solver, who generates responses to the prompts. eva is simple, suits both offline and online RL for LLMs, and sets a new state-of-the-art on challenging benchmarks without extra human prompts: it improves gemma-2-9b-it’s win-rate on Arena-Hard from 51.6% to 60.1% by DPO and 52.6% to 62.4% by RLOO, surpassing claude-3-opus and nearing gemini-1.5-pro, both are orders of magnitude larger. Further ablation studies show eva can induce meaningful learning curriculum, and effectively scale RL for LLMs beyond static human prompts.
Ziyu Ye, Rishabh Agarwal, Tianqi Liu 0002, Rishabh Joshi, Sarmishta Velury, Quoc V. Le, Qijun Tan
ICML1
2025 OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation
abstract
Large Language Model (LLM)-based multi-agent systems show promise for automating real-world tasks but struggle to transfer across domains due to their domain-specific nature. Current approaches face two critical shortcomings: they require complete architectural redesign and full retraining of all components when applied to new domains. We introduce **Workforce**, a hierarchical multi-agent framework that decouples strategic planning from specialized execution through a modular architecture comprising: *(i)* a *domain-agnostic* **Planner** for task decomposition, *(ii)* a **Coordinator** for subtask management, and *(iii)* specialized **Workers** with *domain-specific* tool-calling capabilities. This decoupling enables cross-domain transferability during both inference and training phases: During inference, Workforce seamlessly adapts to new domains by adding or modifying worker agents; For training, we introduce **Optimized Workforce Learning (OWL)**, which improves generalization across domains by optimizing a domain-agnostic planner with reinforcement learning from real-world feedback. To validate our approach, we evaluate Workforce on the GAIA benchmark, covering various realistic, multi-domain agentic tasks. Experimental results demonstrate Workforce achieves open-source state-of-the-art performance (**69.70%**), outperforming commercial systems like OpenAI's Deep Research by **2.34%**. More notably, our OWL-trained 32B model achieves **52.73%** accuracy (**+16.37%**) and demonstrates performance comparable to GPT-4o on challenging tasks. To summarize, by enabling scalable generalization and modular domain transfer, our work establishes a foundation for the next generation of general-purpose AI assistants. *Our code is available at [Anonymous URL](https://anonymous.4open.science/r/annonymous-owl/), and our data is available at [Anonymous URL](https://huggingface.co/anonymous21016).*
Mengkang Hu, Wendong Fan, Yuzhou Nie, Ziyu Ye, Bowei Xia, Zhaoxuan Jin, Yingru Li, Qianshuo Ye, Bernard Ghanem, Ping Luo 0002, Guohao Li 0001
NeurIPS5
2025 A Data Augmentation Approach Using Sentiment Analysis and WordNet-Based Comments Transformation
Ziyu Ye, Zhuowen Zhang, Huakang Li
PRICAI (4)1
2024 Don't Be Pessimistic Too Early: Look K Steps Ahead!
Chaoqi Wang, Ziyu Ye, Kevin Murphy 0002, Yuxin Chen 0001
AISTATS2
2024 Understanding the Role of Equivariance in Self-supervised Learning
abstract
Contrastive learning has been a leading paradigm for self-supervised learning, but it is widely observed that it comes at the price of sacrificing useful features (\eg colors) by being invariant to data augmentations. Given this limitation, there has been a surge of interest in equivariant self-supervised learning (E-SSL) that learns features to be augmentation-aware. However, even for the simplest rotation prediction method, there is a lack of rigorous understanding of why, when, and how E-SSL learns useful features for downstream tasks. To bridge this gap between practice and theory, we establish an information-theoretic perspective to understand the generalization ability of E-SSL. In particular, we identify a critical explaining-away effect in E-SSL that creates a synergy between the equivariant and classification tasks. This synergy effect encourages models to extract class-relevant features to improve its equivariant prediction, which, in turn, benefits downstream tasks requiring semantic features. Based on this perspective, we theoretically analyze the influence of data transformations and reveal several principles for practical designs of E-SSL. Our theory not only aligns well with existing E-SSL methods but also sheds light on new directions by exploring the benefits of model equivariance. We believe that a theoretically grounded understanding on the role of equivariance would inspire more principled and advanced designs in this field. Code is available at https://github.com/kaotty/Understanding-ESSL.
Yifei Wang 0001, Kaiwen Hu, Sharut Gupta, Ziyu Ye, Yisen Wang 0001, Stefanie Jegelka
NeurIPS4
2024 Can Large Language Model Agents Simulate Human Trust Behavior?
abstract
Large Language Model (LLM) agents have been increasingly adopted as simulation tools to model humans in social science and role-playing applications. However, one fundamental question remains: can LLM agents really simulate human behavior? In this paper, we focus on one critical and elemental behavior in human interactions, trust, and investigate whether LLM agents can simulate human trust behavior. We first find that LLM agents generally exhibit trust behavior, referred to as agent trust, under the framework of Trust Games, which are widely recognized in behavioral economics. Then, we discover that GPT-4 agents manifest high behavioral alignment with humans in terms of trust behavior, indicating the feasibility of simulating human trust behavior with LLM agents. In addition, we probe the biases of agent trust and differences in agent trust towards other LLM agents and humans. We also explore the intrinsic properties of agent trust under conditions including external manipulations and advanced reasoning strategies. Our study provides new insights into the behaviors of LLM agents and the fundamental analogy between LLMs and humans beyond value alignment. We further illustrate broader implications of our discoveries for applications where trust is paramount.
Chengxing Xie, Canyu Chen, Feiran Jia, Ziyu Ye, Shiyang Lai, Kai Shu, Jindong Gu, Adel Bibi, Ziniu Hu, David Jurgens, Philip Torr 0001, Bernard Ghanem, Guohao Li 0001
NeurIPS4
2023 Efficient Online Decision Tree Learning with Active Feature Acquisition
abstract
Constructing decision trees online is a classical machine learning problem. Existing works often assume that features are readily available for each incoming data point. However, in many real world applications, both feature values and the labels are unknown a priori and can only be obtained at a cost. For example, in medical diagnosis, doctors have to choose which tests to perform (i.e., making costly feature queries) on a patient in order to make a diagnosis decision (i.e., predicting labels). We provide a fresh perspective to tackle this practical challenge. Our framework consists of an active planning oracle embedded in an online learning scheme for which we investigate several information acquisition functions. Specifically, we employ a surrogate information acquisition function based on adaptive submodularity to actively query feature values with a minimal cost, while using a posterior sampling scheme to maintain a low regret for online prediction. We demonstrate the efficiency and effectiveness of our framework via extensive experiments on various real-world datasets. Our framework also naturally adapts to the challenging setting of online learning with concept drift and is shown to be competitive with baseline models while being more flexible.
Arman Rahbar, Ziyu Ye, Yuxin Chen 0001, Morteza Haghir Chehreghani
IJCAI2
2023 Follow-ups Also Matter: Improving Contextual Bandits via Post-serving Contexts
abstract
Standard contextual bandit problem assumes that all the relevant contexts are observed before the algorithm chooses an arm. This modeling paradigm, while useful, often falls short when dealing with problems in which additional valuable contexts can be observed after arm selection. For example, content recommendation platforms like Youtube, Instagram, Tiktok receive much additional features about a user's reward after the user clicks a content (e.g., how long the user stayed, what is the user's watch speed, etc.). To improve online learning efficiency in these applications, we study a novel contextual bandit problem with post-serving contexts and design a new algorithm, poLinUCB, that achieves tight regret under standard assumptions. Core to our technical proof is a robustified and generalized version of the well-known Elliptical Potential Lemma (EPL), which can accommodate noise in data. Such robustification is necessary for tackling our problem, though we believe it could also be of general interest. Extensive empirical tests on both synthetic and real-world datasets demonstrate the significant benefit of utilitzing post-serving contexts as well as the superior performance of our algorithm over the state-of-the-art approaches.
Chaoqi Wang, Ziyu Ye, Zhe Feng 0004, Ashwinkumar Badanidiyuru
NeurIPS2
2021 Understanding the Effect of Bias in Deep Anomaly Detection
abstract
Anomaly detection presents a unique challenge in machine learning, due to the scarcity of labeled anomaly data. Recent work attempts to mitigate such problems by augmenting training of deep anomaly detection models with additional labeled anomaly samples. However, the labeled data often does not align with the target distribution and introduces harmful bias to the trained model. In this paper, we aim to understand the effect of a biased anomaly set on anomaly detection. Concretely, we view anomaly detection as a supervised learning task where the objective is to optimize the recall at a given false positive rate. We formally study the relative scoring bias of an anomaly detector, defined as the difference in performance with respect to a baseline anomaly detector. We establish the first finite sample rates for estimating the relative scoring bias for deep anomaly detection, and empirically validate our theoretical results on both synthetic and real-world datasets. We also provide an extensive empirical study on how a biased training anomaly set affects the anomaly score function and therefore the detection performance on different anomaly classes. Our study demonstrates scenarios in which the biased anomaly set can be useful or problematic, and provides a solid benchmark for future research.
Ziyu Ye, Yuxin Chen 0001, Haitao Zheng 0001
IJCAI1
2018 Joint Energy Optimization of Video Encoding and Transmission
abstract
Disposable wireless video sensors have many potential applications but are subject to stringent energy constraints. We studied the minimization of end-to-end distortion under an total energy constraint, by means of optimizing FEC code rate, number of source bits, and energy allocation between video encoding and wireless transmission. A two-step approach is employed. First, the FEC rate is optimized by exhaustive search. Then a binary-search-based algorithm is proposed to optimize the energy allocation and number of source bits. Experiments show that the algorithm achieves a PSNR gain up to 1dB over some reasonable baselines. A simpler suboptimal algorithm is also tested and exhibits similar performance.
Ziyu Ye, Rana D. Hegazy, Pamela C. Cosman, Laurence B. Milstein
PCS1