VLDB 2026 Research / reviewers in the wild / expert
Yifan Zhong
dblp:227/0726
· DBLP profile ↗
18ranked-venue papers
5as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DexGraspVLA: A Vision-Language-Action Framework Towards General Dexterous GraspingabstractDexterous grasping remains a fundamental yet challenging problem in robotics. A general-purpose robot must be capable of grasping diverse objects in arbitrary scenarios. However, existing research typically relies on restrictive assumptions, such as single-object settings or limited environments, showing constrained generalization. We present DexGraspVLA, a hierarchical framework for robust generalization in language-guided general dexterous grasping and beyond. It utilizes a pre-trained Vision-Language model as the high-level planner and learns a diffusion-based low-level Action controller. The key insight to achieve generalization lies in iteratively transforming diverse language and visual inputs into domain-invariant representations via foundation models, where imitation learning can be effectively applied due to the alleviation of domain shift. Notably, our method achieves a 90+% dexterous grasping success rate under thousands of challenging unseen cluttered scenes. Empirical analysis confirms the consistency of internal model behavior across environmental variations, validating our design. DexGraspVLA also, for the first time, simultaneously demonstrates free-form long-horizon prompt execution, robustness to adversarial objects and human disturbance, and failure recovery. Extended application to nonprehensile grasping further proves its generality. Yifan Zhong, Xuchuan Huang, Ruochong Li, Ceyao Zhang, Tianrui Guan, Fanlian Zeng, Ka Nam Lui, Yuyao Ye, Yitao Liang, Yaodong Yang 0001, Yuanpei Chen |
AAAI | 1 |
| 2026 | DCNMST: A Deep Contrastive Network With Multiple Self-Supervised Tasks for Diabetic Retinopathy Grading Classification in Internet of Medical Things
Yanfei Sun, Xiangjun Han, Hongyuan Yu, Yifan Zhong, Dongyong Zhang |
IEEE Internet Things J. | 7 |
| 2025 | In-Context Editing: Learning Knowledge from Self-Induced DistributionsabstractIn scenarios where language models must incorporate new information efficiently without extensive retraining, traditional fine-tuning methods are prone to overfitting, degraded generalization, and unnatural language generation. To address these limitations, we introduce Consistent In-Context Editing (ICE), a novel approach leveraging the model's in-context learning capability to optimize towards a contextual distribution rather than a one-hot target. ICE introduces a simple yet effective optimization framework for the model to internalize new knowledge by aligning its output distributions with and without additional context. This method enhances the robustness and effectiveness of gradient-based tuning methods, preventing overfitting and preserving the model's integrity. We analyze ICE across four critical aspects of knowledge editing: accuracy, locality, generalization, and linguistic quality, demonstrating its advantages. Experimental results confirm the effectiveness of ICE and demonstrate its potential for continual editing, ensuring that the integrity of the model is preserved while updating information. Siyuan Qi, Bangcheng Yang, Kailin Jiang, Xiaobo Wang 0004, Jiaqi Li 0021, Yifan Zhong, Yaodong Yang 0001, Zilong Zheng |
ICLR | 6 |
| 2025 | Falcon: Fast Visuomotor Policies via Partial DenoisingabstractDiffusion policies are widely adopted in complex visuomotor tasks for their ability to capture multimodal action distributions. However, the multiple sampling steps required for action generation significantly harm real-time inference efficiency, which limits their applicability in real-time decision-making scenarios. Existing acceleration techniques either require retraining or degrade performance under low sampling steps. Here we propose Falcon, which mitigates this speed-performance trade-off and achieves further acceleration. The core insight is that visuomotor tasks exhibit sequential dependencies between actions. Falcon leverages this by reusing partially denoised actions from historical information rather than sampling from Gaussian noise at each step. By integrating current observations, Falcon reduces sampling steps while preserving performance. Importantly, Falcon is a training-free algorithm that can be applied as a plug-in to further improve decision efficiency on top of existing acceleration techniques. We validated Falcon in 48 simulated environments and 2 real-world robot experiments. demonstrating a 2-7x speedup with negligible performance degradation, offering a promising direction for efficient visuomotor policy design. Haojun Chen, Chengdong Ma, Xiaojian Ma 0001, Zailin Ma, Huimin Wu 0001, Yuanpei Chen, Yifan Zhong, Qing Li 0003, Yaodong Yang 0001 |
ICML | 8 |
| 2024 | Maximum Entropy Heterogeneous-Agent Reinforcement Learningabstract*Multi-agent reinforcement learning* (MARL) has been shown effective for cooperative games in recent years. However, existing state-of-the-art methods face challenges related to sample complexity, training instability, and the risk of converging to a suboptimal Nash Equilibrium. In this paper, we propose a unified framework for learning \emph{stochastic} policies to resolve these issues. We embed cooperative MARL problems into probabilistic graphical models, from which we derive the maximum entropy (MaxEnt) objective for MARL. Based on the MaxEnt framework, we propose *Heterogeneous-Agent Soft Actor-Critic* (HASAC) algorithm. Theoretically, we prove the monotonic improvement and convergence to *quantal response equilibrium* (QRE) properties of HASAC. Furthermore, we generalize a unified template for MaxEnt algorithmic design named *Maximum Entropy Heterogeneous-Agent Mirror Learning* (MEHAML), which provides any induced method with the same guarantees as HASAC. We evaluate HASAC on six benchmarks: Bi-DexHands, Multi-Agent MuJoCo, StarCraft Multi-Agent Challenge, Google Research Football, Multi-Agent Particle Environment, and Light Aircraft Game. Results show that HASAC consistently outperforms strong baselines, exhibiting better sample efficiency, robustness, and sufficient exploration. Jiarong Liu, Yifan Zhong, Siyi Hu 0001, Haobo Fu, Qiang Fu 0016, Xiaojun Chang, Yaodong Yang 0001 |
ICLR | 2 |
| 2024 | CivRealm: A Learning and Reasoning Odyssey in Civilization for Decision-Making AgentsabstractThe generalization of decision-making agents encompasses two fundamental elements: learning from past experiences and reasoning in novel contexts. However, the predominant emphasis in most interactive environments is on learning, often at the expense of complexity in reasoning. In this paper, we introduce CivRealm, an environment inspired by the Civilization game. Civilization’s profound alignment with human society requires sophisticated learning and prior knowledge, while its ever-changing space and action space demand robust reasoning for generalization. Particularly, CivRealm sets up an imperfect-information general-sum game with a changing number of players; it presents a plethora of complex features, challenging the agent to deal with open-ended stochastic environments that require diplomacy and negotiation skills. Within CivRealm, we provide interfaces for two typical agent types: tensor-based agents that focus on learning, and language-based agents that emphasize reasoning. To catalyze further research, we present initial results for both paradigms. The canonical RL-based agents exhibit reasonable performance in mini-games, whereas both RL- and LLM-based agents struggle to make substantial progress in the full game. Overall, CivRealm stands as a unique learning and reasoning challenge for decision-making agents. The code is available at https://github.com/bigai-ai/civrealm. Siyuan Qi, Shuo Chen 0006, Yexin Li, Bangcheng Yang, Pring Wong, Yifan Zhong, Zhaowei Zhang 0001, Nian Liu 0003, Yaodong Yang 0001, Song-Chun Zhu |
ICLR | 8 |
| 2024 | Off-Agent Trust Region Policy Optimization
Ruiqing Chen, Yali Du 0001, Yifan Zhong, Zheng Tian 0002, Fanglei Sun, Yaodong Yang 0001 |
IJCAI | 4 |
| 2024 | Panacea: Pareto Alignment via Preference Adaptation for LLMsabstractCurrent methods for large language model alignment typically use scalar human preference labels. However, this convention tends to oversimplify the multi-dimensional and heterogeneous nature of human preferences, leading to reduced expressivity and even misalignment. This paper presents Panacea, an innovative approach that reframes alignment as a multi-dimensional preference optimization problem. Panacea trains a single model capable of adapting online and Pareto-optimally to diverse sets of preferences without the need for further tuning. A major challenge here is using a low-dimensional preference vector to guide the model's behavior, despite it being governed by an overwhelmingly large number of parameters. To address this, Panacea is designed to use singular value decomposition (SVD)-based low-rank adaptation, which allows the preference vector to be simply injected online as singular values. Theoretically, we prove that Panacea recovers the entire Pareto front with common loss aggregation methods under mild conditions. Moreover, our experiments demonstrate, for the first time, the feasibility of aligning a single LLM to represent an exponentially vast spectrum of human preferences through various optimization methods. Our work marks a step forward in effectively and efficiently aligning models to diverse and intricate human preferences in a controllable and Pareto-optimal manner. Yifan Zhong, Chengdong Ma, Ziran Yang, Haojun Chen, Qingfu Zhang 0001, Siyuan Qi, Yaodong Yang 0001 |
NeurIPS | 1 |
| 2024 | Heterogeneous-Agent Reinforcement LearningabstractThe necessity for cooperation among intelligent machines has popularised cooperative multi-agent reinforcement learning (MARL) in AI research. However, many research endeavours heavily rely on parameter sharing among agents, which confines them to only homogeneous-agent setting and leads to training instability and lack of convergence guarantees. To achieve effective cooperation in the general heterogeneous-agent setting, we propose Heterogeneous-Agent Reinforcement Learning (HARL) algorithms that resolve the aforementioned issues. Central to our findings are the multi-agent advantage decomposition lemma and the sequential update scheme. Based on these, we develop the provably correct Heterogeneous-Agent Trust Region Learning (HATRL), and derive HATRPO and HAPPO by tractable approximations. Furthermore, we discover a novel framework named Heterogeneous-Agent Mirror Learning (HAML), which strengthens theoretical guarantees for HATRPO and HAPPO and provides a general template for cooperative MARL algorithmic designs. We prove that all algorithms derived from HAML inherently enjoy monotonic improvement of joint return and convergence to Nash Equilibrium. As its natural outcome, HAML validates more novel algorithms in addition to HATRPO and HAPPO, including HAA2C, HADDPG, and HATD3, which generally outperform their existing MA-counterparts. We comprehensively test HARL algorithms on six challenging benchmarks and demonstrate their superior effectiveness and stability for coordinating heterogeneous agents compared to strong baselines such as MAPPO and QMIX. Yifan Zhong, Jakub Grudzien Kuba, Xidong Feng, Siyi Hu 0001, Jiaming Ji, Yaodong Yang 0001 |
J. Mach. Learn. Res. | 1 |
| 2024 | Nash Equilibrium Seeking for Multi-Agent Systems Under DoS Attacks and DisturbancesabstractIn this article, the primary focus is on studying a Nash equilibrium (NE) seeking algorithm to maintain system resilience in multi-agent systems (MASs), which are subject to denial-of-service (DoS) attacks and disturbance. Furthermore, we demonstrate the performance of the algorithm in tolerating such attacks. DoS attacks are modeled using Markov processes, and their impact on interagent communication is investigated. The considered n-order MAS is unable to maintain normal communication links when subjected to DoS attacks. To address this issue, stability analysis is conducted to demonstrate the effectiveness of the proposed NE seeking algorithm in achieving secure control of MAS. Conditions for maintaining resilience under attacks are also provided. Finally, numerical simulations are performed on a satellite cluster system, and physical experiments are conducted using a wheeled robot ground platform to validate the effectiveness of the algorithm. Yifan Zhong, Yuan Yuan 0006, Huanhuan Yuan |
IEEE Trans. Ind. Informatics | 1 |
| 2023 | Safety Gymnasium: A Unified Safe Reinforcement Learning BenchmarkabstractArtificial intelligence (AI) systems possess significant potential to drive societal progress. However, their deployment often faces obstacles due to substantial safety concerns. Safe reinforcement learning (SafeRL) emerges as a solution to optimize policies while simultaneously adhering to multiple constraints, thereby addressing the challenge of integrating reinforcement learning in safety-critical scenarios. In this paper, we present an environment suite called Safety-Gymnasium, which encompasses safety-critical tasks in both single and multi-agent scenarios, accepting vector and vision-only input. Additionally, we offer a library of algorithms named Safe Policy Optimization (SafePO), comprising 16 state-of-the-art SafeRL algorithms. This comprehensive library can serve as a validation tool for the research community. By introducing this benchmark, we aim to facilitate the evaluation and comparison of safety performance, thus fostering the development of reinforcement learning for safer, more reliable, and responsible real-world applications. The website of this project can be accessed at https://sites.google.com/view/safety-gymnasium. Jiaming Ji, Borong Zhang, Xuehai Pan, Weidong Huang 0008, Ruiyang Sun, Yiran Geng, Yifan Zhong, Josef Dai, Yaodong Yang 0001 |
NeurIPS | 8 |
| 2023 | MARLlib: A Scalable and Efficient Multi-agent Reinforcement Learning LibraryabstractA significant challenge facing researchers in the area of multi-agent reinforcement learning (MARL) pertains to the identification of a library that can offer fast and compatible development for multi-agent tasks and algorithm combinations, while obviating the need to consider compatibility issues. In this paper, we present MARLlib, a library designed to address the aforementioned challenge by leveraging three key mechanisms: 1) a standardized multi-agent environment wrapper, 2) an agent-level algorithm implementation, and 3) a flexible policy mapping strategy. By utilizing these mechanisms, MARLlib can effectively disentangle the intertwined nature of the multi-agent task and the learning process of the algorithm, with the ability to automatically alter the training strategy based on the current task's attributes. The MARLlib library's source code is publicly accessible on GitHub: https://github.com/Replicable-MARL/MARLlib. Siyi Hu 0001, Yifan Zhong, Minquan Gao, Weixun Wang, Hao Dong 0003, Xiaodan Liang, Zhihui Li 0001, Xiaojun Chang, Yaodong Yang 0001 |
J. Mach. Learn. Res. | 2 |
| 2023 | Adoption of AI in response to COVID-19 - a configurational perspective
Lili Mi, Wei Liu 0135, Yu-Hsi Yuan, Xuefeng Shao, Yifan Zhong |
Pers. Ubiquitous Comput. | 5 |
| 2023 | A design concept of big data analytics model for managers in hospitality industries
Seyedmohammad Mousavian, Shah Jahan Miah, Yifan Zhong |
Pers. Ubiquitous Comput. | 3 |
| 2022 | FormLM: Recommending Creation Ideas for Online Forms by Modelling Semantic and Structural InformationabstractOnline forms are widely used to collect data from human and have a multi-billion market.Many software products provide online services for creating semi-structured forms where questions and descriptions are organized by predefined structures.However, the design and creation process of forms is still tedious and requires expert knowledge.To assist form designers, in this work we present FormLM to model online forms (by enhancing pre-trained language model with form structural information) and recommend form creation ideas (including question / options recommendations and block type suggestion).For model training and evaluation, we collect the first public online form dataset with 62K online forms.Experiment results show that FormLM significantly outperforms general-purpose language models on all tasks, with an improvement by 4.71 on Question Recommendation and 10.6 on Block Type Suggestion in terms of ROUGE-1 and Macro-F1, respectively. Yijia Shao, Mengyu Zhou, Yifan Zhong, Shi Han, Gideon Huang, Dongmei Zhang 0001 |
EMNLP | 3 |
| 2022 | Development and validation of a deep learning model to predict the survival of patients in ICUabstractBACKGROUND: Patients in the intensive care unit (ICU) are often in critical condition and have a high mortality rate. Accurately predicting the survival probability of ICU patients is beneficial to timely care and prioritizing medical resources to improve the overall patient population survival. Models developed by deep learning (DL) algorithms show good performance on many models. However, few DL algorithms have been validated in the dimension of survival time or compared with traditional algorithms. METHODS: Variables from the Early Warning Score, Sequential Organ Failure Assessment Score, Simplified Acute Physiology Score II, Acute Physiology and Chronic Health Evaluation (APACHE) II, and APACHE IV models were selected for model development. The Cox regression, random survival forest (RSF), and DL methods were used to develop prediction models for the survival probability of ICU patients. The prediction performance was independently evaluated in the MIMIC-III Clinical Database (MIMIC-III), the eICU Collaborative Research Database (eICU), and Shanghai Pulmonary Hospital Database (SPH). RESULTS: Forty variables were collected in total for model development. 83 943 participants from 3 databases were included in the study. The New-DL model accurately stratified patients into different survival probability groups with a C-index of >0.7 in the MIMIC-III, eICU, and SPH, performing better than the other models. The calibration curves of the models at 3 and 10 days indicated that the prediction performance was good. A user-friendly interface was developed to enable the model's convenience. CONCLUSIONS: Compared with traditional algorithms, DL algorithms are more accurate in predicting the survival probability during ICU hospitalization. This novel model can provide reliable, individualized survival probability prediction. Hai Tang, Zhuochen Jin, Jiajun Deng, Yunlang She, Yifan Zhong, Weiyan Sun, Yijiu Ren, Nan Cao 0001, Chang Chen 0006 |
J. Am. Medical Informatics Assoc. | 5 |
| 2021 | Cross-Lingual Transfer for Speech Processing Using Acoustic Language SimilarityabstractSpeech processing systems currently do not support the vast majority of languages, in part due to the lack of data in low-resource languages. Cross-lingual transfer offers a compelling way to help bridge this digital divide by incorporating high-resource data into low-resource systems. Current cross-lingual algorithms have shown success in text-based tasks and speech-related tasks over some low-resource languages. However, scaling up speech systems to support hundreds of low-resource languages remains unsolved. To help bridge this gap, we propose a language similarity approach that can efficiently identify acoustic cross-lingual transfer pairs across hundreds of languages. We demonstrate the effectiveness of our approach in language family classification, speech recognition, and speech synthesis tasks. Peter Wu, Jiatong Shi, Yifan Zhong, Shinji Watanabe 0001, Alan W. Black |
ASRU | 3 |
| 2018 | A field study of related video recommendations: newest, most similar, or most relevant?abstractMany video sites recommend videos related to the one a user is watching. These recommendations have been shown to influence what users end up exploring and are an important part of a recommender system. Plenty of methods have been proposed to recommend related videos, but there has been relatively little work that compares competing strategies. We describe a field study of related video recommendations, where we deploy algorithms to recommend related movie trailers. Our results show that recency- and similarity-based algorithms yield the highest click-through rates, and that the recency-based algorithm leads to the most trailer-level engagement. Our findings suggest the potential to design non-personalized yet effective related item recommendation strategies. Yifan Zhong, Tahir Lazaro Sousa Menezes, F. Maxwell Harper |
RecSys | 1 |