EDBT 2026 Demo / reviewers in the wild / expert
Yiqiao Jin
dblp:207/6631
· DBLP profile ↗
16ranked-venue papers
8as first author
16since 2021 · last 2026
0000-0002-6974-5970ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 6 first-author · 14 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reasoning Is Not All You Need: Examining LLMs for Multi-Turn Mental Health ConversationsabstractMohit Chandra, Siddharth Sriraman, Harneet Singh Khanuja, Yiqiao Jin, Munmun De Choudhury. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Mohit Chandra, Siddharth Sriraman, Harneet Singh Khanuja, Yiqiao Jin, Munmun De Choudhury |
ACL (1) | 4 |
| 2026 | SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document UnderstandingabstractMulti-page visual documents such as manuals, brochures, presentations, and posters convey key information through layout, colors, icons, and cross-slide references.While multimodal large language models (MLLMs) offer opportunities in document understanding, current systems struggle with complex, multi-page visual documents, particularly in fine-grained reasoning over elements and pages.We introduce SlideAgent, a versatile agentic framework for understanding multi-modal, multipage, and multi-layout documents, especially slide decks.SlideAgent employs specialized agents and decomposes reasoning into three specialized levels-global, page, and element-to construct a structured, query-agnostic representation that captures both overarching themes and detailed visual or textual cues.During inference, SlideAgent selectively activates specialized agents for multi-level reasoning and integrates their outputs into coherent, contextaware answers.Extensive experiments show that SlideAgent significantly improves accuracy over both proprietary (+7.9%) and opensource models (+9.8%). Yiqiao Jin, Rachneet Kaur, Sumitra Ganesh, Srijan Kumar |
ACL (1) | 1 |
| 2026 | SARA: Selective and Adaptive Retrieval-augmented Generation with Context CompressionabstractYiqiao Jin, Kartik Sharma, Vineeth Rakesh, Yingtong Dou, Menghai Pan, Mahashweta Das, Srijan Kumar. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yiqiao Jin, Kartik Sharma, Vineeth Rakesh, Yingtong Dou, Menghai Pan, Mahashweta Das, Srijan Kumar |
ACL (1) | 1 |
| 2025 | A Survey on Efficient Large Language Model Training: From Data-centric PerspectivesabstractJunyu Luo, Bohan Wu, Xiao Luo, Zhiping Xiao, Yiqiao Jin, Rong-Cheng Tu, Nan Yin, Yifan Wang, Jingyang Yuan, Wei Ju, Ming Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Junyu Luo 0002, Bohan Wu, Xiao Luo 0001, Zhiping Xiao 0001, Yiqiao Jin, Rongcheng Tu, Yifan Wang 0014, Jingyang Yuan, Wei Ju 0001, Ming Zhang 0004 |
ACL (1) | 5 |
| 2024 | Prototypical Reward Network for Data-Efficient RLHFabstractThe reward model for Reinforcement Learning from Human Feedback (RLHF) has proven effective in fine-tuning Large Language Models (LLMs).Notably, collecting human feedback for RLHF can be resource-intensive and lead to scalability issues for LLMs and complex tasks.Our proposed framework Proto-RM leverages prototypical networks to enhance reward models under limited human feedback.By enabling stable and reliable structural learning from fewer samples, Proto-RM significantly enhances LLMs' adaptability and accuracy in interpreting human preferences.Extensive experiments on various datasets demonstrate that Proto-RM significantly improves the performance of reward models and LLMs in human feedback tasks, achieving comparable and usually better results than traditional methods, while requiring significantly less data. in datalimited scenarios.This research offers a promising direction for enhancing the efficiency of reward models and optimizing the fine-tuning of language models under restricted feedback conditions. Jinghan Zhang 0002, Xiting Wang, Yiqiao Jin, Changyu Chen, Xinhao Zhang 0001, Kunpeng Liu 0001 |
ACL (1) | 3 |
| 2024 | Towards Fair Graph Anomaly Detection: Problem, Benchmark Datasets, and EvaluationabstractThe Fair Graph Anomaly Detection (FairGAD) problem aims to accurately detect anomalous nodes in an input graph while avoiding biased predictions against individuals from sensitive subgroups. However, the current literature does not comprehensively discuss this problem, nor does it provide realistic datasets that encompass actual graph structures, anomaly labels, and sensitive attributes. To bridge this gap, we introduce a formal definition of the FairGAD problem and present two novel datasets constructed from the social media platforms Reddit and Twitter. These datasets comprise 1.2 million and 400,000 edges associated with 9,000 and 47,000 nodes, respectively, and leverage political leanings as sensitive attributes and misinformation spreaders as anomaly labels. We demonstrate that our FairGAD datasets significantly differ from the synthetic datasets used by the research community. Using our datasets, we investigate the performance-fairness trade-off in nine existing GAD and non- graph AD methods on five state-of-the-art fairness methods. Code and datasets are available at https://github.com/nigelnnk/FairGAD. Neng Kai Nigel Neo, Yeon-Chang Lee, Yiqiao Jin, Sang-Wook Kim, Srijan Kumar |
CIKM | 3 |
| 2024 | AgentReview: Exploring Peer Review Dynamics with LLM AgentsabstractPeer review is fundamental to the integrity and advancement of scientific publication.Traditional methods of peer review analyses often rely on exploration and statistics of existing peer review data, which do not adequately address the multivariate nature of the process, account for the latent variables, and are further constrained by privacy concerns due to the sensitive nature of the data.We introduce AGENTREVIEW, the first large language model (LLM) based peer review simulation framework, which effectively disentangles the impacts of multiple latent factors and addresses the privacy issue.Our study reveals significant insights, including a notable 37.1% variation in paper decisions due to reviewers' biases, supported by sociological theories such as the social influence theory, altruism fatigue, and authority bias.We believe that this study could offer valuable insights to improve the design of peer review mechanisms.Our code is available at https://github.com/Ahren09/AgentReview. Yiqiao Jin, Qinlin Zhao, Hao Chen 0102, Kaijie Zhu, Yijia Xiao, Jindong Wang 0001 |
EMNLP | 1 |
| 2024 | Large Language Models Can Be Contextual Privacy Protection LearnersabstractYijia Xiao, Yiqiao Jin, Yushi Bai, Yue Wu, Xianjun Yang, Xiao Luo, Wenchao Yu, Xujiang Zhao, Yanchi Liu, Quanquan Gu, Haifeng Chen, Wei Wang, Wei Cheng. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Yijia Xiao, Yiqiao Jin, Yushi Bai, Xianjun Yang, Xiao Luo 0001, Wenchao Yu, Xujiang Zhao, Yanchi Liu, Quanquan Gu, Wei Wang 0010, Wei Cheng 0002 |
EMNLP | 2 |
| 2024 | CompeteAI: Understanding the Competition Dynamics of Large Language Model-based AgentsabstractLarge language models (LLMs) have been widely used as agents to complete different tasks, such as personal assistance or event planning. Although most of the work has focused on cooperation and collaboration between agents, little work explores competition, another important mechanism that promotes the development of society and economy. In this paper, we seek to examine the competition dynamics in LLM-based agents. We first propose a general framework for studying the competition between agents. Then, we implement a practical competitive environment using GPT-4 to simulate a virtual town with two types of agents, including restaurant agents and customer agents. Specifically, the restaurant agents compete with each other to attract more customers, where competition encourages them to transform, such as cultivating new operating strategies. Simulation experiments reveal several interesting findings at the micro and macro levels, which align well with existing market and sociological theories. We hope that the framework and environment can be a promising testbed to study the competition that fosters understanding of society. Code is available at: https://github.com/microsoft/competeai. Qinlin Zhao, Jindong Wang 0001, Yixuan Zhang 0001, Yiqiao Jin, Kaijie Zhu, Hao Chen 0102, Xing Xie 0001 |
ICML | 4 |
| 2024 | Better to Ask in English: Cross-Lingual Evaluation of Large Language Models for Healthcare Queries
Yiqiao Jin, Mohit Chandra, Gaurav Verma 0005, Yibo Hu 0002, Munmun De Choudhury, Srijan Kumar |
WWW | 1 |
| 2023 | Prototypical Fine-Tuning: Towards Robust Performance under Varying Data SizesabstractIn this paper, we move towards combining large parametric models with non-parametric prototypical networks. We propose prototypical fine-tuning, a novel prototypical framework for fine-tuning pretrained language models (LM), which automatically learns a bias to improve predictive performance for varying data sizes, especially low-resource settings. Our prototypical fine-tuning approach can automatically adjust the model capacity according to the number of data points and the model's inherent attributes. Moreover, we propose four principles for effective prototype fine-tuning towards the optimal solution. Experimental results across various datasets show that our work achieves significant performance improvements under various low-resource settings, as well as comparable and usually better performances in high-resource scenarios. Yiqiao Jin, Xiting Wang, Yaru Hao, Yizhou Sun, Xing Xie 0001 |
AAAI | 1 |
| 2023 | Semi-Offline Reinforcement Learning for Optimized Text GenerationabstractExisting reinforcement learning (RL) mainly utilize online or offline settings. The online methods explore the environment with expensive time cost, and the offline methods efficiently obtain reward signals by sacrificing the exploration capability. We propose semi-offline RL, a novel paradigm that can smoothly transit from the offline setting to the online setting, balances the exploration capability and training cost, and provides a theoretical foundation for comparing different RL settings. Based on the semi-offline MDP formulation, we present the RL setting that is optimal in terms of optimization cost, asymptotic error, and overfitting error bound. Extensive experiments show that our semi-offline RL approach is effective in various text generation tasks and datasets, and yields comparable or usually better performance compared with the state-of-the-art methods. Changyu Chen, Xiting Wang, Yiqiao Jin, Victor Ye Dong, Rui Yan 0001 |
ICML | 3 |
| 2023 | Predicting Information Pathways Across Online CommunitiesabstractThe problem of community-level information pathway prediction (CLIPP) aims at predicting the transmission trajectory of content across online communities. A successful solution to CLIPP holds significance as it facilitates the distribution of valuable information to a larger audience and prevents the proliferation of misinfor- mation. Notably, solving CLIPP is non-trivial as inter-community relationships and influence are unknown, information spread is multi-modal, and new content and new communities appear over time. In this work, we address CLIPP by collecting large-scale, multi-modal datasets to examine the diffusion of online YouTube videos on Reddit. We analyze these datasets to construct community influence graphs (CIGs) and develop a novel dynamic graph frame- work, INPAC (Information Pathway Across Online Communities), which incorporates CIGs to capture the temporal variability and multi-modal nature of video propagation across communities. Ex- perimental results in both warm-start and cold-start scenarios show that INPAC outperforms seven baselines in CLIPP. Our code and datasets are available at https://github.com/claws-lab/INPAC Yiqiao Jin, Yeon-Chang Lee, Kartik Sharma, Meng Ye 0002, Karan Sikka, Ajay Divakaran, Srijan Kumar |
KDD | 1 |
| 2023 | Code Recommendation for Open Source Software DevelopersabstractOpen Source Software (OSS) is forming the spines of technology infrastructures, attracting millions of talents to contribute. Notably, it is challenging and critical to consider both the developers’ interests and the semantic features of the project code to recommend appropriate development tasks to OSS developers. In this paper, we formulate the novel problem of code recommendation, whose purpose is to predict the future contribution behaviors of developers given their interaction history, the semantic features of source code, and the hierarchical file structures of projects. We introduce CODER, a novel graph-based CODE Recommendation framework for open source software developers, which accounts for the complex interactions among multiple parties within the system. CODER jointly models microscopic user-code interactions and macroscopic user-project interactions via a heterogeneous graph and further bridges the two levels of information through aggregation on file-structure graphs that reflect the project hierarchy. Moreover, to overcome the lack of reliable benchmarks, we construct three large-scale datasets to facilitate future research in this direction. Extensive experiments show that our CODER framework achieves superior performance under various experimental settings, including intra-project, cross-project, and cold-start recommendation. Yiqiao Jin, Yunsheng Bai, Yanqiao Zhu 0001, Yizhou Sun, Wei Wang 0010 |
WWW | 1 |
| 2022 | Towards Fine-Grained Reasoning for Fake News DetectionabstractThe detection of fake news often requires sophisticated reasoning skills, such as logically combining information by considering word-level subtle clues. In this paper, we move towards fine-grained reasoning for fake news detection by better reflecting the logical processes of human thinking and enabling the modeling of subtle clues. In particular, we propose a fine-grained reasoning framework by following the human’s information-processing model, introduce a mutual-reinforcement-based method for incorporating human knowledge about which evidence is more important, and design a prior-aware bi-channel kernel graph network to model subtle differences between pieces of evidence. Extensive experiments show that our model outperforms the state-of-the-art methods and demonstrate the explainability of our approach. Yiqiao Jin, Xiting Wang, Ruichao Yang, Yizhou Sun, Wei Wang 0010, Hao Liao, Xing Xie 0001 |
AAAI | 1 |
| 2022 | Reinforcement Subgraph Reasoning for Fake News DetectionabstractThe wide spread of fake news has caused serious societal issues. We propose a subgraph reasoning paradigm for fake news detection, which provides a crystal type of explainability by revealing which subgraphs of the news propagation network are the most important for news verification, and concurrently improves the generalization and discrimination power of graph-based detection models by removing task-irrelevant information. In particular, we propose a reinforced subgraph generation method, and perform fine-grained modeling on the generated subgraphs by developing a Hierarchical Path-aware Kernel Graph Attention Network. We also design a curriculum-based optimization method to ensure better convergence and train the two parts in an end-to-end manner. Ruichao Yang, Xiting Wang, Yiqiao Jin, Chaozhuo Li, Jianxun Lian, Xing Xie 0001 |
KDD | 3 |