VLDB 2026 Research / reviewers in the wild / expert
Ye Liu 0006
dblp:96/2615-6
· DBLP profile ↗
17ranked-venue papers
8as first author
13since 2021 · last 2025
0000-0001-7237-7382ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 8 first-author · 11 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Breaking the Batch Barrier (B3) of Contrastive Learning via Smart Batch MiningabstractContrastive learning (CL) is a prevalent technique for training embedding models, which pulls semantically similar examples (positives) closer in the representation space while pushing dissimilar ones (negatives) further apart. A key source of negatives are "in-batch" examples, i.e., positives from other examples in the batch. Effectiveness of such models is hence strongly influenced by the size and quality of training batches. In this work, we propose *Breaking the Batch Barrier* (B3), a novel batch construction strategy designed to curate high-quality batches for CL. Our approach begins by using a pretrained teacher embedding model to rank all examples in the dataset, from which a sparse similarity graph is constructed. A community detection algorithm is then applied to this graph to identify clusters of examples that serve as strong negatives for one another. The clusters are then used to construct batches that are rich in in-batch negatives. Empirical results on the MMEB multimodal embedding benchmark (36 tasks) demonstrate that our method sets a new state of the art, outperforming previous best methods by +1.3 and +2.9 points at the 7B and 2B model scales, respectively. Notably, models trained with B3 surpass existing state-of-the-art results even with a batch size as small as 64, which is 4–16× smaller than that required by other methods. Moreover, experiments show that B3 generalizes well across domains and tasks, maintaining strong performance even when trained with considerably weaker teachers. Raghuveer Thirukovalluru, Ye Liu 0006, Karthikeyan K, Mingyi Su, Ping Nie, Semih Yavuz, Yingbo Zhou 0002, Wenhu Chen, Bhuwan Dhingra |
NeurIPS | 3 |
| 2025 | Can Large Language Models Serve as Evaluators for Code Summarization?abstractCode summarization facilitates program comprehension and software maintenance by converting code snippets into natural-language descriptions. Over the years, numerous methods have been developed for this task, but a key challenge remains: effectively evaluating the quality of generated summaries. While human evaluation is effective for assessing code summary quality, it is labor-intensive and difficult to scale. Commonly used automatic metrics, such as BLEU, ROUGE-L, METEOR, and BERTScore, often fail to align closely with human judgments. In this paper, we explore the potential ofLarge Language Models (LLMs)for evaluating code summarization. We propose CODERPE (Role-Player for Code Summarization Evaluation), a novel method that leverages role-player prompting to assess the quality of generated summaries. Specifically, we prompt LLM-based evaluators to take on diverse roles, such as code reviewer, code author, code editor, and system analyst. Each role evaluates the quality of code summaries across key dimensions, including coherence, consistency, fluency, and relevance. We further explore the robustness of LLMs as evaluators by employing various prompting strategies, including chain-of-thought reasoning, incontext learning, and tailored rating form designs. The results demonstrate that LLMs serve as effective evaluators for code summarization. Notably, our LLM-based evaluator, CODERPE , achieves an 80.18% Spearman correlation with human evaluations, outperforming the existing BERTScore metric by 10.39%. Yang Wu 0010, Yao Wan 0001, Zhaoyang Chu, Wenting Zhao 0006, Ye Liu 0006, Hongyu Zhang 0002, Xuanhua Shi, Hai Jin 0001, Philip S. Yu |
IEEE Trans. Software Eng. | 5 |
| 2024 | CORI: CJKV Benchmark with Romanization Integration - a Step towards Cross-lingual Transfer beyond Textual ScriptsabstractNaively assuming English as a source language may hinder cross-lingual transfer for many languages by failing to consider the importance of language contact. Some languages are more well-connected than others, and target languages can benefit from transferring from closely related languages; for many languages, the set of closely related languages does not include English. In this work, we study the impact of source language for cross-lingual transfer, demonstrating the importance of selecting source languages that have high contact with the target language. We also construct a novel benchmark dataset for close contact Chinese-Japanese-Korean-Vietnamese (CJKV) languages to further encourage in-depth studies of language contact. To comprehensively capture contact between these languages, we propose to integrate Romanized transcription beyond textual scripts via Contrastive Learning objectives, leading to enhanced cross-lingual representations and effective zero-shot cross-lingual transfer. Hoang Nguyen 0006, Ye Liu 0006, Natalie Parde, Eugene Rohrbaugh, Philip S. Yu |
LREC/COLING | 3 |
| 2024 | FOLIO: Natural Language Reasoning with First-Order LogicabstractSimeng Han, Hailey Schoelkopf, Yilun Zhao, Zhenting Qi, Martin Riddell, Wenfei Zhou, James Coady, David Peng, Yujie Qiao, Luke Benson, Lucy Sun, Alexander Wardle-Solano, Hannah Szabó, Ekaterina Zubova, Matthew Burtell, Jonathan Fan, Yixin Liu, Brian Wong, Malcolm Sailor, Ansong Ni, Linyong Nan, Jungo Kasai, Tao Yu, Rui Zhang, Alexander Fabbri, Wojciech Maciej Kryscinski, Semih Yavuz, Ye Liu, Xi Victoria Lin, Shafiq Joty, Yingbo Zhou, Caiming Xiong, Rex Ying, Arman Cohan, Dragomir Radev. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Simeng Han, Hailey Schoelkopf, Yilun Zhao 0001, Zhenting Qi, Martin Riddell, Wenfei Zhou, James Coady, David Peng, Yujie Qiao, Luke Benson, Lucy Sun, Alexander Wardle-Solano, Hannah Szabó, Ekaterina Zubova, Matthew Burtell, Jonathan Fan 0001, Yixin Liu 0003, Malcolm Sailor, Ansong Ni, Linyong Nan, Jungo Kasai, Tao Yu 0009, Rui Zhang 0037, Alexander R. Fabbri, Wojciech Kryscinski, Semih Yavuz, Ye Liu 0006, Xi Victoria Lin, Shafiq R. Joty, Yingbo Zhou 0002, Caiming Xiong, Rex Ying, Arman Cohan, Dragomir R. Radev |
EMNLP | 28 |
| 2024 | kNN-ICL: Compositional Task-Oriented Parsing Generalization with Nearest Neighbor In-Context LearningabstractWenting Zhao, Ye Liu, Yao Wan, Yibo Wang, Qingyang Wu, Zhongfen Deng, Jiangshu Du, Shuaiqi Liu, Yunlong Xu, Philip Yu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Wenting Zhao 0006, Ye Liu 0006, Yao Wan 0001, Yibo Wang 0001, Qingyang Wu, Zhongfen Deng, Jiangshu Du, Shuaiqi Liu 0002, Philip S. Yu |
NAACL-HLT | 2 |
| 2024 | L2CEval: Evaluating Language-to-Code Generation Capabilities of Large Language ModelsabstractAbstract Recently, large language models (LLMs), especially those that are pretrained on code, have demonstrated strong capabilities in generating programs from natural language inputs. Despite promising results, there is a notable lack of a comprehensive evaluation of these models’ language-to-code generation capabilities. Existing studies often focus on specific tasks, model architectures, or learning paradigms, leading to a fragmented understanding of the overall landscape. In this work, we present L2CEval, a systematic evaluation of the language-to-code generation capabilities of LLMs on 7 tasks across the domain spectrum of semantic parsing, math reasoning, and Python programming, analyzing the factors that potentially affect their performance, such as model size, pretraining data, instruction tuning, and different prompting methods. In addition, we assess confidence calibration, and conduct human evaluations to identify typical failures across different tasks and models. L2CEval offers a comprehensive understanding of the capabilities and limitations of LLMs in language-to-code generation. We release the evaluation framework1 and all model outputs, hoping to lay the groundwork for further future research. All future evaluations (e.g., LLaMA-3, StarCoder2, etc) will be updated on the project website: https://l2c-eval.github.io/. Ansong Ni, Yilun Zhao 0001, Martin Riddell, Troy Feng, Stephen Yin, Ye Liu 0006, Semih Yavuz, Caiming Xiong, Shafiq R. Joty, Yingbo Zhou 0002, Dragomir R. Radev, Arman Cohan |
Trans. Assoc. Comput. Linguistics | 8 |
| 2023 | CoF-CoT: Enhancing Large Language Models with Coarse-to-Fine Chain-of-Thought Prompting for Multi-domain NLU TasksabstractWhile Chain-of-Thought prompting is popular in reasoning tasks, its application to Large Language Models (LLMs) in Natural Language Understanding (NLU) is under-explored.Motivated by multi-step reasoning of LLMs, we propose Coarse-to-Fine Chain-of-Thought (CoF-CoT) approach that breaks down NLU tasks into multiple reasoning steps where LLMs can learn to acquire and leverage essential concepts to solve tasks from different granularities.Moreover, we propose leveraging semanticbased Abstract Meaning Representation (AMR) structured knowledge as an intermediate step to capture the nuances and diverse structures of utterances, and to understand connections between their varying levels of granularity.Our proposed approach is demonstrated effective in assisting the LLMs adapt to the multi-grained NLU tasks under both zero-shot and few-shot multi-domain settings 1 . Hoang Nguyen 0006, Ye Liu 0006, Tao Zhang 0055, Philip S. Yu |
EMNLP | 2 |
| 2023 | Slot Induction via Pre-trained Language Model Probing and Multi-level Contrastive LearningabstractRecent advanced methods in Natural Language Understanding for Task-oriented Dialogue (TOD) Systems (e.g., intent detection and slot filling) require a large amount of annotated data to achieve competitive performance.In reality, token-level annotations (slot labels) are time-consuming and difficult to acquire.In this work, we study the Slot Induction (SI) task whose objective is to induce slot boundaries without explicit knowledge of token-level slot annotations.We propose leveraging Unsupervised Pre-trained Language Model (PLM) Probing and Contrastive Learning mechanism to exploit (1) unsupervised semantic knowledge extracted from PLM, and (2) additional sentencelevel intent label signals available from TOD.Our approach is shown to be effective in SI task and capable of bridging the gaps with tokenlevel supervised models on two NLU benchmark datasets.When generalized to emerging intents, our SI objectives also provide enhanced slot label representations, leading to improved performance on the Slot Filling tasks. 1 Hoang Nguyen 0006, Ye Liu 0006, Philip S. Yu |
SIGDIAL | 3 |
| 2023 | Reinforced MOOCs Concept Recommendation in Heterogeneous Information NetworksabstractMassive open online courses (MOOCs), which offer open access and widespread interactive participation through the internet, are quickly becoming the preferred method for online and remote learning. Several MOOC platforms offer the service of course recommendation to users, to improve the learning experience of users. Despite the usefulness of this service, we consider that recommending courses to users directly may neglect their varying degrees of expertise. To mitigate this gap, we examine an interesting problem of concept recommendation in this paper, which can be viewed as recommending knowledge to users in a fine-grained way. We put forward a novel approach, termedHinCRec-RL, forConceptRecommendation in MOOCs, which is based onHeterogeneousInformationNetworks andReinforcementLearning. In particular, we propose to shape the problem of concept recommendation within a reinforcement learning framework to characterize the dynamic interaction between users and knowledge concepts in MOOCs. Furthermore, we propose to form the interactions among users, courses, videos, and concepts into aheterogeneous information network (HIN)to learn the semantic user representations better. We then employ an attentional graph neural network to represent the users in the HIN, based on meta-paths. Extensive experiments are conducted on a real-world dataset collected from a Chinese MOOC platform,XuetangX, to validate the efficacy of our proposed HinCRec-RL. Experimental results and analysis demonstrate that our proposed HinCRec-RL performs well when compared with several state-of-the-art models. Jibing Gong, Yao Wan 0001, Ye Liu 0006, Xuewen Li 0005, Yi Zhao 0029, Cheng Wang 0052, Xiaohan Fang, Wenzheng Feng, Jie Tang 0001 |
ACM Trans. Web | 3 |
| 2022 | Uni-Parser: Unified Semantic Parser for Question Answering on Knowledge Base and DatabaseabstractParsing natural language questions into executable logical forms is a useful and interpretable way to perform question answering on structured data such as knowledge bases (KB) or databases (DB).However, existing approaches on semantic parsing cannot adapt to both modalities, as they suffer from the exponential growth of the logical form candidates and can hardly generalize to unseen data.In this work, we propose Uni-Parser, a unified semantic parser for question answering (QA) on both KB and DB.We introduce the primitive (relation and entity in KB, and table name, column name and cell value in DB) as an essential element in our framework.The number of primitives grows linearly with the number of retrieved relations in KB and DB, preventing us from dealing with exponential logic form candidates.We leverage the generator to predict final logical forms by altering and composing topranked primitives with different operations (e.g.select, where, count).With sufficiently pruned search space by a contrastive primitive ranker, the generator is empowered to capture the composition of primitives enhancing its generalization ability.We achieve competitive results on multiple KB and DB QA benchmarks more efficiently, especially in the compositional and zero-shot settings. Ye Liu 0006, Semih Yavuz, Dragomir R. Radev, Caiming Xiong, Yingbo Zhou 0002 |
EMNLP | 1 |
| 2021 | KG-BART: Knowledge Graph-Augmented BART for Generative Commonsense ReasoningabstractGenerative commonsense reasoning which aims to empower machines to generate sentences with the capacity of reasoning over a set of concepts is a critical bottleneck for text generation. Even the state-of-the-art pre-trained language generation models struggle at this task and often produce implausible and anomalous sentences. One reason is that they rarely consider incorporating the knowledge graph which can provide rich relational information among the commonsense concepts. To promote the ability of commonsense reasoning for text generation, we propose a novel knowledge graph augmented pre-trained language generation model KG-BART, which encompasses the complex relations of concepts through the knowledge graph and produces more logical and natural sentences as output. Moreover, KG-BART can leverage the graph attention to aggregate the rich concept semantics that enhances the model generalization on unseen concept sets. Experiments on benchmark CommonGen dataset verify the effectiveness of our proposed approach by comparing with several strong pre-trained language generation models, particularly KG-BART outperforms BART by 5.80, 4.60, in terms of BLEU-3, 4. Moreover, we also show that the generated context by our model can work as background scenarios to benefit downstream commonsense QA tasks. Ye Liu 0006, Yao Wan 0001, Lifang He 0001, Hao Peng 0001, Philip S. Yu |
AAAI | 1 |
| 2021 | Enriching Non-Autoregressive Transformer with Syntactic and Semantic Structures for Neural Machine TranslationabstractThe non-autoregressive models have boosted the efficiency of neural machine translation through parallelized decoding at the cost of effectiveness, when comparing with the autoregressive counterparts.In this paper, we claim that the syntactic and semantic structures among natural language are critical for non-autoregressive machine translation and can further improve the performance.However, these structures are rarely considered in existing non-autoregressive models.Inspired by this intuition, we propose to incorporate the explicit syntactic and semantic structures of languages into a non-autoregressive Transformer, for the task of neural machine translation.Moreover, we also consider the intermediate latent alignment within target sentences to better learn the long-term token dependencies.Experimental results on two real-world datasets (i.e., WMT14 En-De and WMT16 En-Ro) show that our model achieves a significantly faster speed, as well as keeps the translation quality when compared with several stateof-the-art non-autoregressive models. Ye Liu 0006, Yao Wan 0001, Jianguo Zhang 0005, Wenting Zhao 0006, Philip S. Yu |
EACL | 1 |
| 2021 | HETFORMER: Heterogeneous Transformer with Sparse Attention for Long-Text Extractive SummarizationabstractTo capture the semantic graph structure from raw text, most existing summarization approaches are built on GNNs with a pre-trained model.However, these methods suffer from cumbersome procedures and inefficient computations for long-text documents.To mitigate these issues, this paper proposes HET-FORMER, a Transformer-based pre-trained model with multi-granularity sparse attentions for long-text extractive summarization.Specifically, we model different types of semantic nodes in raw text as a potential heterogeneous graph and directly learn heterogeneous relationships (edges) among nodes by Transformer.Extensive experiments on both single-and multi-document summarization tasks show that HETFORMER achieves stateof-the-art performance in Rouge F1 while using less memory and fewer parameters. Ye Liu 0006, Jianguo Zhang 0005, Yao Wan 0001, Congying Xia, Lifang He 0001, Philip S. Yu |
EMNLP (1) | 1 |
| 2020 | Commonsense Evidence Generation and Injection in Reading ComprehensionabstractHuman tackle reading comprehension not only based on the given context itself but often rely on the commonsense beyond.To empower the machine with commonsense reasoning, in this paper, we propose a Commonsense Evidence Generation and Injection framework in reading comprehension, named CEGI.The framework injects two kinds of auxiliary commonsense evidence into comprehensive reading to equip the machine with the ability of rational thinking.Specifically, we build two evidence generators: one aims to generate textual evidence via a language model; the other aims to extract factual evidence (automatically aligned text-triples) from a commonsense knowledge graph after graph completion.Those evidences incorporate contextual commonsense and serve as the additional inputs to the reasoning model.Thereafter, we propose a deep contextual encoder to extract semantic relationships among the paragraph, question, option, and evidence.Finally, we employ a capsule network to extract different linguistic units (word and phrase) from the relations, and dynamically predict the optimal option based on the extracted units.Experiments on the Cos-mosQA dataset demonstrate that the proposed CEGI model outperforms the current state-ofthe-art approaches and achieves the highest accuracy (83.6%) on the leaderboard. Ye Liu 0006, Tao Yang 0012, Zeyu You, Wei Fan 0001, Philip S. Yu |
SIGdial | 1 |
| 2019 | Generative Question Refinement with Deep Reinforcement Learning in Retrieval-based QA SystemabstractIn real-world question-answering (QA) systems, ill-formed questions, such as wrong words, ill word order and noisy expressions, are common and may prevent the QA systems from understanding and answering them accurately. In order to eliminate the effect of ill-formed questions, we approach the question refinement task and propose a unified model, QREFINE, to refine the ill-formed questions to well-formed question. The basic idea is to learn a Seq2Seq model to generate a new question from the original one. To improve the quality and retrieval performance of the generated questions, we make two major improvements: 1) To better encode the semantics of ill-formed questions, we enrich the representation of questions with character embedding and the recent proposed contextual word embedding such as BERT, besides the traditional context-free word embeddings; 2) To make it capable to generate desired questions, we train the model with deep reinforcement learning techniques that considers an appropriate wording of the generation as an immediate reward and the correlation between generated question and answer as time-delayed long-term rewards. Experimental results on real-world datasets show that the proposed QREFINE method can generate refined questions with more readability but fewer mistakes than the original questions provided by users. Moreover, the refined questions also significantly improve the accuracy of answer retrieval. Ye Liu 0006, Yi Chang 0001, Philip S. Yu |
CIKM | 1 |
| 2018 | Multi-View Multi-Graph Embedding for Brain Network Clustering AnalysisabstractNetwork analysis of human brain connectivity is critically important for understanding brain function and disease states. Embedding a brain network as a whole graph instance into a meaningful low-dimensional representation can be used to investigate disease mechanisms and inform therapeutic interventions. Moreover, by exploiting information from multiple neuroimaging modalities or views, we are able to obtain an embedding that is more useful than the embedding learned from an individual view. Therefore, multi-view multi-graph embedding becomes a crucial task. Currently only a few studies have been devoted to this topic, and most of them focus on vector-based strategy which will cause structural information contained in the original graphs lost. As a novel attempt to tackle this problem, we propose Multi-view Multi-graph Embedding M2E by stacking multi-graphs into multiple partially-symmetric tensors and using tensor techniques to simultaneously leverage the dependencies and correlations among multi-view and multi-graph brain networks. Extensive experiments on real HIV and bipolar disorder brain network datasets demonstrate the superior performance of M2E on clustering brain networks by leveraging the multi-view multi-graph interactions. Ye Liu 0006, Lifang He 0001, Bokai Cao, Philip S. Yu, Ann B. Ragin, Alex D. Leow |
AAAI | 1 |
| 2018 | Data-driven Blockbuster Planning on Online Movie Knowledge LibraryabstractIn the era of big data, logistic planning can be made data-driven to take advantage of accumulated knowledge in the past. While in the movie industry, movie planning can also exploit the existing online movie knowledge library to achieve better results. However, it is ineffective to solely rely on conventional heuristics for movie planning, due to a large number of existing movies and various real-world factors that contribute to the success of each movie, such as the movie genre, available budget, production team (involving actor, actress, director, and writer), etc. In this paper, we study a "Blockbuster Planning" (BP) problem to learn from previous movies and plan for low budget yet high return new movies in a totally data-driven fashion. After a thorough investigation of an online movie knowledge library, a novel movie planning framework "Blockbuster Planning with Maximized Movie Configuration Acquaintance" (BigMovie) is introduced in this paper. From the investment perspective, BigMovie maximizes the estimated gross of the planned movies with a given budget. Meanwhile, from the production team's perspective, BigMovie is able to formulate an optimized team with people/movie genres that team members are acquainted with. We formulate the BP problem as a non-linear binary programming problem and prove its NP-hardness. To solve it in polynomial time, BigMovie relaxes the hard binary constraints and addresses the BP problem as a cubic programming problem. This paper is the short version, and you can move to the full version of the paper to get more information. Ye Liu 0006, Jiawei Zhang 0001, Philip S. Yu |
IEEE BigData | 1 |