VLDB 2026 Research / reviewers in the wild / expert
Luyang Kong
dblp:258/1215
· DBLP profile ↗
6ranked-venue papers
0as first author
6since 2021 · last 2026
0009-0004-3960-4992ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Introduction to the Special Issue on Large Language Models for Recommender SystemsabstractRecommender systems have become pivotal in today’s digital landscape, shaping user experiences across diverse online platforms. Recent advances in Large Language Models (LLMs) such as T5, GPT, LLaMA, and their variants have introduced transformative possibilities for recommender systems. LLMs excel in processing and generating natural language text, offering a unique opportunity to reshape the design and elevate the effectiveness of recommendation algorithms. The main topic of this special issue is to explore the integration of Large Language Models and Recommender Systems, encompassing various facets, including model architectures, recommendation algorithms, evaluation methods, and real-world applications. It provides a dedicated platform for researchers and practitioners to share their insights, innovations, and empirical findings in the realm of LLMs for recommender systems, which helps to promote knowledge exchange, leading to best practices and guidelines for integrating LLMs and recommender systems. Towards this goal, the five articles in this collection span trustworthy issues such as recommendation fairness and diversity with LLMs, as well as classic recommendation problems, including sequential recommendation, click-through rate prediction and bundle recommendation with LLMs. By fostering interdisciplinary collaboration between the natural language processing and recommendation communities, the special issue aspires to advance the state of the art in this evolving field. Yongfeng Zhang 0003, Lei Li 0042, Luyang Kong |
Trans. Recomm. Syst. | 3 |
| 2025 | CSR-Bench: Benchmarking LLM Agents in Deployment of Computer Science Research RepositoriesabstractYijia Xiao, Runhui Wang, Luyang Kong, Davor Golac, Wei Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yijia Xiao, Runhui Wang, Luyang Kong, Davor Golac |
NAACL (Long Papers) | 3 |
| 2024 | Learning from Natural Language Explanations for Generalizable Entity MatchingabstractEntity matching is the task of linking records from different sources that refer to the same real-world entity.Past work has primarily treated entity linking as a standard supervised learning problem.However, supervised entity matching models often do not generalize well to new data, and collecting exhaustive labeled training data is often cost prohibitive.Further, recent efforts have adopted LLMs for this task in few/zero-shot settings, exploiting their general knowledge.But LLMs are prohibitively expensive for performing inference at scale for real-world entity matching tasks.As an efficient alternative, we re-cast entity matching as a conditional generation task as opposed to binary classification.This enables us to "distill" LLM reasoning into smaller entity matching models via natural language explanations.This approach achieves strong performance, especially on out-of-domain generalization tests (↑10.85%F-1) where standalone generative methods struggle.We perform ablations that highlight the importance of explanations, both for performance and model robustness.Explain matching label class given the entity descriptions: Label: Match E_a: Nike Sportswear AF-1 488298-436 MN Navy.E_b: Air Force 1 [BRAND] Somin Wadhwa, Adit Krishnan, Runhui Wang, Byron C. Wallace, Luyang Kong |
EMNLP | 5 |
| 2024 | GRAM: Generative Retrieval Augmented Matching of Data Schemas in the Context of Data SecurityabstractSchema matching constitutes a pivotal phase in the data ingestion process for contemporary database systems. Its objective is to discern pairwise similarities between two sets of attributes, each associated with a distinct data table. This challenge emerges at the initial stages of data analytics, such as when incorporating a third-party table into existing databases to inform business insights. Given its significance in the realm of database systems, schema matching has been under investigation since the 2000s. This study revisits this foundational problem within the context of large language models. Adhering to increasingly stringent data security policies, our focus lies on the zero-shot and few-shot scenarios: the model should analyze only a minimal amount of customer data to execute the matching task, contrasting with the conventional approach of scrutinizing the entire data table. We emphasize that the zero-shot or few-shot assumption is imperative to safeguard the identity and privacy of customer data, even at the potential cost of accuracy. The capability to accurately match attributes under such stringent requirements distinguishes our work from previous literature in this domain. Xuanqing Liu, Runhui Wang, Luyang Kong |
KDD | 4 |
| 2024 | Neural Locality Sensitive Hashing for Entity BlockingabstractLocality-sensitive hashing (LSH) is a fundamental algorithmic technique widely employed in large-scale data processing applications, such as nearest-neighbor search, entity resolution, and clustering. However, its applicability in some real-world scenarios is limited due to the need for careful design of hashing functions that align with specific metrics. Existing LSH-based Entity Blocking solutions primarily rely on generic similarity metrics such as Jaccard similarity, whereas practical use cases often demand complex and customized similarity rules surpassing the capabilities of generic similarity metrics. Consequently, designing LSH functions for these customized similarity rules presents considerable challenges. In this research, we propose a neuralization approach to enhance locality-sensitive hashing by training deep neural networks to serve as hashing functions for complex metrics. We assess the effectiveness of this approach within the context of the entity resolution problem, which frequently involves the use of task-specific metrics in real-world applications. Specifically, we introduce NLSHBlock (Neural-LSH Block), a novel blocking methodology that leverages pre-trained language models, fine-tuned with a novel LSH-based loss function. Through extensive evaluations conducted on a diverse range of real-world datasets, we demonstrate the superiority of NLSHBlock over existing methods, exhibiting significant performance improvements. Furthermore, we showcase the efficacy of NLSHBlock in enhancing the performance of the entity matching phase, particularly within the semi-supervised setting. Runhui Wang, Luyang Kong, Yefan Tao, Andrew Borthwick, Davor Golac, Henrik Johnson, Shadie Hijazi, Dong Deng 0001, Yongfeng Zhang 0003 |
SDM | 2 |
| 2022 | Entity Anchored ICD Coding
Jay DeYoung, Han-Chin Shing, Luyang Kong, Christopher Winestock, Chaitanya P. Shivade |
AMIA | 3 |