VLDB 2026 Research / reviewers in the wild / expert
Yalin Sun
dblp:163/7108
· DBLP profile ↗
6ranked-venue papers
2as first author
4since 2021 · last 2025
0000-0003-1414-8295ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Improving Generated and Retrieved Knowledge Combination Through Zero-shot GenerationabstractOpen-domain Question Answering (QA) has garnered substantial interest by combining the advantages of faithfully retrieved passages and relevant passages generated through Large Language Models (LLMs). However, there is a lack of definitive labels available to pair these sources of knowledge. In order to address this issue, we propose an unsupervised and simple framework called Bi-Reranking for Merging Generated and Retrieved Knowledge (BRMGR), which utilizes re-ranking methods for both retrieved passages and LLM-generated passages. We pair the two types of passages using two separate re-ranking methods and then combine them through greedy matching. We demonstrate that BRMGR is equivalent to employing a bipartite matching loss when assigning each retrieved passage with a corresponding LLM-generated passage. The application of our model yielded experimental results from three datasets, improving their performance by +1.7 and +1.6 on NQ and WebQ datasets, respectively, and obtaining comparable result on TriviaQA dataset when compared to competitive baselines. Xinkai Du, Quanjie Han, Yan Liu 0004, Yalin Sun, Hongbo Shan, Maosong Sun 0001 |
ICASSP | 5 |
| 2025 | Posture-Aware Robust Person Re-Identification via Optimal Transport Calibration
Ruiying Lu, Yalin Sun, Chunlei Peng, Yu Zheng 0006 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Label Dependencies-Aware Set Prediction Networks for Multi-Label Text ClassificationabstractMulti-label text classification involves extracting all relevant labels from a sentence. Given the unordered nature of these labels, we propose approaching the problem as a set prediction task. To address the correlation between labels, we leverage Graph Convolutional Networks and construct an adjacency matrix based on the statistical relations between labels. Additionally, we enhance recall ability by applying the Bhattacharyya distance to the output distributions of the set prediction networks. We evaluate the effectiveness of our approach on two multi-label datasets and demonstrate its superiority over previous baselines through experimental results. Xinkai Du, Quanjie Han, Yalin Sun, Maosong Sun 0001 |
ICASSP | 3 |
| 2024 | Topic-Aware Sensitive Information Detection in Chinese Large Language ModelabstractWith the rapid advancement of deep learning, generative AI models have emerged as a prominent area of focus. However, these developments bring potential security concerns. China’s guiding document of government on generative AI security identifies 31 specific security risks across five categories, including content that violates socialist core values and discriminatory content. Detecting the sensitivity of both user input and model-generated content has therefore become a critical challenge for the security of generative AI models. This paper proposes a robust scheme for detecting sensitive information in the Chinese languages generated from the large language models. Initially, we detect and classify the security risks outlined in the document "Basic Security Requirements for Generative Artificial Intelligence Service" and develop a comprehensive dataset, named Chinese Sensitive Language Detection (CSLD). Specifically, in order to leverage the semantics of languages, we introduce the topic model to pre-analyze text data and construct topic-augmented classification vectors (TACV) that supply effective contextual information for sensitive content detection. Additionally, we propose a topic-infused attention mechanism (TIAM) to provide richer contextual information and relevant topics to guide sensitive information detection. At the same time, the proposed framework is designed to integrate with various classes of Chinese pre-trained models, enabling accurate classification of sensitive content while maintaining low latency and memory usage. Furthermore, the proposed dataset surpasses existing ones in terms of coverage, data volume, and its focus on the security challenges specific to Chinese large language models. Without bells and whistles, our experiments demonstrate that our model outperforms existing models in terms of accuracy and efficiency on the CSLD dataset. Yalin Sun, Ruiying Lu |
TrustCom | 1 |
| 2016 | Individual Differences and Online Health Information Source SelectionabstractOnline information sources are become increasingly diverse. The selection of sources for health information has significant implications for people's healthcare decision-making. However, little is known about how individual differences influence users' selection of online sources for health information. This study intends to fill this gap by exploring the impact of a number of individual characteristics on users' selection of five online sources (search engines, social Q&A sites, social networking sites (SNSs), online health communities, and crowdsourcing sites) for three distinct types of health-related search task (factual, exploratory, and personal experiences). We found that individuals' health literacy and frequency of using a source are the most significant predictors of their source selections across task types. Preference for information has an impact on users' selection of SNSs for exploratory tasks. Extraversion personality has an impact on users' selection of search engines for tasks that seek personal experiences. Nevertheless, demographic factors, including gender, income, and health status, do not predict users' selection of online sources for health information. Yalin Sun, Yan Zhang 0005 |
CHIIR | 1 |
| 2015 | Quality of health information for consumers on the web: A systematic review of indicators, criteria, tools, and evaluation resultsabstractThe quality of online health information for consumers has been a critical issue that concerns all stakeholders in healthcare. To gain an understanding of how quality is evaluated, this systematic review examined 165 articles in which researchers evaluated the quality of consumer‐oriented health information on the web against predefined criteria. It was found that studies typically evaluated quality in relation to the substance and formality of content, as well as to the design of technological platforms. Attention to design, particularly interactivity, privacy, and social and cultural appropriateness is on the rise, which suggests the permeation of a user‐centered perspective into the evaluation of health information systems, and a growing recognition of the need to study these systems from a social‐technical perspective. Researchers used many preexisting instruments to facilitate evaluation of the formality of content; however, only a few were used in multiple studies, and their validity was questioned. The quality of content (i.e., accuracy and completeness) was always evaluated using proprietary instruments constructed based on medical guidelines or textbooks. The evaluation results revealed that the quality of health information varied across medical domains and across websites, and that the overall quality remained problematic. Future research is needed to examine the quality of user‐generated content and to explore opportunities offered by emerging new media that can facilitate the consumer evaluation of health information. Yan Zhang 0005, Yalin Sun, Bo Xie 0001 |
J. Assoc. Inf. Sci. Technol. | 2 |