EDBT 2026 Demo / reviewers in the wild / expert
Weiqing Luo
dblp:313/4765
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2026
0009-0004-8041-3258ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 5 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Utility-Oriented Visual Evidence Selection for Multimodal Retrieval-Augmented GenerationabstractVisual evidence selection is a critical component of multimodal retrieval-augmented generation (RAG), yet existing methods typically rely on semantic relevance or surface-level similarity, which are often misaligned with the actual utility of visual evidence for downstream reasoning.We reformulate multimodal evidence selection from an information-theoretic perspective by defining evidence utility as the information gain induced on a model's output distribution.To overcome the intractability of answer-space optimization, we introduce a latent notion of evidence helpfulness and theoretically show that, under mild assumptions, ranking evidence by information gain on this latent variable is equivalent to answer-space utility.We further propose a training-free, surrogateaccelerated framework that efficiently estimates evidence utility using lightweight multimodal models.Experiments on MRAG-Bench and Visual-RAG across multiple model families demonstrate that our method consistently outperforms state-of-the-art RAG baselines while achieving substantial reductions in computational cost.We release our code at https: //github.com/Hcnaeg/utility-mrag. Weiqing Luo, Zongye Hu, Haofeng Zhang 0008 |
ACL (1) | 1 |
| 2026 | Caddie: A prototype of content-based ad hoc RDF dataset retrievalabstractThe rapid growth of open and structured RDF data on the Web has promoted the development of dataset search as an important research topic. The core function of existing systems is ad hoc dataset retrieval (AHDR) based on the metadata of datasets, which contains limited information and often suffers from quality issues. To overcome the limitations, in this article, we systematically investigate content-based AHDR to exploit the actual RDF data in datasets. We address three main tasks of content-based AHDR with novel methods for handling the large size and complex structure of RDF data to facilitate dataset retrieval, deduplication, and snippet extraction. These methods are integrated into an online and open-source prototype called Caddie . The effectiveness and practicability of its components are evaluated on a public test collection and by a user study. Xiaxia Wang 0001, Qiaosheng Chen, Weiqing Luo, Jeff Z. Pan, Baifan Zhou, Evgeny Kharlamov, Gong Cheng 0001 |
J. Web Semant. | 4 |
| 2025 | TRAWL: External Knowledge-Enhanced Recommendation with LLM AssistanceabstractCombining semantic information with behavioral data is a crucial research area in recommender systems. A promising approach involves leveraging external knowledge to enrich behavioral-based recommender systems with abundant semantic information. However, this approach faces two primary challenges: (1) denoising raw external knowledge and (2) adapting semantic representations. To address these challenges, we propose exTernal knowledge-enhanced RecommendAtion With LLM assistance (TRAWL). This method utilizes large language models to extract relevant recommendation knowledge from raw external data and employs a contrastive learning strategy for adapter training. Experiments on public datasets and real-world online recommender systems validate the effectiveness of our approach. Weiqing Luo, Chonggang Song, Lingling Yi, Gong Cheng 0001 |
CIKM | 1 |
| 2025 | Task-Aware Resolution Optimization for Visual Large Language ModelsabstractReal-world vision-language applications demand varying levels of perceptual granularity.However, most existing visual large language models (VLLMs), such as LLaVA, preassume a fixed resolution for downstream tasks, which leads to subpar performance.To address this problem, we first conduct a comprehensive and pioneering investigation into the resolution preferences of different visionlanguage tasks, revealing a correlation between resolution preferences with ❶ image complexity, and ❷ uncertainty variance of the VLLM at different image input resolutions.Building on this insight, we propose an empirical formula to determine the optimal resolution for a given vision-language task, combining these two factors.Second, based on rigorous experiments, we propose a novel parameter-efficient fine-tuning technique to extend the visual input resolution of pre-trained VLLMs to the identified optimal resolution.Extensive experiments on various vision-language tasks validate the effectiveness of our method. Weiqing Luo, Zhen Tan 0001, Kwonjoon Lee, Behzad Dariush, Tianlong Chen 0001 |
EMNLP | 1 |
| 2025 | Benchmarking Recommendation, Classification, and Tracing Based on Hugging Face Knowledge GraphabstractThe rapid growth of open source machine learning (ML) resources, such as models and datasets, has accelerated IR research. However, existing platforms like Hugging Face do not explicitly utilize structured representations, limiting advanced queries and analyses such as tracing model evolution and recommending relevant datasets. To fill the gap, we construct HuggingKG, the first large-scale knowledge graph built from the Hugging Face community for ML resource management. With 2.6 million nodes and 6.2 million edges, HuggingKG captures domain-specific relations and rich textual attributes. It enables us to further present HuggingBench, a multi-task benchmark with three novel test collections for IR tasks including resource recommendation, classification, and tracing. Our experiments reveal unique characteristics of HuggingKG and the derived tasks. Both resources are publicly available, expected to advance research in open source resource sharing and management. Qiaosheng Chen, Kaijia Huang, Xiao Zhou 0009, Weiqing Luo, Yuanning Cui, Gong Cheng 0001 |
SIGIR | 4 |
| 2024 | ACORDAR 2.0: A Test Collection for Ad Hoc Dataset Retrieval with Densely Pooled Datasets and Question-Style QueriesabstractDataset search, or more specifically, ad hoc dataset retrieval which is a trending specialized IR task, has received increasing attention in both academia and industry. While methods and systems continue evolving, existing test collections for this task exhibit shortcomings, particularly suffering from lexical bias in pooling and limited to keyword-style queries for evaluation. To address these limitations, in this paper, we construct ACORDAR 2.0, a new test collection for this task which is also the largest to date. To reduce lexical bias in pooling, we adapt dense retrieval models to large structured data, using them to find an extended set of semantically relevant datasets to be annotated. To diversify query forms, we employ a large language model to rewrite keyword queries into high-quality question-style queries. We use the test collection to evaluate popular sparse and dense retrieval models to establish a baseline for future studies. The test collection and source code are publicly available. Qiaosheng Chen, Weiqing Luo, Zixian Huang, Tengteng Lin, Xiaxia Wang 0001, Ahmet Soylu, Basil Ell, Baifan Zhou, Evgeny Kharlamov, Gong Cheng 0001 |
SIGIR | 2 |
| 2023 | Dense Re-Ranking with Weak Supervision for RDF Dataset Search
Qiaosheng Chen, Zixian Huang, Weiqing Luo, Tengteng Lin, Gong Cheng 0001 |
ISWC | 4 |