Feng Luo 0005

dblp:181/2672-5 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2026
0009-0005-0448-3462ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Decomposition-Driven Multi-Table Retrieval and Reasoning for Numerical Question Answering
abstract
In this paper, we study the problem of numerical multi-table question answering (MTQA) over large-scale table collections (e.g., online data repositories). This task is essential in many analytical applications. Existing MTQA solutions, such as text-to-SQL or open-domain MTQA methods, are designed for databases and struggle when applied to large-scale table collections. The key limitations include: (1) Limited support for complex table relationships; (2) Ineffective retrieval of relevant tables at scale; (3) Inaccurate answer generation. To overcome these limitations, we propose DMRAL, a Decomposition-driven Multi-table Retrieval and Answering framework for MTQA over large-scale table collections, which consists of: (1) constructing a table relationship graph to capture complex relationships among tables; (2) Table-Aligned Question Decomposer and Coverage-Aware Retriever, which jointly enable the effective identification of relevant tables from large-scale corpora by enhancing the question decomposition quality and maximizing the question coverage of retrieved tables; and (3) Sub-question Guided Reasoner, which produces correct answers by progressively generating and refining the reasoning program based on sub-questions. Experiments on two MTQA datasets demonstrate that DMRAL significantly outperforms existing state-of-the-art MTQA methods, with an average improvement of 24% in table retrieval and 55% in answer accuracy.
Feng Luo 0005, Hui Luo 0001, Zhifeng Bao, Xiaoli Wang 0002, J. Shane Culpepper, Shazia Sadiq
ICDE1
2026 Missing Value Imputation in Tabular Data Lakes Unleashed: A Hybrid Approach
abstract
Abstract Missing values in tabular data lakes can severely impact data analysis and diminish the performance in downstream applications. We highlight that a robust imputation strategy should properly take three aspects of variety into consideration: source of imputed value, the types of tables involved, and the data types of the missing value. Existing imputation methods rely on estimation-based approaches (using a model trained on data from the same table to estimate missing values) or search-based approaches (retrieving values from other tables). Unfortunately, none of these approaches effectively incorporate all three aspects of variety. To address this gap, we propose , a novel framework that uses a C ombination of E stimation-based and S earch-based methods for missing value I mputation in D ata lakes. contains three core modules: (1) the , which efficiently discovers candidate values from tables by exploiting the contextual information; (2) the , which introduces an influence function and a sampling-based exploration strategy to yield accurate estimated values; (3) the , which determines the most suitable method based on table-level and column-level statistics. Extensive experiments conducted on three data lakes demonstrate that effectively and efficiently addresses the missing value problem.
Feng Luo 0005, Hui Luo 0001, Zhifeng Bao, J. Shane Culpepper, Shazia Sadiq, Xiaoli Wang 0002
VLDB J.1
2023 A One-Size-Fits-Three Representation Learning Framework for Patient Similarity Search
abstract
Abstract Patient similarity search is an essential task in healthcare. Recent studies adopted electronic health records (EHRs) to learn patient representations for measuring the clinical similarities. These methods outperformed traditional methods, by capturing more information from various sources consisting of multi-modal EHRs, external knowledge and correlations among medical concepts. They often concerned certain type of data without taking full advantage of various information. We propose a graph representation learning framework, denoted by One-Size-Fits-Three ( OSFT ), that takes into account fusion-attention, neighbor-attention and global-attention from three types of information. Extensive experiments are conducted on two real datasets of MIMIC-III and MIMIC-IV, and the results verified the effectiveness and generality of our framework. When compared with baselines on patient similarity search, our framework achieved good effectiveness and comparative efficiency. The results provide new insights about whether the use of various information can better measure the patient similarity. The source codes are available at https://github.com/emmali808/ADDS/tree/master/EHRDeepHelper .
Yefan Huang, Feng Luo 0005, Xiaoli Wang 0002, Bohan Li 0001
Data Sci. Eng.2
2021 IMAS++: An Intelligent Medical Analysis System Enhanced with Deep Graph Neural Networks
abstract
This paper demonstrates an intelligent medical analysis system. We aim to address two main challenges: 1) medical data often contain heterogeneous information which are usually valuable but difficult to be modeled; 2) medical data are often lacking of large scale labeled data which usually require huge efforts to build. To resolve the first challenge, we propose a novel multi-modal heterogeneous graph model to represent the medical data. Based on this model, graph neural networks can be directly applied to effective medical case clustering. This helps to resolve the second challenge for label assignment in the same cluster. To further evaluate the practical use of the proposed model, the system also proposes an effective similar medical case retrieval framework based on a novel graph similarity learning model. We have implemented the system and the source codes are published at https://github.com/emmali808/ADDS. With our system, users can easily pinpoint valuable historical medical information they are interested in and obtain closely relevant medical cases for further diagnosis.
Feng Luo 0005, Xiaoli Wang 0002
CIKM1