EDBT 2026 Demo / reviewers in the wild / expert
Tengfei Luo
dblp:202/0890
· DBLP profile ↗
8ranked-venue papers
0as first author
8since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 6 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing interpretation of clinical disease-associated copy number variations from multiple sequencing strategies with CNVSeekerabstractMOTIVATION: DNA copy number variations (CNVs) exert a profound impact on major genetic disorders in humans. Although multiple sequencing technologies have become the first line of molecular diagnosis for CNVs, existing tools are unable to resolve the pathogenicity of CNVs directly from raw sequencing data. RESULTS: We developed CNVSeeker, a one-stop and easy-to-use pipeline that provides comprehensive analysis from raw sequencing data to variant interpretation reports, and supports multiple types of sequencing data including short-read data such as whole genome sequencing data and whole exome sequencing data, and long-read sequencing data from Pacific Biosciences HiFi platform or Oxford Nanopore Technologies platform. Through extensive benchmarking, CNVSeeker demonstrated comparable enhancement over the state-of-the-art methods for CNV calling. Moreover, CNVSeeker enables significantly precise variant classification with an accuracy of ∼87%. By applying CNVSeeker to 1946 individuals with autism spectrum disorder (ASD), a total of 133 ASD-associated CNVs in 122 patients were identified, yielding a diagnostic yield of ∼6.3%. Additionally, we have also provided a user-friendly webserver for intuitive visualization of results. This study highlights the potential of CNVSeeker to benefit clinicians and geneticists with limited bioinformatic skill by aiding them interpret CNVs directly from various types of raw sequencing data for auxiliary disease diagnosis. AVAILABILITY AND IMPLEMENTATION: The web server is freely available at https://genemed.tech/cnvseeker and the open-source code can be found at https://github.com/lovelycatZ/CNVSeeker. Xudong Xiang, Xinxin Mao, Tengfei Luo, Chenbin Liu, Bozhao Li, Pei Yu, Dai Wu, Yixiao Zhu, Guihu Zhao, Jinchen Li |
Bioinform. | 3 |
| 2025 | Learning Repetition-Invariant Representations for Polymer InformaticsabstractPolymers are large macromolecules composed of repeating structural units known as monomers and are widely applied in fields such as energy storage, construction, medicine, and aerospace. However, existing graph neural network methods, though effective for small molecules, only model the single unit of polymers and fail to produce consistent vector representations for the true polymer structure with varying numbers of units. To address this challenge, we introduce Graph Repetition Invariance (GRIN), a novel method to learn polymer representations that are invariant to the number of repeating units in their graph representations. GRIN integrates a graph-based maximum spanning tree alignment with repeat-unit augmentation to ensure structural consistency. We provide theoretical guarantees for repetition‐invariance from both model and data perspectives, demonstrating that three repeating units are the minimal augmentation required for optimal invariant representation learning. GRIN outperforms state-of-the-art baselines on both homopolymer and copolymer benchmarks, learning stable, repetition-invariant representations that generalize effectively to polymer chains of unseen sizes. Yihan Zhu, Gang Liu 0025, Eric Inae, Tengfei Luo, Meng Jiang 0001 |
NeurIPS | 4 |
| 2024 | Application of Large Language Models in Chemistry Reaction Data Extraction and CleaningabstractChemical reaction data has existed and still largely exists in unstructured forms. But curating such information into datasets suitable for tasks such as yield and reaction outcome prediction is impractical via manual curation and not possible to automate through programmatic means alone. Large language models (LLMs) have emerged as potent tools, showcasing remarkable capabilities in processing textual information and therefore could be extremely useful in automating this process. To address the challenge of unstructured data, we manually curated a dataset of structured chemical reaction data to fine-tune and evaluate LLMs. We propose a paradigm that leverages prompt-tuning, fine-tuning techniques, and a verifier to check the extracted information. We evaluate the capabilities of various LLMs, including LLAMA-2 and GPT models with different parameter counts, on the data extraction task. Our results show that prompt tuning of GPT-4 yields the best accuracy and evaluation results. Fine-tuning LLAMA-2 models with hundreds of samples does enable them and organize scientific material according to user-defined schemas better though. This workflow shows an adaptable approach for chemical reaction data extraction but also highlights the challenges associated with nuance in chemical information. We open-sourced our code at https://github.com/joker-bruce/LLM_Extraction_Chem. Xiaobao Huang, Mihir Surve, Yuhan Liu 0010, Tengfei Luo, Olaf Wiest, Xiangliang Zhang 0001, Nitesh V. Chawla |
CIKM | 4 |
| 2024 | Graph Diffusion Transformers for Multi-Conditional Molecular GenerationabstractInverse molecular design with diffusion models holds great potential for advancements in material and drug discovery. Despite success in unconditional molecule generation, integrating multiple properties such as synthetic score and gas permeability as condition constraints into diffusion models remains unexplored. We present the Graph Diffusion Transformer (Graph DiT) for multi-conditional molecular generation. Graph DiT has a condition encoder to learn the representation of numerical and categorical properties and utilizes a Transformer-based graph denoiser to achieve molecular graph denoising under conditions. Unlike previous graph diffusion models that add noise separately on the atoms and bonds in the forward diffusion process, we propose a graph-dependent noise model for training Graph DiT, designed to accurately estimate graph-related noise in molecules. We extensively validate the Graph DiT for multi-conditional polymer and small molecule generation. Results demonstrate our superiority across metrics from distribution learning to condition control for molecular properties. A polymer inverse design task for gas separation with feedback from domain experts further demonstrates its practical utility. The code is available at https://github.com/liugangcode/Graph-DiT. Jiaxin Xu, Tengfei Luo |
NeurIPS | 3 |
| 2024 | Rationalizing Graph Neural Networks with Data AugmentationabstractGraph rationales are representative subgraph structures that best explain and support the graph neural network (GNN) predictions. Graph rationalization involves the joint identification of these subgraphs during GNN training, resulting in improved interpretability and generalization. GNN is widely used for node-level tasks such as paper classification and graph-level tasks such as molecular property prediction. However, on both levels, little attention has been given to GNN rationalization and the lack of training examples makes it difficult to identify the optimal graph rationales. In this work, we address the problem by proposing a unified data augmentation framework with two novel operations on environment subgraphs to rationalize GNN prediction. We define the environment subgraph as the remaining subgraph after rationale identification and separation. The framework efficiently performs rationale–environment separation in the representation space for a node’s neighborhood graph or a graph’s complete structure to avoid the high complexity of explicit graph decoding and encoding. We conduct experiments on 17 datasets spanning node classification, graph classification, and graph regression. Results demonstrate that our framework is effective and efficient in rationalizing and enhancing GNNs for different levels of tasks on graphs. Gang Liu 0025, Eric Inae, Tengfei Luo, Meng Jiang 0001 |
ACM Trans. Knowl. Discov. Data | 3 |
| 2023 | Semi-Supervised Graph Imbalanced RegressionabstractData imbalance is easily found in annotated data when the observations of certain continuous label values are difficult to collect for regression tasks. When they come to molecule and polymer property predictions, the annotated graph datasets are often small because labeling them requires expensive equipment and effort. To address the lack of examples of rare label values in graph regression tasks, we propose a semi-supervised framework to progressively balance training data and reduce model bias via self-training. The training data balance is achieved by (1) pseudo-labeling more graphs for under-represented labels with a novel regression confidence measurement and (2) augmenting graph examples in latent space for remaining rare labels after data balancing with pseudo-labels. The former is to identify quality examples from unlabeled data whose labels are confidently predicted and sample a subset of them with a reverse distribution from the imbalanced annotated data. The latter collaborates with the former to target a perfect balance using a novel label-anchored mixup algorithm. We perform experiments in seven regression tasks on graph datasets. Results demonstrate that the proposed framework significantly reduces the error of predicted graph properties, especially in under-represented label areas. Gang Liu 0025, Tong Zhao 0003, Eric Inae, Tengfei Luo, Meng Jiang 0001 |
KDD | 4 |
| 2023 | Data-Centric Learning from Unlabeled Graphs with Diffusion ModelabstractGraph property prediction tasks are important and numerous. While each task offers a small size of labeled examples, unlabeled graphs have been collected from various sources and at a large scale. A conventional approach is training a model with the unlabeled graphs on self-supervised tasks and then fine-tuning the model on the prediction tasks. However, the self-supervised task knowledge could not be aligned or sometimes conflicted with what the predictions needed. In this paper, we propose to extract the knowledge underlying the large set of unlabeled graphs as a specific set of useful data points to augment each property prediction model. We use a diffusion model to fully utilize the unlabeled graphs and design two new objectives to guide the model's denoising process with each task's labeled data to generate task-specific graph examples and their labels. Experiments demonstrate that our data-centric approach performs significantly better than fifteen existing various methods on fifteen tasks. The performance improvement brought by unlabeled data is visible as the generated labeled examples unlike the self-supervised learning. Gang Liu 0025, Eric Inae, Tong Zhao 0003, Jiaxin Xu, Tengfei Luo, Meng Jiang 0001 |
NeurIPS | 5 |
| 2022 | Graph Rationalization with Environment-based AugmentationsabstractRationale is defined as a subset of input features that best explains or supports the prediction by machine learning models. Rationale identification has improved the generalizability and interpretability of neural networks on vision and language data. In graph applications such as molecule and polymer property prediction, identifying representative subgraph structures named as graph rationales plays an essential role in the performance of graph neural networks. Existing graph pooling and/or distribution intervention methods suffer from the lack of examples to learn to identify optimal graph rationales. In this work, we introduce a new augmentation operation called environment replacement that automatically creates virtual data examples to improve rationale identification. We propose an efficient framework that performs rationale-environment separation and representation learning on the real and augmented examples in latent spaces to avoid the high complexity of explicit graph decoding and encoding. Comparing against recent techniques, experiments on seven molecular and four polymer datasets demonstrate the effectiveness and efficiency of the proposed augmentation-based graph rationalization framework. Data and the implementation of the proposed framework are publicly available https://github.com/liugangcode/GREA. Gang Liu 0025, Tong Zhao 0003, Jiaxin Xu, Tengfei Luo, Meng Jiang 0001 |
KDD | 4 |