VLDB 2026 Research / reviewers in the wild / expert
Shengfei Lyu
dblp:268/5763
· DBLP profile ↗
12ranked-venue papers
2as first author
11since 2021 · last 2025
0000-0002-1843-6836ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Generation-Augmented and Embedding Fusion in Document-Level Event Argument ExtractionabstractDocument-level event argument extraction is a crucial task that aims to extract arguments from the entire document, beyond sentence-level analysis. Prior classification-based models still fail to explicitly capture significant relationships and heavily relies on large-scale datasets. In this study, we propose a novel approach called Generation-Augmented and Embedding Fusion. This approach first uses predefined templates and generative language models to produce an embedding capturing role relationship information, then integrates it into the foundational embedding derived from a classification model through a noval embedding fusion mechanism. We conduct the extensive experiments on the RAMS and WikiEvents datasets to demonstrate that our approach is more effective than the baselines, and that it is also data-efficient in low-resource scenarios. Xingjian Lin, Shengfei Lyu, Xin Wang 0179, Qiuju Chen, Huanhuan Chen 0001 |
COLING | 2 |
| 2025 | ESGenius: Benchmarking LLMs on Environmental, Social, and Governance (ESG) and Sustainability KnowledgeabstractChaoyue He, Xin Zhou, Yi Wu, Xinjia Yu, Yan Zhang, Lei Zhang, Di Wang, Shengfei Lyu, Hong Xu, Wang Xiaoqiao, Wei Liu, Chunyan Miao. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Chaoyue He, Xin Zhou 0008, Xinjia Yu, Lei Zhang 0199, Di Wang 0004, Shengfei Lyu, Hong Xu 0004, Xiaoqiao Wang, Chunyan Miao |
EMNLP | 8 |
| 2025 | Variational Counterfactual Intervention Planning to Achieve Target OutcomesabstractA key challenge in personalized healthcare is identifying optimal intervention sequences to guide temporal systems toward target outcomes, a novel problem we formalize as counterfactual target achievement. In addressing this problem, directly adopting counterfactual estimation methods face compounding errors due to the unobservability of counterfactuals. To overcome this, we propose Variational Counterfactual Intervention Planning (VCIP), which reformulates the problem by modeling the conditional likelihood of achieving target outcomes, implemented through variational inference. By leveraging the g-formula to bridge the gap between interventional and observational log-likelihoods, VCIP enables reliable training from observational data. Experiments on both synthetic and real-world datasets show that VCIP significantly outperforms existing methods in target achievement accuracy. Xin Wang 0179, Shengfei Lyu, Chi Luo, Xiren Zhou, Huanhuan Chen 0001 |
ICML | 2 |
| 2025 | Enhancing Counterfactual Estimation: A Focus on Temporal TreatmentsabstractIn the medical field, treatment sequences significantly influence future outcomes through complex temporal interactions. Therefore, highlighting the role of temporal treatments within the model is crucial for accurate counterfactual estimation, which is often overlooked in current methods. To address this, we employ Koopman theory, known for its capability to model complex dynamic systems, and introduce a novel model named the Counterfactual Temporal Dynamics Network via Neural Koopman Operators (CTD-NKO). This model utilizes Koopman operators to encapsulate sequential treatment data, aiming to capture the causal dynamics within the system induced by temporal interactions between treatments. Moreover, CTD-NKO implements a weighting strategy that aligns joint and marginal distributions of the system state and the current treatment to mitigate time-varying confounding bias. This deviates from the balanced representation strategy employed by existing methods, as we demonstrate that such a strategy may suffer from the potential information loss of historical treatments. These designs allow CTD-NKO to exploit treatment information more thoroughly and effectively, resulting in superior performance on both synthetic and real-world datasets. Xin Wang 0179, Shengfei Lyu, Kangyang Luo, Lishan Yang 0004, Huanhuan Chen 0001, Chunyan Miao |
IJCAI | 2 |
| 2024 | A Dual-module Framework for Counterfactual Estimation over TimeabstractEfficiently and effectively estimating counterfactuals over time is crucial for optimizing treatment strategies. We present the Adversarial Counterfactual Temporal Inference Network (ACTIN), a novel framework with dual modules to enhance counterfactual estimation. The balancing module employs a distribution-based adversarial method to learn balanced representations, extending beyond the limitations of current classification-based methods to mitigate confounding bias across various treatment types. The integrating module adopts a novel Temporal Integration Predicting (TIP) strategy, which has a wider receptive field of treatments and balanced representations from the beginning to the current time for a more profound level of analysis. TIP goes beyond the established Direct Predicting (DP) strategy, which only relies on current treatments and representations, by empowering the integrating module to effectively capture long-range dependencies and temporal treatment interactions. ACTIN exceeds the confines of specific base models, and when implemented with simple base models, consistently delivers state-of-the-art performance and efficiency across both synthetic and real-world datasets. Xin Wang 0179, Shengfei Lyu, Lishan Yang 0004, Yibing Zhan, Huanhuan Chen 0001 |
ICML | 2 |
| 2023 | Temporal knowledge graph embedding via sparse transfer matrix
Xin Wang 0179, Shengfei Lyu, Xiangyu Wang 0016, Huanhuan Chen 0001 |
Inf. Sci. | 2 |
| 2022 | Generalization Bounds for Estimating Causal Effects of Continuous TreatmentsabstractWe focus on estimating causal effects of continuous treatments (e.g., dosage in medicine), also known as dose-response function. Existing methods in causal inference for continuous treatments using neural networks are effective and to some extent reduce selection bias, which is introduced by non-randomized treatments among individuals and might lead to covariate imbalance and thus unreliable inference. To theoretically support the alleviation of selection bias in the setting of continuous treatments, we exploit the re-weighting schema and the Integral Probability Metric (IPM) distance to derive an upper bound on the counterfactual loss of estimating the average dose-response function (ADRF), and herein the IPM distance builds a bridge from a source (factual) domain to an infinite number of target (counterfactual) domains. We provide a discretized approximation of the IPM distance with a theoretical guarantee in the practical implementation. Based on the theoretical analyses, we also propose a novel algorithm, called Average Dose- response estiMatIon via re-weighTing schema (ADMIT). ADMIT simultaneously learns a re-weighting network, which aims to alleviate the selection bias, and an inference network, which makes factual and counterfactual estimations. In addition, the effectiveness of ADMIT is empirically demonstrated in both synthetic and semi-synthetic experiments by outperforming the existing benchmarks. Xin Wang 0179, Shengfei Lyu, Huanhuan Chen 0001 |
NeurIPS | 2 |
| 2022 | Contextual-Aware Information Extractor with Adaptive Objective for Chinese Medical DialoguesabstractElectronic Medical Records ( EMRs ) are the foundation of modern medical information systems. Despite the benefits of EMRs, the exhausting process of constructing EMRs decreases the efficiency of medical consultation. Therefore, it becomes an emerging research field to automatically extract EMRs from medical dialogues. In Chinese medical dialogues, the phenomena of omission and reference are extremely common, leading to strong contextual relevance among utterances. However, recent studies on converting Chinese medical dialogues to EMRs lack a reliable mechanism to effectively exploit the contextual relevance information among utterances. Moreover, they neglect the frequency imbalance of different items and treat these items indiscriminately, which eventually degrade the overall system performance. In this article, we proposed a Contextual-Aware Information Extractor ( CANE ), which employs a local-to-global mechanism over utterances to model the contextual relevance among utterances. Furthermore, an adaptive objective is introduced to alleviate the frequency imbalance of items by dynamically assigning weights to each sample. Experimental results indicate that CANE outperforms previous state-of-the-arts with considerable improvements (+6.11% and +3.39% on F1-score). Gangqiang Hu, Shengfei Lyu, Jinlong Li 0001, Huanhuan Chen 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 2 |
| 2022 | Mapping the Buried Cable by Ground Penetrating Radar and Gaussian-Process RegressionabstractWith the rapid expansion of urban areas and the increasing use of electricity, the need for locating buried cables is becoming urgent. In this paper, a novel method to locate underground cables based on Ground Penetrating Radar (GPR) and Gaussian-process regression is proposed. Firstly, the coordinate system of the detected area is conducted, and the input and output of locating buried cables are determined. The GPR is moved along the established parallel detection lines, and the hyperbolic signatures generated by buried cables are identified and fitted, thus the positions and depths of some points on the cable could be derived. On the basis of the established coordinate system and the derived points on the cable, the clustering method and cable fitting algorithm based on Gaussian-process regression are proposed to find the most likely locations of the underground cables. Furthermore, the confidence intervals of the cables’ locations are also obtained. Both the position and depth noises are taken into account in our method, ensuring the robustness in different environments and equipment. Experiments on real-world datasets are conducted, and the obtained results demonstrate the effectiveness of the proposed method. Xiren Zhou, Qiuju Chen, Shengfei Lyu, Huanhuan Chen 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Convolutional Recurrent Neural Networks for Text ClassificationabstractRecurrent neural network (RNN) and convolutional neural network (CNN) are two prevailing architectures used in text classification. Traditional approaches combine the strengths of these two networks by straightly streamlining them or linking features extracted from them. In this article, a novel approach is proposed to maintain the strengths of RNN and CNN to a great extent. In the proposed approach, a bi-directional RNN encodes each word into forward and backward hidden states. Then, a neural tensor layer is used to fuse bi-directional hidden states to get word representations. Meanwhile, a convolutional neural network is utilized to learn the importance of each word for text classification. Empirical experiments are conducted on several datasets for text classification. The superior performance of the proposed approach confirms its effectiveness. Shengfei Lyu |
J. Database Manag. | 1 |
| 2021 | Short isometric shapelet transform for binary time series classificationabstractIn the research area of time series classification, the ensemble shapelet transform algorithm is one of state-of-the-art algorithms for classification. However, its high time complexity is an issue to hinder its application since its base classifier shapelet transform includes a high time complexity of a distance calculation and shapelet selection. Therefore, in this paper we introduce a novel algorithm, i.e. short isometric shapelet transform, which contains two strategies to reduce the time complexity. The first strategy of SIST fixes the length of shapelet based on a simplified distance calculation, which largely reduces the number of shapelet candidates as well as speeds up the distance calculation in the ensemble shapelet transform algorithm. The second strategy is to train a single linear classifier in the feature space instead of an ensemble classifier. The theoretical evidences of these two strategies are presented to guarantee a near-lossless accuracy under some preconditions while reducing the time complexity. Furthermore, empirical experiments demonstrate the superior performance of the proposed algorithm. Weibo Shu, Yaqiang Yao, Shengfei Lyu, Jinlong Li 0001, Huanhuan Chen 0001 |
Knowl. Inf. Syst. | 3 |
| 2020 | Multiclass Probabilistic Classification Vector MachineabstractThe probabilistic classification vector machine (PCVM) synthesizes the advantages of both the support vector machine and the relevant vector machine, delivering a sparse Bayesian solution to classification problems. However, the PCVM is currently only applicable to binary cases. Extending the PCVM to multiclass cases via heuristic voting strategies such as one-vs-rest or one-vs-one often results in a dilemma where classifiers make contradictory predictions, and those strategies might lose the benefits of probabilistic outputs. To overcome this problem, we extend the PCVM and propose a multiclass PCVM (mPCVM). Two learning algorithms, i.e., one top-down algorithm and one bottom-up algorithm, have been implemented in the mPCVM. The top-down algorithm obtains the maximum a posteriori (MAP) point estimates of the parameters based on an expectation-maximization algorithm, and the bottom-up algorithm is an incremental paradigm by maximizing the marginal likelihood. The superior performance of the mPCVMs, especially when the investigated problem has a large number of classes, is extensively evaluated on the synthetic and benchmark data sets. Shengfei Lyu, Xing Tian, Yang Li 0066, Bingbing Jiang 0001, Huanhuan Chen 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |