VLDB 2026 Research / reviewers in the wild / expert
Li-Wei H. Lehman
dblp:87/2340
· DBLP profile ↗
13ranked-venue papers
6as first author
4since 2021 · last 2025
0000-0002-3782-9977ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Generative modeling · 33% Learning paradigms · 19% Probabilistic and Bayesian machine learning · 19% | |
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Medical and health informatics · 66% Computational science and engineering · 34% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 18 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling › diffusion model › conditional diffusion model
conditional denoising diffusion |
0.7 | 1 | 2023 | A Diffusion Model with Contrastive Learning for ICU False Arrhythmia Alarm Reduction · IJCAI 2023 |
Machine learning › Generative modeling
diffusion model |
0.7 | 1 | 2023 | A Diffusion Model with Contrastive Learning for ICU False Arrhythmia Alarm Reduction · IJCAI 2023 |
Medical and health informatics
clinical monitoring |
0.7 | 1 | 2023 | VTaC: A Benchmark Dataset of Ventricular Tachycardia Alarms from ICU Monitors · NeurIPS 2023 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.6 | 1 | 2022 | Knowledge Distillation via Constrained Variational Inference · AAAI 2022 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference |
0.6 | 1 | 2022 | Knowledge Distillation via Constrained Variational Inference · AAAI 2022 |
Computational science and engineering › partial differential equations
partial differential equation discovery |
0.4 | 1 | 2020 | Robust Low-Rank Discovery of Data-Driven Partial Differential Equations · AAAI 2020 |
Mathematical optimization › statistical estimation › regression
sparse regression |
0.4 | 1 | 2020 | Robust Low-Rank Discovery of Data-Driven Partial Differential Equations · AAAI 2020 |
Computer vision › 3D vision
feature matching |
0.4 | 1 | 2019 | Retaining Privileged Information for Multi-Task Learning · KDD 2019 |
Machine learning › Learning paradigms › supervised learning
learning using privileged information |
0.4 | 1 | 2019 | Retaining Privileged Information for Multi-Task Learning · KDD 2019 |
Machine learning › Learning paradigms
multi-task learning |
0.4 | 1 | 2019 | Retaining Privileged Information for Multi-Task Learning · KDD 2019 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
0.2 | 1 | 2022 | Knowledge Distillation via Constrained Variational Inference · AAAI 2022 |
Natural language and speech › Information extraction and text analysis
topic model |
0.2 | 1 | 2022 | Knowledge Distillation via Constrained Variational Inference · AAAI 2022 |
Machine learning and data management › matrix recovery
low-rank matrix recovery |
0.1 | 1 | 2020 | Robust Low-Rank Discovery of Data-Driven Partial Differential Equations · AAAI 2020 |
Internet architecture and protocols › future internet architecture
active networks |
0.0 | 1 | 1998 | Active Reliable Multicast · INFOCOM 1998 |
Internet of things and sensor networks › wireless sensor network
in-network processing |
0.0 | 1 | 1998 | Active Reliable Multicast · INFOCOM 1998 |
Transport protocols and congestion control
loss recovery |
0.0 | 1 | 1998 | Active Reliable Multicast · INFOCOM 1998 |
Internet architecture and protocols
multicast |
0.0 | 1 | 1998 | Active Reliable Multicast · INFOCOM 1998 |
Internet architecture and protocols › multicast
reliable multicast |
0.0 | 1 | 1998 | Active Reliable Multicast · INFOCOM 1998 |
Methods — techniques the papers use, named apart from their topics
self-attention · 1.3residual links · 1.3contrastive learning · 1.3threshold ridge regression · 1.3nuclear norm minimization · 1.3l1/l0 regularization · 1.3supervised learning · 0.7semi-supervised learning · 0.7generative model · 0.7deep learning · 0.7variational inference · 0.6knowledge distillation · 0.6automatic differentiation variational inference · 0.6sample complexity analysis · 0.4feature matching · 0.4simulation · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Creative style transfer for image stylization via learning neural permutation
Zedong Zhang, Gan Sun, Li-Wei H. Lehman, Jian Yang 0003, Jun Li 0027 |
Knowl. Based Syst. | 4 |
| 2023 | A Diffusion Model with Contrastive Learning for ICU False Arrhythmia Alarm ReductionabstractThe high rate of false arrhythmia alarms in intensive care units (ICUs) can negatively impact patient care and lead to slow staff response time due to alarm fatigue. To reduce false alarms in ICUs, previous works proposed conventional supervised learning methods which have inherent limitations in dealing with high-dimensional, sparse, unbalanced, and limited data. We propose a deep generative approach based on the conditional denoising diffusion model to detect false arrhythmia alarms in the ICUs. Conditioning on past waveform data of a patient, our approach generates waveform predictions of the patient during an actual arrhythmia event, and uses the distance between the generated and the observed samples to classify the alarm. We design a network with residual links and self-attention mechanism to capture long-term dependencies in signal sequences, and leverage the contrastive learning mechanism to maximize distances between true and false arrhythmia alarms. We demonstrate the effectiveness of our approach on the MIMIC II arrhythmia dataset for detecting false alarms in both retrospective and real-time settings. Guoshuai Zhao 0001, Xueming Qian, Li-Wei H. Lehman |
IJCAI | 4 |
| 2023 | VTaC: A Benchmark Dataset of Ventricular Tachycardia Alarms from ICU MonitorsabstractFalse arrhythmia alarms in intensive care units (ICUs) are a continuing problem despite considerable effort from industrial and academic algorithm developers. Of all life-threatening arrhythmias, ventricular tachycardia (VT) stands out as the most challenging arrhythmia to detect reliably. We introduce a new annotated VT alarm database, VTaC (Ventricular Tachycardia annotated alarms from ICUs) consisting of over 5,000 waveform recordings with VT alarms triggered by bedside monitors in the ICU. Each VT alarm waveform in the dataset has been labeled by at least two independent human expert annotators. The dataset encompasses data collected from ICUs in two major US hospitals and includes data from three leading bedside monitor manufacturers, providing a diverse and representative collection of alarm waveform data. Each waveform recording comprises at least two electrocardiogram (ECG) leads and one or more pulsatile waveforms, such as photoplethysmogram (PPG or PLETH) and arterial blood pressure (ABP) waveforms. We demonstrate the utility of this new benchmark dataset for the task of false arrhythmia alarm reduction, and present performance of multiple machine learning approaches, including conventional supervised machine learning, deep learning, semi-supervised learning, and generative approaches for the task of VT false alarm reduction. Li-Wei H. Lehman, Benjamin Moody, Harsh Deep, Hasan Saeed, Lucas McCullum, Diane Perry, Tristan Struja, Qiao Li 0011, Gari D. Clifford, Roger G. Mark |
NeurIPS | 1 |
| 2022 | Knowledge Distillation via Constrained Variational InferenceabstractKnowledge distillation has been used to capture the knowledge of a teacher model and distill it into a student model with some desirable characteristics such as being smaller, more efficient, or more generalizable. In this paper, we propose a framework for distilling the knowledge of a powerful discriminative model such as a neural network into commonly used graphical models known to be more interpretable (e.g., topic models, autoregressive Hidden Markov Models). Posterior of latent variables in these graphical models (e.g., topic proportions in topic models) is often used as feature representation for predictive tasks. However, these posterior-derived features are known to have poor predictive performance compared to the features learned via purely discriminative approaches. Our framework constrains variational inference for posterior variables in graphical models with a similarity preserving constraint. This constraint distills the knowledge of the discriminative model into the graphical model by ensuring that input pairs with (dis)similar representation in the teacher model also have (dis)similar representation in the student model. By adding this constraint to the variational inference scheme, we guide the graphical model to be a reasonable density model for the data while having predictive features which are as close as possible to those of a discriminative model. To make our framework applicable to a wide range of graphical models, we build upon the Automatic Differentiation Variational Inference (ADVI), a black-box inference framework for graphical models. We demonstrate the effectiveness of our framework on two real-world tasks of disease subtyping and disease trajectory modeling. Ardavan Saeedi, Yuria Utsumi, Li Sun 0010, Kayhan Batmanghelich, Li-Wei H. Lehman |
AAAI | 5 |
| 2020 | Robust Low-Rank Discovery of Data-Driven Partial Differential EquationsabstractPartial differential equations (PDEs) are essential foundations to model dynamic processes in natural sciences. Discovering the underlying PDEs of complex data collected from real world is key to understanding the dynamic processes of natural laws or behaviors. However, both the collected data and their partial derivatives are often corrupted by noise, especially from sparse outlying entries, due to measurement/process noise in the real-world applications. Our work is motivated by the observation that the underlying data modeled by PDEs are in fact often low rank. We thus develop a robust low-rank discovery framework to recover both the low-rank data and the sparse outlying entries by integrating double low-rank and sparse recoveries with a (group) sparse regression method, which is implemented as a minimization problem using mixed nuclear norms with ℓ1 and ℓ0 norms. We propose a low-rank sequential (grouped) threshold ridge regression algorithm to solve the minimization problem. Results from several experiments on seven canonical models (i.e., four PDEs and three parametric PDEs) verify that our framework outperforms the state-of-art sparse and group sparse regression methods. Code is available at https://github.com/junli2019/Robust-Discovery-of-PDEs Jun Li 0027, Gan Sun, Guoshuai Zhao 0001, Li-Wei H. Lehman |
AAAI | 4 |
| 2020 | Is Deep Reinforcement Learning Ready for Practical Applications in Healthcare? A Sensitivity Analysis of Duel-DDQN for Hemodynamic Management in Sepsis Patients
Mingyu Lu, Zach Shahn, Daby M. Sow, Finale Doshi-Velez, Li-Wei H. Lehman |
AMIA | 5 |
| 2019 | Retaining Privileged Information for Multi-Task LearningabstractKnowledge transfer has been of great interest in current machine learning research, as many have speculated its importance in modeling the human ability to rapidly generalize learned models to new scenarios. Particularly in cases where training samples are limited, knowledge transfer shows improvement on both the learning speed and generalization performance of related tasks. Recently, Learning Using Privileged Information (LUPI) has presented a new direction in knowledge transfer by modeling the transfer of prior knowledge as a Teacher-Student interaction process. Under LUPI, a Teacher model uses Privileged Information (PI) that is only available at training time to improve the sample complexity required to train a Student learner for a given task. In this work, we present a LUPI formulation that allows privileged information to be retained in a multi-task learning setting. We propose a novel feature matching algorithm that projects samples from the original feature space and the privilege information space into a joint latent space in a way that informs similarity between training samples. Our experiments show that useful knowledge from PI is maintained in the latent space and greatly improves the sample efficiency of other related learning tasks. We also provide an analysis of sample complexity of the proposed LUPI method, which under some favorable assumptions can achieve a greater sample efficiency than brute force methods. Fengyi Tang, Cao Xiao, Fei Wang 0001, Li-Wei H. Lehman |
KDD | 5 |
| 2018 | Improving Sepsis Treatment Strategies by Combining Deep and Kernel-Based Reinforcement Learning
Xuefeng Peng, David Wihl, Omer Gottesman, Matthieu Komorowski, Li-Wei H. Lehman, Andrew Slavin Ross, A. Aldo Faisal, Finale Doshi-Velez |
AMIA | 6 |
| 2018 | A Model-Based Machine Learning Approach to Probing Autonomic Regulation From Nonstationary Vital-Sign Time SeriesabstractPhysiological variables, such as heart rate (HR), blood pressure (BP) and respiration (RESP), are tightly regulated and coupled under healthy conditions, and a break-down in the coupling has been associated with aging and disease. We present an approach that incorporates physiological modeling within a switching linear dynamical systems (SLDS) framework to assess the various functional components of the autonomic regulation through transfer function analysis of nonstationary multivariate time series of vital signs. We validate our proposed SLDS-based transfer function analysis technique in automatically capturing 1) changes in baroreflex gain due to postural changes in a tilt-table study including ten subjects, and 2) the effect of aging on the autonomic control using HR/RESP recordings from 40 healthy adults. Next, using HR/BP time series of more than 450 adult ICU patients, we show that our technique can be used to reveal coupling changes associated with severe sepsis (AUC = 0.74, sensitivity = 0.74, specificity = 0.60). Our findings indicate that reduced HR/BP coupling is significantly associated with severe sepsis even after adjusting for clinical interventions (P 0.001). These results demonstrate the utility of our approach in phenotyping complex vital-sign dynamics, and in providing mechanistic hypotheses in terms of break-down of autoregulatory systems under healthy and disease conditions. Li-Wei H. Lehman, Roger G. Mark, Shamim Nemati |
IEEE J. Biomed. Health Informatics | 1 |
| 2015 | A Physiological Time Series Dynamics-Based Approach to Patient Monitoring and Outcome PredictionabstractCardiovascular variables such as heart rate (HR) and blood pressure (BP) are regulated by an underlying control system, and therefore, the time series of these vital signs exhibit rich dynamical patterns of interaction in response to external perturbations (e.g., drug administration), as well as pathological states (e.g., onset of sepsis and hypotension). A question of interest is whether "similar" dynamical patterns can be identified across a heterogeneous patient cohort, and be used for prognosis of patients' health and progress. In this paper, we used a switching vector autoregressive framework to systematically learn and identify a collection of vital sign time series dynamics, which are possibly recurrent within the same patient and may be shared across the entire cohort. We show that these dynamical behaviors can be used to characterize the physiological "state" of a patient. We validate our technique using simulated time series of the cardiovascular system, and human recordings of HR and BP time series from an orthostatic stress study with known postural states. Using the HR and BP dynamics of an intensive care unit (ICU) cohort of over 450 patients from the MIMIC II database, we demonstrate that the discovered cardiovascular dynamics are significantly associated with hospital mortality (dynamic modes 3 and 9, p=0.001, p=0.006 from logistic regression after adjusting for the APACHE scores). Combining the dynamics of BP time series and SAPS-I or APACHE-III provided a more accurate assessment of patient survival/mortality in the hospital than using SAPS-I and APACHE-III alone (p=0.005 and p=0.045). Our results suggest that the discovered dynamics of vital sign time series may contain additional prognostic value beyond that of the baseline acuity measures, and can potentially be used as an independent predictor of outcomes in the ICU. Li-Wei H. Lehman, Ryan P. Adams, Louis Mayaud, George B. Moody, Atul Malhotra, Roger G. Mark, Shamim Nemati |
IEEE J. Biomed. Health Informatics | 1 |
| 2012 | Risk Stratification of ICU Patients Using Topic Models Inferred from Unstructured Progress Notes
Li-Wei H. Lehman, Mohammed Saeed 0001, William J. Long, Joon Lee, Roger G. Mark |
AMIA | 1 |
| 2004 | PCoord: Network Position Estimation Using Peer-to-Peer MeasurementsabstractSeveral recently emerged Internet services make use of application-level or overlay networks. Examples of such services include overlay multicast, structured peer-to-peer lookup services, and peer-to-peer file sharing. Many of these services could benefit from enabling participating end hosts to estimate their relative network locations within the overlay. We present PCoord, a peer-to-peer network coordinate system for overlay topology discovery and distance prediction. The goal of PCoord is to allow participating peer nodes in an overlay network to collaboratively construct an accurate geometric model of the overlay network topology in a completely decentralized peer-to-peer fashion. We evaluate the PCoord approach through extensive simulations using both real network measurements and simulated topologies. Our results indicate that the constructed geometric model can give accurate pair-wise distance prediction and nearest neighbor discovery. In particular, using a simulated overlay network consisting of over 3,400 peer nodes, our results indicate that over 90% of the peers can predict their closest peers by probing only a small fraction of the global peer population. Li-Wei H. Lehman, Steven Lerman |
NCA | 1 |
| 1998 | Active Reliable MulticastabstractThis paper presents a novel loss recovery scheme, active reliable multicast (ARM), for large scale reliable multicast. ARM is "active" in that routers in the multicast tree play an active role in loss recovery. Additionally, ARM utilizes soft-state storage within the network to improve performance and scalability. In the upstream direction, routers suppress duplicate NACKs from multiple receivers to control the implosion problem. By suppressing duplicate NACKs, ARM also lessens the traffic that propagates back through the network, In the downstream direction, routers limit the delivery of repair packets to receivers experiencing loss, thereby reducing network bandwidth consumption. Finally, to reduce wide-area recovery latency and to distribute the retransmission load, routers cache multicast data on a "best effort" basis. ARM is flexible and robust in that it does not require all nodes to be active, nor does it require any specific router or receiver to perform loss recovery. Analysis and simulation results show that ARM yields significant benefits even when less than half the routers within the multicast tree can perform ARM processing. Li-Wei H. Lehman, Stephen J. Garland, David L. Tennenhouse |
INFOCOM | 1 |