VLDB 2026 Research / reviewers in the wild / expert
Yizhe Xu
dblp:204/4745
· DBLP profile ↗
8ranked-venue papers
5as first author
6since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Deep learning architectures and training · 87% Transfer learning and domain adaptation · 13% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 100% |
Topics — the 4 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training
foundation model |
0.8 | 1 | 2024 | MOTOR: A Time-to-Event Foundation Model For Structured Medical Records · ICLR 2024 |
Machine learning › Deep learning architectures and training › foundation model
medical foundation model |
0.8 | 1 | 2024 | MOTOR: A Time-to-Event Foundation Model For Structured Medical Records · ICLR 2024 |
Medical and health informatics
electronic health records |
0.8 | 1 | 2024 | MOTOR: A Time-to-Event Foundation Model For Structured Medical Records · ICLR 2024 |
Medical and health informatics › clinical prediction
time-to-event prediction |
0.8 | 1 | 2024 | MOTOR: A Time-to-Event Foundation Model For Structured Medical Records · ICLR 2024 |
Methods — techniques the papers use, named apart from their topics
survival analysis · 1.5self-supervised pretraining · 1.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | MOTOR: A Time-to-Event Foundation Model For Structured Medical RecordsabstractWe present a self-supervised, time-to-event (TTE) foundation model called MOTOR (Many Outcome Time Oriented Representations) which is pretrained on timestamped sequences of events in electronic health records (EHR) and health insurance claims. TTE models are used for estimating the probability distribution of the time until a specific event occurs, which is an important task in medical settings. TTE models provide many advantages over classification using fixed time horizons, including naturally handling censored observations, but are challenging to train with limited labeled data. MOTOR addresses this challenge by pretraining on up to 55M patient records (9B clinical events). We evaluate MOTOR's transfer learning performance on 19 tasks, across 3 patient databases (a private EHR system, MIMIC-IV, and Merative claims data). Task-specific models adapted from MOTOR improve time-dependent C statistics by 4.6\% over state-of-the-art, improve label efficiency by up to 95\%, and are more robust to temporal distributional shifts. We further evaluate cross-site portability by adapting our MOTOR foundation model for six prediction tasks on the MIMIC-IV dataset, where it outperforms all baselines. MOTOR is the first foundation model for medical TTE predictions and we release a 143M parameter pretrained model for research use at https://huggingface.co/StanfordShahLab/motor-t-base. Ethan Steinberg, Jason Alan Fries, Yizhe Xu, Nigam H. Shah |
ICLR | 3 |
| 2023 | Hovering Control of Flapping Wings in Tandem with Multi-RotorsabstractThis work briefly covers our efforts to stabilize the flight dynamics of Northeatern's tailless bat-inspired micro aerial vehicle, Aerobat. Flapping robots are not new. A plethora of examples is mainly dominated by insect-style design paradigms that are passively stable. However, Aerobat, in addition for being tailless, possesses morphing wings that add to the inherent complexity of flight control. The robot can dynamically adjust its wing platform configurations during gaitcycles, increasing its efficiency and agility. We employ a guard design with manifold small thrusters to stabilize Aerobat's position and orientation in hovering, a flapping system in tandem with a multi-rotor. For flight control purposes, we take an approach based on assuming the guard cannot observe Aeroat's states. Then, we propose an observer to estimate the unknown states of the guard which are then used for closed-loop hovering control of the Guard-Aerobat platform. Aniket Dhole, Bibek Gupta, Adarsh Salagame, Xuejian Niu, Yizhe Xu, Kaushik Venkatesh Krishnamurthy, Paul Ghanem, Ioannis Mandralis, Eric Sihite, Alireza Ramezani |
IROS | 5 |
| 2023 | Clinical utility gains from incorporating comorbidity and geographic location information into risk estimation equations for atherosclerotic cardiovascular diseaseabstractOBJECTIVE: There are over 363 customized risk models of the American College of Cardiology and the American Heart Association (ACC/AHA) pooled cohort equations (PCE) in the literature, but their gains in clinical utility are rarely evaluated. We build new risk models for patients with specific comorbidities and geographic locations and evaluate whether performance improvements translate to gains in clinical utility. MATERIALS AND METHODS: We retrain a baseline PCE using the ACC/AHA PCE variables and revise it to incorporate subject-level information of geographic location and 2 comorbidity conditions. We apply fixed effects, random effects, and extreme gradient boosting (XGB) models to handle the correlation and heterogeneity induced by locations. Models are trained using 2 464 522 claims records from Optum©'s Clinformatics® Data Mart and validated in the hold-out set (N = 1 056 224). We evaluate models' performance overall and across subgroups defined by the presence or absence of chronic kidney disease (CKD) or rheumatoid arthritis (RA) and geographic locations. We evaluate models' expected utility using net benefit and models' statistical properties using several discrimination and calibration metrics. RESULTS: The revised fixed effects and XGB models yielded improved discrimination, compared to baseline PCE, overall and in all comorbidity subgroups. XGB improved calibration for the subgroups with CKD or RA. However, the gains in net benefit are negligible, especially under low exchange rates. CONCLUSIONS: Common approaches to revising risk calculators incorporating extra information or applying flexible models may enhance statistical performance; however, such improvement does not necessarily translate to higher clinical utility. Thus, we recommend future works to quantify the consequences of using risk calculators to guide clinical decisions. Yizhe Xu, Agata Foryciarz, Ethan Steinberg, Nigam H. Shah |
J. Am. Medical Informatics Assoc. | 1 |
| 2023 | Principled estimation and evaluation of treatment effect heterogeneity: A case study application to dabigatran for patients with atrial fibrillationabstractOBJECTIVE: To apply the latest guidance for estimating and evaluating heterogeneous treatment effects (HTEs) in an end-to-end case study of the Long-term Anticoagulation Therapy (RE-LY) trial, and summarize the main takeaways from applying state-of-the-art metalearners and novel evaluation metrics in-depth to inform their applications to personalized care in biomedical research. METHODS: Based on the characteristics of the RE-LY data, we selected four metalearners (S-learner with Lasso, X-learner with Lasso, R-learner with random survival forest and Lasso, and causal survival forest) to estimate the HTEs of dabigatran. For the outcomes of (1) stroke or systemic embolism and (2) major bleeding, we compared dabigatran 150 mg, dabigatran 110 mg, and warfarin. We assessed the overestimation of treatment heterogeneity by the metalearners via a global null analysis and their discrimination and calibration ability using two novel metrics: rank-weighted average treatment effects (RATE) and estimated calibration error for treatment heterogeneity. Finally, we visualized the relationships between estimated treatment effects and baseline covariates using partial dependence plots. RESULTS: The RATE metric suggested that either the applied metalearners had poor performance of estimating HTEs or there was no treatment heterogeneity for either the stroke/SE or major bleeding outcome of any treatment comparison. Partial dependence plots revealed that several covariates had consistent relationships with the treatment effects estimated by multiple metalearners. The applied metalearners showed differential performance across outcomes and treatment comparisons, and the X- and R-learners yielded smaller calibration errors than the others. CONCLUSIONS: HTE estimation is difficult, and a principled estimation and evaluation process is necessary to provide reliable evidence and prevent false discoveries. We have demonstrated how to choose appropriate metalearners based on specific data properties, applied them using the off-the-shelf implementation tool survlearners, and evaluated their performance using recently defined formal metrics. We suggest that clinical implications should be drawn based on the common trends across the applied metalearners. Yizhe Xu, Katelyn K. Bechler, Alison Callahan, Nigam H. Shah |
J. Biomed. Informatics | 1 |
| 2023 | EI-HCR: An Efficient End-to-End Hybrid Consistency Regularization Algorithm for Semisupervised Remote Sensing Image SegmentationabstractRecently, remote sensing image (RSI) semantic segmentation technology has advanced greatly, with the fully supervised process achieving particularly strong performance. However, the technology depends heavily on dataset labels, leading to high annotation costs. To alleviate this problem, we propose a novel efficient end-to-end hybrid consistency regularization algorithm (EI-HCR) for the semi-supervised semantic segmentation of RSI, wherein only a few labeled images and a large number of unlabeled images are effectively used. First, we devise data perturbation (DP) consistency regularization (CR), which includes a data mix-up method to combine unlabeled and labeled images. Then, we employ teacher and student networks to conduct model perturbation (MP) CR. Both segmentation results are regarded as pseudo-labels for each other. In the end, the semi-supervised loss is composed of DP and MP consistency loss, and supervises network training along with the fully supervised loss. More importantly, we first combine the characteristics of knowledge distillation to make the student network more lightweight, efficiently reducing the model inference time. Experimental results demonstrate the effectiveness of EI-HCR on the ISPRS Vaihingen and Massachusetts Buildings datasets. With only 5% of the labeled images, EI-HCR can achieve the same accuracy as the fully supervised training with 50% of the labeled images, and the number of student model parameters is only 9.64 M, indicating the method’s great advantages over other algorithms. Yizhe Xu, Liangliang Yan, Jie Jiang 0005 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Calibration Error for Heterogeneous Treatment EffectsabstractRecently, many researchers have advanced data-driven methods for modeling heterogeneous treatment effects (HTEs). Even still, estimation of HTEs is a difficult task–these methods frequently over- or under-estimate the treatment effects, leading to poor calibration of the resulting models. However, while many methods exist for evaluating the calibration of prediction and classification models, formal approaches to assess the calibration of HTE models are limited to the calibration slope. In this paper, we define an analogue of the (L2) expected calibration error for HTEs, and propose a robust estimator. Our approach is motivated by doubly robust treatment effect estimators, making it unbiased, and resilient to confounding, overfitting, and high-dimensionality issues. Furthermore, our method is straightforward to adapt to many structures under which treatment effects can be identified, including randomized trials, observational studies, and survival analysis. We illustrate how to use our proposed metric to evaluate the calibration of learned HTE models through the application to the CRITEO-UPLIFT Trial. Yizhe Xu, Steve Yadlowsky |
AISTATS | 1 |
| 2014 | Distributed and autonomous control of the FREEDM system: A power electronics based distribution systemabstractA truly distributed and autonomous control strategy is proposed for the FREEDM System-a power electronics based distribution grid. The proposed control strategy requires no communication between and among all grid assets. By utilizing available local quantities (V, f) as a way to communicate among all connected devices, the strategy contains a novel dual droop control loops between the medium voltage feeder and dispatchable resources such as energy sources, storages. The proposed control is put forward to achieve intelligent power management of the FREEDM system under all modes of operation. This is a pre-requisite to use FREEDM system as a transformative platform for plug-and-play of energy (aka Energy Internet). Simulation results verify the autonomous nature of the proposed control. Alex Q. Huang, Yizhe Xu, Fei Wang 0045, Wensong Yu |
IECON | 3 |
| 2014 | Five-level bidirectional converter for renewable power generation systemabstractA novel five-level DC/AC bidirectional converter is developed and applied for interfacing renewable energy into the grid. The converter has reduced switching losses, voltage stress, harmonic distortion and electromagnetic interference caused by switching operation of power devices. The bidirectional operation feature allows it to interface with parallel dc-dc PV optimizers that requires a start-up dc link voltage of 200-250Vdc. Two DC decoupled capacitors, a high-frequency three-level converter, a grid-frequency unfolding bridge and a filter compose the proposed converter. The high-frequency converter offers a three level positive voltage output, while the unfolding bridge is responsible for switching the flow direction every half grid cycle. With the proposed control strategy, the converter operates via two modes, i.e. the inverter and rectifier mode. In inverter mode, the output current of the converter generates a sinusoidal current in phase with the grid voltage, while in rectifier mode, grid voltage acts as the source and dc link supplies the PV optimizer with a threshold voltage to avoid false tripping of the optimizer. A hardware prototype is further developed to verify the performance of the topology and control strategy. Yizhe Xu, Yen-mo Chen, Alex Q. Huang |
IECON | 1 |