Yizhe Xu

dblp:204/4745 · DBLP profile ↗
← Back
8ranked-venue papers
5as first author
6since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Deep learning architectures and training · 87% Transfer learning and domain adaptation · 13%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Medical and health informatics · 100%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
foundation model
0.812024
MOTOR: A Time-to-Event Foundation Model For Structured Medical Records · ICLR 2024
Machine learning › Deep learning architectures and training › foundation model
medical foundation model
0.812024
MOTOR: A Time-to-Event Foundation Model For Structured Medical Records · ICLR 2024
Medical and health informatics
electronic health records
0.812024
MOTOR: A Time-to-Event Foundation Model For Structured Medical Records · ICLR 2024
Medical and health informatics › clinical prediction
time-to-event prediction
0.812024
MOTOR: A Time-to-Event Foundation Model For Structured Medical Records · ICLR 2024

Methods — techniques the papers use, named apart from their topics

survival analysis · 1.5self-supervised pretraining · 1.5
YearPublicationVenuePosition
2024 MOTOR: A Time-to-Event Foundation Model For Structured Medical Records
abstract
We present a self-supervised, time-to-event (TTE) foundation model called MOTOR (Many Outcome Time Oriented Representations) which is pretrained on timestamped sequences of events in electronic health records (EHR) and health insurance claims. TTE models are used for estimating the probability distribution of the time until a specific event occurs, which is an important task in medical settings. TTE models provide many advantages over classification using fixed time horizons, including naturally handling censored observations, but are challenging to train with limited labeled data. MOTOR addresses this challenge by pretraining on up to 55M patient records (9B clinical events). We evaluate MOTOR's transfer learning performance on 19 tasks, across 3 patient databases (a private EHR system, MIMIC-IV, and Merative claims data). Task-specific models adapted from MOTOR improve time-dependent C statistics by 4.6\% over state-of-the-art, improve label efficiency by up to 95\%, and are more robust to temporal distributional shifts. We further evaluate cross-site portability by adapting our MOTOR foundation model for six prediction tasks on the MIMIC-IV dataset, where it outperforms all baselines. MOTOR is the first foundation model for medical TTE predictions and we release a 143M parameter pretrained model for research use at https://huggingface.co/StanfordShahLab/motor-t-base.
Ethan Steinberg, Jason Alan Fries, Yizhe Xu, Nigam H. Shah
ICLR3
2023 Hovering Control of Flapping Wings in Tandem with Multi-Rotors
abstract
This work briefly covers our efforts to stabilize the flight dynamics of Northeatern's tailless bat-inspired micro aerial vehicle, Aerobat. Flapping robots are not new. A plethora of examples is mainly dominated by insect-style design paradigms that are passively stable. However, Aerobat, in addition for being tailless, possesses morphing wings that add to the inherent complexity of flight control. The robot can dynamically adjust its wing platform configurations during gaitcycles, increasing its efficiency and agility. We employ a guard design with manifold small thrusters to stabilize Aerobat's position and orientation in hovering, a flapping system in tandem with a multi-rotor. For flight control purposes, we take an approach based on assuming the guard cannot observe Aeroat's states. Then, we propose an observer to estimate the unknown states of the guard which are then used for closed-loop hovering control of the Guard-Aerobat platform.
Aniket Dhole, Bibek Gupta, Adarsh Salagame, Xuejian Niu, Yizhe Xu, Kaushik Venkatesh Krishnamurthy, Paul Ghanem, Ioannis Mandralis, Eric Sihite, Alireza Ramezani
IROS5
2023 Clinical utility gains from incorporating comorbidity and geographic location information into risk estimation equations for atherosclerotic cardiovascular disease
abstract
OBJECTIVE: There are over 363 customized risk models of the American College of Cardiology and the American Heart Association (ACC/AHA) pooled cohort equations (PCE) in the literature, but their gains in clinical utility are rarely evaluated. We build new risk models for patients with specific comorbidities and geographic locations and evaluate whether performance improvements translate to gains in clinical utility. MATERIALS AND METHODS: We retrain a baseline PCE using the ACC/AHA PCE variables and revise it to incorporate subject-level information of geographic location and 2 comorbidity conditions. We apply fixed effects, random effects, and extreme gradient boosting (XGB) models to handle the correlation and heterogeneity induced by locations. Models are trained using 2 464 522 claims records from Optum©'s Clinformatics® Data Mart and validated in the hold-out set (N = 1 056 224). We evaluate models' performance overall and across subgroups defined by the presence or absence of chronic kidney disease (CKD) or rheumatoid arthritis (RA) and geographic locations. We evaluate models' expected utility using net benefit and models' statistical properties using several discrimination and calibration metrics. RESULTS: The revised fixed effects and XGB models yielded improved discrimination, compared to baseline PCE, overall and in all comorbidity subgroups. XGB improved calibration for the subgroups with CKD or RA. However, the gains in net benefit are negligible, especially under low exchange rates. CONCLUSIONS: Common approaches to revising risk calculators incorporating extra information or applying flexible models may enhance statistical performance; however, such improvement does not necessarily translate to higher clinical utility. Thus, we recommend future works to quantify the consequences of using risk calculators to guide clinical decisions.
Yizhe Xu, Agata Foryciarz, Ethan Steinberg, Nigam H. Shah
J. Am. Medical Informatics Assoc.1
2023 Principled estimation and evaluation of treatment effect heterogeneity: A case study application to dabigatran for patients with atrial fibrillation
abstract
OBJECTIVE: To apply the latest guidance for estimating and evaluating heterogeneous treatment effects (HTEs) in an end-to-end case study of the Long-term Anticoagulation Therapy (RE-LY) trial, and summarize the main takeaways from applying state-of-the-art metalearners and novel evaluation metrics in-depth to inform their applications to personalized care in biomedical research. METHODS: Based on the characteristics of the RE-LY data, we selected four metalearners (S-learner with Lasso, X-learner with Lasso, R-learner with random survival forest and Lasso, and causal survival forest) to estimate the HTEs of dabigatran. For the outcomes of (1) stroke or systemic embolism and (2) major bleeding, we compared dabigatran 150 mg, dabigatran 110 mg, and warfarin. We assessed the overestimation of treatment heterogeneity by the metalearners via a global null analysis and their discrimination and calibration ability using two novel metrics: rank-weighted average treatment effects (RATE) and estimated calibration error for treatment heterogeneity. Finally, we visualized the relationships between estimated treatment effects and baseline covariates using partial dependence plots. RESULTS: The RATE metric suggested that either the applied metalearners had poor performance of estimating HTEs or there was no treatment heterogeneity for either the stroke/SE or major bleeding outcome of any treatment comparison. Partial dependence plots revealed that several covariates had consistent relationships with the treatment effects estimated by multiple metalearners. The applied metalearners showed differential performance across outcomes and treatment comparisons, and the X- and R-learners yielded smaller calibration errors than the others. CONCLUSIONS: HTE estimation is difficult, and a principled estimation and evaluation process is necessary to provide reliable evidence and prevent false discoveries. We have demonstrated how to choose appropriate metalearners based on specific data properties, applied them using the off-the-shelf implementation tool survlearners, and evaluated their performance using recently defined formal metrics. We suggest that clinical implications should be drawn based on the common trends across the applied metalearners.
Yizhe Xu, Katelyn K. Bechler, Alison Callahan, Nigam H. Shah
J. Biomed. Informatics1
2023 EI-HCR: An Efficient End-to-End Hybrid Consistency Regularization Algorithm for Semisupervised Remote Sensing Image Segmentation
abstract
Recently, remote sensing image (RSI) semantic segmentation technology has advanced greatly, with the fully supervised process achieving particularly strong performance. However, the technology depends heavily on dataset labels, leading to high annotation costs. To alleviate this problem, we propose a novel efficient end-to-end hybrid consistency regularization algorithm (EI-HCR) for the semi-supervised semantic segmentation of RSI, wherein only a few labeled images and a large number of unlabeled images are effectively used. First, we devise data perturbation (DP) consistency regularization (CR), which includes a data mix-up method to combine unlabeled and labeled images. Then, we employ teacher and student networks to conduct model perturbation (MP) CR. Both segmentation results are regarded as pseudo-labels for each other. In the end, the semi-supervised loss is composed of DP and MP consistency loss, and supervises network training along with the fully supervised loss. More importantly, we first combine the characteristics of knowledge distillation to make the student network more lightweight, efficiently reducing the model inference time. Experimental results demonstrate the effectiveness of EI-HCR on the ISPRS Vaihingen and Massachusetts Buildings datasets. With only 5% of the labeled images, EI-HCR can achieve the same accuracy as the fully supervised training with 50% of the labeled images, and the number of student model parameters is only 9.64 M, indicating the method’s great advantages over other algorithms.
Yizhe Xu, Liangliang Yan, Jie Jiang 0005
IEEE Trans. Geosci. Remote. Sens.1
2022 Calibration Error for Heterogeneous Treatment Effects
abstract
Recently, many researchers have advanced data-driven methods for modeling heterogeneous treatment effects (HTEs). Even still, estimation of HTEs is a difficult task–these methods frequently over- or under-estimate the treatment effects, leading to poor calibration of the resulting models. However, while many methods exist for evaluating the calibration of prediction and classification models, formal approaches to assess the calibration of HTE models are limited to the calibration slope. In this paper, we define an analogue of the (L2) expected calibration error for HTEs, and propose a robust estimator. Our approach is motivated by doubly robust treatment effect estimators, making it unbiased, and resilient to confounding, overfitting, and high-dimensionality issues. Furthermore, our method is straightforward to adapt to many structures under which treatment effects can be identified, including randomized trials, observational studies, and survival analysis. We illustrate how to use our proposed metric to evaluate the calibration of learned HTE models through the application to the CRITEO-UPLIFT Trial.
Yizhe Xu, Steve Yadlowsky
AISTATS1
2014 Distributed and autonomous control of the FREEDM system: A power electronics based distribution system
abstract
A truly distributed and autonomous control strategy is proposed for the FREEDM System-a power electronics based distribution grid. The proposed control strategy requires no communication between and among all grid assets. By utilizing available local quantities (V, f) as a way to communicate among all connected devices, the strategy contains a novel dual droop control loops between the medium voltage feeder and dispatchable resources such as energy sources, storages. The proposed control is put forward to achieve intelligent power management of the FREEDM system under all modes of operation. This is a pre-requisite to use FREEDM system as a transformative platform for plug-and-play of energy (aka Energy Internet). Simulation results verify the autonomous nature of the proposed control.
Alex Q. Huang, Yizhe Xu, Fei Wang 0045, Wensong Yu
IECON3
2014 Five-level bidirectional converter for renewable power generation system
abstract
A novel five-level DC/AC bidirectional converter is developed and applied for interfacing renewable energy into the grid. The converter has reduced switching losses, voltage stress, harmonic distortion and electromagnetic interference caused by switching operation of power devices. The bidirectional operation feature allows it to interface with parallel dc-dc PV optimizers that requires a start-up dc link voltage of 200-250Vdc. Two DC decoupled capacitors, a high-frequency three-level converter, a grid-frequency unfolding bridge and a filter compose the proposed converter. The high-frequency converter offers a three level positive voltage output, while the unfolding bridge is responsible for switching the flow direction every half grid cycle. With the proposed control strategy, the converter operates via two modes, i.e. the inverter and rectifier mode. In inverter mode, the output current of the converter generates a sinusoidal current in phase with the grid voltage, while in rectifier mode, grid voltage acts as the source and dc link supplies the PV optimizer with a threshold voltage to avoid false tripping of the optimizer. A hardware prototype is further developed to verify the performance of the topology and control strategy.
Yizhe Xu, Yen-mo Chen, Alex Q. Huang
IECON1