Yanbo Xu

dblp:25/6978 · DBLP profile ↗
← Back
24ranked-venue papers
11as first author
7since 2021 · last 2025
0000-0003-4129-7376ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 6 first-authorGraphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021Computer networks · 2Human-computer interaction and ubiquitous computing · 2Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Efficient and distributed learning · 21% Probabilistic and Bayesian machine learning · 18% Generative modeling · 16%
Interdisciplinary, comprehensive, and emerging computing
3 papers
Medical and health informatics · 54% Computational social science and digital humanities · 46%
Computer graphics and multimedia
2 papers
Visual content generation and editing · 54% Rendering · 46%
Network and information security
1 paper
Cyber-physical and IoT security · 77% Systems and software security · 23%

Topics — the 22 heaviest of 24, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Medical and health informatics
clinical decision support
0.822020
HOLMES: Health OnLine Model Ensemble Serving for Deep Learning Models in Intensive Care Units · KDD 2020
RAIM: Recurrent Attentive and Intensive Model of Multimodal Patient Monitoring Data · KDD 2018
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes › point process › temporal point process
hawkes process
0.712023
SMURF-THP: Score Matching-based UnceRtainty quantiFication for Transformer Hawkes Process · ICML 2023
Machine learning › Probabilistic and Bayesian machine learning › stochastic processes
point process
0.712023
SMURF-THP: Score Matching-based UnceRtainty quantiFication for Transformer Hawkes Process · ICML 2023
Machine learning › Generative modeling
score matching
0.712023
SMURF-THP: Score Matching-based UnceRtainty quantiFication for Transformer Hawkes Process · ICML 2023
Machine learning › Trustworthy machine learning
uncertainty estimation
0.712023
SMURF-THP: Score Matching-based UnceRtainty quantiFication for Transformer Hawkes Process · ICML 2023
Rendering
neural radiance fields
0.712023
FaceDNeRF: Semantics-Driven Face Reconstruction, Prompt Editing and Relighting with Diffusion Models · NeurIPS 2023
Machine learning › Efficient and distributed learning › inference efficiency
cost-aware inference
0.612022
UnfoldML: Cost-Aware and Uncertainty-Based Dynamic 2D Prediction for Multi-Stage Classification · NeurIPS 2022
Machine learning › Time series and sequential data › time series modeling
dynamic prediction
0.612022
UnfoldML: Cost-Aware and Uncertainty-Based Dynamic 2D Prediction for Multi-Stage Classification · NeurIPS 2022
Machine learning › Generative modeling
generative adversarial network
0.612022
TransEditor: Transformer-Based Dual-Space GAN for Highly Controllable Facial Editing · CVPR 2022
Machine learning › Efficient and distributed learning › adaptive computation
model cascading
0.612022
UnfoldML: Cost-Aware and Uncertainty-Based Dynamic 2D Prediction for Multi-Stage Classification · NeurIPS 2022
Computer vision › 3D vision › 3d object detection › image-based 3d object detection
monocular 3d object detection
0.612022
Semi-supervised Monocular 3D Object Detection by Multi-view Consistency · ECCV (8) 2022
Machine learning › Kernel, tree and ensemble methods › classifier combination
multi-stage classification
0.612022
UnfoldML: Cost-Aware and Uncertainty-Based Dynamic 2D Prediction for Multi-Stage Classification · NeurIPS 2022
Computer vision › 3D vision › 3d reconstruction
multi-view reconstruction
0.612022
Semi-supervised Monocular 3D Object Detection by Multi-view Consistency · ECCV (8) 2022
Visual content generation and editing
face editing
0.612022
TransEditor: Transformer-Based Dual-Space GAN for Highly Controllable Facial Editing · CVPR 2022
Computational social science and digital humanities
causal inference
0.512021
Split-Treatment Analysis to Rank Heterogeneous Causal Effects for Prospective Interventions · WSDM 2021
Computational social science and digital humanities › causal inference
heterogeneous treatment effect
0.512021
Split-Treatment Analysis to Rank Heterogeneous Causal Effects for Prospective Interventions · WSDM 2021
Machine learning › Efficient and distributed learning
inference serving
0.412020
HOLMES: Health OnLine Model Ensemble Serving for Deep Learning Models in Intensive Care Units · KDD 2020
Machine learning › Deep learning architectures and training
recurrent neural network
0.312018
RAIM: Recurrent Attentive and Intensive Model of Multimodal Patient Monitoring Data · KDD 2018
Medical and health informatics › clinical monitoring
patient monitoring
0.312018
RAIM: Recurrent Attentive and Intensive Model of Multimodal Patient Monitoring Data · KDD 2018
Systems and software security › vulnerability discovery
static analysis
0.312025
FirmProj: Detecting Firmware Leakage in IoT Update Processes via Companion App Analysis · ASE 2025
Visual content generation and editing
3D GAN inversion
0.212023
FaceDNeRF: Semantics-Driven Face Reconstruction, Prompt Editing and Relighting with Diffusion Models · NeurIPS 2023
Machine learning › Kernel, tree and ensemble methods
model ensemble
0.112020
HOLMES: Health OnLine Model Ensemble Serving for Deep Learning Models in Intensive Care Units · KDD 2020

Methods — techniques the papers use, named apart from their topics

transformer · 1.1inversion strategy · 1.1dual-space editing · 1.1static analysis · 0.9large language model · 0.9ensemble selection · 0.9score matching · 0.7multimodal pretraining · 0.7latent diffusion model · 0.7illumination and identity preserving loss · 0.7confidence interval computation · 0.7uncertainty estimation · 0.6semi-supervised learning · 0.6multi-view consistency · 0.6cascade of classifiers · 0.6split-treatment analysis · 0.5sensitivity analysis · 0.5latency-aware scheduling · 0.4
YearPublicationVenuePosition
2025 FirmProj: Detecting Firmware Leakage in IoT Update Processes via Companion App Analysis
abstract
The rapid growth of the Internet of Things (IoT) has led to the widespread use of companion apps for device management. However, these apps expose a critical vulnerability in the IoT ecosystem: insufficient verification procedures during device firmware updates (DFU), often resulting in firmware leakage. Once leaked, the firmware reveals sensitive design details, creating a straightforward path for attackers to reverse-engineer devices. To address this issue, we designed an automated analysis tool called FirmProj. It systematically evaluates firmware leakage risks by examining IoT companion apps. FirmProj combines advanced static analysis techniques with large language models to identify DFU modules, extract firmware files, and detect security vulnerabilities. In a large-scale study involving 10,047 IoT companion apps, FirmProj successfully retrieved 3,434 firmware files, uncovering severe flaws in DFU implementations that can lead to firmware leakage. These findings resulted in the assignment of 35 CVE IDs. Our results highlight the urgent need to strengthen firmware protection mechanisms throughout the IoT ecosystem.
Wenzhi Li, Jialong Guo, Jiongyi Chen, Yujie Xing, Yanbo Xu, Shishuai Yang, Wenrui Diao
ASE6
2023 SMURF-THP: Score Matching-based UnceRtainty quantiFication for Transformer Hawkes Process
abstract
Transformer Hawkes process models have shown to be successful in modeling event sequence data. However, most of the existing training methods rely on maximizing the likelihood of event sequences, which involves calculating some intractable integral. Moreover, the existing methods fail to provide uncertainty quantification for model predictions, e.g., confidence interval for the predicted event’s arrival time. To address these issues, we propose SMURF-THP, a score-based method for learning Transformer Hawkes process and quantifying prediction uncertainty. Specifically, SMURF-THP learns the score function of the event’s arrival time based on a score-matching objective that avoids the intractable computation. With such a learnt score function, we can sample arrival time of events from the predictive distribution. This naturally allows for the quantification of uncertainty by computing confidence intervals over the generated samples. We conduct extensive experiments in both event type prediction and uncertainty quantification on time of arrival. In all the experiments, SMURF-THP outperforms existing likelihood-based methods in confidence calibration while exhibiting comparable prediction accuracy.
Zichong Li, Yanbo Xu, Simiao Zuo, Haoming Jiang, Chao Zhang 0014, Tuo Zhao, Hongyuan Zha
ICML2
2023 FaceDNeRF: Semantics-Driven Face Reconstruction, Prompt Editing and Relighting with Diffusion Models
abstract
The ability to create high-quality 3D faces from a single image has become increasingly important with wide applications in video conferencing, AR/VR, and advanced video editing in movie industries. In this paper, we propose Face Diffusion NeRF (FaceDNeRF), a new generative method to reconstruct high-quality Face NeRFs from single images, complete with semantic editing and relighting capabilities. FaceDNeRF utilizes high-resolution 3D GAN inversion and expertly trained 2D latent-diffusion model, allowing users to manipulate and construct Face NeRFs in zero-shot learning without the need for explicit 3D data. With carefully designed illumination and identity preserving loss, as well as multi-modal pre-training, FaceDNeRF offers users unparalleled control over the editing process enabling them to create and edit face NeRFs using just single-view images, text prompts, and explicit target lighting. The advanced features of FaceDNeRF have been designed to produce more impressive results than existing 2D editing approaches that rely on 2D segmentation maps for editable attributes. Experiments show that our FaceDNeRF achieves exceptionally realistic results and unprecedented flexibility in editing compared with state-of-the-art 3D face reconstruction and editing methods. Our code will be available at https://github.com/BillyXYB/FaceDNeRF.
Hao Zhang 0106, Tianyuan Dai, Yanbo Xu, Yu-Wing Tai, Chi-Keung Tang
NeurIPS3
2022 TransEditor: Transformer-Based Dual-Space GAN for Highly Controllable Facial Editing
abstract
Recent advances like StyleGAN have promoted the growth of controllable facial editing. To address its core challenge of attribute decoupling in a single latent space, attempts have been made to adopt dual-space GAN for better disentanglement of style and content representations. Nonetheless, these methods are still incompetent to obtain plausible editing results with high controllability, especially for complicated attributes. In this study, we highlight the importance of interaction in a dual-space GAN for more controllable editing. We propose TransEditor, a novel Transformer-based framework to enhance such interaction. Besides, we develop a new dual-space editing and inversion strategy to provide additional editing flexibility. Extensive experiments demonstrate the superiority of the proposed framework in image quality and editing capability, suggesting the effectiveness of TransEditor for highly controllable facial editing. Code and models are publicly available at https://github.com/BillyXYB/TransEditor.
Yanbo Xu, Yueqin Yin, Liming Jiang 0001, Qianyi Wu, Chengyao Zheng, Chen Change Loy, Bo Dai 0002, Wayne Wu
CVPR1
2022 Semi-supervised Monocular 3D Object Detection by Multi-view Consistency
Qing Lian, Yanbo Xu, Weilong Yao, Ying-Cong Chen, Tong Zhang 0001
ECCV (8)2
2022 UnfoldML: Cost-Aware and Uncertainty-Based Dynamic 2D Prediction for Multi-Stage Classification
abstract
Machine Learning (ML) research has focused on maximizing the accuracy of predictive tasks. ML models, however, are increasingly more complex, resource intensive, and costlier to deploy in resource-constrained environments. These issues are exacerbated for prediction tasks with sequential classification on progressively transitioned stages with “happens-before” relation between them.We argue that it is possible to “unfold” a monolithic single multi-class classifier, typically trained for all stages using all data, into a series of single-stage classifiers. Each single- stage classifier can be cascaded gradually from cheaper to more expensive binary classifiers that are trained using only the necessary data modalities or features required for that stage. UnfoldML is a cost-aware and uncertainty-based dynamic 2D prediction pipeline for multi-stage classification that enables (1) navigation of the accuracy/cost tradeoff space, (2) reducing the spatio-temporal cost of inference by orders of magnitude, and (3) early prediction on proceeding stages. UnfoldML achieves orders of magnitude better cost in clinical settings, while detecting multi- stage disease development in real time. It achieves within 0.1% accuracy from the highest-performing multi-class baseline, while saving close to 20X on spatio- temporal cost of inference and earlier (3.5hrs) disease onset prediction. We also show that UnfoldML generalizes to image classification, where it can predict different level of labels (from coarse to fine) given different level of abstractions of a image, saving close to 5X cost with as little as 0.4% accuracy reduction.
Yanbo Xu, Alind Khare, Glenn Matlin, Monish Ramadoss, Rishikesan Kamaleswaran, Chao Zhang 0014, Alexey Tumanov
NeurIPS1
2021 Split-Treatment Analysis to Rank Heterogeneous Causal Effects for Prospective Interventions
abstract
For many kinds of interventions, such as a new advertisement, marketing intervention, or feature recommendation, it is important to target a specific subset of people for maximizing its benefits at minimum cost or potential harm. However, a key challenge is that no data is available about the effect of such a prospective intervention since it has not been deployed yet. In this work, we propose a split-treatment analysis that ranks the individuals most likely to be positively affected by a prospective intervention using past observational data. Unlike standard causal inference methods, the split-treatment method does not need any observations of the target treatments themselves. Instead it relies on observations of a proxy treatment that is caused by the target treatment. Under reasonable assumptions, we show that the ranking of heterogeneous causal effect based on the proxy treatment is the same as the ranking based on the target treatment's effect. In the absence of any interventional data for cross-validation, Split-Treatment uses sensitivity analyses for unobserved confounding to eliminate unreliable models. We apply Split-Treatment to simulated data and a large-scale, real-world targeting task and validate our discovered rankings via a randomized experiment for the latter.
Yanbo Xu, Divyat Mahajan, Liz Manrao, Amit Sharma 0007, Emre Kiciman
WSDM1
2020 HOLMES: Health OnLine Model Ensemble Serving for Deep Learning Models in Intensive Care Units
abstract
Deep learning models have achieved expert-level performance in healthcare with an exclusive focus on training accurate models. However, in many clinical environments such as intensive care unit (ICU), real-time model serving is equally if not more important than accuracy, because in ICU patient care is simultaneously more urgent and more expensive. Clinical decisions and their timeliness, therefore, directly affect both the patient outcome and the cost of care. To make timely decisions, we argue the underlying serving system must be latency-aware. To compound the challenge, health analytic applications often require a combination of models instead of a single model, to better specialize individual models for different targets, multi-modal data, different prediction windows, and potentially personalized predictions. To address these challenges, we propose HOLMES---an online model ensemble serving framework for healthcare applications. HOLMES dynamically identifies the best performing set of models to ensemble for highest accuracy, while also satisfying sub-second latency constraints on end-to-end prediction. We demonstrate that HOLMES is able to navigate the accuracy/latency tradeoff efficiently, compose the ensemble, and serve the model ensemble pipeline, scaling to simultaneously streaming data from 100 patients, each producing waveform data at 250~Hz. HOLMES outperforms the conventional offline batch-processed inference for the same clinical task in terms of accuracy and latency (by order of magnitude). HOLMES is tested on risk prediction task on pediatric cardio ICU data with above 95% prediction accuracy and sub-second latency on 64-bed simulation.
Shenda Hong, Yanbo Xu, Alind Khare, Satria Priambada, Kevin O. Maher, Alaa Aljiffry, Jimeng Sun 0001, Alexey Tumanov
KDD2
2018 RAIM: Recurrent Attentive and Intensive Model of Multimodal Patient Monitoring Data
abstract
With the improvement of medical data capturing, vast amount of continuous patient monitoring data, e.g., electrocardiogram (ECG), real-time vital signs and medications, become available for clinical decision support at intensive care units (ICUs). However, it becomes increasingly challenging to model such data, due to high density of the monitoring data, heterogeneous data types and the requirement for interpretable models.
Yanbo Xu, Siddharth Biswal, Shriprasad R. Deshpande, Kevin O. Maher, Jimeng Sun 0001
KDD1
2017 Predicting Changes in Pediatric Medical Complexity using Large Longitudinal Health Records
Yanbo Xu, Mohammad Taha Bahadori, Elizabeth Searles, Javier Tejedor-Sojo, Jimeng Sun 0001
AMIA1
2016 Improving efficacy attribution in a self-directed learning environment using prior knowledge individualization
abstract
Models of learning in EDM and LAK are pushing the boundaries of what can be measured from large quantities of historical data. When controlled randomization is present in the learning platform, such as randomized ordering of problems within a problem set, natural quasi-randomized controlled studies can be conducted, post-hoc. Difficulty and learning gain attribution are among factors of interest that can be studied with secondary analyses under these conditions. However, much of the content that we might like to evaluate for learning value is not administered as a random stimulus to students but instead is being self-selected, such as a student choosing to seek help in the discussion forums, wiki pages, or other pedagogically relevant material in online courseware. Help seekers, by virtue of their motivation to seek help, tend to be the ones who have the least knowledge. When presented with a cohort of students with a bi-modal or uniform knowledge distribution, this can present problems with model interpretability when a single point estimation is used to represent cohort prior knowledge. Since resource access is indicative of a low knowledge student, a model can tend towards attributing the resources with low or negative learning gain in order to better explain performance given the higher average prior point estimate. In this paper we present several individualized prior strategies and demonstrate how learning efficacy attribution validity and prediction accuracy improve as a result. Level of education attained, relative past assessment performance, and the prior per student cold start heuristic were employed and compared as prior knowledge individualization strategies.
Zachary A. Pardos, Yanbo Xu
LAK2
2015 Exemplar-based large vocabulary speech recognition using k-nearest neighbors
abstract
This paper describes a large scale exemplar-based acoustic modeling approach for large vocabulary continuous speech recognition. We construct an index of labeled training frames using high-level features extracted from the bottleneck layer of a deep neural network as indexing features. At recognition time, each test frame is turned into a query and a set of k-nearest neighbor frames is retrieved from the index. This set is further filtered using majority voting and the remaining frames are used to derive an estimate of the context-dependent state posteriors of the query, which can then be used for recognition. Using an approximate nearest neighbor search approach based on asymmetric hashing, we are able to construct an index on over 25,000 hours of training data. We present both frame classification and recognition experiments on a Voice Search task.
Yanbo Xu, Olivier Siohan, David Simcha, Sanjiv Kumar, Hank Liao
ICASSP1
2014 Using EEG in Knowledge Tracing
Yanbo Xu, Kai-min Chang, Yueran Yuan, Jack Mostow
EDM1
2014 Doing More with Less: Student Modeling and Performance Prediction with Reduced Content Models
Yun Huang 0002, Yanbo Xu, Peter Brusilovsky
UMAP2
2013 Using Item Response Theory to Refine Knowledge Tracing
Yanbo Xu, Jack Mostow
EDM1
2013 Speech enhancement using convolutive nonnegative matrix factorization with cosparsity regularization
Majid Mirbagheri, Yanbo Xu, Sahar Akram, Shihab A. Shamma
INTERSPEECH2
2012 Comparison of methods to trace multiple subskills: Is LR-DBN best?
Yanbo Xu, Jack Mostow
EDM1
2012 Dimension reduction in regression using Gaussian Mixture Models
abstract
Linear-Nonlinear regression models play a fundamental role in characterizing nonlinear systems. In this paper, we propose a method to estimate the linear transform in such models equivalent to a subspace of a small dimension in the input space that is relevant for eliciting response. The novel aspect of this work is the formulation of the mutual information between the transformed inputs and output as a closed-form function of the parameters of their joint density in the form of Gaussian Mixture Models and we subsequently maximize this measure to find relevant dimensions. Instead of a commonly used mutual information measure based on Kullback-Leibler divergence, we use a measure called Quadratic Euclidean Mutual Information. Through experiments on both synthesized data and real MEG recordings, the effectiveness of the proposed method is demonstrated.
Majid Mirbagheri, Yanbo Xu, Shihab A. Shamma
ICASSP2
2012 CAR: Contour-based routing in wireless sensor networks
abstract
MAP is a connectivity-based routing protocol aimed at improving the load balance performance of traditional geographical routing methods. It attempts to find parallel routing paths by taking advantage of the concept of skeleton in the continuous domain. However, MAP suffers seriously from overloading the sensor nodes that are close to the skeleton. In this paper, we propose a contour-based routing protocol, CAR, that does not require geographical information, produces short routing paths, and achieves outstanding load balancing. Our experimental results show that CAR outperforms MAP in terms of both load balancing and routing path length.
Jie Cheng 0003, Qiang Ye 0001, Lei Zhang 0066, Yanbo Xu, Hongbo Jiang 0001, Hongwei Du 0001
ICC4
2011 Desperately Seeking Subscripts: Towards Automated Model Parameterization
Jack Mostow, Yanbo Xu, Md. Ahaduzzaman Munna
EDM2
2011 Using Logistic Regression to Trace Multiple Sub-skills in a Dynamic Bayes Net
Yanbo Xu, Jack Mostow
EDM1
2011 Logistic Regression in a Dynamic Bayes Net Models Multiple Subskills Better!
Yanbo Xu, Jack Mostow
EDM1
2010 Efficient Mobile Content Delivery Based on Co-Route Prediction in Urban Transport
abstract
Routing is one of the most challenging open problems in pocket-switched-networks (PSN). In this paper, we propose a novel co-route media content forwarding scheme (CRMF), in which new contact opportunities are created for occasionally disconnected mobile users. Our study is inspired by two observations: one is that many people tend to make regular journeys to the same place, so their trajectories show a high degree of temporal and spatial regularity. The other is that the number of repeated journeys for an individual commuter is greater than that of the repeated contacts with another commuter who possess similar seasonal movement patterns. Our main contributions include: we properly install store-and-forward routers based on vehicle mobility patterns and human regular movement behaviors; we also propose a router-centric prediction scheme that collects passenger historical trajectory information to determine the delivery scheme. The simulation results demonstrate that this approach improves delivery ratio and also reduces the delivery latency compared to memory (history)-less delivery scheme.
Le Shu, Hongbo Jiang 0001, Xiaoqiang Ma, Lanchao Liu, Kai Peng 0001, Bo Liu 0104, Jie Cheng 0003, Yanbo Xu
GLOBECOM8
2008 Computing Stable Skeletons with Particle Filters
Xiang Bai, Xingwei Yang, Longin Jan Latecki, Yanbo Xu, Wenyu Liu 0001
PRICAI4