EDBT 2026 Demo / reviewers in the wild / expert
Raed Kontar
dblp:216/2976 · also Raed Al Kontar
· DBLP profile ↗
19ranked-venue papers
3as first author
16since 2021 · last 2025
0000-0002-4546-324XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 2 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FCOM: A Federated Collaborative Online Monitoring Framework via Representation LearningabstractMonitoring a large population of dynamic processes with limited resources presents a significant challenge across various industrial sectors. This is due to 1) the inherent disparity between the available monitoring resources and the extensive number of processes to be monitored and 2) the unpredictable and heterogeneous dynamics inherent in the progression of these processes. Online learning approaches, commonly referred to as bandit methods, have demonstrated notable potential in addressing this issue by dynamically allocating resources and effectively balancing the exploitation of high-reward processes and the exploration of uncertain ones. However, most online learning algorithms are designed for 1) a centralized setting that requires data sharing across processes for accurate predictions or 2) a homogeneity assumption that estimates a single global model from decentralized data. To overcome these limitations and enable online learning in a heterogeneous population under a decentralized setting, we propose a federated collaborative online monitoring method. Our approach utilizes representation learning to capture the latent representative models within the population and introduces a novel federated collaborative UCB algorithm to estimate these models from sequentially observed decentralized data. This strategy facilitates informed monitoring of resource allocation. The efficacy of our method is demonstrated through theoretical analysis, simulation studies, and its application to decentralized cognitive degradation monitoring in Alzheimer’s disease. Tanapol Kosolwattana, Huazheng Wang, Raed Kontar |
AAAI | 3 |
| 2025 | Inv-Entropy: A Fully Probabilistic Framework for Uncertainty Quantification in Language ModelsabstractLarge language models (LLMs) have transformed natural language processing, but their reliable deployment requires effective uncertainty quantification (UQ). Existing UQ methods are often heuristic and lack a fully probabilistic foundation. This paper begins by providing a theoretical justification for the role of perturbations in UQ for LLMs. We then introduce a dual random walk perspective, modeling input–output pairs as two Markov chains with transition probabilities defined by semantic similarity. Building on this, we propose a fully probabilistic framework based on an inverse model, which quantifies uncertainty by evaluating the diversity of the input space conditioned on a given output through systematic perturbations. Within this framework, we define a new uncertainty measure, Inv-Entropy. A key strength of our framework is its flexibility: it supports various definitions of uncertainty measures, embeddings, perturbation strategies, and similarity metrics. We also propose GAAP, a perturbation algorithm based on genetic algorithms, which enhances the diversity of sampled inputs. In addition, we introduce a new evaluation metric, Temperature Sensitivity of Uncertainty (TSU), which directly assesses uncertainty without relying on correctness as a proxy. Extensive experiments demonstrate that Inv-Entropy outperforms existing semantic UQ methods. Haoyi Song, Ruihan Ji, Naichen Shi, Fan Lai 0001, Raed Kontar |
NeurIPS | 5 |
| 2025 | Diffusion-Based Surrogate Modeling and Multi-Fidelity CalibrationabstractPhysics simulations have become fundamental tools to study myriad engineering systems. As physics simulations often involve simplifications, their outputs should be calibrated using real-world data. In this paper, we present a diffusion-based surrogate (DBS) that calibrates multi-fidelity physics simulations with diffusion generative processes. DBS categorizes multi-fidelity physics simulations into inexpensive and expensive simulations, depending on the computational costs. The inexpensive simulations, which can be obtained with low latency, directly inject contextual information into diffusion models. Furthermore, when results from expensive simulations are available, DBS refines the quality of generated samples via a guided diffusion process. This design circumvents the need for large amounts of expensive physics simulations to train denoising diffusion models, thus lending flexibility to practitioners. DBS builds on Bayesian probabilistic models and is equipped with a theoretical guarantee that provides upper bounds on the Wasserstein distance between the sample and underlying true distribution. The probabilistic nature of DBS also provides a convenient approach for uncertainty quantification in prediction. Our models excel in cases where physics simulations are imperfect and sometimes inaccessible. We use a numerical simulation in fluid dynamics and a case study in laser-based metal powder deposition additive manufacturing to demonstrate how DBS calibrates multi-fidelity physics simulations with observations to obtain surrogates with superior predictive performance. Naichen Shi, Shenghan Guo, Raed Kontar |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2025 | Collaborative and Distributed Bayesian Optimization via ConsensusabstractOptimal design is a critical yet challenging task within many applications. This challenge arises from the need for extensive trial and error, often done through simulations or running field experiments. Fortunately, sequential optimal design, also referred to as Bayesian optimization when using surrogates with a Bayesian flavor, has played a key role in accelerating the design process through efficient sequential sampling strategies. However, a key opportunity exists nowadays. The increased connectivity of edge devices sets forth a new collaborative paradigm for Bayesian optimization. A paradigm whereby different clients collaboratively borrow strength from each other by effectively distributing their experimentation efforts to improve and fast-track their optimal design process. To this end, we bring the notion of consensus to Bayesian optimization, where clients agree (i.e., reach a consensus) on their next-to-sample designs. Our approach provides a generic and flexible framework that can incorporate different collaboration mechanisms. In lieu of this, we propose transitional collaborative mechanisms where clients initially rely more on each other to maneuver through the early stages with scant data, then, at the late stages, focus on their own objectives to get client-specific solutions. Theoretically, we show the sub-linear growth in regret for our proposed framework. Empirically, through simulated datasets and a real-world collaborative sensor design experiment, we show that our framework can effectively accelerate and improve the optimal design process and benefit all participants. Note to Practitioners—The proposed algorithm allows multiple clients to collaboratively distribute their trial-and-error efforts to fast-track and improve the optimal design process. In the algorithm, each client performs a test locally and then shares the results with an orchestrator. Using the information from all clients, the orchestrator then finds the best new experiment that each client should undertake and sends those back for the next round of experiments. Through this process, all clients can leverage each other’s strengths and optimize their designs with far fewer experiments than each client operating in isolation. This is confirmed through many simulation examples, along with a real-life sensor design experiment where multiple collaborating agents seqeuntially coordinate their experimentation efforts. The goal is to rapidly discover the biosensor design and measurement format parameters that find the maximum amount of captured target analyte. Xubo Yue, Yang Liu 0479, Albert S. Berahas, Blake N. Johnson, Raed Kontar |
IEEE Trans Autom. Sci. Eng. | 5 |
| 2024 | Collaborative and Federated Black-box Optimization: A Bayesian Optimization PerspectiveabstractWe focus on collaborative and federated black-box optimization (BBOpt), where agents optimize their heterogeneous black-box functions through collaborative sequential experimentation. From a Bayesian optimization perspective, we address the fundamental challenges of distributed experimentation, heterogeneity, and privacy within BBOpt, and propose three unifying frameworks to tackle these issues: (i) a global framework where experiments are centrally coordinated, (ii) a local framework that allows agents to make decisions conditioned on shared information, and (iii) a predictive framework that enhances local surrogates through collaboration to improve decision-making. We categorize existing methods within these frameworks and highlight key open questions to unlock the full potential of federated BBOpt. Our overarching goal is to shift federated learning from its predominantly descriptive/predictive paradigm to a prescriptive one, particularly in the context of BBOpt —an inherently sequential decision-making problem. Raed Kontar |
IEEE Big Data | 1 |
| 2024 | Triple Component Matrix Factorization: Untangling Global, Local, and Noisy ComponentsabstractIn this work, we study the problem of common and unique feature extraction from noisy data. When we have $N$ observation matrices from $N$ different and associated sources corrupted by sparse and potentially gross noise, can we recover the common and unique components from these noisy observations? This is a challenging task as the number of parameters to estimate is approximately thrice the number of observations. Despite the difficulty, we propose an intuitive alternating minimization algorithm called triple component matrix factorization (TCMF) to recover the three components exactly. TCMF is distinguished from existing works in literature thanks to two salient features. First, TCMF is a principled method to separate the three components given noisy observations provably. Second, the bulk of the computation in TCMF can be distributed. On the technical side, we formulate the problem as a constrained nonconvex nonsmooth optimization problem. Despite the intricate nature of the problem, we provide a Taylor series characterization of its solution by solving the corresponding Karush–Kuhn–Tucker conditions. Using this characterization, we can show that the alternating minimization algorithm makes significant progress at each iteration and converges into the ground truth at a linear rate. Numerical experiments in video segmentation and anomaly detection highlight the superior feature extraction abilities of TCMF. Naichen Shi, Salar Fattahi, Raed Kontar |
J. Mach. Learn. Res. | 3 |
| 2024 | Personalized PCA: Decoupling Shared and Unique FeaturesabstractIn this paper, we tackle a significant challenge in PCA: heterogeneity. When data are collected from different sources with heterogeneous trends while still sharing some congruency, it is critical to extract shared knowledge while retaining the unique features of each source. To this end, we propose personalized PCA (PerPCA), which uses mutually orthogonal global and local principal components to encode both unique and shared features. We show that, under mild conditions, both unique and shared features can be identified and recovered by a constrained optimization problem, even if the covariance matrices are immensely different. Also, we design a fully federated algorithm inspired by distributed Stiefel gradient descent to solve the problem. The algorithm introduces a new group of operations called generalized retractions to handle orthogonality constraints, and only requires global PCs to be shared across sources. We prove the linear convergence of the algorithm under suitable assumptions. Comprehensive numerical experiments highlight PerPCA's superior performance in feature extraction and prediction from heterogeneous datasets. As a systematic approach to decouple shared and unique features from heterogeneous datasets, PerPCA finds applications in several tasks, including video segmentation, topic extraction, and feature clustering. Naichen Shi, Raed Kontar |
J. Mach. Learn. Res. | 2 |
| 2024 | Federated Gaussian Process: Convergence, Automatic Personalization and Multi-Fidelity ModelingabstractIn this paper, we propose FGPR: a Federated Gaussian process ( GP) regression framework that uses an averaging strategy for model aggregation and stochastic gradient descent for local computations. Notably, the resulting global model excels in personalization as FGPR jointly learns a shared prior across all devices. The predictive posterior is then obtained by exploiting this shared prior and conditioning on local data, which encodes personalized features from a specific dataset. Theoretically, we show that FGPR converges to a critical point of the full log-marginal likelihood function, subject to statistical errors. This result offers standalone value as it brings federated learning theoretical results to correlated paradigms. Through extensive case studies on several regression tasks, we show that FGPR excels in a wide range of applications and is a promising approach for privacy-preserving multi-fidelity data modeling. Xubo Yue, Raed Kontar |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Fed-ensemble: Ensemble Models in Federated Learning for Improved Generalization and Uncertainty QuantificationabstractThe increase in the computational power of edge devices has opened up the possibility of processing some of the data at the edge and distributing model learning. This paradigm is often called federated learning (FL), where edge devices exploit their local computational resources to train models collaboratively. Though FL has seen recent success, it is unclear how to characterize uncertainties in FL predictions. In this paper, we proposeFed-ensemble: a simple approach that brings model ensembling to FL. Instead of aggregating local models to update a single global model,Fed-ensembleuses random permutations to update a group of$K$models and then obtains predictions through model averaging.Fed-ensemblecan be readily utilized within established FL methods and does not impose a computational overhead compared with single-model methods. Empirical results show that our model has superior performance over several FL algorithms on a wide range of data sets and excels in heterogeneous settings often encountered in FL applications. Also, by carefully choosing client-dependent weights in the inference stage,Fed-ensemblebecomes personalized and yields even better performance. Theoretically, we show that predictions on new data from all$K$models belong to the same predictive posterior distribution under a neural tangent kernel regime. This result, in turn, sheds light on the generalization advantages of model averaging and justifies the uncertainty quantification capability. We also illustrate thatFed-ensemblehas an elegant Bayesian interpretation.Note to Practitioners—provides an algorithm that extracts a set of$K$solutions without imposing any additional communication overhead in FL. Given multiple solutions,Fed-ensemblecan be exploited to personalize inference as well as quantify uncertainty. Such capabilities may be beneficial within multiple practical systems that require uncertainty-aware decision-making. Further,Fed-ensemblemay be useful for model validation and hypothesis testing. Naichen Shi, Fan Lai 0001, Raed Kontar, Mosharaf Chowdhury |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2024 | SALR: Sharpness-Aware Learning Rate Scheduler for Improved GeneralizationabstractIn an effort to improve generalization in deep learning and automate the process of learning rate scheduling, we propose SALR: a sharpness-aware learning rate update technique designed to recover flat minimizers. Our method dynamically updates the learning rate of gradient-based optimizers based on the local sharpness of the loss function. This allows optimizers to automatically increase learning rates at sharp valleys to increase the chance of escaping them. We demonstrate the effectiveness of SALR when adopted by various algorithms over a broad range of networks. Our experiments indicate that SALR improves generalization, converges faster, and drives solutions to significantly flatter regions. Xubo Yue, Maher Nouiehed, Raed Kontar |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Federated Condition Monitoring Signal Prediction With Improved GeneralizationabstractRevolutionary advances in Internet of Things technologies have paved the way for a significant increase in computational resources at edge devices that collect condition monitoring (CM) data. This poses a significant opportunity for federated analytics (FA), which exploits edge computing resources to distribute model learning, reduce communication traffic, and circumvent the need to share raw data. In this article, we study CM signal prediction where operating units that have data storage and computational capabilities jointly learn models without sharing their collected CM signals. The key challenge we aim to address is learning effective FA models in the presence of heterogeneity, which is often intrinsic to CM signals. To this end, we first introduce a federated framework for CM signal prediction that tries to improve generalization by encouraging flat solutions through distributed computations. Then, a personalization approach is proposed to adapt the learned model to new clients without losing old knowledge. We examine our proposed framework on CM signals from aircraft turbofan engines under three realistic federated CM scenarios. Experimental results highlight the capability of our model to decentralize model inference while improving generalization and robustness to heterogeneity across CM signals. Seokhyun Chung, Raed Kontar |
IEEE Trans. Reliab. | 2 |
| 2023 | Personalized Dictionary Learning for Heterogeneous DatasetsabstractWe introduce a relevant yet challenging problem named Personalized Dictionary Learning (PerDL), where the goal is to learn sparse linear representations from heterogeneous datasets that share some commonality. In PerDL, we model each dataset's shared and unique features as global and local dictionaries. Challenges for PerDL not only are inherited from classical dictionary learning(DL), but also arise due to the unknown nature of the shared and unique features. In this paper, we rigorously formulate this problem and provide conditions under which the global and local dictionaries can be provably disentangled. Under these conditions, we provide a meta-algorithm called Personalized Matching and Averaging (PerMA) that can recover both global and local dictionaries from heterogeneous datasets. PerMA is highly efficient; it converges to the ground truth at a linear rate under suitable conditions. Moreover, it automatically borrows strength from strong learners to improve the prediction of weak learners. As a general framework for extracting global and local dictionaries, we show the application of PerDL in different learning tasks, such as training with imbalanced datasets and video surveillance. Geyu Liang, Naichen Shi, Raed Kontar, Salar Fattahi |
NeurIPS | 3 |
| 2022 | Gaussian Process Parameter Estimation Using Mini-batch Stochastic Gradient Descent: Convergence Guarantees and Empirical BenefitsabstractStochastic gradient descent (SGD) and its variants have established themselves as the go-to algorithms for large-scale machine learning problems with independent samples due to their generalization performance and intrinsic computational advantage. However, the fact that the stochastic gradient is a biased estimator of the full gradient with correlated samples has led to the lack of theoretical understanding of how SGD behaves under correlated settings and hindered its use in such cases. In this paper, we focus on hyperparmeter estimation for the Gaussian process (GP) and take a step forward towards breaking the barrier by proving minibatch SGD converges to a critical point of the full log-likelihood loss function, and recovers model hyperparameters with rate $O(\frac{1}{K})$ for $K$ iterations, up to a statistical error term depending on the minibatch size. Our theoretical guarantees hold provided that the kernel functions exhibit exponential or polynomial eigendecay which is satisfied by a wide range of kernels commonly used in GPs. Numerical studies on both simulated and real datasets demonstrate that minibatch SGD has better generalization over state-of-the-art GP methods while reducing the computational burden and opening a new, previously unexplored, data size regime for GPs. Hao Chen 0067, Raed Kontar, Garvesh Raskutti |
J. Mach. Learn. Res. | 3 |
| 2022 | A Multi-Stage Approach for Knowledge-Guided Predictions With Application to Additive ManufacturingabstractInspired by sequential additive manufacturing operations, we consider prediction tasks arising in processes that comprise of sequential sub-operations and propose a multi-stage inference procedure that exploits prior knowledge of the operational sequence. Our approach decomposes a data-driven model into several easier problems each corresponding to a sub-operation and then introduces a Bayesian inference procedure to quantify and propagate uncertainty across operational stages. We also complement our model with an approach to incorporate physical knowledge of the output of a sub-operation which is often more practical in reality relative to understanding the physics of the entire process. Comprehensive simulations and two case studies on additive manufacturing show that the proposed framework provides well-quantified uncertainties and superior predictive accuracy compared to a single-stage predictive approach.Note to Practitioners—This paper is motivated by sequential operations that often occur in manufacturing processes. For example, several additive manufacturing processes consist of multiple sequential steps, e.g., printing, washing, and curing in stereolithography, or printing, debinding, and sintering in binder jetting. In such settings, a complex data-driven model that blindly throws all given data into a single predictive model might not be optimal. To this end, we propose a multi-stage inference procedure that decomposes the problem into easier sub-problem using the prior knowledge of the operational sequence, and propagates uncertainty across stages using Bayesian neural networks. Here we note that even if sequential operations are not existent in reality, one may conceptually decompose a complex system into simpler pieces and exploit our procedure. Also, our approach is able to incorporate physical knowledge of the output of a sub-operation. Seokhyun Chung, Cheng-Hao Chou, Xiaozhu Fang, Raed Kontar, Chinedum Emmanuel Okwudire |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2021 | Minimizing Negative Transfer of Knowledge in Multivariate Gaussian Processes: A Scalable and Regularized ApproachabstractRecently there has been an increasing interest in the multivariate Gaussian process (MGP) which extends the Gaussian process (GP) to deal with multiple outputs. One approach to construct the MGP and account for non-trivial commonalities amongst outputs employs a convolution process (CP). The CP is based on the idea of sharing latent functions across several convolutions. Despite the elegance of the CP construction, it provides new challenges that need yet to be tackled. First, even with a moderate number of outputs, model building is extremely prohibitive due to the huge increase in computational demands and number of parameters to be estimated. Second, the negative transfer of knowledge may occur when some outputs do not share commonalities. In this paper we address these issues. We propose a regularized pairwise modeling approach for the MGP established using CP. The key feature of our approach is to distribute the estimation of the full multivariate model into a group of bivariate GPs which are individually built. Interestingly pairwise modeling turns out to possess unique characteristics, which allows us to tackle the challenge of negative transfer through penalizing the latent function that facilitates information sharing in each bivariate model. Predictions are then made through combining predictions from the bivariate models within a Bayesian framework. The proposed method has excellent scalability when the number of outputs is large and minimizes the negative transfer of knowledge between uncorrelated outputs. Statistical guarantees for the proposed method are studied and its advantageous features are demonstrated through numerical studies. Raed Kontar, Garvesh Raskutti |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Functional Principal Component Analysis for Extrapolating Multistream Longitudinal DataabstractThe advance of modern sensor technologies enables collection of multi-stream longitudinal data where multiple signals from different units are collected in real-time. In this article, we present a non-parametric approach to predict the evolution of multi-stream longitudinal data for an in-service unit through borrowing strength from other historical units. Our approach first decomposes each stream into a linear combination of eigenfunctions and their corresponding functional principal component (FPC) scores. A Gaussian process prior for the FPC scores is then established based on a functional semi-metric that measures similarities between streams of historical units and the in-service unit. Finally, an empirical Bayesian updating strategy is derived to update the established prior using real-time stream data obtained from the in-service unit. Experiments on synthetic and real world data show that the proposed framework outperforms state-of-the-art approaches and can effectively account for heterogeneity as well as achieve high predictive accuracy. Seokhyun Chung, Raed Kontar |
IEEE Trans. Reliab. | 2 |
| 2020 | Why Non-myopic Bayesian Optimization is Promising and How Far Should We Look-ahead? A Study via RolloutabstractLookahead, also known as non-myopic, Bayesian optimization (BO) aims to find optimal sampling policies through solving a dynamic programming (DP) formulation that maximizes a long-term reward over a rolling horizon. Though promising, lookahead BO faces the risk of error propagation through its increased dependence on a possibly mis-specified model. In this work we focus on the rollout approximation for solving the intractable DP. We first prove the improving nature of rollout in tackling lookahead BO and provide a sufficient condition for the used heuristic to be rollout improving. We then provide both a theoretical and practical guideline to decide on the rolling horizon stagewise. This guideline is built on quantifying the negative effect of a mis-specified model. To illustrate our idea, we provide case studies on both single and multi-information source BO. Empirical results show the advantageous properties of our method over several myopic and non-myopic BO algorithms. Xubo Yue, Raed Kontar |
AISTATS | 2 |
| 2020 | Stochastic Gradient Descent in Correlated Settings: A Study on Gaussian ProcessesabstractStochastic gradient descent (SGD) and its variants have established themselves as the go-to algorithms for large-scale machine learning problems with independent samples due to their generalization performance and intrinsic computational advantage. However, the fact that the stochastic gradient is a biased estimator of the full gradient with correlated samples has led to the lack of theoretical understanding of how SGD behaves under correlated settings and hindered its use in such cases. In this paper, we focus on the Gaussian process (GP) and take a step forward towards breaking the barrier by proving minibatch SGD converges to a critical point of the full loss function, and recovers model hyperparameters with rate $O(\frac{1}{K})$ up to a statistical error term depending on the minibatch size. Numerical studies on both simulated and real datasets demonstrate that minibatch SGD has better generalization over state-of-the-art GP methods while reducing the computational burden and opening a new, previously unexplored, data size regime for GPs. Hao Chen 0067, Raed Kontar, Garvesh Raskutti |
NeurIPS | 3 |
| 2018 | Nonparametric-Condition-Based Remaining Useful Life Prediction Incorporating External FactorsabstractThe use of condition monitoring (CM) signals to predict the remaining useful life of in-service units plays a critical role in reliability engineering. Many models assume that CM signals behave under similar external conditions or that external factors have no effect on the evolution of these signals. These assumptions might not hold in real-life applications. In this paper, we propose a nonparametric framework for modeling the evolution of CM signals under different external factors. The unique feature of our model is that it does not assume any functional form for CM signals and is able to incorporate the effect of external factors through a reparametrization technique called hypersphere decomposition. Through extensive numerical studies and a case study on automotive lead-acid batteries, we demonstrate the advantageous features of our proposed method specifically when the evolution of CM signals is impacted by external factors. Raed Kontar, Chaitanya Sankavaram, Yilu Zhang |
IEEE Trans. Reliab. | 1 |