VLDB 2026 Research / reviewers in the wild / expert
Anahita Khojandi
dblp:149/8048
· DBLP profile ↗
12ranked-venue papers
1as first author
8since 2021 · last 2025
0000-0001-6818-2048ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Theory of computation · 3 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Unsupervised Machine Learning for Cybersecurity Anomaly Detection in Traditional and Software-Defined Networking EnvironmentsabstractCybersecurity has become a field of increasing importance within the past years, with the National Academy of Engineering most recently designating securing cyberspace as one of the fourteen Grand Challenges in Engineering in the 21st Century. Henceforth, it is imperative to design a robust anomaly detection and response approach that can identify and mitigate anomalous Internet traffic. In this study, we present several unsupervised/semi-supervised machine learning models to combat prolific anomalous data on a computer network. Specifically, we employ five unsupervised machine learning models, including a Generative Adversarial Network (GAN), Deep Belief Network (DBN), Restricted Boltzmann Machine (RBM), One-Class Support Vector Machine (OCSVM), and Isolation Forest (I-Forest). We use these models separately and, when applicable, combined together to examine their anomaly detection performance on three prominent traditional networking datasets, namely the KDD-Cup 99, NSL-KDD, and CIC-IDS2017 dataset, and implement these models within a software-defined networking and industrial Internet-of-Things environment using the DNP3 intrusion detection dataset. Furthermore, we investigate the generalizability of the models across the two datasets. Our results suggest I-Forest and DBN overall perform better than other models in traditional and software-defined networking environments; our GAN manages to outperform some benchmark models on the CIC-IDS2017 dataset. Curtis Rookard, Anahita Khojandi |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2024 | RRIoT: Recurrent reinforcement learning for cyber threat detection on IoT devices
Curtis Rookard, Anahita Khojandi |
Comput. Secur. | 2 |
| 2023 | Determining the Most Significant Metadata Features to Indicate Defective Software CommitsabstractDefects are largely inevitable in the software development life cycle. Since we cannot avoid them during the development process, we can only desire to fight back with our limited resources in terms of time and monetary investment. Like in many other fields, machine learning models can be of help to mitigate the problem of defects by predicting both bug frequency and defective modules at different granularity levels. However, machine learning models are as good as the quality of the pre-selected set of features under consideration. Therefore, importance must be given while selecting only the necessary features from the original set of features. In this study, we compared various machine learning models with varying feature selection techniques and found the superiority of random forest-based machine learning techniques with wrapper methods. Random forest-based models with the wrapper method were able to detect all the buggy classes successfully on the validation data set. Rupam Kumar Dey, Anahita Khojandi, Kalyan S. Perumalla |
SERA | 2 |
| 2022 | A Machine Learning-Enabled Partially Observable Markov Decision Process Framework for Early Sepsis PredictionabstractSepsis is a life-threatening condition, caused by the body’s extreme response to an infection. In the United States, 1.7 million cases of sepsis occur annually, resulting in 265,000 deaths. Delayed diagnosis and treatment are associated with higher mortality rates. An exponential rise in the availability of medical data has allowed for the development of sophisticated machine learning algorithms to predict sepsis earlier than the onset. However, these models often underperform, as the training data are retrospective and do not fully capture the uncertain future. In this study, we develop a novel framework, which we refer to as MLePOMDP, to leverage and combine the underlying, high-level knowledge about sepsis progression and machine learning (ML) for classification. Specifically, we use a hidden Markov model to describe sepsis development at a high level, where the ML model makes the higher-order “observations” from temporal data. Consequently, a partially observable Markov decision process (POMDP) model is developed to make classification decisions. We analytically establish that the optimal policy is of threshold-type, which we exploit to efficiently optimize MLePOMDP. MLePOMDP is calibrated and tested using high-frequency physiological data collected from bedside monitors. Different from past POMDP-based frameworks, MLePOMDP is developed for a prediction task using a very small state definition, produces highly interpretable results, and accounts for a novel and clinically meaningful action space. Our results show that MLePOMDP outperforms machine learning–based benchmarks by up to 8% in precision. Importantly, MLePOMDP is able to reduce false alarms by up to 28%. An additional experiment is conducted to show the generalizability of MLePOMDP to different patient cohorts. Summary of Contribution: This study develops a novel real-time decision support framework for early sepsis prediction by integrating well-known machine learning models (random forest and neural networks) with a well-established sequential decision-making model, namely, a partially observable Markov decision process (POMDP). The structural properties of the optimal policy are further explored and a threshold-type structure is established, which is then leveraged to develop a customized algorithm to solve the problem more efficiently. The resulting framework demonstrates the benefit of applying POMDPs to augment machine learning outputs. Specifically, the framework results in the reduction of false alarms in sepsis predictions where decisions are made in real time, hence improving the overall prediction precision. Zeyu Liu 0002, Anahita Khojandi, Xueping Li 0002, Akram Mohammed, Robert L. Davis, Rishikesan Kamaleswaran |
INFORMS J. Comput. | 2 |
| 2022 | A Dynamic Deep Reinforcement Learning-Bayesian Framework for Anomaly DetectionabstractTo assure the successful operation of connected and automated vehicles, it is critical to detect and isolate anomalous and/or faulty information in a timely manner. To do so, anomaly detection techniques should be implemented in real-time where if the probability of anomalous information exceeds a certain threshold, the information is dealt with accordingly. Traditionally, the threshold for judging whether the data is anomalous is fixed and determined a priori. However, not only does this approach fail to account for the feedback obtained during a trip on the performance of the algorithms, but it also fails to respond to potential changes in rates of anomalies. Hence, it is important to develop an approach that can dynamically alter this threshold in response to exogenous factors to assure reliable and robust system operation. We develop a mathematical framework which utilizes a dynamic threshold for an anomaly classification algorithm in order to maximize the safety of a trip. Specifically, we develop and pair an anomaly classification algorithm based on convolutional neural networks (CNN), with a partially observable Markov decision process (POMDP) model. We solve the resulting POMDP model using the asynchronous advantage actor critic (A3C) deep reinforcement learning algorithm. The prescribed policy determines the anomaly classification threshold in real-time that maximizes the performance. Our numerical experiments show that the POMDP model outperforms state-of-the-art benchmarks, especially under more difficult to detect anomaly profiles. Jeremy Watts, Franco van Wyk, Shahrbanoo Rezaei, Yiyang Wang 0002, Neda Masoud, Anahita Khojandi |
IEEE Trans. Intell. Transp. Syst. | 6 |
| 2021 | Hidden Markov models as recurrent neural networks: An application to Alzheimer's diseaseabstractHidden Markov models (HMMs) are commonly used for disease progression modeling when the true patient health state is not fully known. Since HMMs typically have multiple local optima, incorporating additional patient covariates can improve parameter estimation and predictive performance. To allow for this, we develop hidden Markov recurrent neural networks (HMRNNs), a special case of recurrent neural networks that combine neural networks' flexibility with HMMs' interpretability. The HMRNN can be reduced to a standard HMM, with an identical likelihood function and parameter interpretations, but it can also combine an HMM with other predictive neural networks that take patient information as input. The HMRNN estimates all parameters simultaneously via gradient descent. Using a dataset of Alzheimer's disease patients, we demonstrate how the HMRNN can combine an HMM with other predictive neural networks to improve disease forecasting and to offer a novel clinical interpretation compared with a standard HMM trained via expectation-maximization. Matthew Baucum, Anahita Khojandi, Theodore Papamarkou |
BIBE | 2 |
| 2021 | Improving Deep Reinforcement Learning With Transitional Variational Autoencoders: A Healthcare ApplicationabstractReinforcement learning is a powerful tool for developing personalized treatment regimens from healthcare data. Yet training reinforcement learning agents through direct interactions with patients is often impractical for ethical reasons. One solution is to train reinforcement learning agents using an 'environment model,' which is learned from retrospective patient data, and can simulate realistic patient trajectories. In this study, we propose transitional variational autoencoders (tVAE), a generative neural network architecture that learns a direct mapping between distributions over clinical measurements at adjacent time points. Unlike other models, the tVAE requires few distributional assumptions, and benefits from identical training, and testing architectures. This model produces more realistic patient trajectories than state-of-the-art sequential decision-making models, and generative neural networks, and can be used to learn effective treatment policies. Matthew Baucum, Anahita Khojandi, Rama K. Vasudevan |
IEEE J. Biomed. Health Informatics | 2 |
| 2021 | Real-Time Sensor Anomaly Detection and Recovery in Connected Automated Vehicle SensorsabstractIn this paper we propose a novel observer-based method to improve the safety and security of connected and automated vehicle (CAV) transportation. The proposed method combines model-based signal filtering and anomaly detection methods. Specifically, we use adaptive extended Kalman filter (AEKF) to smooth sensor readings of a CAV based on a nonlinear car-following model. Using the car-following model the subject vehicle (i.e., the following vehicle) utilizes the leading vehicle's information to detect sensor anomalies by employing previously-trained One Class Support Vector Machine (OCSVM) models. This approach allows the AEKF to estimate the state of a vehicle not only based on the vehicle's location and speed, but also by taking into account the state of the surrounding traffic. A communication time delay factor is considered in the car-following model to make it more suitable for real-world applications. Our experiments show that compared with the AEKF with a traditional x2-detector, our proposed method achieves a better anomaly detection performance. We also demonstrate that a larger time delay factor has a negative impact on the overall detection performance. Yiyang Wang 0002, Neda Masoud, Anahita Khojandi |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2020 | On the k-Strong Roman Domination Problem
Zeyu Liu 0002, Xueping Li 0002, Anahita Khojandi |
Discret. Appl. Math. | 3 |
| 2020 | Real-Time Sensor Anomaly Detection and Identification in Automated VehiclesabstractConnected and automated vehicles (CAVs) are expected to revolutionize the transportation industry, mainly through allowing for a real-time and seamless exchange of information between vehicles and roadside infrastructure. Although connectivity and automation are projected to bring about a vast number of benefits, they can give rise to new challenges in terms of safety, security, and privacy. To navigate roadways, CAVs need to heavily rely on their sensor readings and the information received from other vehicles and roadside units. Hence, anomalous sensor readings caused by either malicious cyber attacks or faulty vehicle sensors can result in disruptive consequences and possibly lead to fatal crashes. As a result, before the mass implementation of CAVs, it is important to develop methodologies that can detect anomalies and identify their sources seamlessly and in real time. In this paper, we develop an anomaly detection approach through combining a deep learning method, namely convolutional neural network (CNN), with a well-established anomaly detection method, and Kalman filtering with a χ2-detector, to detect and identify anomalous behavior in CAVs. Our numerical experiments demonstrate that the developed approach can detect anomalies and identify their sources with high accuracy, sensitivity, and F1 score. In addition, this developed approach outperforms the anomaly detection and identification capabilities of both CNNs and Kalman filtering with a χ2-detector method alone. It is envisioned that this research will contribute to the development of safer and more resilient CAV systems that implement a holistic view toward intelligent transportation system (ITS) concepts. Franco van Wyk, Yiyang Wang 0002, Anahita Khojandi, Neda Masoud |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2019 | Improving Prediction Performance Using Hierarchical Analysis of Real-Time Data: A Sepsis Case StudyabstractThis paper presents a novel method for hierarchical analysis of machine learning algorithms to improve predictions of at risk patients, thus further enabling prompt therapy. Specifically, we develop a multi-layer machine learning approach to analyze continuous, high-frequency data. We illustrate the capabilities of this approach for early identification of patients at risk of sepsis, a potentially life-threatening complication of an infection, using high-frequency (minute-by-minute) physiological data collected from bedside monitors. In our analysis of a cohort of 586 patients, the model obtained from analyzing the output of a previously developed sepsis prediction model resulted in improved outcomes. Specifically, the original model failed to predict 11.76 ± 4.26% of sepsis patients earlier than Systemic Inflammatory Response Syndrome (SIRS) criteria, commonly used to identify patients at risk for rapid physiological deterioration resulting from sepsis. In contrast, the multi-layer model only failed to predict 3.21 ± 3.11% of sepsis patients earlier than SIRS. In addition, sepsis patients were predicted on average 204.87 ± 7.90 minutes earlier than SIRS criteria using the multi-layer model, which can potentially help reduce mortality and morbidity if implemented in the ICU. Franco van Wyk, Anahita Khojandi, Rishikesan Kamaleswaran |
IEEE J. Biomed. Health Informatics | 2 |
| 2014 | Optimal Implantable Cardioverter Defibrillator (ICD) Generator ReplacementabstractImplantable cardioverter defibrillators (ICDs) include small, battery-powered generators, the longevity of which depends on a patient's rate of consumption. Generator replacement, however, involves risks, including death. Hence, a trade-off exists between prematurely exposing the patient to these risks and allowing for the possibility that the device is unable to deliver therapy when needed. Currently, replacements are performed using a one-size-fits-all approach. Here, we develop a Markov decision process model to determine patient-specific optimal replacement policies as a function of patient age and the remaining battery capacity. We analytically establish that the optimal policy is of threshold-type in the remaining capacity, but not necessarily in patient age. Based on clinical data, we conduct a large computational study that suggests that under the optimal policy, patients undergoing initial implantation at age 30–40, 41–60, and 61–80 see an approximate decrease in the total expected number of replacements of 8%–14%, 8%–15% and 8%–19%, respectively, while achieving the same or greater expected lifetime. Anahita Khojandi, Lisa M. Maillart, Oleg A. Prokopyev, Mark S. Roberts, Timothy Brown, William W. Barrington |
INFORMS J. Comput. | 1 |