Morteza Noshad

dblp:28/10329 · also Morteza Noshad Iranzad · DBLP profile ↗
← Back
15ranked-venue papers
6as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Universal Training of Neural Networks to Achieve Bayes Optimal Classification Accuracy
abstract
This work invokes the notion of f-divergence to introduce a novel upper bound on the Bayes error rate of a general classification task. We show that the proposed bound can be computed by sampling from the output of a parameterized model. Using this practical interpretation, we introduce the Bayes optimal learning threshold (BOLT) loss whose minimization enforces a classification model to achieve the Bayes error rate. We validate the proposed loss for image and text classification tasks, considering MNIST, Fashion-MNIST, CIFAR10, and IMDb datasets. Numerical experiments demonstrate that models trained with BOLT achieve performance on par with or exceeding that of cross-entropy, particularly on challenging datasets. This highlights the potential of BOLT in improving generalization.
Mohammadreza Tavasoli Naeini, Ali Bereyhi, Morteza Noshad, Ben Liang 0001, Alfred O. Hero III
ICASSP3
2023 Graph-based clinical recommender: Predicting specialists procedure orders using graph representation learning
Sajjad Fouladvand, Federico Reyes Gomez, Hamed Nilforoshan, Matthew Schwede, Morteza Noshad, Olivia Jee, Jiaxuan You, Rok Sosic, Jure Leskovec, Jonathan H. Chen
J. Biomed. Informatics5
2022 Leveraging EHR Audit Log Data to Unlock New Insights into Care Processes and Outcomes
Christian Rose, Robert Thombley, Morteza Noshad, Ron Li, Wendy Lu, Heather A Clancy, David Schlessinger, Vincent X. Liu, Jonathan H. Chen, Julia Adler-Milstein
AMIA3
2022 Team is brain: leveraging EHR audit log data for new insights into acute care processes
abstract
OBJECTIVE: To determine whether novel measures of contextual factors from multi-site electronic health record (EHR) audit log data can explain variation in clinical process outcomes. MATERIALS AND METHODS: We selected one widely-used process outcome: emergency department (ED)-based team time to deliver tissue plasminogen activator (tPA) to patients with acute ischemic stroke (AIS). We evaluated Epic audit log data (that tracks EHR user-interactions) for 3052 AIS patients aged 18+ who received tPA after presenting to an ED at three Northern California health systems (Stanford Health Care, UCSF Health, and Kaiser Permanente Northern California). Our primary outcome was door-to-needle time (DNT) and we assessed bivariate and multivariate relationships with six audit log-derived measures of treatment team busyness and prior team experience. RESULTS: Prior team experience was consistently associated with shorter DNT; teams with greater prior experience specifically on AIS cases had shorter DNT (minutes) across all sites: (Site 1: -94.73, 95% CI: -129.53 to 59.92; Site 2: -80.93, 95% CI: -130.43 to 31.43; Site 3: -42.95, 95% CI: -62.73 to 23.17). Teams with greater prior experience across all types of cases also had shorter DNT at two sites: (Site 1: -6.96, 95% CI: -14.56 to 0.65; Site 2: -19.16, 95% CI: -36.15 to 2.16; Site 3: -11.07, 95% CI: -17.39 to 4.74). Team busyness was not consistently associated with DNT across study sites. CONCLUSIONS: EHR audit log data offers a novel, scalable approach to measure key contextual factors relevant to clinical process outcomes across multiple sites. Audit log-based measures of team experience were associated with better process outcomes for AIS care, suggesting opportunities to study underlying mechanisms and improve care through deliberate training, team-building, and scheduling to maximize team experience.
Christian Rose, Robert Thombley, Morteza Noshad, Heather A Clancy, David Schlessinger, Ron C. Li, Vincent X. Liu, Jonathan H. Chen, Julia Adler-Milstein
J. Am. Medical Informatics Assoc.3
2022 Signal from the noise: A mixed graphical and quantitative process mining approach to evaluate care pathways applied to emergency stroke care
Morteza Noshad, Christian Rose, Jonathan H. Chen
J. Biomed. Informatics1
2021 Machine Learning Predictability of Clinical Next Generation Sequencing for Hematologic Malignancies to Guide High-Value Precision Medicine
Grace Y. E. Kim, Morteza Noshad, Henning Stehr, Rebecca Rojansky, Dita Gratzinger, Jean Oak, Rondeep Brar, David Iberri, Christina Kong, James L. Zehnder, Jonathan H. Chen
AMIA2
2021 Signal from the Noise: Quantitative Measures of Conformity and Variability From Process Mining Maps
Christian Rose, Morteza Noshad, Jonathan H. Chen
AMIA2
2021 NIR Iris Challenge Evaluation in Non-cooperative Environments: Segmentation and Localization
abstract
For iris recognition in non-cooperative environments, iris segmentation has been regarded as the first most important challenge still open to the biometric community, affecting all downstream tasks from normalization to recognition. In recent years, deep learning technologies have gained significant popularity among various computer vision tasks and also been introduced in iris biometrics, especially iris segmentation. To investigate recent developments and attract more interest of researchers in the iris segmentation method, we organized the 2021 NIR Iris Challenge Evaluation in Non-cooperative Environments: Segmentation and Localization (NIR-ISL 2021) at the 2021 International Joint Conference on Biometrics (IJCB 2021). The challenge was used as a public platform to assess the performance of iris segmentation and localization methods on Asian and African NIR iris images captured in non-cooperative environments. The three best-performing entries achieved solid and satisfactory iris segmentation and localization results in most cases, and their code and models have been made publicly available for reproducibility research.
Caiyong Wang, Yunlong Wang 0003, Kunbo Zhang, Jawad Muhammad, Qi Zhang 0015, Qichuan Tian, Zhaofeng He 0001, Zhenan Sun, Tianbao Liu, Wei Yang 0006, Dongliang Wu, Yingfeng Liu, Ruiye Zhou, Huihai Wu, Junbao Wang, Wantong Xiong, Xueyu Shi, Shao Zeng, Peihua Li, Huijie Wu, Xinhui Zhang, Menghan Zhang, Fadi Boutros, Naser Damer, Arjan Kuijper, Juan E. Tapia, Andres Valenzuela, Christoph Busch 0001, Gourav Gupta, Kiran B. Raja, Xi Wu 0004, Xiaojie Li 0001, Jingfu Yang, Hongyan Jing, Xin Wang 0045, Bin Kong 0001, Youbing Yin, Qi Song 0001, Siwei Lyu, Shu Hu 0001, Leon Premk, Matej Vitek, Vitomir Struc, Peter Peer, Jalil Nourmohammadi-Khiarak, Farhang Jaryani, Samaneh Salehi Nasab, Seyed Naeim Moafinejad, Yasin Amini, Morteza Noshad
IJCB63
2020 Context is Key: Using the Audit Log to Capture Contextual Factors Affecting Stroke Care Processes
Morteza Noshad, Christian Rose, Robert Thombley, Jonathan Chiang, Conor K. Corbin, Vincent X. Liu, Julia Adler-Milstein, Jonathan H. Chen
AMIA1
2019 Scalable Mutual Information Estimation Using Dependence Graphs
abstract
The Mutual Information (MI) is an often used measure of dependency between two random variables utilized in information theory, statistics and machine learning. Recently several MI estimators have been proposed that can achieve parametric MSE convergence rate. However, most of the previously proposed estimators have high computational complexity of at least O(N2). We propose a unified method for empirical non-parametric estimation of general MI function between random vectors in d based on N i.i.d. samples. The reduced complexity MI estimator, called the ensemble dependency graph estimator (EDGE), combines randomized locality sensitive hashing (LSH), dependency graphs, and ensemble bias-reduction methods. We prove that EDGE achieves optimal computational complexity O(N), and can achieve the optimal parametric MSE rate of O(1/N) if the density is d times differentiable. To the best of our knowledge EDGE is the first non-parametric MI estimator that can achieve parametric MSE rates with linear time complexity. We illustrate the utility of EDGE for the analysis of the information plane (IP) in deep learning. Using EDGE we shed light on a controversy on whether or not the compression property of information bottleneck (IB) in fact holds for ReLu and other rectification functions in deep neural networks (DNN).
Morteza Noshad, Alfred O. Hero III
ICASSP1
2018 Scalable Hash-Based Estimation of Divergence Measures
abstract
We propose a scalable divergence estimation method based on hashing. Consider two continuous random variables $X$ and $Y$ whose densities have bounded support. We consider a particular locality sensitive random hashing, and consider the ratio of samples in each hash bin having non-zero numbers of Y samples. We prove that the weighted average of these ratios over all of the hash bins converges to f-divergences between the two samples sets. We derive the MSE rates for two families of smooth functions; the Hölder smoothness class and differentiable functions. In particular, it is proved that if the density functions have bounded derivatives up to the order $d$, where $d$ is the dimension of samples, the optimal parametric MSE rate of $O(1/N)$ can be achieved. The computational complexity is shown to be $O(N)$, which is optimal. To the best of our knowledge, this is the first empirical divergence estimator that has optimal computational complexity and can achieve the optimal parametric MSE estimation rate of $O(1/N)$.
Morteza Noshad, Alfred O. Hero III
AISTATS1
2018 Rate-Optimal Meta Learning of Classification Error
abstract
Meta learning of optimal classifier error rates allows an experimenter to empirically estimate the intrinsic ability of any estimator to discriminate between two populations, circumventing the difficult problem of estimating the optimal Bayes classifier. To this end we propose a weighted nearest neighbor (WNN) graph estimator for a tight bound on the Bayes classification error; the Henze-Penrose (HP) divergence. Similar to recently proposed HP estimators [1], the proposed estimator is non-parametric and does not require density estimation. However, unlike previous approaches the proposed estimator is rate-optimal, i.e., its mean squared estimation error (MSEE) decays to zero at the fastest possible rate of O(1/M+1/N) where M, N are the sample sizes of the respective populations. We illustrate the proposed WNN meta estimator for several simulated and real data sets.
Morteza Noshad, Alfred O. Hero III
ICASSP1
2017 Information theoretic structure learning with confidence
abstract
Information theoretic measures (e.g. the Kullback Liebler divergence and Shannon mutual information) have been used for exploring possibly nonlinear multivariate dependencies in high dimension. If these dependencies are assumed to follow a Markov factor graph model, this exploration process is called structure discovery. For discrete-valued samples, estimates of the information divergence over the parametric class of multinomial models lead to structure discovery methods whose mean squared error achieves parametric convergence rates as the sample size grows. However, a naive application of this method to continuous nonparametric multivariate models converges much more slowly. In this paper we introduce a new method for nonparametric structure discovery that uses weighted ensemble divergence estimators that achieve parametric convergence rates and obey an asymptotic central limit theorem that facilitates hypothesis testing and other types of statistical validation.
Kevin R. Moon, Morteza Noshad, Salimeh Yasaei Sekeh, Alfred O. Hero III
ICASSP2
2017 Direct estimation of information divergence using nearest neighbor ratios
abstract
We propose a direct estimation method for Rényi and f-divergence measures based on a new graph theoretical interpretation. Suppose that we are given two sample sets X and Y, respectively with N and M samples, where η := M/N is a constant value. Considering the k-nearest neighbor (k-NN) graph of Y in the joint data set (X, Y), we show that the average powered ratio of the number of X points to the number of Y points among all k-NN points is proportional to Rényi divergence of X and Y densities. A similar method can also be used to estimate f-divergence measures. We derive bias and variance rates, and show that for the class of γ-Hölder smooth functions, the estimator achieves the MSE rate of O(N-2γ/(γ+d)). Furthermore, by using a weighted ensemble estimation technique, for density functions with continuous and bounded derivatives of up to the order d, and some extra conditions at the support set boundary, we derive an ensemble estimator that achieves the parametric MSE rate of O(1/N). Our estimator requires no boundary correction, and remarkably, the boundary issues do not show up. Our approach is also more computationally tractable than other competing estimators, which makes them appealing in many practical applications.
Morteza Noshad, Kevin R. Moon, Salimeh Yasaei Sekeh, Alfred O. Hero III
ISIT1
2016 Low-complexity stochastic Generalized Belief Propagation
abstract
The generalized belief propagation (GBP), introduced by Yedidia et al., is an extension of the belief propagation (BP) algorithm, which is widely used in different problems involved in calculating exact or approximate marginals of probability distributions. In many problems, it has been observed that the accuracy of GBP outperforms that of BP considerably. However, due to its generally higher complexity compared to BP, its application is limited in practice. In this paper, we introduce a stochastic version of GBP called stochastic generalized belief propagation (SGBP) that can be considered as an extension to the stochastic BP (SBP) algorithm introduced by Noorshams et al. They have shown that SBP reduces the complexity per iteration of BP by an order of magnitude in alphabet size. In contrast to SBP, SGBP can reduce the computation complexity if certain topological conditions are met by the region graph associated to a graphical model. However, this reduction can be larger than only one order of magnitude in alphabet size. In this paper, we characterize these conditions and the amount of complexity gain that one can obtain by using SGBP. Finally, using similar proof techniques employed by Noorshams et al., for general graphical models satisfy contraction conditions, we prove the asymptotic convergence of SGBP to the unique GBP fixed point, as well as providing non-asymptotic upper bounds on the mean square error and on the high probability error.
Farzin Haddadpour, Mahdi Jafari Siavoshani, Morteza Noshad
ISIT3