Saeed Amizadeh

dblp:48/8399 · DBLP profile ↗
← Back
16ranked-venue papers
7as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 7 first-author · 2 since 2021Databases, data management, data science and information retrieval · 5 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
5 papers
Query processing and optimization · 41% Machine learning and data management · 30% Data mining · 29%
Artificial intelligence
7 papers
Vision and language · 42% Probabilistic and Bayesian machine learning · 20% Knowledge representation and reasoning · 15%
Computer graphics and multimedia
1 paper
Audio and music processing · 100%
Theoretical computer science
1 paper
Automated reasoning and model checking · 50% Computational complexity · 50%

Topics — the 23 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
cross-modal supervision
0.812024
Weakly-supervised Audio Separation via Bi-modal Semantic Similarity · ICLR 2024
Audio and music processing
source separation
0.812024
Weakly-supervised Audio Separation via Bi-modal Semantic Similarity · ICLR 2024
Machine learning and data management
machine learning pipeline
0.512021
WindTunnel: Towards Differentiable ML Pipelines Beyond a Single Modele · Proc. VLDB Endow. 2021
Knowledge, reasoning and agents › Knowledge representation and reasoning
neuro-symbolic reasoning
0.412020
Neuro-Symbolic Visual Reasoning: Disentangling "Visual" from "Reasoning" · ICML 2020
Computational complexity › circuit complexity › boolean circuits
circuit satisfiability
0.412019
Learning To Solve Circuit-SAT: An Unsupervised Differentiable Approach · ICLR (Poster) 2019
Automated reasoning and model checking
satisfiability
0.412019
Learning To Solve Circuit-SAT: An Unsupervised Differentiable Approach · ICLR (Poster) 2019
Query processing and optimization
cardinality estimation
0.312018
Towards a Learning Optimizer for Shared Clouds · Proc. VLDB Endow. 2018
Query processing and optimization › cardinality estimation
learned cardinality estimation
0.312018
Towards a Learning Optimizer for Shared Clouds · Proc. VLDB Endow. 2018
Query processing and optimization › query optimization › learned query optimization
learned query optimizer
0.312018
Towards a Learning Optimizer for Shared Clouds · Proc. VLDB Endow. 2018
Query processing and optimization
query optimization
0.312018
Towards a Learning Optimizer for Shared Clouds · Proc. VLDB Endow. 2018
Machine learning › Representation and self-supervised learning
contrastive learning
0.212024
Weakly-supervised Audio Separation via Bi-modal Semantic Similarity · ICLR 2024
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model
hidden markov model
0.212015
Inertial Hidden Markov Models: Modeling Change in Multivariate Time Series · AAAI 2015
Data mining
anomaly detection
0.212015
Generic and Scalable Framework for Automated Time-series Anomaly Detection · KDD 2015
Data mining
time series analysis
0.212015
Inertial Hidden Markov Models: Modeling Change in Multivariate Time Series · AAAI 2015
Data mining › anomaly detection
time series anomaly detection
0.212015
Generic and Scalable Framework for Automated Time-series Anomaly Detection · KDD 2015
Data mining › time series analysis
time series segmentation
0.212015
Inertial Hidden Markov Models: Modeling Change in Multivariate Time Series · AAAI 2015
Machine learning › Deep learning architectures and training
backpropagation
0.112021
WindTunnel: Towards Differentiable ML Pipelines Beyond a Single Modele · Proc. VLDB Endow. 2021
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
markov random field
0.112010
Latent Variable Model for Learning in Pairwise Markov Networks · AAAI 2010
Machine learning › Probabilistic and Bayesian machine learning › statistical inference
parameter estimation
0.112010
Latent Variable Model for Learning in Pairwise Markov Networks · AAAI 2010
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
structure learning
0.112010
Latent Variable Model for Learning in Pairwise Markov Networks · AAAI 2010
Machine learning and data management
learned database components
0.112018
Towards a Learning Optimizer for Shared Clouds · Proc. VLDB Endow. 2018
Data mining › predictive modeling
forecasting
0.112015
Generic and Scalable Framework for Automated Time-series Anomaly Detection · KDD 2015
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models
0.012010
Latent Variable Model for Learning in Pairwise Markov Networks · AAAI 2010

Methods — techniques the papers use, named apart from their topics

weak supervision · 1.5bi-modal semantic similarity · 1.5CLAP · 1.5gradient boosting trees · 1.0backpropagation · 1.0unsupervised differentiable approach · 0.8graph neural network · 0.8subgraph template learning · 0.7machine learning · 0.7expectation-maximization · 0.4neuro-symbolic reasoning · 0.4disentangled representation learning · 0.4state transition regularization · 0.2
YearPublicationVenuePosition
2025 CorrGAN: Simultaneous Learning of Speech Enhancement and Perceptual Quality Loss Functions
abstract
Deep-learning models have allowed effective end-to-end SE systems in the Speech Enhancement (SE) field. Most of these methods are trained using a fixed reconstruction loss in a supervised setting. Often these losses do not perfectly represent the desired perceptual quality metrics, resulting in sub-optimal performance. Recently, there have been efforts to learn the behavior of those metrics directly via neural nets for training SE models. However, an accurate estimation of the true metric function introduces statistical complexity for training because it attempts to capture the exact value of the metric. We propose an adversarial training strategy based on statistical correlation that avoids the complexity of estimating the SE metric while learning to mimic its overall behavior. We call this framework CorrGAN and show its significant improvement over standard losses of the SOTA baselines and achieve SOTA performance on the VoiceBank+DEMAND dataset.
Vasily Zadorozhnyy, Saeed Amizadeh, Qiang Ye 0003, Kazuhito Koishida
ICASSP2
2025 Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale Problems
abstract
Transformers and their attention mechanism have been revolutionary in the field of Machine Learning. While originally proposed for the language data, they quickly found their way to the image, video, graph, etc. data modalities with various signal geometries. Despite this versatility, generalizing the attention mechanism to scenarios where data is presented at different scales from potentially different modalities is not straightforward. The attempts to incorporate hierarchy and multi-modality within transformers are largely based on ad hoc heuristics, which are not seamlessly generalizable to similar problems with potentially different structures. To address this problem, in this paper, we take a fundamentally different approach: we first propose a mathematical construct to represent multi-modal, multi-scale data. We then mathematically derive the neural attention mechanics for the proposed construct from the first principle of entropy minimization. We show that the derived formulation is optimal in the sense of being the closest to the standard Softmax attention while incorporating the inductive biases originating from the hierarchical/geometric information of the problem. We further propose an efficient algorithm based on dynamic programming to compute our derived attention mechanism. By incorporating it within transformers, we show that the proposed hierarchical attention mechanism not only can be employed to train transformer models in hierarchical/multi-modal settings from scratch, but it can also be used to inject hierarchical information into classical, pre-trained transformer models post training, resulting in more efficient models in zero-shot manner.
Saeed Amizadeh, Sara Abdali, Yinheng Li, Kazuhito Koishida
NeurIPS1
2024 Weakly-supervised Audio Separation via Bi-modal Semantic Similarity
abstract
Conditional sound separation in multi-source audio mixtures without having access to single source sound data during training is a long standing challenge. Existing mix-and-separate based methods suffer from significant performance drop with multi-source training mixtures due to the lack of supervision signal for single source separation cases during training. However, in the case of language-conditional audio separation, we do have access to corresponding text descriptions for each audio mixture in our training data, which can be seen as (rough) representations of the audio samples in the language modality. That raises the curious question of how to generate supervision signal for single-source audio extraction by leveraging the fact that single-source sounding language entities can be easily extracted from the text description. To this end, in this paper, we propose a generic bi-modal separation framework which can enhance the existing unsupervised frameworks to separate single-source signals in a target modality (i.e., audio) using the easily separable corresponding signals in the conditioning modality (i.e., language), without having access to single-source samples in the target modality during training. We empirically show that this is well within reach if we have access to a pretrained joint embedding model between the two modalities (i.e., CLAP). Furthermore, we propose to incorporate our framework into two fundamental scenarios to enhance separation performance. First, we show that our proposed methodology significantly improves the performance of purely unsupervised baselines by reducing the distribution shift between training and test samples. In particular, we show that our framework can achieve 71% boost in terms of Signal-to-Distortion Ratio (SDR) over the baseline, reaching 97.5% of the supervised learning performance. Second, we show that we can further improve the performance of the supervised learning itself by 17% if we augment it by our proposed weakly-supervised framework. Our framework achieves this by making large corpora of unsupervised data available to the supervised learning model as well as utilizing a natural, robust regularization mechanism through weak supervision from the language modality, and hence enabling a powerful semi-supervised framework for audio separation. Code is released at https://github.com/microsoft/BiModalAudioSeparation.
Tanvir Mahmud, Saeed Amizadeh, Kazuhito Koishida, Diana Marculescu
ICLR2
2021 WindTunnel: Towards Differentiable ML Pipelines Beyond a Single Modele
abstract
While deep neural networks (DNNs) have shown to be successful in several domains like computer vision, non-DNN models such as linear models and gradient boosting trees are still considered state-of-the-art over tabular data. When using these models, data scientists often author machine learning (ML) pipelines: DAG of ML operators comprising data transforms and ML models, whereby each operator is sequentially trained one-at-a-time. Conversely, when training DNNs, layers composing the neural networks are simultaneously trained using backpropagation. In this paper, we argue that the training scheme of ML pipelines is sub-optimal because it tries to optimize a single operator at a time thus losing the chance of global optimization. We therefore propose WindTunnel: a system that translates a trained ML pipeline into a pipeline of neural network modules and jointly optimizes the modules using backpropagation. We also suggest translation methodologies for several non-differentiable operators such as gradient boosting trees and categorical feature encoders. Our experiments show that fine-tuning of the translated WindTunnel pipelines is a promising technique able to increase the final accuracy.
Gyeong-In Yu, Saeed Amizadeh, Sehoon Kim 0001, Artidoro Pagnoni, Ce Zhang 0001, Byung-Gon Chun, Markus Weimer, Matteo Interlandi
Proc. VLDB Endow.2
2020 Neuro-Symbolic Visual Reasoning: Disentangling "Visual" from "Reasoning"
Saeed Amizadeh, Hamid Palangi, Oleksandr Polozov, Kazuhito Koishida
ICML1
2019 Learning To Solve Circuit-SAT: An Unsupervised Differentiable Approach
Saeed Amizadeh, Sergiy Matusevych, Markus Weimer
ICLR (Poster)1
2019 Coded Elastic Computing
abstract
Cloud providers have recently introduced new offerings whereby spare computing resources are accessible at discounts compared to on-demand computing. Exploiting such opportunity is challenging inasmuch as such resources are accessed with low-priority and therefore can elastically leave (through preemption) and join the computation at any time. In this paper, we design a new technique called coded elastic computing enabling distributed computations over elastic resources. The proposed technique allows machines to leave the computation without sacrificing the algorithm-level performance, and, at the same time, flexibly reduce the workload at existing machines when new ones join the computation. Leveraging coded redundancy, our approach is able to achieve similar computational cost as the original (uncoded) method when all machines are present; the cost gracefully increases when machines are preempted and reduces when machines join. The performance of the proposed technique is evaluated on matrix-vector multiplication and linear regression tasks, and shows improvements over existing techniques.
Yaoqing Yang 0002, Matteo Interlandi, Pulkit Grover, Soummya Kar, Saeed Amizadeh, Markus Weimer
ISIT5
2019 Machine Learning at Microsoft with ML.NET
abstract
Machine Learning is transitioning from an art and science into a technology available to every developer. In the near future, every application on every platform will incorporate trained models to encode data-based decisions that would be impossible for developers to author. This presents a significant engineering challenge, since currently data science and modeling are largely decoupled from standard software development processes. This separation makes incorporating machine learning capabilities inside applications unnecessarily costly and difficult, and furthermore discourage developers from embracing ML in first place. In this paper we present ML.NET, a framework developed at Microsoft over the last decade in response to the challenge of making it easy to ship machine learning models in large software applications. We present its architecture, and illuminate the application demands that shaped it. Specifically, we introduce DataView, the core data abstraction of ML.NET which allows it to capture full predictive pipelines efficiently and consistently across training and inference lifecycles. We close the paper with a surprisingly favorable performance study of ML.NET compared to more recent entrants, and a discussion of some lessons learned.
Saeed Amizadeh, Mikhail Bilenko, Rogan Carr, Wei-Sheng Chin, Yael Dekel, Xavier Dupré, Vadim Eksarevskiy, Senja Filipi, Tom Finley, Abhishek Goswami, Monte Hoover, Scott Inglis, Matteo Interlandi, Najeeb Kazmi, Gleb Krivosheev, Pete Luferenko, Ivan Matantsev, Sergiy Matusevych, Shahab Moradi, Gani Nazirov, Justin Ormont, Gal Oshri, Artidoro Pagnoni, Jignesh Parmar, Prabhat Roy, Mohammad Zeeshan Siddiqui, Markus Weimer, Shauheen Zahirazami
KDD2
2018 Towards a Learning Optimizer for Shared Clouds
abstract
Query optimizers are notorious for inaccurate cost estimates, leading to poor performance. The root of the problem lies in inaccurate cardinality estimates, i.e., the size of intermediate (and final) results in a query plan. These estimates also determine the resources consumed in modern shared cloud infrastructures. In this paper, we present C ARD L EARNER , a machine learning based approach to learn cardinality models from previous job executions and use them to predict the cardinalities in future jobs. The key intuition in our approach is that shared cloud workloads are often recurring and overlapping in nature, and so we could learn cardinality models for overlapping subgraph templates. We discuss various learning approaches and show how learning a large number of smaller models results in high accuracy and explainability. We further present an exploration technique to avoid learning bias by considering alternate join orders and learning cardinality models over them. We describe the feedback loop to apply the learned models back to future job executions. Finally, we show a detailed evaluation of our models (up to 5 orders of magnitude less error), query plans (60% applicability), performance (up to 100% faster, 3x fewer resources), and exploration (optimal in few 10s of executions).
Chenggang Wu 0001, Alekh Jindal, Saeed Amizadeh, Hiren Patel, Wangchao Le, Shi Qiao 0001, Sriram Rao
Proc. VLDB Endow.3
2015 Inertial Hidden Markov Models: Modeling Change in Multivariate Time Series
abstract
Faced with the problem of characterizing systematic changes in multivariate time series in an unsupervised manner, we derive and test two methods of regularizing hidden Markov models for this task. Regularization on state transitions provides smooth transitioning among states, such that the sequences are split into broad, contiguous segments. Our methods are compared with a recent hierarchical Dirichlet process hidden Markov model (HDP-HMM) and a baseline standard hidden Markov model, of which the former suffers from poor performance on moderate-dimensional data and sensitivity to parameter settings, while the latter suffers from rapid state transitioning, over-segmentation and poor performance on a segmentation task involving human activity accelerometer data from the UCI Repository. The regularized methods developed here are able to perfectly characterize change of behavior in the human activity data for roughly half of the real-data test cases, with accuracy of 94% and low variation of information. In contrast to the HDP-HMM, our methods provide simple, drop-in replacements for standard hidden Markov model update rules, allowing standard expectation maximization (EM) algorithms to be used for learning.
George D. Montañez, Saeed Amizadeh, Nikolay Laptev
AAAI2
2015 Generic and Scalable Framework for Automated Time-series Anomaly Detection
abstract
This paper introduces a generic and scalable framework for automated anomaly detection on large scale time-series data. Early detection of anomalies plays a key role in maintaining consistency of person's data and protects corporations against malicious attackers. Current state of the art anomaly detection approaches suffer from scalability, use-case restrictions, difficulty of use and a large number of false positives. Our system at Yahoo, EGADS, uses a collection of anomaly detection and forecasting models with an anomaly filtering layer for accurate and scalable anomaly detection on time-series. We compare our approach against other anomaly detection systems on real and synthetic data with varying time-series characteristics. We found that our framework allows for 50-60% improvement in precision and recall for a variety of use-cases. Both the data and the framework are being open-sourced. The open-sourcing of the data, in particular, represents the first of its kind effort to establish the standard benchmark for anomaly detection.
Nikolay Laptev, Saeed Amizadeh, Ian Flint
KDD2
2013 The Bregman Variational Dual-Tree Framework
Saeed Amizadeh, Bo Thiesson, Milos Hauskrecht
UAI1
2012 Sampling Strategies to Evaluate the Performance of Unknown Predictors
abstract
The focus of this paper is on how to select a small sample of examples for labeling that can help us to evaluate many different classification models unknown at the time of sampling. We are particularly interested in studying the sampling strategies for problems in which the prevalence of the two classes is highly biased toward one of the classes. The evaluation measures of interest we want to estimate as accurately as possible are those obtained from the contingency table. We provide a careful theoretical analysis on sensitivity, specificity, and precision and show how sampling strategies should be adapted to the rate of skewness in data in order to effectively compute the three aforementioned evaluation measures.
Hamed Valizadegan, Saeed Amizadeh, Milos Hauskrecht
SDM2
2012 Variational Dual-Tree Framework for Large-Scale Transition Matrix Approximation
Saeed Amizadeh, Bo Thiesson, Milos Hauskrecht
UAI1
2011 An Efficient Framework for Constructing Generalized Locally-Induced Text Metrics
Saeed Amizadeh, Shuguang Wang, Milos Hauskrecht
IJCAI1
2010 Latent Variable Model for Learning in Pairwise Markov Networks
abstract
Pairwise Markov Networks (PMN) are an important class of Markov networks which, due to their simplicity, are widely used in many applications such as image analysis, bioinformatics, sensor networks, etc. However, learning of Markov networks from data is a challenging task; there are many possible structures one must consider and each of these structures comes with its own parameters making it easy to overfit the model with limited data. To deal with the problem, recent learning methods build upon the L1 regularization to express the bias towards sparse network structures. In this paper, we propose a new and more flexible framework that let us bias the structure, that can, for example, encode the preference to networks with certain local substructures which as a whole exhibit some special global structure. We experiment with and show the benefit of our framework on two types of problems: learning of modular networks and learning of traffic networks models.
Saeed Amizadeh, Milos Hauskrecht
AAAI1