Vahid Noroozi

dblp:18/9437 · DBLP profile ↗
← Back
18ranked-venue papers
5as first author
6since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorTheory of computation · 1
YearPublicationVenuePosition
2026 Scaling Test-Time Compute to Achieve IOI Gold Medal with Open-Weight Models
abstract
Mehrzad Samadi, Aleksander Ficek, Sean Narenthiran, Siddhartha Jain, Wasi Uddin Ahmad, Somshubra Majumdar, Vahid Noroozi, Boris Ginsburg. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Mehrzad Samadi, Aleksander Ficek, Sean Narenthiran, Siddhartha Jain 0001, Wasi Uddin Ahmad, Somshubra Majumdar, Vahid Noroozi, Boris Ginsburg
ACL (1)7
2024 Investigating End-to-End ASR Architectures for Long Form Audio Transcription
abstract
This paper presents an overview and evaluation of some of the end-to-end ASR models on long-form audio. We study three categories of Automatic Speech Recognition(ASR) models based on their core architecture: (1) convolutional, (2) convolutional with squeeze-and-excitation, and (3) convolutional models with attention. We selected one ASR model from each category and evaluated the Word Error Rate, maximum audio length and real-time factor for each model on a variety of long audio benchmarks: Earnings-21 and 22, CORAAL, and TED-LIUM3. The model from the category of self-attention with local attention and global token has the best accuracy compared to other architectures. We also compared models with CTC and RNNT decoders and showed that CTC-based models are more robust and efficient than RNNT on long form audio.
Nithin Rao Koluguri, Samuel Kriman, Georgy Zelenfroind, Somshubra Majumdar, Dima Rekesh, Vahid Noroozi, Jagadeesh Balam, Boris Ginsburg
ICASSP6
2024 Stateful Conformer with Cache-Based Inference for Streaming Automatic Speech Recognition
abstract
In this paper, we propose an efficient and accurate streaming speech recognition model based on the FastConformer architecture. We adapted the FastConformer architecture for streaming applications through: (1) constraining both the look-ahead and past contexts in the encoder, and (2) introducing an activation caching mechanism to enable the non-autoregressive encoder to operate autoregressively during inference. The proposed model is thoughtfully designed in a way to eliminate the accuracy disparity between the train and inference time which is common for many streaming models. Furthermore, our proposed encoder works with various decoder configurations including Connectionist Temporal Classification (CTC) and RNN-Transducer (RNNT) decoders. We evaluate the proposed model and demonstrate that it can achieve better accuracy with lower latency and inference time compared to a conventional buffered streaming model baseline.
Vahid Noroozi, Somshubra Majumdar, Jagadeesh Balam, Boris Ginsburg
ICASSP1
2024 Instruction Data Generation and Unsupervised Adaptation for Speech Language Models
Vahid Noroozi, Zhehuai Chen, Somshubra Majumdar, Steve Huang, Jagadeesh Balam, Boris Ginsburg
INTERSPEECH1
2023 Fast Conformer With Linearly Scalable Attention For Efficient Speech Recognition
abstract
Conformer-based models have become the dominant end-to-end architecture for speech processing tasks. With the objective of enhancing the conformer architecture for efficient training and inference, we carefully redesigned Conformer with a novel downsampling schema. The proposed model, named Fast Conformer(FC), is 2.8 × faster than the original Conformer, supports scaling to Billion parameters without any changes to the core architecture and also achieves state-of-the-art accuracy on Automatic Speech Recognition benchmarks. To enable transcription of long-form speech up to 11 hours, we replaced global attention with limited context attention post-training, while also improving accuracy through fine-tuning with the addition of a global token. Fast Conformer, when combined with a Transformer decoder also outperforms the original Conformer in accuracy and in speed for Speech Translation and Spoken Language Understanding.
Dima Rekesh, Nithin Rao Koluguri, Samuel Kriman, Somshubra Majumdar, Vahid Noroozi, He Huang 0012, Oleksii Hrinchuk, Krishna C. Puvvada, Jagadeesh Balam, Boris Ginsburg
ASRU5
2021 SPGISpeech: 5, 000 Hours of Transcribed Financial Audio for Fully Formatted End-to-End Speech Recognition
abstract
In the English speech-to-text (STT) machine learning task, acoustic models are conventionally trained on uncased Latin characters, and any necessary orthography (such as capitalization, punctuation, and denormalization of non-standard words) is imputed by separate post-processing models.This adds complexity and limits performance, as many formatting tasks benefit from semantic information present in the acoustic signal but absent in transcription.Here we propose a new STT task: endto-end neural transcription with fully formatted text for target labels.We present baseline Conformer-based models trained on a corpus of 5,000 hours of professionally transcribed earnings calls, achieving a CER of 1.7.As a contribution to the STT research community, we release the corpus free for noncommercial use. 1
Patrick K. O'Neill, Vitaly Lavrukhin, Somshubra Majumdar, Vahid Noroozi, Yuekai Zhang, Oleksii Kuchaiev, Jagadeesh Balam, Yuliya Dovzhenko, Keenan Freyberg, Michael D. Shulman, Boris Ginsburg, Shinji Watanabe 0001, Georg Kucsko
Interspeech4
2020 Real-World Multi-Domain Data Applications for Generalizations to Clinical Settings
abstract
With promising results of machine learning based models in computer vision, applications on medical imaging data have been increasing exponentially. Deep learning models perform well when trained on standardized datasets from artificial settings. However, generalization and translation to real-world clinical settings are challenging. The complexity of real-world applications in healthcare emanates from different data distributions across multiple device domains, variations in image resolution, human errors, and the lack of manual grading. Moreover, healthcare applications not only suffer from scarcity in labeled data, but also face limited access to un-labeled data. These limitations pose additional challenges to developing translatable applications for clinical care. In this paper, we utilize self-supervised representation learning methods, formulated effectively in transfer learning settings, to address limited data availability and assess the importance of real-world data for generalizations to clinical settings. We show that by employing a self-supervised approach with transfer learning on a multi-domain real-world dataset, we can achieve 16% relative improvement on a standardized dataset over supervised baselines.
Nooshin Mojab, Vahid Noroozi, Darvin Yi, Manoj Prabhakar Nallabothula, Abdullah Aleem, Philip S. Yu, Joelle A. Hallak
ICMLA2
2019 MARS: Memory Attention-Aware Recommender System
abstract
In this paper, we study the problem of modeling users' diverse interests. Previous methods usually learn a fixed user representation, which has a limited ability to represent distinct interests of a user. In order to model users' various interests, we propose a Memory Attention-aware Recommender System (MARS). MARS utilizes a memory component and a novel attentional mechanism to learn deep adaptive user representations. Trained in an end-to-end fashion, MARS adaptively summarizes users' interests. In the experiments, MARS outperforms seven state-of-the-art methods on three real-world datasets in terms of recall and mean average precision. We also demonstrate that MARS has a great interpretability to explain its recommendation results, which is important in many recommendation scenarios.
Lei Zheng 0001, Chun-Ta Lu, Lifang He 0001, Sihong Xie, He Huang 0008, Chaozhuo Li, Vahid Noroozi, Philip S. Yu
DSAA7
2019 Leveraging Semi-Supervised Learning for Fairness using Neural Networks
abstract
There has been a growing concern about the fairness of decision-making systems based on machine learning. The shortage of labeled data has been always a challenging problem facing machine learning based systems. In such scenarios, semi-supervised learning has shown to be an effective way of exploiting unlabeled data to improve upon the performance of model. Notably, unlabeled data do not contain label information which itself can be a significant source of bias in training machine learning systems. This inspired us to tackle the challenge of fairness by formulating the problem in a semi-supervised framework. In this paper, we propose a semi-supervised algorithm using neural networks benefiting from unlabeled data to not just improve the performance but also improve the fairness of the decision-making process. The proposed model, called SSFair, exploits the information in the unlabeled data to mitigate the bias in the training data.
Vahid Noroozi, Sara Bahaadini, Samira Sheikhi, Nooshin Mojab, Philip S. Yu
ICMLA1
2018 DeepFP: A Deep Learning Framework For User Fingerprinting via Mobile Motion Sensors
abstract
In this paper, we propose a deep learning framework for user fingerprinting via mobile motion sensors, DeepFP, which can identify and track users based on their behavioral patterns while interacting with the smartphone. Existing machine learning techniques for user identification are classification-oriented and thus are not amenable easily to large-scale, real world deployment. They need to be trained on all the users whom they want to identify. DeepFP exploits metric learning techniques and deep neural networks to address the challenges of current user identification techniques. We leverage feature embedding to directly extract informative features and map input samples to a discriminative lower-dimensional space, where recurrent neural networks are used to model the temporal information of data. DeepFP does not need to re-train to identify new users which makes it feasible to be used in real world scenarios with a huge number of users, without needing a large number of training samples. Experiments on a publicly available mobile sensors dataset and comparison with other embedding methods depict the effectiveness of DeepFP.
Sara Amini, Vahid Noroozi, Sara Bahaadini, Philip S. Yu, Chris Kanich
IEEE BigData2
2018 Semi-supervised Deep Representation Learning for Multi-View Problems
abstract
While neural networks for learning representation of multi-view data have been previously proposed as one of the state-of-the-art multi-view dimension reduction techniques, how to make the representation discriminative with only a small amount of labeled data is not well-studied. We introduce a semi-supervised neural network model, named Multi-view Discriminative Neural Network (MDNN), for multi-view problems. MDNN finds nonlinear view-specific mappings by projecting samples to a common feature space using multiple coupled deep networks. It is capable of leveraging both labeled and unlabeled data to project multi-view data so that samples from different classes are separated and those from the same class are clustered together. It also uses the inter-view correlation between views to exploit the available information in both the labeled and unlabeled data. Extensive experiments conducted on four datasets demonstrate the effectiveness of the proposed algorithm for multi-view semi-supervised learning.
Vahid Noroozi, Sara Bahaadini, Lei Zheng 0001, Sihong Xie, Weixiang Shao, Philip S. Yu
IEEE BigData1
2018 DeepAuth: A Framework for Continuous User Re-authentication in Mobile Apps
abstract
With the increasing volume of transactions taking place online, mobile fraud has also increased. Mobile applications often authenticate the user only at install time. The user may then remain logged in for hours or weeks. Any unauthorized access may lead to financial, criminal or privacy losses. In this work, we leverage currently available built-in motion sensors in smartphones to learn users' behavioral characteristics while interacting with the mobile device to provide an implicit re-authentication mechanism that enables a frictionless and secure user experience in the application. This approach improves the generality as well as power efficiency of the authentication mechanism compared to using the camera feed which involves (a) specific hardware, (b) higher battery usage and (c) privacy concerns. We present DeepAuth as a generic framework for re-authenticating users in a mobile app. In our approach, we use time and frequency domain features extracted from motion sensors and a long short-term memory (LSTM) model with negative sampling to build a re-authentication framework. The framework is able to re-authenticate a user with 96.70% accuracy in 20 seconds from a set of data collected from 47 volunteers.
Sara Amini, Vahid Noroozi, Amit Pande, Satyajit Gupte, Philip S. Yu, Chris Kanich
CIKM2
2018 Direct: Deep Discriminative Embedding for Clustering of Ligo Data
abstract
In this paper, benefiting from the strong ability of deep neural network in estimating non-linear functions, we propose a discriminative embedding function to be used as a feature extractor for clustering tasks. The trained embedding function transfers knowledge from the domain of a labeled set of morphologically-distinct images, known as classes, to a new domain within which new classes can potentially be isolated and identified. Our target application in this paper is the Gravity Spy Project, which is an effort to characterize transient, non-Gaussian noise present in data from the Advanced Laser Interferometer Gravitational-wave Observatory, or LIGO. Accumulating large, labeled sets of noise features and identifying of new classes of noise lead to a better understanding of their origin, which makes their removal from the data and/or detectors possible.
Sara Bahaadini, Neda Rohani, Aggelos K. Katsaggelos, Vahid Noroozi, Scott Coughlin, Michael Zevin
ICIP4
2018 Machine learning for Gravity Spy: Glitch classification and dataset
Sara Bahaadini, Vahid Noroozi, Neda Rohani, Scott Coughlin, Michael Zevin, Joshua R. Smith 0003, Vicky Kalogera, Aggelos K. Katsaggelos
Inf. Sci.2
2017 Hierarchical collaborative embedding for context-aware recommendations
abstract
In a variety of recommender systems, items, such as news or articles, are associated with text. Most of previous recommender systems learn item embeddings from the textual content by utilizing the bag-of-words technique. However, due to its limited ability to capture semantic meanings in the text, these methods lead to the shallow modeling of items. Recently proposed deep learning based methods try to overcome the limitation by leveraging Recurrent Neural Networks (RNN). Suffering from the problem of modeling long sequences for RNN, these methods are unable to effectively model items based on their textual content as well. In this paper, in order to overcome aforementioned limitations and accurately capture semantic meanings within the textual content, we propose Hierarchical Collaborative Embedding (HCE). HCE tightly couples a Hierarchical Recurrent Network (HRN) with Probabilistic Matrix Factorization (PMF) to provide top-N ranking lists of items for users. In the experiments, we show that HCE beats strong baselines by a wide margin on three real-world datasets.
Lei Zheng 0001, Bokai Cao, Vahid Noroozi, Philip S. Yu, Nianzu Ma
IEEE BigData3
2017 SEVEN: Deep Semi-supervised Verification Networks
abstract
Verification determines whether two samples belong to the same class or not, and has important applications such as face and fingerprint verification, where thousands or millions of categories are present but each category has scarce labeled examples, presenting two major challenges for existing deep learning models. We propose a deep semi-supervised model named SEmi-supervised VErification Network (SEVEN) to address these challenges. The model consists of two complementary components. The generative component addresses the lack of supervision within each category by learning general salient structures from a large amount of data across categories. The discriminative component exploits the learned general features to mitigate the lack of supervision within categories, and also directs the generative component to find more informative structures of the whole data manifold. The two components are tied together in SEVEN to allow an end-to-end training of the two components. Extensive experiments on four verification tasks demonstrate that SEVEN significantly outperforms other state-of-the-art deep semi-supervised techniques when labeled data are in short supply. Furthermore, SEVEN is competitive with fully supervised baselines trained with a larger amount of labeled data. It indicates the importance of the generative component in SEVEN.
Vahid Noroozi, Lei Zheng 0001, Sara Bahaadini, Sihong Xie, Philip S. Yu
IJCAI1
2017 Joint Deep Modeling of Users and Items Using Reviews for Recommendation
abstract
A large amount of information exists in reviews written by users. This source of information has been ignored by most of the current recommender systems while it can potentially alleviate the sparsity problem and improve the quality of recommendations. In this paper, we present a deep model to learn item properties and user behaviors jointly from review text. The proposed model, named Deep Cooperative Neural Networks (DeepCoNN), consists of two parallel neural networks coupled in the last layers. One of the networks focuses on learning user behaviors exploiting reviews written by the user, and the other one learns item properties from the reviews written for the item. A shared layer is introduced on the top to couple these two networks together. The shared layer enables latent factors learned for users and items to interact with each other in a manner similar to factorization machine techniques. Experimental results demonstrate that DeepCoNN significantly outperforms all baseline recommender systems on a variety of datasets.
Lei Zheng 0001, Vahid Noroozi, Philip S. Yu
WSDM2
2012 Two phased cellular PSO: A new collaborative cellular algorithm for optimization in dynamic environments
abstract
Many real world optimization problems are dynamic in which the fitness landscape is time dependent and the optima change over time such as dynamic economic modeling, dynamic resource scheduling, and dynamic vehicle routing. Such problems challenge traditional optimization methods as well as conventional evolutionary optimization algorithms. For such environments, optimization algorithms not only have to find the global optimum but also closely track its trajectory. In this paper, we propose a collaborative version of cellular PSO, named Two Phased cellular PSO to address dynamic optimization problems. The proposed algorithm introduces two search phases in order to create a more efficient balance between exploration and exploitation in cellular PSO. The conventional PSO in cellular PSO is replaced by a proposed PSO to increase the exploration capability and an exploitation phase is added to increase exploitation is the promising cells. Moreover, the cell capacity threshold which is a key parameter of cellular PSO is eliminated due to these modifications. To demonstrate the performance and robustness of the proposed algorithm, it is evaluated in various dynamic environment modeled by Moving Peaks Benchmark. The results show that for all the experimented dynamic environments, TP-CPSO outperforms all compared algorithms including cellular PSO.
Ali Sharifi, Vahid Noroozi, Masoud Bashiri, Ali B. Hashemi, Mohammad Reza Meybodi
IEEE Congress on Evolutionary Computation2