VLDB 2026 Research / reviewers in the wild / expert
Masoud Faraki
dblp:143/9779
· DBLP profile ↗
18ranked-venue papers
10as first author
8since 2021 · last 2025
0000-0002-0304-993XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 7 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 7 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PaVeRL-SQL: Text-to-SQL via Partial-Match Rewards and Verbal Reinforcement Learning
Heng Hao, Oxana Verkholyak, D. Ataee Tarzanagh, Baruch Gutow, Sima Didari, Masoud Faraki, Hankyu Moon, Seungjai Min |
IEEE Big Data | 7 |
| 2023 | Domain Generalization Guided by Gradient Signal to Noise Ratio of ParametersabstractOverfitting to the source domain is a common issue in gradient-based training of deep neural networks. To compensate for the over-parameterized models, numerous regularization techniques have been introduced such as those based on dropout. While these methods achieve significant improvements on classical benchmarks such as ImageNet, their performance diminishes with the introduction of domain shift in the test set i.e. when the unseen data comes from a significantly different distribution. In this paper, we move away from the classical approach of Bernoulli sampled dropout mask construction and propose to base the selection on gradient-signal-to-noise ratio (GSNR) of network’s parameters. Specifically, at each training step, parameters with high GSNR will be discarded. Furthermore, we alleviate the burden of manually searching for the optimal dropout ratio by leveraging a meta-learning approach. We evaluate our method on standard domain generalization benchmarks and achieve competitive results on classification and face anti-spoofing problems. Mateusz Michalkiewicz, Masoud Faraki, Xiang Yu 0002, Manmohan Krishna Chandraker, Mahsa Baktash |
ICCV | 2 |
| 2023 | Split to Learn: Gradient Split for Multi-Task Human Image AnalysisabstractThis paper presents an approach to train a unified deep network that simultaneously solves multiple human-related tasks. A multi-task framework is favorable for sharing information across tasks under restricted computational resources. However, tasks not only share information but may also compete for resources and conflict with each other, making the optimization of shared parameters difficult and leading to suboptimal performance. We propose a simple but effective training scheme called GradSplit that alleviates this issue by utilizing asymmetric inter-task relations. Specifically, at each convolution module, it splits features into T groups for T tasks and trains each group only using the gradient back-propagated from the task losses with which it does not have conflicts. During training, we apply GradSplit to a series of convolution modules. As a result, each module is trained to generate a set of task-specific features using the shared features from the previous module. This enables a network to use complementary information across tasks while circumventing gradient conflicts. Experimental results show that GradSplit achieves a better accuracy-efficiency trade-off than existing methods. It minimizes accuracy drop caused by task conflicts while significantly saving compute resources in terms of both FLOPs and memory at inference. We further show that GradSplit achieves higher cross-dataset accuracy compared to single-task and other multi-task networks. Weijian Deng, Yumin Suh, Xiang Yu 0002, Masoud Faraki, Liang Zheng 0001, Manmohan Krishna Chandraker |
WACV | 4 |
| 2022 | Learning to Learn across Diverse Data Biases in Deep Face RecognitionabstractConvolutional Neural Networks have achieved remarkable success in face recognition, in part due to the abundant availability of data. However, the data used for training CNNs is often imbalanced. Prior works largely focus on the long-tailed nature of face datasets in data volume per identity, or focus on single bias variation. In this paper, we show that many bias variations such as ethnicity, head pose, occlusion and blur can jointly affect the accuracy significantly. We propose a sample level weighting approach termed Multi-variation Cosine Margin (MvCoM), to simultaneously consider the multiple variation factors, which orthogonally enhances the face recognition losses to incorporate the importance of training samples. Further, we leverage a learning to learn approach, guided by a held-out meta learning set and use an additive modeling to predict the MvCoM. Extensive experiments on challenging face recognition benchmarks demonstrate the advantages of our method in jointly handling imbalances due to multiple variations. Chang Liu 0022, Xiang Yu 0002, Yi-Hsuan Tsai, Masoud Faraki, Ramin Moslemi, Manmohan Krishna Chandraker, Yun Fu 0001 |
CVPR | 4 |
| 2022 | Controllable Dynamic Multi-Task ArchitecturesabstractMulti-task learning commonly encounters competition for resources among tasks, specifically when model capac-ity is limited. This challenge motivates models which al-low control over the relative importance of tasks and total compute cost during inference time. In this work, we pro-pose such a controllable multi-task network that dynami-cally adjusts its architecture and weights to match the de-sired task preference as well as the resource constraints. In contrast to the existing dynamic multi-task approaches that adjust only the weights within a fixed architecture, our approach affords the flexibility to dynamically control the total computational cost and match the user-preferred task importance better. We propose a disentangled training of two hype rnetwo rks, by exploiting task affinity and a novel branching regularized loss, to take input prefer-ences and accordingly predict tree-structured models with adapted weights. Experiments on three multi-task bench-marks, namely PASCAL-Context, NYU-v2, and CIFAR-100, show the efficacy of our approach. Project page is available at https://www.nec-labs.com/-mas/DYMU. Dripta S. Raychaudhuri, Yumin Suh, Samuel Schulter, Xiang Yu 0002, Masoud Faraki, Amit K. Roy-Chowdhury, Manmohan Krishna Chandraker |
CVPR | 5 |
| 2022 | On Generalizing Beyond Domains in Cross-Domain Continual LearningabstractHumans have the ability to accumulate knowledge of new tasks in varying conditions, but deep neural networks of-ten suffer from catastrophic forgetting of previously learned knowledge after learning a new task. Many recent methods focus on preventing catastrophic forgetting under the assumption of train and test data following similar distributions. In this work, we consider a more realistic scenario of continual learning under domain shifts where the model must generalize its inference to an unseen domain. To this end, we encourage learning semantically meaningful features by equipping the classifier with class similarity metrics as learning parameters which are obtained through Mahalanobis similarity computations. Learning of the backbone representation along with these extra parameters is done seamlessly in an end-to-end manner. In addition, we propose an approach based on the exponential moving average of the parameters for better knowledge distillation. We demonstrate that, to a great extent, existing continual learning algorithms fail to handle the forgetting issue under multiple distributions, while our proposed approach learns new tasks under domain shift with accuracy boosts up to 10% on challenging datasets such as DomainNet and OfficeHome. Christian Simon, Masoud Faraki, Yi-Hsuan Tsai, Xiang Yu 0002, Samuel Schulter, Yumin Suh, Mehrtash Harandi, Manmohan Krishna Chandraker |
CVPR | 2 |
| 2022 | Learning Semantic Segmentation from Multiple Datasets with Label Shifts
Dongwan Kim, Yi-Hsuan Tsai, Yumin Suh, Masoud Faraki, Sparsh Garg, Manmohan Krishna Chandraker, Bohyung Han |
ECCV (28) | 4 |
| 2021 | Cross-Domain Similarity Learning for Face Recognition in Unseen DomainsabstractFace recognition models trained under the assumption of identical training and test distributions often suffer from poor generalization when faced with unknown variations, such as a novel ethnicity or unpredictable individual make-ups during test time. In this paper, we introduce a novel cross-domain metric learning loss, which we dub Cross-Domain Triplet (CDT) loss, to improve face recognition in unseen domains. The CDT loss encourages learning semantically meaningful features by enforcing compact feature clusters of identities from one domain, where the compactness is measured by underlying similarity metrics that belong to another training domain with different statistics. Intuitively, it discriminatively correlates explicit metrics derived from one domain, with triplet samples from another domain in a unified loss function to be minimized within a network, which leads to better alignment of the training domains. The network parameters are further enforced to learn generalized features under domain shift, in a model-agnostic learning pipeline. Unlike the recent work of Meta Face Recognition [18], our method does not require careful hard-pair sample mining and filtering strategy during training. Extensive experiments on various face recognition benchmarks show the superiority of our method in handling variations, compared to baseline and the state-of-the-art methods. Masoud Faraki, Xiang Yu 0002, Yi-Hsuan Tsai, Yumin Suh, Manmohan Krishna Chandraker |
CVPR | 1 |
| 2019 | Learning Factorized Representations for Open-Set Domain Adaptation
Mahsa Baktash, Masoud Faraki, Tom Drummond, Mathieu Salzmann |
ICLR (Poster) | 2 |
| 2018 | Large-Scale Metric Learning: A Voyage From Shallow to DeepabstractDespite its attractive properties, the performance of the recently introduced Keep It Simple and Straightforward MEtric learning (KISSME) method is greatly dependent on principal component analysis as a preprocessing step. This dependence can lead to difficulties, e.g., when the dimensionality is not meticulously set. To address this issue, we devise a unified formulation for joint dimensionality reduction and metric learning based on the KISSME algorithm. Our joint formulation is expressed as an optimization problem on the Grassmann manifold, and hence enjoys the properties of Riemannian optimization techniques. Following the success of deep learning in recent years, we also devise end-to-end learning of a generic deep network for metric learning using our derivation. Masoud Faraki, Mehrtash Harandi, Fatih Porikli |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2018 | A Comprehensive Look at Coding Techniques on Riemannian ManifoldsabstractCore to many learning pipelines is visual recognition such as image and video classification. In such applications, having a compact yet rich and informative representation plays a pivotal role. An underlying assumption in traditional coding schemes [e.g., sparse coding (SC)] is that the data geometrically comply with the Euclidean space. In other words, the data are presented to the algorithm in vector form and Euclidean axioms are fulfilled. This is of course restrictive in machine learning, computer vision, and signal processing, as shown by a large number of recent studies. This paper takes a further step and provides a comprehensive mathematical framework to perform coding in curved and non-Euclidean spaces, i.e., Riemannian manifolds. To this end, we start by the simplest form of coding, namely, bag of words. Then, inspired by the success of vector of locally aggregated descriptors in addressing computer vision problems, we will introduce its Riemannian extensions. Finally, we study Riemannian form of SC, locality-constrained linear coding, and collaborative coding. Through rigorous tests, we demonstrate the superior performance of our Riemannian coding schemes against the state-of-the-art methods on several visual classification tasks, including head pose classification, video-based face recognition, and dynamic scene recognition. Masoud Faraki, Mehrtash Harandi, Fatih Porikli |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2017 | No fuss metric learning, a Hilbert space scenario
Masoud Faraki, Mehrtash Harandi, Fatih Porikli |
Pattern Recognit. Lett. | 1 |
| 2016 | Image set classification by symmetric positive semi-definite matricesabstractRepresenting images and videos by covariance descriptors and leveraging the inherent manifold structure of Symmetric Positive Definite (SPD) matrices leads to enhanced performances in various visual recognition tasks. However, when covariance descriptors are used to represent image sets, the result is often rank-deficient. Thus, most existing approaches adhere to blind perturbation with predefined regularizers just to be able to employ inference tools. To overcome this problem, we introduce novel similarity measures specifically designed for rank-deficient covariance descriptors, i.e., symmetric positive semi-definite matrices. In particular, we derive positive definite kernels that can be decomposed into the kernels on the cone of SPD matrices and kernels on the Grassmann manifolds. Our experiments evidence that, our method achieves superior results for image set classification on various recognition tasks including hand gesture classification, face recognition from video sequences, and dynamic scene categorization. Masoud Faraki, Mehrtash Harandi, Fatih Porikli |
WACV | 1 |
| 2015 | More about VLAD: A leap from Euclidean to Riemannian manifoldsabstractThis paper takes a step forward in image and video coding by extending the well-known Vector of Locally Aggregated Descriptors (VLAD) onto an extensive space of curved Riemannian manifolds. We provide a comprehensive mathematical framework that formulates the aggregation problem of such manifold data into an elegant solution. In particular, we consider structured descriptors from visual data, namely Region Covariance Descriptors and linear subspaces that reside on the manifold of Symmetric Positive Definite matrices and the Grassmannian manifolds, respectively. Through rigorous experimental validation, we demonstrate the superior performance of this novel Riemannian VLAD descriptor on several visual classification tasks including video-based face recognition, dynamic scene recognition, and head pose classification. Masoud Faraki, Mehrtash Harandi, Fatih Porikli |
CVPR | 1 |
| 2015 | Approximate infinite-dimensional Region Covariance Descriptors for image classificationabstractWe introduce methods to estimate infinite-dimensional Region Covariance Descriptors (RCovDs) by exploiting two feature mappings, namely random Fourier features and the Nyström method. In general, infinite-dimensional RCovDs offer better discriminatory power over their low-dimensional counterparts. However, the underlying Riemannian structure, i.e., the manifold of Symmetric Positive Definite (SPD) matrices, is out of reach to great extent for infinite-dimensional RCovDs. To overcome this difficulty, we propose to approximate the infinite-dimensional RCovDs by making use of the aforementioned explicit mappings. We will empirically show that the proposed finite-dimensional approximations of infinite-dimensional RCovDs consistently outperform the low-dimensional RCovDs for image classification task, while enjoying the Riemannian structure of the SPD manifolds. Moreover, our methods achieve the state-of-the-art performance on three different image classification tasks. Masoud Faraki, Mehrtash Harandi, Fatih Porikli |
ICASSP | 1 |
| 2015 | Material Classification on Symmetric Positive Definite ManifoldsabstractThis paper tackles the problem of categorizing materials and textures by exploiting the second order statistics. To this end, we introduce the Extrinsic Vector of Locally Aggregated Descriptors (E-VLAD), a method to combine local and structured descriptors into a unified vector representation where each local descriptor is a Covariance Descriptor (CovD). In doing so, we make use of an accelerated method of obtaining a visual codebook where each atom is itself a CovD. We will then introduce an efficient way of aggregating local CovDs into a vector representation. Our method could be understood as an extrinsic extension of the highly acclaimed method of Vector of Locally Aggregated Descriptors [17] (or VLAD) to CovDs. We will show that the proposed method is extremely powerful in classifying materials/ textures and can outperform complex machineries even with simple classifiers. Masoud Faraki, Mehrtash Harandi, Fatih Porikli |
WACV | 1 |
| 2015 | Log-Euclidean bag of words for human action recognitionabstractRepresenting videos by densely extracted local space–time features has recently become a popular approach for analysing actions. In this study, the authors tackle the problem of categorising human actions by devising bag of words (BoWs) models based on covariance matrices of spatiotemporal features, with the features formed from histograms of optical flow. Since covariance matrices form a special type of Riemannian manifold, the space of symmetric positive definite (SPD) matrices, non‐Euclidean geometry should be taken into account while discriminating between covariance matrices. To this end, the authors propose to embed SPD manifolds to Euclidean spaces via a diffeomorphism and extend the BoW approach to its Riemannian version. The proposed BoW approach takes into account the manifold geometry of SPD matrices during the generation of the codebook and histograms. Experiments on challenging human action datasets show that the proposed method obtains notable improvements in discrimination accuracy, in comparison with several state‐of‐the‐art methods. Masoud Faraki, Maziar Palhang, Conrad Sanderson |
IET Comput. Vis. | 1 |
| 2014 | Fisher tensors for classifying human epithelial cells
Masoud Faraki, Mehrtash Harandi, Arnold Wiliem, Brian C. Lovell |
Pattern Recognit. | 1 |