EDBT 2026 Demo / reviewers in the wild / expert
Mahsa Baktash
dblp:119/1507 · also Mahsa Baktashmotlagh
· DBLP profile ↗
61ranked-venue papers
8as first author
36since 2021 · last 2027
0000-0001-5255-8194ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 42 · 6 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 35 · 4 first-author · 19 since 2021Databases, data management, data science and information retrieval · 7 · 6 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | Probing post-hoc reasoning in LLMs over multi-step pathfinding tasksabstractLarge language models (LLMs) are increasingly evaluated not only on the correctness of their outputs but also on the explanations they provide for their answers. To study LLM output behaviour in a controlled setting, including both final answers and post-hoc explanations, we use shortest-path routing as a testbed: it provides unambiguous ground truth while allowing systematic variation in topology and edge-weight distributions. We construct a large suite of graphs spanning five canonical topologies, grid, fat-tree, jellyfish, Erdős-Rényi, and Barabási-Albert, with fixed, uniform, discrete, and lognormal weights and evaluate both answer accuracy and the linguistic form of post-hoc explanations across multiple models. Two post-hoc explanation styles emerge consistently. An analytical style approximates algorithmic exploration, produces concise explanations, and remains accurate across graph families. A verbose explanation style narrates step-by-step simulations of procedures such as Dijkstra. In our results, this style is often associated with suboptimal paths or invalid edges, especially under irregular structure and non-uniform weights. We further identify recurring failure modes, including hallucinated edges that break path continuity, loops with repeated nodes, and paths that fail to connect the correct source and destination, that characterise this breakdown. All answers and explanations are produced without code execution; the observed behaviours therefore reflect text-only problem solving and retrospective verbalisation. Our findings show that post-hoc explanation style, not only final-answer accuracy, is associated with robustness within the evaluated model set, and they support shortest-path routing as a controlled diagnostic probe for evaluating LLM outputs and post-hoc explanations on structured graph tasks. Yaying Chen, Siamak Layeghy, Mahsa Baktash, Marius Portmann |
Expert Syst. Appl. | 3 |
| 2026 | Domain Generalizing DINO for Visual Regression via Latent Distractor Subspace ConsistencyabstractVision Foundation Models, such as DINO [20], have demonstrated remarkable generalization in classification; however, their application to out-of-domain visual regression tasks remains a significant and underexplored challenge. Unlike classification, domain generalization in regression poses distinct challenges: regression produces continuous outputs and is particularly sensitive to high-variance, label-irrelevant factors (e.g., illumination, blur, or contrast). These factors can entangle with task-relevant features and induce spurious correlations. While recent regression methods [11], [15], [24], [38], [39] have shown promise, they often rely on CNN backbones and require the pre-specification of known distractors. This demands significant domain expertise and fails to address spurious correlations that emerge during training. To address these challenges, we propose LDSC, a Latent Distractor Subspace Consistency framework that disentangles intermediate feature representation into task-relevant and latent distractor subspaces, and regularizes the latter under photometric perturbations to suppress spurious correlations while preserving discriminative features during training. Our proposed method, LDSC, is the first to effectively adapt the powerful DINO backbone for domain generalized visual regression. LDSC achieves state-of-the-art results on seven benchmark regression datasets, demonstrating its strong performance in domain generalization for visual regression with percentage improvements of (41.75%, 20.12%, 52.05%, 8.27%, 22.21%, 3.55%) over state-of-the-art DG regression methods, respectively. Project page is available: ldsc-iitd.github.io. Nikhil Reddy, Chetan Arora 0001, Mahsa Baktash |
WACV | 3 |
| 2025 | PEFTDiff: Diffusion-Guided Transferability Estimation for Parameter-Efficient Fine-Tuning
Prafful Kumar Khoba, Zijian Wang 0009, Chetan Arora 0001, Mahsa Baktash |
ICCV | 4 |
| 2025 | Not All Views Are Created Equal: Analyzing Viewpoint Instabilities in Vision Foundation Models
Mateusz Michalkiewicz, Sheena Bai, Mahsa Baktash, Varun Jampani, Guha Balakrishnan |
ICCV | 3 |
| 2025 | MOS: Model Synergy for Test-Time Adaptation on LiDAR-Based 3D Object DetectionabstractLiDAR-based 3D object detection is crucial for various applications but often experiences performance degradation in real-world deployments due to domain shifts. While most studies focus on cross-dataset shifts, such as changes in environments and object geometries, practical corruptions from sensor variations and weather conditions remain underexplored. In this work, we propose a novel online test-time adaptation framework for 3D detectors that effectively tackles these shifts, including a challenging $\textit{cross-corruption}$ scenario where cross-dataset shifts and corruptions co-occur. By leveraging long-term knowledge from previous test batches, our approach mitigates catastrophic forgetting and adapts effectively to diverse shifts. Specifically, we propose a Model Synergy (MOS) strategy that dynamically selects historical checkpoints with diverse knowledge and assembles them to best accommodate the current test batch. This assembly is directed by our proposed Synergy Weights (SW), which perform a weighted averaging of the selected checkpoints, minimizing redundancy in the composite model. The SWs are computed by evaluating the similarity of predicted bounding boxes on the test data and the independence of features between checkpoint pairs in the model bank. To maintain an efficient and informative model bank, we discard checkpoints with the lowest average SW scores, replacing them with newly updated models. Our method was rigorously tested against existing test-time adaptation strategies across three datasets and eight types of corruptions, demonstrating superior adaptability to dynamic scenes and conditions. Notably, it achieved a 67.3% improvement in a challenging cross-corruption scenario, offering a more comprehensive benchmark for adaptation. Source code: https://github.com/zhuoxiao-chen/MOS. Zhuoxiao Chen, Junjie Meng, Mahsa Baktash, Yonggang Zhang 0003, Zi Huang, Yadan Luo |
ICLR | 3 |
| 2025 | Improving Out-of-Distribution Detection via Dynamic Covariance CalibrationabstractOut-of-Distribution (OOD) detection is essential for the trustworthiness of AI systems. Methods using prior information (i.e., subspace-based methods) have shown effective performance by extracting information geometry to detect OOD data with a more appropriate distance metric. However, these methods fail to address the geometry distorted by ill-distributed samples, due to the limitation of statically extracting information geometry from the training distribution. In this paper, we argue that the influence of ill-distributed samples can be corrected by dynamically adjusting the prior geometry in response to new data. Based on this insight, we propose a novel approach that dynamically updates the prior covariance matrix using real-time input features, refining its information. Specifically, we reduce the covariance along the direction of real-time input features and constrain adjustments to the residual space, thus preserving essential data characteristics and avoiding effects on unintended directions in the principal space. We evaluate our method on two pre-trained models for the CIFAR dataset and five pre-trained models for ImageNet-1k, including the self-supervised DINO model. Extensive experiments demonstrate that our approach significantly enhances OOD detection across various models. The code is released at https://github.com/workerbcd/ooddcc. Kaiyu Guo, Zijian Wang 0009, Tan Pan, Brian C. Lovell, Mahsa Baktash |
ICML | 5 |
| 2025 | Shape-Space Deformer: Unified Visuo-Tactile Representations for Robotic Manipulation of Deformable ObjectsabstractAccurate modelling of object deformations is crucial for a wide range of robotic manipulation tasks, where interacting with soft or deformable objects is essential. Current methods struggle to generalise to unseen forces or adapt to new objects, limiting their utility in real-world applications. We propose Shape-Space Deformer, a unified representation for encoding a diverse range of object deformations using template augmentation to achieve robust, fine-grained reconstructions that are resilient to outliers and unwanted artefacts. Our method improves generalization to unseen forces and can rapidly adapt to novel objects, significantly outperforming existing approaches. We perform extensive experiments to test a range of force generalisation settings and evaluate our method's ability to reconstruct unseen deformations. Our results demonstrate significant improvements in reconstruction accuracy and robustness. Our approach is suitable for real-time performance, making it ready for downstream manipulation applications. Sean M. V. Collins, Brendan Tidd, Mahsa Baktash, Peyman Moghadam |
ICRA | 3 |
| 2025 | WisWheat: A Three-Tiered Vision-Language Dataset for Wheat ManagementabstractWheat management strategies play a critical role in determining yield. Traditional management decisions often rely on labour-intensive expert inspections, which are expensive, subjective and difficult to scale. Recently, Vision-Language Models (VLMs) have emerged as a promising solution to enable scalable, data-driven management support. However, due to a lack of domain-specific knowledge, directly applying VLMs to wheat management tasks results in poor quantification and reasoning capabilities, ultimately producing vague or even misleading management recommendations. In response, we propose WisWheat, a wheat-specific dataset with a three-layered design to enhance VLM performance on wheat management tasks: (1) a foundational pretraining dataset of 47,871 image-caption pairs for coarsely adapting VLMs to wheat morphology; (2) a quantitative dataset comprising 7,263 VQA-style image-question-answer triplets for quantitative trait measuring tasks; and (3) an Instruction Fine-tuning dataset with 4,888 samples targeting biotic and abiotic stress diagnosis and management plan for different phenological stages. Extensive experimental results demonstrate that fine-tuning open-source VLMs (e.g., Qwen2.5 7B) on our dataset leads to significant performance improvements. Specifically, the Qwen2.5 VL 7B fine-tuned on our wheat instruction dataset achieves accuracy scores of 79.2% and 84.6% on wheat stress and growth stage conversation tasks respectively, surpassing even general-purpose commercial models such as GPT-4o by a margin of 11.9% and 34.6%. Selena Song, Javier Fernandez, Yadan Luo, Mahsa Baktash, Zijian Wang 0009 |
ACM Multimedia | 5 |
| 2025 | Minimal Semantic Sufficiency Meets Unsupervised Domain GeneralizationabstractThe generalization ability of deep learning has been extensively studied in supervised settings, yet it remains less explored in unsupervised scenarios. Recently, the Unsupervised Domain Generalization (UDG) task has been proposed to enhance the generalization of models trained with prevalent unsupervised learning techniques, such as Self-Supervised Learning (SSL). UDG confronts the challenge of distinguishing semantics from variations without category labels. Although some recent methods have employed domain labels to tackle this issue, such domain labels are often unavailable in real-world contexts. In this paper, we address these limitations by formalizing UDG as the task of learning a Minimal Sufficient Semantic Representation: a representation that (i) preserves all semantic information shared across augmented views (sufficiency), and (ii) maximally removes information irrelevant to semantics (minimality). We theoretically ground these objectives from the perspective of information theory, demonstrating that optimizing representations to achieve sufficiency and minimality directly reduces out-of-distribution risk. Practically, we implement this optimization through Minimal-Sufficient UDG (MS-UDG), a learnable model by integrating (a) an InfoNCE-based objective to achieve sufficiency; (b) two complementary components to promote minimality: a novel semantic-variation disentanglement loss and a reconstruction-based mechanism for capturing adequate variation. Empirically, MS-UDG sets a new state-of-the-art on popular unsupervised domain-generalization benchmarks, consistently outperforming existing SSL and UDG methods, without category or domain labels during representation learning. Tan Pan, Kaiyu Guo, Dongli Xu, Zhaorui Tan, Chen Jiang 0006, Deshu Chen, Xin Guo 0010, Brian C. Lovell, Limei Han, Mahsa Baktash |
NeurIPS | 11 |
| 2025 | Spectral Distribution Alignment for Enhanced Generalization in Regression
Kaiyu Guo, Zijian Wang 0009, Brian C. Lovell, Mahsa Baktash |
ECML/PKDD (6) | 4 |
| 2025 | Leveraging Gradient Information for Out-of-Domain Performance Estimations
Ekaterina Khramtsova, Mahsa Baktash, Guido Zuccon, Xi Wang 0021, Mathieu Salzmann |
ECML/PKDD (6) | 2 |
| 2025 | Feature Space Perturbation: A Panacea to Enhanced Transferability EstimationabstractLeveraging a transferability estimation metric facilitates the non-trivial challenge of selecting the optimal model for the downstream task from a pool of pre-trained models. Most existing metrics primarily focus on identifying the statistical relationship between feature embeddings and the corresponding labels within the target dataset, but over-look crucial aspect of model robustness. This oversight may limit their effectiveness in accurately ranking pre-trained models. To address this limitation, we introduce a feature perturbation method that enhances the transferability estimation process by systematically altering the feature space. Our method includes a Spread operation that increases intra-class variability, adding complexity within classes, and an Attract operation that minimizes the distances between different classes, thereby blurring the class boundaries. Through extensive experimentation, we demonstrate the efficacy of our feature perturbation method in providing a more precise and robust estimation of model transferability. Notably, the existing LogMe method exhibited a significant improvement, showing a 28.84% increase in performance after applying our feature perturbation method. The implementation is available at https://github.com/prafful-kumar/enhancing_TE.git Prafful Kumar Khoba, Zijian Wang 0009, Chetan Arora 0001, Mahsa Baktash |
WACV | 4 |
| 2024 | Source-Free Domain-Invariant Performance Prediction
Ekaterina Khramtsova, Mahsa Baktash, Guido Zuccon, Xi Wang 0021, Mathieu Salzmann |
ECCV (80) | 2 |
| 2024 | DiPEx: Dispersing Prompt Expansion for Class-Agnostic Object DetectionabstractClass-agnostic object detection (OD) can be a cornerstone or a bottleneck for many downstream vision tasks. Despite considerable advancements in bottom-up and multi-object discovery methods that leverage basic visual cues to identify salient objects, consistently achieving a high recall rate remains difficult due to the diversity of object types and their contextual complexity. In this work, we investigate using vision-language models (VLMs) to enhance object detection via a self-supervised prompt learning strategy. Our initial findings indicate that manually crafted text queries often result in undetected objects, primarily because detection confidence diminishes when the query words exhibit semantic overlap. To address this, we propose a Dispersing Prompt Expansion (DiPEx) approach. DiPEx progressively learns to expand a set of distinct, non-overlapping hyperspherical prompts to enhance recall rates, thereby improving performance in downstream tasks such as out-of-distribution OD. Specifically, DiPEx initiates the process by self-training generic parent prompts and selecting the one with the highest semantic uncertainty for further expansion. The resulting child prompts are expected to inherit semantics from their parent prompts while capturing more fine-grained semantics. We apply dispersion losses to ensure high inter-class discrepancy among child prompts while preserving semantic consistency between parent-child prompt pairs. To prevent excessive growth of the prompt sets, we utilize the maximum angular coverage (MAC) of the semantic space as a criterion for early termination. We demonstrate the effectiveness of DiPEx through extensive class-agnostic OD and OOD-OD experiments on MS-COCO and LVIS, surpassing other prompting methods by up to 20.1% in AR and achieving a 21.3% AP improvement over SAM. Jia Syuen Lim, Zhuoxiao Chen, Zhi Chen 0010, Mahsa Baktash, Xin Yu 0002, Zi Huang, Yadan Luo |
NeurIPS | 4 |
| 2024 | Color-Oriented Redundancy Reduction in Dataset DistillationabstractDataset Distillation (DD) is designed to generate condensed representations of extensive image datasets, enhancing training efficiency. Despite recent advances, there remains considerable potential for improvement, particularly in addressing the notable redundancy within the color space of distilled images. In this paper, we propose a two-fold optimization strategy to minimize color redundancy at the individual image and overall dataset levels, respectively. At the image level, we employ a palette network, a specialized neural network, to dynamically allocate colors from a reduced color space to each pixel. The palette network identifies essential areas in synthetic images for model training, and consequently assigns more unique colors to them. At the dataset level, we develop a color-guided initialization strategy to minimize redundancy among images. Representative images with the least replicated color patterns are selected based on the information gain. A comprehensive performance study involving various datasets and evaluation scenarios is conducted, demonstrating the superior performance of our proposed color-aware DD compared to existing DD methods. Zijian Wang 0009, Mahsa Baktash, Yadan Luo, Zi Huang |
NeurIPS | 3 |
| 2024 | Embark on DenseQuest: A System for Selecting the Best Dense Retriever for a Custom CollectionabstractIn this demo we present a web-based application for selecting an effective pre-trained dense retriever to use on a private collection. Our system, DenseQuest, provides unsupervised selection and ranking capabilities to predict the best dense retriever among a pool of available dense retrievers, tailored to an uploaded target collection. DenseQuest implements a number of existing approaches, including a recent, highly effective method powered by Large Language Models (LLMs), which requires neither queries nor relevance judgments. The system is designed to be intuitive and easy to use for those information retrieval engineers and researchers who need to identify a general-purpose dense retrieval model to encode or search a new private target collection. Our demonstration illustrates conceptual architecture and the different use case scenarios of the system implemented on the cloud, enabling universal access and use. DenseQuest is available at https://densequest.ielab.io. Ekaterina Khramtsova, Teerapong Leelanupab, Shengyao Zhuang, Mahsa Baktash, Guido Zuccon |
SIGIR | 4 |
| 2024 | Leveraging LLMs for Unsupervised Dense Retriever RankingabstractIn this paper we present Large Language Model Assisted Retrieval Model Ranking (LARMOR), an effective unsupervised approach that leverages LLMs for selecting which dense retriever to use on a test corpus (target). Dense retriever selection is crucial for many IR applications that rely on using dense retrievers trained on public corpora to encode or search a new, private target corpus. This is because when confronted with domain shift, where the downstream corpora, domains, or tasks of the target corpus differ from the domain/task the dense retriever was trained on, its performance often drops. Furthermore, when the target corpus is unlabeled, e.g., in a zero-shot scenario, the direct evaluation of the model on the target corpus becomes unfeasible. Unsupervised selection of the most effective pre-trained dense retriever becomes then a crucial challenge. Current methods for dense retriever selection are insufficient in handling scenarios with domain shift. Ekaterina Khramtsova, Shengyao Zhuang, Mahsa Baktash, Guido Zuccon |
SIGIR | 3 |
| 2024 | Domain-Aware Knowledge Distillation for Continual Model GeneralizationabstractGeneralization on unseen domains is critical for Deep Neural Networks (DNNs) to perform well in real-world applications such as autonomous navigation. However, catastrophic forgetting limits the ability of domain generalization and unsupervised domain adaption approaches to adapt to constantly changing target domains. To overcome these challenges, We propose DoSe framework, a Domain-aware Self-Distillation method based on batch normalization prototypes to facilitate continual model generalization across varying target domains. Specifically, we enforce the consistency of batch normalization statistics between two batches of images sampled from the same target domain distribution between the student and teacher models. To alleviate catastrophic forgetting, we introduce a novel exemplar-based replay buffer to identify difficult samples for the model to retain the knowledge. Specifically, we demonstrate that identifying difficult samples and updating the model periodically using them can help in preserving knowledge learned from previously seen domains. We conduct extensive experiments on two real-world datasets ACDC, C-Driving, and one synthetic dataset SHIFT to verify the efficiency of the proposed DoSe framework. On ACDC, our method outperforms existing SOTA in Domain Generalization, Unsupervised Domain Adaptation, and Daytime settings by 26%, 14%, and 70% respectively. Nikhil Reddy, Mahsa Baktash, Chetan Arora 0001 |
WACV | 2 |
| 2023 | Convolutional Persistence as a Remedy to Neural Model AnalysisabstractWhile deep neural networks are proven to be effective learning systems, their analysis is complex due to the high-dimensionality of their weight space. Persistent topological properties can be used as an additional descriptor, providing insights on how the network weights evolve during training. In this paper, we focus on convolutional neural networks, and define the topology of the space, populated by convolutional filters (i.e., kernels). We perform an extensive analysis of topological properties of the convolutional filters. Specifically, we define a metric based on persistent homology, namely, Convolutional Topology Representation, to determine an important factor in neural networks training: the generalizability of the model to the test set. We further analyse how various training methods affect the topology of convolutional layers. Ekaterina Khramtsova, Guido Zuccon, Xi Wang 0021, Mahsa Baktash |
AISTATS | 4 |
| 2023 | Revisiting Domain-Adaptive 3D Object Detection by Reliable, Diverse and Class-balanced Pseudo-LabelingabstractUnsupervised domain adaptation (DA) with the aid of pseudo labeling techniques has emerged as a crucial approach for domain-adaptive 3D object detection. While effective, existing DA methods suffer from a substantial drop in performance when applied to a multi-class training setting, due to the co-existence of low-quality pseudo labels and class imbalance issues. In this paper, we address this challenge by proposing a novel ReDB framework tailored for learning to detect all classes at once. Our approach produces Reliable, Diverse, and class-Balanced pseudo 3D boxes to iteratively guide the self-training on a distributionally different target domain. To alleviate disruptions caused by the environmental discrepancy (e.g., beam numbers), the proposed cross-domain examination (CDE) assesses the correctness of pseudo labels by copy-pasting target instances into a source environment and measuring the prediction consistency. To reduce computational overhead and mitigate the object shift (e.g., scales and point densities), we design an overlapped boxes counting (OBC) metric that allows to uniformly downsample pseudo-labeled objects across different geometric characteristics. To confront the issue of inter-class imbalance, we progressively augment the target point clouds with a class-balanced set of pseudo-labeled target instances and source objects, which boosts recognition accuracies on both frequently appearing and rare classes. Experimental results on three benchmark datasets using both voxel-based (i.e., SECOND) and point-based 3D detectors (i.e., PointRCNN) demonstrate that our proposed ReDB approach outperforms existing 3D domain adaptation methods by a large margin, improving 23.15% mAP on the nuScenes → KITTI task. The code is available at https://github.com/zhuoxiao-chen/ReDB-DA-3Ddet. Zhuoxiao Chen, Yadan Luo, Zheng Wang 0044, Mahsa Baktash, Zi Huang |
ICCV | 4 |
| 2023 | Kecor: Kernel Coding Rate Maximization for Active 3D Object DetectionabstractAchieving a reliable LiDAR-based object detector in autonomous driving is paramount, but its success hinges on obtaining large amounts of precise 3D annotations. Active learning (AL) seeks to mitigate the annotation burden through algorithms that use fewer labels and can attain performance comparable to fully supervised learning. Although AL has shown promise, current approaches prioritize the selection of unlabeled point clouds with high uncertainty and/or diversity, leading to the selection of more instances for labeling and reduced computational efficiency. In this paper, we resort to a novel kernel coding rate maximization (Kecor) strategy which aims to identify the most informative point clouds to acquire labels through the lens of information theory. Greedy search is applied to seek desired point clouds that can maximize the minimal number of bits required to encode the latent features. To determine the uniqueness and informativeness of the selected samples from the model perspective, we construct a proxy network of the 3D detector head and compute the outer product of Jacobians from all proxy layers to form the empirical neural tangent kernel (NTK) matrix. To accommodate both one-stage (i.e., Second) and two-stage detectors (i.e., Pv-rcnn), we further incorporate the classification entropy maximization and well trade-off between detection performance and the total number of bounding boxes selected for annotation. Extensive experiments conducted on two 3D benchmarks and a 2D detection dataset evidence the superiority and versatility of the proposed approach. Our results show that approximately 44% box-level annotation costs and 26% computational time are reduced compared to the state-of-the-art AL method, without compromising detection performance. Source code: https://github.com/Luoyadan/KECOR-active-3Ddet. Yadan Luo, Zhuoxiao Chen, Zhen Fang 0001, Zheng Zhang 0006, Mahsa Baktash, Zi Huang |
ICCV | 5 |
| 2023 | Domain Generalization Guided by Gradient Signal to Noise Ratio of ParametersabstractOverfitting to the source domain is a common issue in gradient-based training of deep neural networks. To compensate for the over-parameterized models, numerous regularization techniques have been introduced such as those based on dropout. While these methods achieve significant improvements on classical benchmarks such as ImageNet, their performance diminishes with the introduction of domain shift in the test set i.e. when the unseen data comes from a significantly different distribution. In this paper, we move away from the classical approach of Bernoulli sampled dropout mask construction and propose to base the selection on gradient-signal-to-noise ratio (GSNR) of network’s parameters. Specifically, at each training step, parameters with high GSNR will be discarded. Furthermore, we alleviate the burden of manually searching for the optimal dropout ratio by leveraging a meta-learning approach. We evaluate our method on standard domain generalization benchmarks and achieve competitive results on classification and face anti-spoofing problems. Mateusz Michalkiewicz, Masoud Faraki, Xiang Yu 0002, Manmohan Krishna Chandraker, Mahsa Baktash |
ICCV | 5 |
| 2023 | How Far Pre-trained Models Are from Neural Collapse on the Target Dataset Informs their TransferabilityabstractThis paper focuses on model transferability estimation, i.e., assessing the performance of pre-trained models on a downstream task without performing fine-tuning. Motivated by neural collapse (NC) [25] that reveals particular feature geometry at the terminal stage of training, we consider model transferability as how far the target activations obtained by pre-trained models are from their hypothetical state in the terminal phase of the model fine-tuned on the target domain. We propose a metric that measures this proximity based on three phenomena of NC: within-class variability collapse, simplex encoded label interpolation geometry structure is formed, and the nearest center classifier becomes optimal on training data. Through experiments on 11 datasets, we confirm none of the three NC proxies are dispensable, which allows us to obtain very competitive transferability estimation accuracy with approximately 10× wall-clock time speed up compared to state-of-the-art approaches. Zijian Wang 0009, Yadan Luo, Liang Zheng 0001, Zi Huang, Mahsa Baktash |
ICCV | 5 |
| 2023 | Exploring Active 3D Object Detection from a Generalization Perspective
Yadan Luo, Zhuoxiao Chen, Zijian Wang 0009, Xin Yu 0002, Zi Huang, Mahsa Baktash |
ICLR | 6 |
| 2023 | Center-aware Adversarial Augmentation for Single Domain GeneralizationabstractDomain generalization (DG) aims to learn a model from multiple training (i.e., source) domains that can generalize well to the unseen test (i.e., target) data coming from a different distribution. Single domain generalization (Single-DG) has recently emerged to tackle a more challenging, yet realistic setting, where only one source domain is available at training time. The existing Single-DG approaches typically are based on data augmentation strategies and aim to expand the span of source data by augmenting out-of-domain samples. Generally speaking, they aim to generate hard examples to confuse the classifier. While this may make the classifier robust to small perturbation, the generated samples are typically not diverse enough to mimic a large domain shift, resulting in sub-optimal generalization performance. To alleviate this, we propose a center-aware adversarial augmentation technique that expands the source distribution by altering the source samples so as to push them away from the class centers via a novel angular center loss. We conduct extensive experiments to demonstrate the effectiveness of our approach on several benchmark datasets for Single-DG and show that our method outperforms the state-of-the-art in most cases. Mahsa Baktash, Zijian Wang 0009, Mathieu Salzmann |
WACV | 2 |
| 2023 | FFM: Injecting Out-of-Domain Knowledge via Factorized Frequency ModificationabstractThis work investigates the Single Domain Generalization (SDG) problem and aims to generalize a model from a single source (i.e., training) domain to multiple target (i.e., test) domains coming from different distributions. Most of the existing SDG approaches focus on generating out-of-domain samples by either transforming the source images into different styles or optimizing adversarial noise perturbations applied on the source images. In this paper, we show that generating images with diverse styles can be complementary to creating hard samples when handling the SDG task, and propose our approach of Factorized Frequency Modification (FFM) to fulfill this requirement. Specifically, we design a unified framework consisting of a style transformation module, an adversarial perturbation module, and a dynamic frequency selection module. We seamlessly equip the framework with iterative adversarial training that facilitates learning discriminative features from hard and diverse augmented samples. Extensive experiments are performed on four image recognition benchmark datasets of Digits, CIFAR-10-C, CIFAR-100-C, and PACS, which demonstrates that our method outperforms existing state-of-the-art approaches. Zijian Wang 0009, Yadan Luo, Zi Huang, Mahsa Baktash |
WACV | 4 |
| 2023 | DI-NIDS: Domain invariant network intrusion detection systemabstractThe performance of machine learning based network intrusion detection systems (NIDSs) severely degrades when deployed on a network with significantly different feature distributions from the ones of the training dataset. In various applications, such as computer vision, domain adaptation techniques have been successful in mitigating the gap between the distributions of the training and test data. In the case of network intrusion detection however, the state-of-the-art domain adaptation approaches have had limited success. According to recent studies, as well as our own results, the performance of an NIDS considerably deteriorates when the ‘unseen’ test dataset does not follow the training dataset distribution. In order to enhance the generalisability of machine learning based network intrusion detection systems, we propose to extract domain invariant features using adversarial domain adaptation from multiple network domains, and then apply an unsupervised technique for recognising abnormalities, i.e., intrusions. More specifically, we train a domain adversarial neural network on labelled source domains, extract the domain invariant features, and train a One-Class SVM (OSVM) model to detect anomalies. At test time, we feedforward the unlabelled test data to the feature extractor network to project it into a domain invariant space, and then apply OSVM on the extracted features to achieve our final goal of detecting intrusions. Our extensive experiments on the NIDS benchmark datasets of NFv2-CIC-2018 and NFv2-UNSW-NB15 show that our proposed setup demonstrates superior cross-domain performance in comparison to the previous approaches. Siamak Layeghy, Mahsa Baktash, Marius Portmann |
Knowl. Based Syst. | 2 |
| 2023 | Source-Free Progressive Graph Learning for Open-Set Domain AdaptationabstractOpen-set domain adaptation (OSDA) aims to transfer knowledge from a label-rich source domain to a label-scarce target domain while addressing disturbances from irrelevant target classes not present in the source data. However, most OSDA approaches are limited due to the lack of essential theoretical analysis of generalization bound, reliance on the coexistence of source and target data during adaptation, and failure to accurately estimate model predictions' uncertainty. To address these limitations, the Progressive Graph Learning (PGL) framework is proposed. PGL decomposes the target hypothesis space into shared and unknown subspaces and progressively pseudo-labels the most confident known samples from the target domain for hypothesis adaptation. PGL guarantees a tight upper bound of the target error by integrating a graph neural network with episodic training and leveraging adversarial learning to close the gap between the source and target distributions. The proposed approach also tackles a more realistic source-free open-set domain adaptation (SF-OSDA) setting that makes no assumptions about the coexistence of source and target domains. In a two-stage framework, the SF-PGL model' uniformly selects the most confident target instances from each category at a fixed ratio, and the confidence thresholds in each class weigh the classification loss in the adaptation step. The proposed methods are evaluated on benchmark image classification and action recognition datasets, where they demonstrate superiority and flexibility in recognizing both shared and unknown categories. Additionally, balanced pseudo-labeling plays a significant role in improving calibration, making the trained model less prone to over- or under-confident predictions on the target data. Yadan Luo, Zijian Wang 0009, Zhuoxiao Chen, Zi Huang, Mahsa Baktash |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Interpretable Signed Link Prediction With Signed Infomax Hyperbolic GraphabstractSigned link prediction in social networks aims to reveal the underlying relationships (i.e., links) among users (i.e., nodes) given their existing positive and negative interactions observed. Most of the prior efforts are devoted to learning node embeddings with graph neural networks (GNNs), which preserve the signed network topology by message-passing along edges to facilitate the downstream link prediction task. Nevertheless, the existing graph-based approaches could hardly provide human-intelligible explanations for the following three questions: (1) which neighbors to aggregate, (2) which path to propagate along, and (3) which social theory to follow in the learning process. To answer the aforementioned questions, in this paper, we investigate how to reconcile thebalanceandstatussocial rules with information theory and develop a unified framework, termed as Signed Infomax Hyperbolic Graph (SIHG). By maximizing the mutual information between edge polarities and node embeddings, one can identify the most representative neighboring nodes that support the inference of edge sign. Different from existing GNNs that could only group features of friends in the subspace, the proposed SIHG incorporates the signed attention module, which is also capable of pushing hostile users far away from each other to preserve the geometry of antagonism. The polarity of the learned edge attention maps, in turn, provides interpretations of the social theories used in each aggregation. In order to model high-order user relations and complex hierarchies, the node embeddings are projected and measured in a hyperbolic space with a lower distortion. Extensive experiments on four signed network benchmarks demonstrate that the proposed SIHG framework significantly outperforms the state-of-the-arts in signed link prediction. Yadan Luo, Zi Huang, Hongxu Chen 0002, Yang Yang 0002, Hongzhi Yin, Mahsa Baktash |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2022 | Master of All: Simultaneous Generalization of Urban-Scene Segmentation to All Adverse Weather Conditions
Nikhil Reddy, Abhinav Singhal, Mahsa Baktash, Chetan Arora 0001 |
ECCV (39) | 4 |
| 2022 | Contrastive Class-aware Adaptation for Domain GeneralizationabstractDomain generalization (DG) tackles the problem of learning a model that generalizes to data drawn from a target domain that was unseen during training. A major trend in this area consists of learning a domain-invariant representation by minimizing the discrepancy across multiple source domains. This strategy, however, does not apply to the challenging yet realistic single-source scenario. In this paper, in contrast to existing methods that focus on domain discrepancy, we exploit the fact that discrepancies also arise across samples from the same class. We therefore develop a unified framework for both multisource and single-source DG that exploits contrastive learning to maximize the gap between samples from the same class, either from different domains or from the same one, while separating the samples from different classes. Our results on standard multisource and single-source DG benchmark datasets demonstrate the benefits of our method over the state-of-the-art ones in both settings. Mahsa Baktash, Mathieu Salzmann |
ICPR | 2 |
| 2022 | Learning to Generate the Unknowns as a Remedy to the Open-Set Domain ShiftabstractIn many situations, the data one has access to at test time follows a different distribution from the training data. Over the years, this problem has been tackled by closed-set domain adaptation techniques. Recently, open-set domain adaptation has emerged to address the more realistic scenario where additional unknown classes are present in the target data. In this setting, existing techniques focus on the challenging task of isolating the unknown target samples, so as to avoid the negative transfer resulting from aligning the source feature distributions with the broader target one that encompasses the additional unknown classes. Here, we propose a simpler and more effective solution consisting of complementing the source data distribution and making it comparable to the target one by enabling the model to generate source samples corresponding to the unknown target classes. We formulate this as a general module that can be incorporated into any existing closed-set approach and show that this strategy allows us to outperform the state of the art on open-set domain adaptation benchmark datasets. Mahsa Baktash, Mathieu Salzmann |
WACV | 1 |
| 2021 | Learning to Diversify for Single Domain GeneralizationabstractDomain generalization (DG) aims to generalize a model trained on multiple source (i.e., training) domains to a distributionally different target (i.e., test) domain. In contrast to the conventional DG that strictly requires the availability of multiple source domains, this paper considers a more realistic yet challenging scenario, namely Single Domain Generalization (Single-DG), where only one source domain is available for training. In this scenario, the limited diversity may jeopardize the model generalization on unseen target domains. To tackle this problem, we propose a style-complement module to enhance the generalization power of the model by synthesizing images from diverse distributions that are complementary to the source ones. More specifically, we adopt a tractable upper bound of mutual information (MI) between the generated and source samples and perform a two-step optimization iteratively: (1) by minimizing the MI upper bound approximation for each sample pair, the generated images are forced to be diversified from the source samples; (2) subsequently, we maximize the MI between the samples from the same semantic category, which assists the network to learn discriminative features from diverse-styled images. Extensive experiments on three benchmark datasets demonstrate the superiority of our approach, which surpasses the state-of-the-art single-DG methods by up to 25.14%. The code will be publicly available at https://github.com/BUserName/Learning_to_diversify Zijian Wang 0009, Yadan Luo, Ruihong Qiu, Zi Huang, Mahsa Baktash |
ICCV | 5 |
| 2021 | Semi-supervised Keypoint Localization
Olga Moskvyak, Frédéric Maire, Feras Dayoub, Mahsa Baktash |
ICLR | 4 |
| 2021 | Conditional Extreme Value Theory for Open Set Video Domain AdaptationabstractWith the advent of media streaming, video action recognition has become progressively important for various applications, yet at the high expense of requiring large-scale data labelling. To overcome the problem of expensive data labelling, domain adaptation techniques have been proposed, which transfer knowledge from fully labelled data (i.e., source domain) to unlabelled data (i.e., target domain). The majority of video domain adaptation algorithms are proposed for closed-set scenarios in which all the classes are shared among the domains. In this work, we propose an open-set video domain adaptation approach to mitigate the domain discrepancy between the source and target data, allowing the target data to contain additional classes that do not belong to the source domain. Different from previous works, which only focus on improving accuracy for shared classes, we aim to jointly enhance the alignment of the shared classes and recognition of unknown samples. Towards this goal, class-conditional extreme value theory is applied to enhance the unknown recognition. Specifically, the entropy values of target samples are modelled as generalised extreme value distributions, which allows separating unknown samples lying in the tail of the distribution. To alleviate the negative transfer issue, weights computed by the distance from the sample entropy to the threshold are leveraged in adversarial learning in the sense that confident source and target samples are aligned, and unconfident samples are pushed away. The proposed method has been thoroughly evaluated on both small-scale and large-scale cross-domain video datasets and achieved the state-of-the-art performance. Zhuoxiao Chen, Yadan Luo, Mahsa Baktash |
MMAsia | 3 |
| 2021 | Keypoint-Aligned Embeddings for Image Retrieval and Re-identificationabstractLearning embeddings that are invariant to the pose of the object is crucial in visual image retrieval and re-identification. The existing approaches for person, vehicle, or animal re-identification tasks suffer from high intra-class variance due to deformable shapes and different camera viewpoints. To overcome this limitation, we propose to align the image embedding with a predefined order of the key-points. The proposed keypoint aligned embeddings model (KAE-Net) learns part-level features via multi-task learning which is guided by keypoint locations. More specifically, KAE-Net extracts channels from a feature map activated by a specific keypoint through learning the auxiliary task of heatmap reconstruction for this keypoint. The KAE-Net is compact, generic and conceptually simple. It achieves state of the art performance on the benchmark datasets of CUB-200-2011, Cars196 and VeRi-776 for retrieval and re-identification tasks. Olga Moskvyak, Frédéric Maire, Feras Dayoub, Mahsa Baktash |
WACV | 4 |
| 2020 | Learning from the Past: Continual Meta-Learning with Bayesian Graph Neural NetworksabstractMeta-learning for few-shot learning allows a machine to leverage previously acquired knowledge as a prior, thus improving the performance on novel tasks with only small amounts of data. However, most mainstream models suffer from catastrophic forgetting and insufficient robustness issues, thereby failing to fully retain or exploit long-term knowledge while being prone to cause severe error accumulation. In this paper, we propose a novel Continual Meta-Learning approach with Bayesian Graph Neural Networks (CML-BGNN) that mathematically formulates meta-learning as continual learning of a sequence of tasks. With each task forming as a graph, the intra- and inter-task correlations can be well preserved via message-passing and history transition. To remedy topological uncertainty from graph initialization, we utilize Bayes by Backprop strategy that approximates the posterior distribution of task-specific parameters with amortized inference networks, which are seamlessly integrated into the end-to-end edge learning. Extensive experiments conducted on the miniImageNet and tieredImageNet datasets demonstrate the effectiveness and efficiency of the proposed method, improving the performance by 42.8% compared with state-of-the-art on the miniImageNet 5-way 1-shot classification task. Yadan Luo, Zi Huang, Zheng Zhang 0006, Ziwei Wang 0003, Mahsa Baktash, Yang Yang 0002 |
AAAI | 5 |
| 2020 | A Simple and Scalable Shape Representation for 3D Reconstruction
Mateusz Michalkiewicz, Eugene Belilovsky, Mahsa Baktash, Anders P. Eriksson |
BMVC | 3 |
| 2020 | CosMo: Conditional Seq2Seq-based Mixture Model for Zero-Shot Commonsense Question AnsweringabstractCommonsense reasoning refers to the ability of evaluating a social situation and acting accordingly.Identification of the implicit causes and effects of a social context is the driving capability which can enable machines to perform commonsense reasoning.The dynamic world of social interactions requires context-dependent on-demand systems to infer such underlying information.However, current approaches in this realm lack the ability to perform commonsense reasoning upon facing an unseen situation, mostly due to incapability of identifying a diverse range of implicit social relations.Hence they fail to estimate the correct reasoning path.In this paper, we present Conditional SEQ2SEQ-based Mixture model (COSMO), which provides us with the capabilities of dynamic and diverse content generation.We use COSMO to generate context-dependent clauses, which form a dynamic Knowledge Graph (KG) on-the-fly for commonsense reasoning.To show the adaptability of our model to context-dependant knowledge generation, we address the task of zero-shot commonsense question answering.The empirical results indicate an improvement of up to +5.2% over the state-of-the-art models. Farhad Moghimifar, Lizhen Qu, Terry Yue Zhuo, Mahsa Baktash, Gholamreza Haffari |
COLING | 4 |
| 2020 | Few-Shot Single-View 3-D Object Reconstruction with Compositional Priors
Mateusz Michalkiewicz, Sarah Parisot, Stavros Tsogkas, Mahsa Baktash, Anders P. Eriksson, Eugene Belilovsky |
ECCV (25) | 4 |
| 2020 | Progressive Graph Learning for Open-Set Domain AdaptationabstractDomain shift is a fundamental problem in visual recognition which typically arises when the source and target data follow different distributions. The existing domain adaptation approaches which tackle this problem work in the "closed-set" setting with the assumption that the source and the target data share exactly the same classes of objects. In this paper, we tackle a more realistic problem of the "open-set" domain shift where the target data contains additional classes that were not present in the source data. More specifically, we introduce an end-to-end Progressive Graph Learning (PGL) framework where a graph neural network with episodic training is integrated to suppress underlying conditional shift and adversarial learning is adopted to close the gap between the source and target distributions. Compared to the existing open-set adaptation approaches, our approach guarantees to achieve a tighter upper bound of the target error. Extensive experiments on three standard open-set benchmarks evidence that our approach significantly outperforms the state-of-the-arts in open-set domain adaptation. Yadan Luo, Zijian Wang 0009, Zi Huang, Mahsa Baktash |
ICML | 4 |
| 2020 | Adversarial Bipartite Graph Learning for Video Domain AdaptationabstractDomain adaptation techniques, which focus on adapting models between distributionally different domains, are rarely explored in the video recognition area due to the significant spatial and temporal shifts across the source (i.e. training) and target (i.e. test) domains. As such, recent works on visual domain adaptation which leverage adversarial learning to unify the source and target video representations and strengthen the feature transferability are not highly effective on the videos. To overcome this limitation, in this paper, we learn a domain-agnostic video classifier instead of learning domain-invariant representations, and propose an Adversarial Bipartite Graph (ABG) learning framework which directly models the source-target interactions with a network topology of the bipartite graph. Specifically, the source and target frames are sampled as heterogeneous vertexes while the edges connecting two types of nodes measure the affinity among them. Through message-passing, each vertex aggregates the features from its heterogeneous neighbors, forcing the features coming from the same class to be mixed evenly. Explicitly exposing the video classifier to such cross-domain representations at the training and test stages makes our model less biased to the labeled source data, which in-turn results in achieving a better generalization on the target domain. The proposed framework is agnostic to the choices of frame aggregation, and therefore, four different aggregation functions are investigated for capturing appearance and temporal dynamics. To further enhance the model capacity and testify the robustness of the proposed architecture on difficult transfer tasks, we extend our model to work in a semi-supervised setting using an additional video-level bipartite graph. Extensive experiments conducted on four benchmark datasets evidence the effectiveness of the proposed approach over the state-of-the-art methods on the task of video recognition. Yadan Luo, Zi Huang, Zijian Wang 0009, Zheng Zhang 0006, Mahsa Baktash |
ACM Multimedia | 5 |
| 2020 | Prototype-Matching Graph Network for Heterogeneous Domain AdaptationabstractEven though the multimedia data is ubiquitous on the web, the scarcity of the annotated data and variety of data modalities hinder their usage by multimedia applications. Heterogeneous domain adaptation (HDA) has therefore arisen to address such limitations by facilitating the knowledge transfer between heterogeneous domains. Existing HDA methods only focus on aligning the cross-domain feature distributions and ignore the importance of maximizing the margin among different classes, which may lead to a sub-optimal classification performance. To tackle this problem, in this paper, we propose the Prototype-Matching Graph Network (PMGN), which gradually explores the domain-invariant class prototype representations. Specifically, we build an end-to-end Graph Prototypical Network, which computes the class prototypes through multiple layers of edge learning, node aggregation, and discrepancy minimization. Our framework utilizes the Swap training strategy to provide adequate supervision for training the edge learning component. Moreover, the proposed PMGN can be equipped with the clustering module that utilises the KL-divergence as a distance metric to reduce the distribution difference between the source and target data. Extensive experiments on three HDA tasks (i.e. object recognition, text-to-image classification, and text categorization) demonstrate the superiority of our approach over the state-of-the-art HDA methods. Zijian Wang 0009, Yadan Luo, Zi Huang, Mahsa Baktash |
ACM Multimedia | 4 |
| 2020 | Correlation-aware adversarial domain adaptation and generalization
Mohammad Mahfujur Rahman, Clinton Fookes, Mahsa Baktash, Sridha Sridharan |
Pattern Recognit. | 3 |
| 2019 | Implicit Surface Representations As Layers in Neural NetworksabstractImplicit shape representations, such as Level Sets, provide a very elegant formulation for performing computations involving curves and surfaces. However, including implicit representations into canonical Neural Network formulations is far from straightforward. This has consequently restricted existing approaches to shape inference, to significantly less effective representations, perhaps most commonly voxels occupancy maps or sparse point clouds. To overcome this limitation we propose a novel formulation that permits the use of implicit representations of curves and surfaces, of arbitrary topology, as individual layers in Neural Network architectures with end-to-end trainability. Specifically, we propose to represent the output as an oriented level set of a continuous and discretised embedding function. We investigate the benefits of our approach on the task of 3D shape prediction from a single image; and demonstrate its ability to produce a more accurate reconstruction compared to voxel-based representations. We further show that our model is flexible and can be applied to a variety of shape inference problems. Mateusz Michalkiewicz, Jhony K. Pontes, Dominic Jack, Mahsa Baktash, Anders P. Eriksson |
ICCV | 4 |
| 2019 | Learning Factorized Representations for Open-Set Domain Adaptation
Mahsa Baktash, Masoud Faraki, Tom Drummond, Mathieu Salzmann |
ICLR (Poster) | 1 |
| 2019 | Multi-Component Image Translation for Deep Domain GeneralizationabstractDomain adaption (DA) and domain generalization (DG) are two closely related methods which are both concerned with the task of assigning labels to an unlabeled data set. The only dissimilarity between these approaches is that DA can access the target data during the training phase, while the target data is totally unseen during the training phase in DG. The task of DG is challenging as we have no earlier knowledge of the target samples. If DA methods are applied directly to DG by a simple exclusion of the target data from training, poor performance will result for a given task. In this paper, we tackle the domain generalization challenge in two ways. In our first approach, we propose a novel deep domain generalization architecture utilizing synthetic data generated by a Generative Adversarial Network (GAN). The discrepancy between the generated images and synthetic images is minimized using existing domain discrepancy metrics such as maximum mean discrepancy or correlation alignment. In our second approach, we introduce a protocol for applying DA methods to a DG scenario by excluding the target data from the training phase, splitting the source data to training and validation parts, and treating the validation data as target data for DA. We conduct extensive experiments on four cross-domain benchmark datasets. Experimental results signify our proposed model outperforms the current state-of-the-art methods for DG. Mohammad Mahfujur Rahman, Clinton Fookes, Mahsa Baktash, Sridha Sridharan |
WACV | 3 |
| 2017 | From Shared Subspaces to Shared Landmarks: A Robust Multi-Source Classification ApproachabstractTraining machine leaning algorithms on augmented data fromdifferent related sources is a challenging task. This problemarises in several applications, such as the Internet of Things(IoT), where data may be collected from devices with differentsettings. The learned model on such datasets can generalizepoorly due to distribution bias. In this paper we considerthe problem of classifying unseen datasets, given several labeledtraining samples drawn from similar distributions. Weexploit the intrinsic structure of samples in a latent subspaceand identify landmarks, a subset of training instances fromdifferent sources that should be similar. Incorporating subspacelearning and landmark selection enhances generalizationby alleviating the impact of noise and outliers, as well asimproving efficiency by reducing the size of the data. However,since addressing the two issues simultaneously resultsin an intractable problem, we relax the objective functionby leveraging the theory of nonlinear projection and solve atractable convex optimisation. Through comprehensive analysis,we show that our proposed approach outperforms stateof-the-art results on several benchmark datasets, while keepingthe computational complexity low. Sarah M. Erfani, Mahsa Baktash, Masud Moshtaghi, Vinh Nguyen 0003, Christopher Leckie, James Bailey 0001, Kotagiri Ramamohanarao |
AAAI | 2 |
| 2017 | Deep discovery of facial motions using a shallow embedding layerabstractUnique encoding of the dynamics of facial actions has potential to provide a spontaneous facial expression recognition system. The most promising existing approaches rely on deep learning of facial actions. However, current approaches are often computationally intensive and require a great deal of memory/processing time, and typically the temporal aspect of facial actions are often ignored, despite the potential wealth of information available from the spatial dynamic movements and their temporal evolution over time from neutral state to apex state. To tackle aforementioned challenges, we propose a deep learning framework by using the 3D convolutional filters to extract spatio-temporal features, followed by the LSTM network which is able to integrate the dynamic evolution of short-duration of spatio-temporal features as an emotion progresses from the neutral state to the apex state. In order to reduce the redundancy of parameters and accelerate the learning of the recurrent neural network, we propose a shallow embedding layer to reduce the number of parameters in the LSTM by up to 98% without sacrificing recognition accuracy. As the fully connected layer approximately contains 95% of the parameters in the network, we decrease the number of parameters in this layer before passing features to the LSTM network, which significantly improves training speed and enables the possibility of deploying a state of the art deep network on real-time applications. We evaluate our proposed framework on the DISFA and UNBC-McMaster Shoulder pain datasets. Afsane Ghasemi, Mahsa Baktash, Simon Denman, Sridha Sridharan, Dung Nguyen Tien, Clinton Fookes |
ICIP | 2 |
| 2016 | Robust Domain Generalisation by Enforcing Distribution Invariance
Sarah M. Erfani, Mahsa Baktash, Masud Moshtaghi, Vinh Nguyen 0003, Christopher Leckie, James Bailey 0001, Kotagiri Ramamohanarao |
IJCAI | 2 |
| 2016 | R1STM: One-class Support Tensor Machine with Randomised KernelabstractIdentifying unusual or anomalous patterns in an underlying dataset is an important but challenging task in many applications. The focus of the unsupervised anomaly detection literature has mostly been on vectorised data. However, many applications are more naturally described using higher-order tensor representations. Approaches that vectorise tensorial data can destroy the structural information encoded in the high-dimensional space, and lead to the problem of the curse of dimensionality. In this paper we present the first unsupervised tensorial anomaly detection method, along with a randomised version of our method. Our anomaly detection method, the One-class Support Tensor Machine (1STM), is a generalisation of conventional one-class Support Vector Machines to higher-order spaces. 1STM preserves the multiway structure of tensor data, while achieving significant improvement in accuracy and efficiency over conventional vectorised methods. We then leverage the theory of nonlinear random projections to propose the Randomised 1STM (R1STM). Our empirical analysis on several real and synthetic datasets shows that our R1STM algorithm delivers comparable or better accuracy to a state-of-the-art deep learning method and traditional kernelised approaches for anomaly detection, while being approximately 100 times faster in training and testing. Sarah M. Erfani, Mahsa Baktash, Sutharshan Rajasegarar, Vinh Nguyen 0003, Christopher Leckie, James Bailey 0001, Kotagiri Ramamohanarao |
SDM | 2 |
| 2016 | Distribution-Matching Embedding for Visual Domain AdaptationabstractDomain-invariant representations are key to addressing the domain shift problem where the training and test examples follow different distributions. Existing techniques that have attempted to match the distributions of the source and target domains typically compare these distributions in the original feature space. This space, however, may not be directly suitable for such a comparison, since some of the features may have been distorted by the domain shift, or may be domain specific. In this paper, we introduce a Distribution-Matching Embedding approach: An unsupervised domain adaptation method that overcomes this issue by mapping the data to a latent space where the distance between the empirical distributions of the source and target examples is minimized. In other words, we seek to extract the information that is invariant across the source and target data. In particular, we study two different distances to compare the source and target distributions: the Maximum Mean Discrepancy and the Hellinger distance. Furthermore, we show that our approach allows us to learn either a linear embedding, or a nonlinear one. We demonstrate the benefits of our approach on the tasks of visual object recognition, text categorization, and WiFi localization. Mahsa Baktash, Mehrtash Harandi, Mathieu Salzmann |
J. Mach. Learn. Res. | 1 |
| 2015 | R1SVM: A Randomised Nonlinear Approach to Large-Scale Anomaly DetectionabstractThe problem of unsupervised anomaly detection arises in awide variety of practical applications. While one-class sup-port vector machines have demonstrated their effectiveness asan anomaly detection technique, their ability to model largedatasets is limited due to their memory and time complexityfor training. To address this issue for supervised learning ofkernel machines, there has been growing interest in randomprojection methods as an alternative to the computationallyexpensive problems of kernel matrix construction and sup-port vector optimisation. In this paper we leverage the theoryof nonlinear random projections and propose the RandomisedOne-class SVM (R1SVM), which is an efficient and scalableanomaly detection technique that can be trained on large-scale datasets. Our empirical analysis on several real-life andsynthetic datasets shows that our randomised 1SVM algo-rithm achieves comparable or better accuracy to deep autoen-coder and traditional kernelised approaches for anomaly de-tection, while being approximately 100 times faster in train-ing and testing Sarah M. Erfani, Mahsa Baktash, Sutharshan Rajasegarar, Shanika Karunasekera, Christopher Leckie |
AAAI | 2 |
| 2015 | Beyond Gauss: Image-Set Matching on the Riemannian Manifold of PDFsabstractState-of-the-art image-set matching techniques typically implicitly model each image-set with a Gaussian distribution. Here, we propose to go beyond these representations and model image-sets as probability distribution functions (PDFs) using kernel density estimators. To compare and match image-sets, we exploit Csiszar f-divergences, which bear strong connections to the geodesic distance defined on the space of PDFs, i.e., the statistical manifold. Furthermore, we introduce valid positive definite kernels on the statistical manifolds, which let us make use of more powerful classification schemes to match image-sets. Finally, we introduce a supervised dimensionality reduction technique that learns a latent space where f-divergences reflect the class labels of the data. Our experiments on diverse problems, such as video-based face recognition and dynamic texture classification, evidence the benefits of our approach over the state-of-the-art image-set matching methods. Mehrtash Harandi, Mathieu Salzmann, Mahsa Baktash |
ICCV | 3 |
| 2014 | Domain Adaptation on the Statistical ManifoldabstractIn this paper, we tackle the problem of unsupervised domain adaptation for classification. In the unsupervised scenario where no labeled samples from the target domain are provided, a popular approach consists in transforming the data such that the source and target distributions become similar. To compare the two distributions, existing approaches make use of the Maximum Mean Discrepancy (MMD). However, this does not exploit the fact that probability distributions lie on a Riemannian manifold. Here, we propose to make better use of the structure of this manifold and rely on the distance on the manifold to compare the source and target distributions. In this framework, we introduce a sample selection method and a subspace-based method for unsupervised domain adaptation, and show that both these manifold-based techniques outperform the corresponding approaches based on the MMD. Furthermore, we show that our subspace-based approach yields state-of-the-art results on a standard object recognition benchmark. Mahsa Baktash, Mehrtash Harandi, Brian C. Lovell, Mathieu Salzmann |
CVPR | 1 |
| 2014 | Discriminative Non-Linear Stationary Subspace Analysis for Video ClassificationabstractLow-dimensional representations are key to the success of many video classification algorithms. However, the commonly-used dimensionality reduction techniques fail to account for the fact that only part of the signal is shared across all the videos in one class. As a consequence, the resulting representations contain instance-specific information, which introduces noise in the classification process. In this paper, we introduce non-linear stationary subspace analysis: a method that overcomes this issue by explicitly separating the stationary parts of the video signal (i.e., the parts shared across all videos in one class), from its non-stationary parts (i.e., the parts specific to individual videos). Our method also encourages the new representation to be discriminative, thus accounting for the underlying classification problem. We demonstrate the effectiveness of our approach on dynamic texture recognition, scene classification and action recognition. Mahsa Baktash, Mehrtash Harandi, Brian C. Lovell, Mathieu Salzmann |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2013 | Unsupervised Domain Adaptation by Domain Invariant ProjectionabstractDomain-invariant representations are key to addressing the domain shift problem where the training and test examples follow different distributions. Existing techniques that have attempted to match the distributions of the source and target domains typically compare these distributions in the original feature space. This space, however, may not be directly suitable for such a comparison, since some of the features may have been distorted by the domain shift, or may be domain specific. In this paper, we introduce a Domain Invariant Projection approach: An unsupervised domain adaptation method that overcomes this issue by extracting the information that is invariant across the source and target domains. More specifically, we learn a projection of the data to a low-dimensional latent space where the distance between the empirical distributions of the source and target examples is minimized. We demonstrate the effectiveness of our approach on the task of visual object recognition and show that it outperforms state-of-the-art methods on a standard domain adaptation benchmark dataset. Mahsa Baktash, Mehrtash Harandi, Brian C. Lovell, Mathieu Salzmann |
ICCV | 1 |
| 2013 | Non-Linear Stationary Subspace Analysis with Application to Video ClassificationabstractLow-dimensional representations are key to the success of many video classification algorithms. However, the commonly-used dimensionality reduction techniques fail to account for the fact that only part of the signal is shared across all the videos in one class. As a consequence, the resulting representations contain instance-specific information, which introduces noise in the classification process. In this paper, we introduce Non-Linear Stationary Subspace Analysis: A method that overcomes this issue by explicitly separating the stationary parts of the video signal (i.e., the parts shared across all videos in one class), from its non-stationary parts (i.e., specific to individual videos). We demonstrate the effectiveness of our approach on action recognition, dynamic texture classification and scene recognition. Mahsa Baktash, Mehrtash Harandi, Abbas Bigdeli, Brian C. Lovell, Mathieu Salzmann |
ICML (3) | 1 |
| 2012 | Directional Space-Time Oriented Gradients for 3D Visual Pattern Analysis
Ehsan Norouznezhad, Mehrtash Harandi, Abbas Bigdeli, Mahsa Baktash, Adam Postula, Brian C. Lovell |
ECCV (3) | 4 |
| 2012 | A wireless mesh sensor network for hazard and safety monitoring at the Port of BrisbaneabstractA wireless sensor network (WSN) was designed and implemented to provide reliable long-term hazard monitoring at the Port of Brisbane, Australia. The proposed system consists of four sensor nodes, a wireless gateway and a central monitoring computer. The sensor nodes are capable of measuring a range of hazardous events along with the time and location of events in maritime environment. Each sensor node is also equipped with a Global Positioning System (GPS) module and a ZigBee module. The monitoring server is a personal computer application server with Internet connectivity. Sensor nodes utilize smart algorithm to save energy and AES-128 encryption to encode data prior to sending the data packet through the wireless ZigBee protocol to gateway. The gateway collects and decrypts the data packets and forwards them to the monitoring computer through wireless connection. A database server running on the monitoring computer stores the captured data for visualization and further analysis. The monitoring server is interfaced to Google Maps to overlay real-time data from the sensor nodes onto map in the correct corresponding locations. The WSN system was successfully deployed and tested at the Port of Brisbane, Queensland, Australia. Amin Ahmadi, Abbas Bigdeli, Mahsa Baktash, Brian C. Lovell |
LCN | 3 |
| 2011 | Dynamic resource aware sensor networks: Integration of sensor cloud and ERPsabstractToday's work in the sensor networks community focuses on collecting and processing data from specific networks with associated base stations. One of the most important requirements in these networks is minimizing resource usage such as processing power and storage size on sensor nodes. Resource constraints in the sensor nodes can be divided into four categories: energy, communication, storage and computational power. In this paper, we present an efficient deployment of sensors with possibility of accessing most recent data through information obtained from ERP (enterprise resource planning) systems' re-configuration models. In this scheme the probability of losing any precious data or events would be minimized. In other words, our main focus in this paper is using ERP or high level distributed decision making systems' processed data or prediction models to reduce resource usage needed on sensor nodes. In the area of integration of sensors and ERPs or distributed decision making systems, very little work has been reported. However, those reported in the literature, mostly address relaying data from sensors to ERPs and tend to largely ignore issues that come with sensor resource constraints. Also they don't make use of the information processed and generated by ERPs in the cloud environment to optimize power consumption and transmission frequency of sensors, which is our aim. Mahsa Baktash, Abbas Bigdeli, Brian C. Lovell |
AVSS | 1 |