VLDB 2026 Research / reviewers in the wild / expert
Mehrtash Harandi
dblp:92/5921 · also Mehrtash Tafazzoli Harandi
· DBLP profile ↗
164ranked-venue papers
23as first author
82since 2021 · last 2026
0000-0002-6937-6300ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 125 · 17 first-author · 64 since 2021Graphics, computer vision, multimedia, augmented reality and games · 104 · 15 first-author · 48 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PCGS: Progressive Compression of 3D Gaussian Splattingabstract3D Gaussian Splatting (3DGS) achieves impressive rendering fidelity and speed for novel view synthesis. However, its substantial data size poses a significant challenge for practical applications. While many compression techniques have been proposed, they fail to efficiently utilize existing bitstreams in on-demand applications due to their lack of progressivity, leading to a waste of resource. To address this issue, we propose PCGS (Progressive Compression of 3D Gaussian Splatting), which adaptively controls both the quantity and quality of Gaussians (or anchors) to enable effective progressivity for on-demand applications. For quantity, we introduce a progressive masking strategy that incrementally incorporates new anchors while refining existing ones to enhance fidelity. For quality, we propose a progressive quantization approach that gradually reduces quantization step sizes to achieve finer modeling of Gaussian attributes. Furthermore, to compact the incremental bitstreams, we leverage existing quantization results to refine probability prediction, improving entropy coding efficiency across progressive levels. PCGS achieves progressivity while maintaining compression performance comparable to SoTA non-progressive methods. Yihang Chen 0002, Qianyi Wu, Weiyao Lin, Mehrtash Harandi, Jianfei Cai 0001 |
AAAI | 5 |
| 2026 | DIET: Machine Unlearning on a Data-DietabstractMachine Unlearning (MU) aims to remove the influence of specific knowledge from a pretrained model. Existing methods often rely on retained training data to preserve utility; such dependence is impractical due to privacy and scalability constraints. A further complication arises when unlearning is applied to vision-language models (VLMs), where entangled multimodal representations make targeted forgetting especially challenging. We propose DIET, a principled retain-data-free unlearning method for VLMs that addresses these challenges by leveraging the geometry of hyperbolic space. The core idea is to push forget embeddings toward class-mismatched prototypes located at the boundary of the hyperbolic space. In hyperbolic geometry, points near the boundary become infinitely distant from interior points. As a result, moving forget embeddings to the boundary makes their influence on the model asymptotically negligible. To formalize this, we guide the forgetting process using the Busemann function, which quantifies directional distance to the boundary. We further develop an adaptive scheme based on optimal transport that selects mismatched prototypes for each forget embedding, enabling flexible unlearning dynamics. Extensive experiments on fine-grained datasets such as Flowers102, OxfordPets, and StanfordCars show that DIET achieves an average forget accuracy of 8.06%, while preserving 69.04% utility using only 16 samples per concept, significantly outperforming the best retain-free baselines with a 117.5% improvement in model utility, and showing competitive performance to retain-data baselines with only a 3.79% drop Nilakshan Kunananthaseelan, Jing Wu 0021, Trung Le 0001, Gholamreza Haffari, Mehrtash Harandi |
AAAI | 5 |
| 2026 | Subspace-Guided Knowledge Distillation for Efficient Model TransferabstractCompact models can be effectively trained via Knowledge Distillation (KD), where a lightweight student model learns to replicate the behavior of a larger, high-performing teacher. A persistent challenge in KD lies in the misalignment between the representational spaces of teacher and student networks, especially when they differ in architecture or capacity. To address this, we propose Subspace-Driven Knowledge Distillation (SDMD), a novel framework that mitigates representational disparity by projecting features into an indefinite inner product space. This relaxation from traditional Hilbert spaces enables more flexible geometric alignment, capturing transformations such as rotations and reflections that are often necessary for accurate knowledge transfer. By learning a subspace that bridges the semantic gap between teacher and student, SDMD facilitates more effective distillation without increasing model complexity. We validate SDMD through extensive experiments on large-scale image classification (ImageNet-1K) and object detection (COCO), where it consistently outperforms existing distillation methods. Notably, SDMD-trained models not only achieve state-of-the-art results in distilled settings but also surpass the performance of equivalent models trained from scratch, highlighting the strength of our subspace-based alignment strategy. Zeeshan Hayder, Ali Cheraghian, Lars Petersson, Mehrtash Harandi |
WACV | 4 |
| 2026 | PointCaM: Cut-and-Mix for open-set point cloud learning
Shi Qiu 0001, Weihao Li 0005, Saeed Anwar, Mehrtash Harandi, Nick Barnes, Lars Petersson |
Comput. Vis. Image Underst. | 5 |
| 2026 | Feedforward Compression of Static and Streamable 3D Gaussian SplattingabstractRecent advances in 3D Gaussian Splatting (3DGS) have enabled real-time, high-fidelity novel view synthesis, yet their substantial storage cost remains a major barrier to practical deployment. Although several compression techniques have been explored, they share a common limitation:each existing 3DGS requires per-scene optimization to achieve compression, making the compressionslow and inefficient. In this work, we present Fast Compression of 3D Gaussian Splatting (FCGS), an optimization-free approach that compresses existing 3DGS in a single feed-forward pass, reducing compression time from minutes to seconds. To enhance compression efficiency, we design a multi-path entropy module that routes Gaussian attributes through separate entropy-constrained paths, achieving a better trade-off between size and fidelity. Furthermore, we introduce both inter- and intra-Gaussian context models to effectively remove redundancies for the unstructured Gaussian representation. Experimental results show that FCGS achieves over 20× compression while maintaining high fidelity, outperforming most State-of-The-Art (SoTA) per-scene optimization-based methods. Beyond static scenes, we further extend FCGS to a streamable setting which eliminates redundant temporal information, demonstrating its strong potential for compressing streamable 3DGS data. Yihang Chen 0002, Qianyi Wu, Weiyao Lin, Junhui Hou, Mehrtash Harandi, Jianfei Cai 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Riemannian Implicit Differentiation via a Fixed-Point Equation for Riemannian Bilevel OptimizationabstractVarious Riemannian optimization tasks, such as Riemannian metaoptimization (RMO) and Riemannian metalearning, can be formulated as Riemannian bilevel optimization problems (i.e., the inner-level and outer-level optimization). Implicit differentiation has shown effectiveness in solving RMO, which decouples the computation of outer gradients from the inner-level process, avoiding huge computational burdens. However, extending implicit differentiation to other Riemannian bilevel optimization tasks is nontrivial because it requires much expert involvement for case-by-case derivations. In this article, we propose a Riemannian implicit differentiation method that provides a unified expression for outer gradients, leading to flexible application to other tasks with less expert involvement. Specifically, we formulate the inner-level optimization as a root-finding process of a fixed-point equation, through which the inner-level optimization among different tasks is formulated in a unified way. By differentiating the fixed-point equation, we derive a unified expression for outer gradients, circumventing the case-by-case derivations for different tasks. Then, we present convergence analysis and approximation error analysis, which guarantee the effectiveness of our method in various Riemannian optimization tasks. We further conduct experiments on multiple Riemannian optimization tasks, and the experimental results confirm the effectiveness. Xiaomeng Fan, Yuwei Wu 0001, Zhi Gao 0002, Zhipeng Lu 0003, Mehrtash Harandi, Yunde Jia |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Erasing Undesirable Influence in Diffusion ModelsabstractDiffusion models are highly effective at generating high-quality images but pose risks, such as the unintentional generation of NSFW (not safe for work) content. Although various techniques have been proposed to mitigate unwanted influences in diffusion models while preserving overall performance, achieving a balance between these goals remains challenging. In this work, we introduce EraseDiff, an algorithm designed to preserve the utility of the diffusion model on retained data while removing the unwanted information associated with the data to be forgotten. Our approach formulates this task as a constrained optimization problem using the value function, resulting in a natural first-order algorithm for solving the optimization problem. By altering the generative process to deviate away from the groundtruth denoising trajectory, we update parameters for preservation while controlling constraint reduction to ensure effective erasure, striking an optimal trade-off. Extensive experiments and thorough comparisons with state-of-the-art algorithms demonstrate that EraseDiff effectively preserves the model’s utility, efficacy, and efficiency.$\color{Red} {\text {WARNING}}$: This paper contains sexually explicit imagery that may be offensive in nature. Jing Wu 0021, Trung Le 0001, Munawar Hayat, Mehrtash Harandi |
CVPR | 4 |
| 2025 | A Good Teacher Adapts Their Knowledge for Distillation
Chengyao Qian, Trung Le 0001, Mehrtash Harandi |
ICCV | 3 |
| 2025 | MUNBa: Machine Unlearning Via Nash BargainingabstractMachine Unlearning (MU) aims to selectively erase harmful behaviors from models while retaining the overall utility of the model. As a multi-task learning problem, MU involves balancing objectives related to forgetting specific concepts/data and preserving general performance. A naive integration of these forgetting and preserving objectives can lead to gradient conflicts and dominance, impeding MU algorithms from reaching optimal solutions. To address the gradient conflict and dominance issue, we reformulate MU as a two-player cooperative game, where the two players, namely, the forgetting player and the preservation player, contribute via their gradient proposals to maximize their overall gain and balance their contributions. To this end, inspired by the Nash bargaining theory, we derive a closed-form solution to guide the model toward the Pareto stationary point. Our formulation of MU guarantees an equilibrium solution, where any deviation from the final state would lead to a reduction in the overall objectives for both players, ensuring optimality in each objective. We evaluate our algorithm's effectiveness on a diverse set of tasks across image classification and image generation. Extensive experiments with ResNet, vision-language model CLIP, and text-to-image diffusion models demonstrate that our method outperforms state-of-the-art MU algorithms, achieving a better trade-off between forgetting and preserving. Our results also highlight improvements in forgetting precision, preservation of generalization, and robustness against adversarial attacks. Jing Wu 0021, Mehrtash Harandi |
ICCV | 2 |
| 2025 | Fast Feedforward 3D Gaussian Splatting CompressionabstractWith 3D Gaussian Splatting (3DGS) advancing real-time and high-fidelity rendering for novel view synthesis, storage requirements pose challenges for their widespread adoption. Although various compression techniques have been proposed, previous art suffers from a common limitation: for any existing 3DGS, per-scene optimization is needed to achieve compression, making the compression sluggish and slow. To address this issue, we introduce Fast Compression of 3D Gaussian Splatting (FCGS), an optimization-free model that can compress 3DGS representations rapidly in a single feed-forward pass, which significantly reduces compression time from minutes to seconds. To enhance compression efficiency, we propose a multi-path entropy module that assigns Gaussian attributes to different entropy constraint paths for balance between size and fidelity. We also carefully design both inter- and intra-Gaussian context models to remove redundancies among the unstructured Gaussian blobs. Overall, FCGS achieves a compression ratio of over 20X while maintaining fidelity, surpassing most per-scene SOTA optimization-based methods. Code: github.com/YihangChen-ee/FCGS. Yihang Chen 0002, Qianyi Wu, Weiyao Lin, Mehrtash Harandi, Jianfei Cai 0001 |
ICLR | 5 |
| 2025 | Unleashing Diffusion Transformers for Visual Correspondence by Modulating Massive ActivationsabstractPre-trained stable diffusion models (SD) have shown great advances in visual correspondence.
In this paper, we investigate the capabilities of Diffusion Transformers (DiTs) for accurate dense correspondence. Distinct from SD, DiTs exhibit a critical phenomenon in which very few feature activations exhibit significantly larger values than others, known as massive activations, leading to uninformative representations and significant performance degradation for DiTs.
The massive activations consistently concentrate at very few fixed dimensions across all image patch tokens, holding little local information.
We analyze these dimension-concentrated massive activations and uncover that their concentration is inherently linked to the Adaptive Layer Normalization (AdaLN) in DiTs.
Building on these findings, we propose the Diffusion Transformer Feature (DiTF), a training-free AdaLN-based framework that extracts semantically discriminative features from DiTs.
Specifically, DiTF leverages AdaLN to adaptively localize and normalize massive activations through channel-wise modulation.
Furthermore, a channel discard strategy is introduced to mitigate the adverse effects of massive activations.
Experimental results demonstrate that our DiTF outperforms both DINO and SD-based models and establishes a new state-of-the-art performance for DiTs in different visual correspondence tasks (e.g., with +9.4\% on Spair-71k and +4.4\% on AP-10K-C.S.). Chaofan Gan, Yuanpeng Tu, Tieyuan Chen, Yuxi Li 0009, Mehrtash Harandi, Weiyao Lin |
NeurIPS | 6 |
| 2025 | Token-Level Self-Play with Importance-Aware Guidance for Large Language ModelsabstractLeveraging the power of Large Language Models (LLMs) through preference optimization is crucial for aligning model outputs with human values. Direct Preference Optimization (DPO) has recently emerged as a simple yet effective method by directly optimizing on preference data without the need for explicit reward models. However, DPO typically relies on human-labeled preference data, which can limit its scalability. Self-Play Fine-Tuning (SPIN) addresses this by allowing models to generate their own rejected samples, reducing the dependence on human annotations. Nevertheless, SPIN uniformly applies learning signals across all tokens, ignoring the fine-grained quality variations within responses. As the model improves, rejected samples increasingly contain high-quality tokens, making the uniform treatment of tokens suboptimal. In this paper, we propose SWIFT (Self-Play Weighted Fine-Tuning), a fine-grained self-refinement method that assigns token-level importance weights estimated from a stronger teacher model. Beyond alignment, we also demonstrate that SWIFT serves as an effective knowledge distillation strategy by using the teacher not for logits matching, but for reward-guided token weighting. Extensive experiments on diverse benchmarks and settings demonstrate that SWIFT consistently surpasses both existing alignment approaches and conventional knowledge distillation methods. Tue Le, Hoang Tran Vuong, Quyen Tran, Ngo Van Linh 0001, Mehrtash Harandi, Trung Le 0001 |
NeurIPS | 5 |
| 2025 | Unveiling m-Sharpness Through the Structure of Stochastic Gradient NoiseabstractSharpness-aware minimization (SAM) has emerged as a highly effective technique to improve model generalization, but its underlying principles are not fully understood. We investigate m-sharpness, where SAM performance improves monotonically as the micro-batch size for computing perturbations decreases, a phenomenon critical for distributed training yet lacking rigorous explanation. We leverage an extended Stochastic Differential Equation (SDE) framework and analyze stochastic gradient noise (SGN) to characterize the dynamics of SAM variants, including n-SAM and m-SAM. Our analysis reveals that stochastic perturbations induce an implicit variance-based sharpness regularization whose strength increases as m decreases. Motivated by this insight, we propose Reweighted SAM (RW-SAM), which employs sharpness-weighted sampling to mimic the generalization benefits of m-SAM while remaining parallelizable. Comprehensive experiments validate our theory and method. Haocheng Luo, Mehrtash Harandi, Dinh Q. Phung, Trung Le 0001 |
NeurIPS | 2 |
| 2025 | Geometry-Aware Collaborative Multi-Solutions Optimizer for Model Fine-Tuning with Parameter EfficiencyabstractWe propose a framework grounded in gradient flow theory and informed by geometric structure that provides multiple diverse solutions for a given task, ensuring collaborative results that enhance performance and adaptability across different tasks. This framework enables flexibility, allowing for efficient task-specific fine-tuning while preserving the knowledge of the pre-trained foundation models. Extensive experiments across transfer learning, few-shot learning, and domain generalization show that our proposed approach consistently outperforms existing Bayesian methods, delivering strong performance with affordable computational overhead and offering a practical solution by updating only a small subset of parameters. Van-Anh Nguyen, Trung Le 0001, Mehrtash Harandi, Ehsan Abbasnejad, Thanh-Toan Do, Dinh Q. Phung |
NeurIPS | 3 |
| 2025 | SeCo-INR: Semantically Conditioned Implicit Neural Representations for Improved Medical Image Super-ResolutionabstractImplicit Neural Representations (INRs) have recently advanced the field of deep learning due to their ability to learn continuous representations of signals without the need for large training datasets. Although INR methods have been studied for medical image super-resolution, their adaptability to localized priors in medical images has not been extensively explored. Medical images contain rich anatomical divisions that could provide valuable local prior information to enhance the accuracy and robustness of INRs. In this work, we propose a novel framework, referred to as the Semantically Conditioned INR (SeCo-INR), that conditions an INR using local priors from a medical image, enabling accurate model fitting and interpolation capabilities to achieve super-resolution. Our framework learns a continuous representation of the semantic segmentation features of a medical image and utilizes it to derive the optimal INR for each semantic region of the image. We tested our framework using several medical imaging modalities and achieved higher quantitative scores and more realistic super-resolution outputs compared to state-of-the-art methods. Mevan Ekanayake, Gary F. Egan, Mehrtash Harandi, Zhaolin Chen |
WACV | 4 |
| 2025 | HVQ-VAE: Variational auto-encoder with hyperbolic vector quantization
Shangyu Chen, Pengfei Fang, Mehrtash Harandi, Trung Le 0001, Jianfei Cai 0001, Dinh Q. Phung |
Comput. Vis. Image Underst. | 3 |
| 2025 | Curvature Learning for Generalization of Hyperbolic Neural Networks
Xiaomeng Fan, Yuwei Wu 0001, Zhi Gao 0002, Mehrtash Harandi, Yunde Jia |
Int. J. Comput. Vis. | 4 |
| 2025 | Co-Manifold learning for semi-supervised medical image segmentationabstractIn this study, we investigate jointly learning Hyperbolic and Euclidean space representations and match the consistency for semi-supervised medical image segmentation. We argue that for complex medical volumetric data, hyperbolic spaces are beneficial to model data inductive biases. We propose an approach incorporating the two geometries to co-train a variational encoder–decoder model with a Hyperbolic probabilistic latent space and a separate variational encoder–decoder model with a Euclidean probabilistic latent space with complementary representations, thereby bridging the gap of co-training across manifolds (Co-Manifold learning) in a principled manner. To capture complementary information and hierarchical relationships, we propose a Latent Space Loss aimed at maximizing disagreement between embeddings across manifolds. Additionally, we employ adversarial learning to enhance segmentation performance by guiding the network in hyperbolic latent space using confident regions identified by the network in Euclidean space. Conversely, the network in Euclidean space is informed by hyperbolic uncertainty, creating a dual uncertainty-aware framework that enables the two spaces to collaboratively learn confident regions from each other. Our proposed method achieves competitive results on two benchmarks for semi-supervised medical image segmentation on medical scans. The code is publicly available at: https://github.com/himashi92/Co-Manifold . • A novel co-training approach based on two different geometrical spaces. • Introduces bijective functions to enforce prediction consistency. • Develops a knowledge distillation method with dual uncertainty-aware loss. Himashi Peiris, Zhaolin Chen, Gary F. Egan, Mehrtash Harandi |
Neurocomputing | 4 |
| 2025 | Flashbacks to harmonize stability and plasticity in continual learningabstractWe introduce Flashback Learning (FL), a novel method designed to harmonize the stability and plasticity of models in Continual Learning (CL). Unlike prior approaches that primarily focus on regularizing model updates to preserve old information while learning new concepts, FL explicitly balances this trade-off through a bidirectional form of regularization. This approach effectively guides the model to swiftly incorporate new knowledge while actively retaining its old knowledge. FL operates through a two-phase training process and can be seamlessly integrated into various CL methods, including replay, parameter regularization, distillation, and dynamic architecture techniques. In designing FL, we use two distinct knowledge bases: one to enhance plasticity and another to improve stability. FL ensures a more balanced model by utilizing both knowledge bases to regularize model updates. Theoretically, we analyze how the FL mechanism enhances the stability-plasticity balance. Empirically, FL demonstrates tangible improvements over baseline methods within the same training budget. By integrating FL into at least one representative baseline from each CL category, we observed an average accuracy improvement of up to 4.91% in Class-Incremental and 3.51% in Task-Incremental settings on standard image classification benchmarks. Additionally, measurements of the stability-to-plasticity ratio confirm that FL effectively enhances this balance. FL also outperforms state-of-the-art CL methods on more challenging datasets like ImageNet. The codes of this article will be available at https://github.com/csiro-robotics/Flashback-Learning. Leila Mahmoodi, Peyman Moghadam, Munawar Hayat, Christian Simon, Mehrtash Harandi |
Neural Networks | 5 |
| 2025 | HAC++: Towards 100X Compression of 3D Gaussian Splattingabstract3D Gaussian Splatting (3DGS) has emerged as a promising representation for novel view synthesis, boosting rapid rendering speed with high fidelity. However, the substantial Gaussians and their associated attributes necessitate effective compression techniques. Nevertheless, the sparse and unorganized nature of the point cloud of Gaussians (or anchors in our paper) presents challenges for compression. In this paper, we propose HAC++, which explicitly minimizes the representation's entropy during optimization, enabling efficient arithmetic coding after training for compressed storage. Specifically, to reduce entropy, HAC++ leverages the relationships between unorganized anchors and a structured hash grid, utilizing their mutual information for context modeling. Additionally, HAC++ captures intra-anchor contextual relationships to further enhance compression performance. To facilitate entropy coding, we utilize Gaussian distributions to precisely estimate the probability of each quantized attribute, where an adaptive quantization module is proposed to enable high-precision quantization of these attributes for improved fidelity restoration. Moreover, we incorporate an adaptive masking strategy to eliminate non-effective Gaussians and anchors. Overall, HAC++ achieves a remarkable size reduction of over $100\times$100× compared to vanilla 3DGS when averaged on all datasets, while simultaneously improving fidelity. It also delivers more than $20\times$20× size reduction compared to Scaffold-GS. Yihang Chen 0002, Qianyi Wu, Weiyao Lin, Mehrtash Harandi, Jianfei Cai 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | LaViP: Language-Grounded Visual PromptingabstractWe introduce a language-grounded visual prompting method to adapt the visual encoder of vision-language models for downstream tasks. By capitalizing on language integration, we devise a parameter-efficient strategy to adjust the input of the visual encoder, eliminating the need to modify or add to the model's parameters. Due to this design choice, our algorithm can operate even in black-box scenarios, showcasing adaptability in situations where access to the model's parameters is constrained. We will empirically demonstrate that, compared to prior art, grounding visual prompts with language enhances both the accuracy and speed of adaptation. Moreover, our algorithm excels in base-to-novel class generalization, overcoming limitations of visual prompting and exhibiting the capacity to generalize beyond seen classes. We thoroughly assess and evaluate our method across a variety of image recognition datasets, such as EuroSAT, UCF101, DTD, and CLEVR, spanning different learning situations, including few-shot adaptation, base-to-novel class generalization, and transfer learning. Nilakshan Kunananthaseelan, Jing Zhang 0052, Mehrtash Harandi |
AAAI | 3 |
| 2024 | Concealing Sensitive Samples against Gradient Leakage in Federated LearningabstractFederated Learning (FL) is a distributed learning paradigm that enhances users' privacy by eliminating the need for clients to share raw, private data with the server. Despite the success, recent studies expose the vulnerability of FL to model inversion attacks, where adversaries reconstruct users’ private data via eavesdropping on the shared gradient information. We hypothesize that a key factor in the success of such attacks is the low entanglement among gradients per data within the batch during stochastic optimization. This creates a vulnerability that an adversary can exploit to reconstruct the sensitive data. Building upon this insight, we present a simple, yet effective defense strategy that obfuscates the gradients of the sensitive data with concealed samples. To achieve this, we propose synthesizing concealed samples to mimic the sensitive data at the gradient level while ensuring their visual dissimilarity from the actual sensitive data. Compared to the previous art, our empirical evaluations suggest that the proposed technique provides the strongest protection while simultaneously maintaining the FL performance. Code is located at https://github.com/JingWu321/DCS-2. Jing Wu 0021, Munawar Hayat, Mingyi Zhou, Mehrtash Harandi |
AAAI | 4 |
| 2024 | How Far can we Compress Instant-NGP-Based NeRF?abstractIn recent years, Neural Radiance Field (NeRF) has demonstrated remarkable capabilities in representing 3D scenes. To expedite the rendering process, learnable explicit representations have been introduced for combination with implicit NeRF representation, which however results in a large storage space requirement. In this paper, we introduce the Context-based NeRF Compression (CNC) framework, which leverages highly efficient context models to provide a storage-friendly NeRF representation. Specifically, we excavate both level-wise and dimension-wise context dependencies to enable probability prediction for information entropy reduction. Additionally, we exploit hash collision and occupancy grids as strong prior knowledge for better context modeling. To the best of our knowledge, we are the first to construct and exploit context models for NeRF compression. We achieve a size reduction of 100x and 70× with improved fidelity against the baseline Instant-NGP on Synthesic-NeRF and Tanks and Temples datasets, respectively. Additionally, we attain 86.7% and 82.3% storage size reduction against the SOTA NeRF compression method BiRF. Our code is available here: https://github.com/YihangChen-ee/CNC. Yihang Chen 0002, Qianyi Wu, Mehrtash Harandi, Jianfei Cai 0001 |
CVPR | 3 |
| 2024 | Text-Enhanced Data-Free Approach for Federated Class-Incremental LearningabstractFederated Class-Incremental Learning (FCIL) is an underexplored yet pivotal issue, involving the dynamic addition of new classes in the context of federated learning. In this field, Data-Free Knowledge Transfer (DFKT) plays a crucial role in addressing catastrophic forgetting and data privacy problems. However, prior approaches lack the crucial synergy between DFKT and the model training phases, causing DFKT to encounter difficulties in generating high-quality data from a non-anchored latent space of the old task model. In this paper, we introduce LANDER (Label Text Centered Data-Free Knowledge Transfer) to address this issue by utilizing label text embeddings (LTE) produced by pretrained language models. Specifically, during the model training phase, our approach treats LTE as anchor points and constrains the feature embeddings of corresponding training samples around them, enriching the surrounding area with more meaningful information. In the DFKT phase, by using these LTE anchors, LANDER can synthesize more meaningful samples, thereby effectively addressing the forgetting problem. Additionally, instead of tightly constraining embeddings toward the anchor, the Bounding Loss is introduced to encourage sample embeddings to remain flexible within a defined radius. This approach preserves the natural differences in sample embeddings and mitigates the embedding overlap caused by heterogeneous federated settings. Extensive experiments conducted on CIFAR100, Tiny-ImageNet, and ImageNet demonstrate that LANDER significantly outperforms previous methods and achieves state-of-the-art performance in FCIL. The code is available at https://github.com/tmtuan1307/lander. Minh-Tuan Tran, Trung Le 0001, Xuan-May Le, Mehrtash Harandi, Dinh Q. Phung |
CVPR | 4 |
| 2024 | NAYER: Noisy Layer Data Generation for Efficient and Effective Data-free Knowledge DistillationabstractData-Free Knowledge Distillation (DFKD) has made significant recent strides by transferring knowledge from a teacher neural network to a student neural network without accessing the original data. Nonetheless, existing approaches encounter a significant challenge when attempting to generate samples from random noise inputs, which inherently lack meaningful information. Consequently, these models struggle to effectively map this noise to the ground-truth sample distribution, resulting in prolonging training times and low-quality outputs. In this paper, we propose a novel Noisy Layer Generation method (NAYER) which re-locates the random source from the input to a noisy layer and utilizes the meaningful constant label-text embedding (LTE) as the input. LTE is generated by using the language model once, and then it is stored in memory for all subsequent training processes. The significance of LTE lies in its ability to contain substantial meaningful inter-class information, enabling the generation of high-quality samples with only a few training steps. Simultaneously, the noisy layer plays a key role in addressing the issue of diversity in sample generation by preventing the model from overemphasizing the constrained label information. By reinitializing the noisy layer in each iteration, we aim to facilitate the generation of diverse samples while still retaining the method's efficiency, thanks to the ease of learning provided by LTE. Experiments carried out on multiple datasets demonstrate that our NAYER not only outperforms the state-of-the-art methods but also achieves speeds 5 to 15 times faster than previous approaches. The code is available at https://github.com/tmtuan1307/nayer. Minh-Tuan Tran, Trung Le 0001, Xuan-May Le, Mehrtash Harandi, Quan Hung Tran, Dinh Q. Phung |
CVPR | 4 |
| 2024 | Backpropagation-free Network for 3D Test-time AdaptationabstractReal-world systems often encounter new data over time, which leads to experiencing target domain shifts. Existing Test- Time Adaptation (TTA) methods tend to apply computationally heavy and memory-intensive backpropagation-based approaches to handle this. Here, we propose a novel method that uses a backpropagation-free approach for TTA for the specific case of 3D data. Our model uses a two-stream architecture to maintain knowledge about the source domain as well as complementary target-domain-specific information. The backpropagation-free property of our model helps address the well-known forgetting prob-lem and mitigates the error accumulation issue. The pro-posed method also eliminates the need for the usually noisy process of pseudo-labeling and reliance on costly self-supervised training. Moreover, our method leverages sub-space learning, effectively reducing the distribution vari-ance between the two domains. Furthermore, the source-domain-specific and the target-domain-specific streams are aligned using a novel entropy-based adaptive fusion strat-egy. Extensive experiments on popular benchmarks demon-strate the effectiveness of our method. The code will be available at https://github.com/abie-e/BFTT3D. Yanshuo Wang, Ali Cheraghian, Zeeshan Hayder, Sameera Ramasinghe, Shafin Rahman, David Ahmedt-Aristizabal, Xuesong Li 0001, Lars Petersson, Mehrtash Harandi |
CVPR | 10 |
| 2024 | HAC: Hash-Grid Assisted Context for 3D Gaussian Splatting Compression
Yihang Chen 0002, Qianyi Wu, Weiyao Lin, Mehrtash Harandi, Jianfei Cai 0001 |
ECCV (7) | 4 |
| 2024 | Canonical Shape Projection Is All You Need for 3D Few-Shot Class Incremental Learning
Ali Cheraghian, Zeeshan Hayder, Sameera Ramasinghe, Shafin Rahman, Javad Jafaryahya, Lars Petersson, Mehrtash Harandi |
ECCV (41) | 7 |
| 2024 | Scissorhands: Scrub Data Influence via Connection Sensitivity in Networks
Jing Wu 0021, Mehrtash Harandi |
ECCV (47) | 2 |
| 2024 | Stereographic Projection for Embedding Hierarchical Structures in Hyperbolic Space
Shangyu Chen, Xiaohao Yang, Pengfei Fang, Mehrtash Harandi, Dinh Q. Phung, Jianfei Cai 0001 |
ICPR (9) | 4 |
| 2024 | A CNN system for segmenting tropical cyclones neighborhoods in geostationary imagesabstractWe present progress towards an automated system that can segment tropical cyclones (TCs) and their neighboring regions from geostationary infrared brightness temperature images. The purpose of these segmentation maps is to provide an area that can be used to compute the contribution of TCs to the upwelling radiation budget. Our previous work has identified regions of TC clouds, but it is known that TCs impact a larger area than just those covered by their clouds. Hence it is necessary to properly label both the TC clouds and the associated clear-sky regions in the vicinity of the TCs. Here we present a convolutional neural network method that can be used to reproduce cloud masks generated with our earlier, first-principles algorithm. We also discuss our efforts to create an extended training set of TC masks that include both clouds and clear sky. Joshua May, Mehrtash Harandi, J. Scott Tyo, Elizabeth A Ritchie-Tyo |
IGARSS | 2 |
| 2024 | FIRE: A Dataset for Feedback Integration and Refinement Evaluation of Multimodal ModelsabstractVision language models (VLMs) have achieved impressive progress in diverse applications, becoming a prevalent research direction. In this paper, we build FIRE, a feedback-refinement dataset, consisting of 1.1M multi-turn conversations that are derived from 27 source datasets, empowering VLMs to spontaneously refine their responses based on user feedback across diverse tasks. To scale up the data collection, FIRE is collected in two components: FIRE-100K and FIRE-1M, where FIRE-100K is generated by GPT-4V, and FIRE-1M is freely generated via models trained on FIRE-100K. Then, we build FIRE-Bench, a benchmark to comprehensively evaluate the feedback-refining capability of VLMs, which contains 11K feedback-refinement conversations as the test data, two evaluation settings, and a model to provide feedback for VLMs. We develop the FIRE-LLaVA model by fine-tuning LLaVA on FIRE-100K and FIRE-1M, which shows remarkable feedback-refining capability on FIRE-Bench and outperforms untrained VLMs by 50%, making more efficient user-agent interactions and underscoring the significance of the FIRE dataset. Pengxiang Li 0002, Zhi Gao 0002, Bofei Zhang, Yuwei Wu 0001, Mehrtash Harandi, Yunde Jia, Song-Chun Zhu, Qing Li 0003 |
NeurIPS | 6 |
| 2024 | Explicit Eigenvalue Regularization Improves Sharpness-Aware MinimizationabstractSharpness-Aware Minimization (SAM) has attracted significant attention for its effectiveness in improving generalization across various tasks. However, its underlying principles remain poorly understood. In this work, we analyze SAM’s training dynamics using the maximum eigenvalue of the Hessian as a measure of sharpness and propose a third-order stochastic differential equation (SDE), which reveals that the dynamics are driven by a complex mixture of second- and third-order terms. We show that alignment between the perturbation vector and the top eigenvector is crucial for SAM’s effectiveness in regularizing sharpness, but find that this alignment is often inadequate in practice, which limits SAM's efficiency. Building on these insights, we introduce Eigen-SAM, an algorithm that explicitly aims to regularize the top Hessian eigenvalue by aligning the perturbation vector with the leading eigenvector. We validate the effectiveness of our theory and the practical advantages of our proposed approach through comprehensive experiments. Code is available at https://github.com/RitianLuo/EigenSAM. Haocheng Luo, Tuan Truong, Tung Pham 0001, Mehrtash Harandi, Dinh Q. Phung, Trung Le 0001 |
NeurIPS | 4 |
| 2024 | Continual Test-time Domain Adaptation via Dynamic Sample SelectionabstractThe objective of Continual Test-time Domain Adaptation (CTDA) is to gradually adapt a pre-trained model to a sequence of target domains without accessing the source data. This paper proposes a Dynamic Sample Selection (DSS) method for CTDA. DSS consists of dynamic thresholding, positive learning, and negative learning processes. Traditionally, models learn from unlabeled unknown environment data and equally rely on all samples’ pseudo-labels to update their parameters through self-training. However, noisy predictions exist in these pseudo-labels, so all samples are not equally trustworthy. Therefore, in our method, a dynamic thresholding module is first designed to select suspected low-quality from high-quality samples. The selected low-quality samples are more likely to be wrongly predicted. Therefore, we apply joint positive and negative learning on both high- and low-quality samples to reduce the risk of using wrong information. We conduct extensive experiments that demonstrate the effectiveness of our proposed method for CTDA in the image domain, outperforming the state-of-the-art results. Furthermore, our approach is also evaluated in the 3D point cloud domain, showcasing its versatility and potential for broader applicability. Yanshuo Wang, Ali Cheraghian, Shafin Rahman, David Ahmedt-Aristizabal, Lars Petersson, Mehrtash Harandi |
WACV | 7 |
| 2024 | Automated Segmentation of Tropical Cyclone Clouds in Geostationary Infrared ImagesabstractWe demonstrate that a convolutional neural network (CNN) based on the U-Net architecture can be used to create a cloud mask data set that accurately identifies the clouds associated with tropical cyclones (TCs). The CNN can be trained using a single year of cloud masks produced by an earlier first-principles algorithm, and the results are insensitive to the specific year of training data used. These masks were originally created in order to compute the upwelling radiation due to TC clouds, and we show that the predicted masks result in both pixel areas and radiation calculations that are nearly identical to those computed using the earlier masks. Joshua May, Elizabeth A. Ritchie, Mehrtash Harandi, J. Scott Tyo |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2024 | Curved Geometric Networks for Visual Anomaly RecognitionabstractLearning a latent embedding to understand the underlying nature of data distribution is often formulated in Euclidean spaces with zero curvature. However, the success of the geometry constraints, posed in the embedding space, indicates that curved spaces might encode more structural information, leading to better discriminative power and hence richer representations. In this work, we investigate the benefits of the curved space for analyzing anomalous, open-set, or out-of-distribution (OOD) objects in data. This is achieved by considering embeddings via three geometry constraints, namely, spherical geometry (with positive curvature), hyperbolic geometry (with negative curvature), or mixed geometry (with both positive and negative curvatures). Three geometric constraints can be chosen interchangeably in a unified design, given the task at hand. Tailored for the embeddings in the curved space, we also formulate functions to compute the anomaly score. Two types of geometric modules (i.e., geometric-in-one (GiO) and geometric-in-two (GiT) models) are proposed to plug in the original Euclidean classifier, and anomaly scores are computed from the curved embeddings. We evaluate the resulting designs under a diverse set of visual recognition scenarios, including image detection (multiclass OOD detection and one-class anomaly detection) and segmentation (multiclass anomaly segmentation and one-class anomaly segmentation). The empirical results show the effectiveness of our proposal through consistent improvement over various scenarios. The code is made available at https://github.com/JHome1/GiO-GiT. Pengfei Fang, Weihao Li 0005, Junlin Han, Lars Petersson, Mehrtash Harandi |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | GOSS: towards generalized open-set semantic segmentationabstractAbstract In this paper, we extend Open-set Semantic Segmentation (OSS) into a new image segmentation task called Generalized Open-set Semantic Segmentation (GOSS). Previously, with well-known OSS, the intelligent agents only detect unknown regions without further processing, limiting their perception capacity of the environment. It stands to reason that further analysis of the detected unknown pixels would be beneficial for agents’ decision-making. Therefore, we propose GOSS, which holistically unifies the abilities of two well-defined segmentation tasks, i.e. OSS and generic segmentation. Specifically, GOSS classifies pixels as belonging to known classes, and clusters (or groups) of pixels of unknown class are labelled as such. We propose a metric that balances the pixel classification and clustering aspects to evaluate this newly expanded task. Moreover, we build benchmark tests on existing datasets and propose neural architectures as baselines. Our experiments on multiple benchmarks demonstrate the effectiveness of our baselines. Code is made available at https://github.com/JHome1/GOSS_Segmentor . Weihao Li 0005, Junlin Han, Jiyang Zheng, Pengfei Fang, Mehrtash Harandi, Lars Petersson |
Vis. Comput. | 6 |
| 2024 | Publisher Correction: GOSS: towards generalized open-set semantic segmentation
Weihao Li 0005, Junlin Han, Jiyang Zheng, Pengfei Fang, Mehrtash Harandi, Lars Petersson |
Vis. Comput. | 6 |
| 2023 | Exploring Data Geometry for Continual LearningabstractContinual learning aims to efficiently learn from a non-stationary stream of data while avoiding forgetting the knowledge of old data. In many practical applications, data complies with non-Euclidean geometry. As such, the commonly used Euclidean space cannot gracefully capture non-Euclidean geometric structures of data, leading to in-ferior results. In this paper, we study continual learning from a novel perspective by exploring data geometry for the non-stationary stream of data. Our method dynamically expands the geometry of the underlying space to match growing geometric structures induced by new data, and pre-vents forgetting by keeping geometric structures of old data into account. In doing so, making use of the mixed cur-vature space, we propose an incremental search scheme, through which the growing geometric structures are en-coded. Then, we introduce an angular-regularization loss and a neighbor-robustness loss to train the model, capa-ble of penalizing the change of global geometric structures and local geometric structures. Experiments show that our method achieves better performance than baseline methods designed in Euclidean space. Zhi Gao 0002, Yunde Jia, Mehrtash Harandi, Yuwei Wu 0001 |
CVPR | 5 |
| 2023 | Energy-based Self-Training and Normalization for Unsupervised Domain AdaptationabstractWe propose an Unsupervised Domain Adaptation (UDA) method by making use of Energy-Based Learning (EBL) and demonstrate 1. EBL can be used to improve the instance selection for a self-training task on the unlabelled target domain, and 2. alignment and normalizing energy scores can learn domain-invariant representations. For the former, we show that an energy-based selection criterion can be used to model instance selections by mimicking the joint distribution between data and predictions in the target domain. As per learning domain invariant representations, we show that stable domain alignment can be achieved by a combined energy alignment and an energy normalization process. We implement our method in consistent with the vision-transformer (ViT) backbone and show that our proposed method can outperform state-of-the-art ViT based UDA methods on diverse benchmarks (DomainNet, Office-Home, and VISDA2017). Samitha Herath, Basura Fernando, Ehsan Abbasnejad, Munawar Hayat, Shahram Khadivi, Mehrtash Harandi, Seyed Hamid Rezatofighi, Gholamreza Haffari |
ICCV | 6 |
| 2023 | Hyperbolic Audio-visual Zero-shot LearningabstractAudio-visual zero-shot learning aims to classify samples consisting of a pair of corresponding audio and video sequences from classes that are not present during training. An analysis of the audio-visual data reveals a large degree of hyperbolicity, indicating the potential benefit of using a hyperbolic transformation to achieve curvature-aware geometric learning, with the aim of exploring more complex hierarchical data structures for this task. The proposed approach employs a novel loss function that incorporates cross-modality alignment between video and audio features in the hyperbolic space. Additionally, we explore the use of multiple adaptive curvatures for hyperbolic projections. The experimental results on this very challenging task demonstrate that our proposed hyperbolic approach for zero-shot learning outperforms the SOTA method on three datasets: VGGSound-GZSL, UCF-GZSL, and ActivityNet-GZSL achieving a harmonic mean (HM) improvement of around 3.0%, 7.0%, and 5.3%, respectively. Zeeshan Hayder, Junlin Han, Pengfei Fang, Mehrtash Harandi, Lars Petersson |
ICCV | 5 |
| 2023 | Can we Distill Knowledge from Powerful Teachers Directly?abstractKnowledge distillation efficiently improves a small model’s performance by mimicking the teacher model’s behavior. Most existing methods assume that distilling from a large and accurate teacher model leads to better student models. However, several studies show the difficulty of distillation from large teacher models and opt for heuristics to address this. In this work, we demonstrate that large teacher models can still be effective in knowledge distillation. We show that the spurious features learned by large models are the cause of difficulty in knowledge distillation for small students. To overcome this issue, we propose employing ℓ1regularization to prevent teacher models from learning an excessive number of spurious features. Our method alleviates the poor learning for small students when there is a significant disparity in size between the teachers and students. We achieve substantial improvement on various architectures, e.g., ResNet, WideResNet, and VGG, on diverse datasets, including CIFAR-100, Tiny-ImageNet, and ImageNet. Chengyao Qian, Munawar Hayat, Mehrtash Harandi |
ICIP | 3 |
| 2023 | Vector Quantized Wasserstein Auto-EncoderabstractLearning deep discrete latent presentations offers a promise of better symbolic and summarized abstractions that are more useful to subsequent downstream tasks. Inspired by the seminal Vector Quantized Variational Auto-Encoder (VQ-VAE), most of work in learning deep discrete representations has mainly focused on improving the original VQ-VAE form and none of them has studied learning deep discrete representations from the generative viewpoint. In this work, we study learning deep discrete representations from the generative viewpoint. Specifically, we endow discrete distributions over sequences of codewords and learn a deterministic decoder that transports the distribution over the sequences of codewords to the data distribution via minimizing a WS distance between them. We develop further theories to connect it with the clustering viewpoint of WS distance, allowing us to have a better and more controllable clustering solution. Finally, we empirically evaluate our method on several well-known benchmarks, where it achieves better qualitative and quantitative performances than the other VQ-VAE variants in terms of the codebook utilization and image reconstruction/generation. Long Tung Vuong, Trung Le 0001, He Zhao 0001, Chuanxia Zheng, Mehrtash Harandi, Jianfei Cai 0001, Dinh Q. Phung |
ICML | 5 |
| 2023 | L3DMC: Lifelong Learning Using Distillation via Mixed-Curvature Space
Kaushik Roy 0008, Peyman Moghadam, Mehrtash Harandi |
MICCAI (2) | 3 |
| 2023 | EndoSurf: Neural Surface Reconstruction of Deformable Tissues with Stereo Endoscope Videos
Ruyi Zha, Xuelian Cheng, Hongdong Li, Mehrtash Harandi, ZongYuan Ge |
MICCAI (9) | 4 |
| 2023 | LAVA:Label-efficient Visual Learning and AdaptationabstractWe present LAVA, a simple yet effective method for multi-domain visual transfer learning with limited data. LAVA builds on a few recent innovations to enable adapting to partially labelled datasets with class and domain shifts. First, LAVA learns self-supervised visual representations on the source dataset and ground them using class label semantics to overcome transfer collapse problems associated with supervised pretraining. Secondly, LAVA maximises the gains from unlabelled target data via a novel method which uses multi-crop augmentations to obtain highly robust pseudo-labels. By combining these ingredients, LAVA achieves a new state-of-the-art on ImageNet semi-supervised protocol, as well as on 7 out of 10 datasets in multi-domain few-shot learning on the Meta-dataset.1 Islam Nassar, Munawar Hayat, Ehsan Abbasnejad, Seyed Hamid Rezatofighi, Mehrtash Harandi, Gholamreza Haffari |
WACV | 5 |
| 2023 | Poincaré Kernels for Hyperbolic Representations
Pengfei Fang, Mehrtash Harandi, Zhen-Zhong Lan, Lars Petersson |
Int. J. Comput. Vis. | 2 |
| 2023 | Guest Editorial : Learning with Manifolds in Computer Vision
Mohamed Daoudi, Mehrtash Harandi, Vittorio Murino |
Image Vis. Comput. | 2 |
| 2023 | Subspace distillation for continual learningabstractAn ultimate objective in continual learning is to preserve knowledge learned in preceding tasks while learning new tasks. To mitigate forgetting prior knowledge, we propose a novel knowledge distillation technique that takes into the account the manifold structure of the latent/output space of a neural network in learning novel tasks. To achieve this, we propose to approximate the data manifold up-to its first order, hence benefiting from linear subspaces to model the structure and maintain the knowledge of a neural network while learning novel concepts. We demonstrate that the modeling with subspaces provides several intriguing properties, including robustness to noise and therefore effective for mitigating Catastrophic Forgetting in continual learning. We also discuss and show how our proposed method can be adopted to address both classification and segmentation problems. Empirically, we observe that our proposed method outperforms various continual learning methods on several challenging datasets including Pascal VOC, and Tiny-Imagenet. Furthermore, we show how the proposed method can be seamlessly combined with existing learning approaches to improve their performances. The codes of this article will be available at https://github.com/csiro-robotics/SDCL. Kaushik Roy 0008, Christian Simon, Peyman Moghadam, Mehrtash Harandi |
Neural Networks | 4 |
| 2023 | Learning to Optimize on Riemannian ManifoldsabstractMany learning tasks are modeled as optimization problems with nonlinear constraints, such as principal component analysis and fitting a Gaussian mixture model. A popular way to solve such problems is resorting to Riemannian optimization algorithms, which yet heavily rely on both human involvement and expert knowledge about Riemannian manifolds. In this paper, we propose a Riemannian meta-optimization method to automatically learn a Riemannian optimizer. We parameterize the Riemannian optimizer by a novel recurrent network and utilize Riemannian operations to ensure that our method is faithful to the geometry of manifolds. The proposed method explores the distribution of the underlying data by minimizing the objective of updated parameters, and hence is capable of learning task-specific optimizations. We introduce a Riemannian implicit differentiation training scheme to achieve efficient training in terms of numerical stability and computational cost. Unlike conventional meta-optimization training schemes that need to differentiate through the whole optimization trajectory, our training scheme is only related to the final two optimization steps. In this way, our training scheme avoids the exploding gradient problem, and significantly reduces the computational load and memory footprint. We discuss experimental results across various constrained problems, including principal component analysis on Grassmann manifolds, face recognition, person re-identification, and texture image classification on Stiefel manifolds, clustering and similarity learning on symmetric positive definite manifolds, and few-shot learning on hyperbolic manifolds. Zhi Gao 0002, Yuwei Wu 0001, Xiaomeng Fan, Mehrtash Harandi, Yunde Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Curvature-Adaptive Meta-Learning for Fast Adaptation to Manifold DataabstractMeta-learning methods are shown to be effective in quickly adapting a model to novel tasks. Most existing meta-learning methods represent data and carry out fast adaptation in euclidean space. In fact, data of real-world applications usually resides in complex and various Riemannian manifolds. In this paper, we propose a curvature-adaptive meta-learning method that achieves fast adaptation to manifold data by producing suitable curvature. Specifically, we represent data in the product manifold of multiple constant curvature spaces and build a product manifold neural network as the base-learner. In this way, our method is capable of encoding complex manifold data into discriminative and generic representations. Then, we introduce curvature generation and curvature updating schemes, through which suitable product manifolds for various forms of data manifolds are constructed via few optimization steps. The curvature generation scheme identifies task-specific curvature initialization, leading to a shorter optimization trajectory. The curvature updating scheme automatically produces appropriate learning rate and search direction for curvature, making a faster and more adaptive optimization paradigm compared to hand-designed optimization schemes. We evaluate our method on a broad set of problems including few-shot classification, few-shot regression, and reinforcement learning tasks. Experimental results show that our method achieves substantial improvements as compared to meta-learning methods ignoring the geometry of the underlying space. Zhi Gao 0002, Yuwei Wu 0001, Mehrtash Harandi, Yunde Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | RMAML: Riemannian meta-learning with orthogonality constraints
Hadi Tabealhojeh, Peyman Adibi, Hossein Karshenas, Soumava Kumar Roy, Mehrtash Harandi |
Pattern Recognit. | 5 |
| 2023 | Domain Neural AdaptationabstractDomain adaptation is concerned with the problem of generalizing a classification model to a target domain with little or no labeled data, by leveraging the abundant labeled data from a related source domain. The source and target domains possess different joint probability distributions, making it challenging for model generalization. In this article, we introduce domain neural adaptation (DNA): an approach that exploits nonlinear deep neural network to 1) match the source and target joint distributions in the network activation space and 2) learn the classifier in an end-to-end manner. Specifically, we employ the relative chi-square divergence to compare the two joint distributions, and show that the divergence can be estimated via seeking the maximal value of a quadratic functional over the reproducing kernel hilbert space. The analytic solution to this maximization problem enables us to explicitly express the divergence estimate as a function of the neural network mapping. We optimize the network parameters to minimize the estimated joint distribution divergence and the classification loss, yielding a classification model that generalizes well to the target domain. Empirical results on several visual datasets demonstrate that our solution is statistically better than its competitors. Sentao Chen, Zijie Hong, Mehrtash Harandi, Xiaowei Yang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Efficient Riemannian Meta-Optimization by Implicit DifferentiationabstractTo solve optimization problems with nonlinear constrains, the recently developed Riemannian meta-optimization methods show promise, which train neural networks as an optimizer to perform optimization on Riemannian manifolds. A key challenge is the heavy computational and memory burdens, because computing the meta-gradient with respect to the optimizer involves a series of time-consuming derivatives, and stores large computation graphs in memory. In this paper, we propose an efficient Riemannian meta-optimization method that decouples the complex computation scheme from the meta-gradient. We derive Riemannian implicit differentiation to compute the meta-gradient by establishing a link between Riemannian optimization and the implicit function theorem. As a result, the updating our optimizer is only related to the final two iterations, which in turn speeds up our method and reduces the memory footprint significantly. We theoretically study the computational load and memory footprint of our method for long optimization trajectories, and conduct an empirical study to demonstrate the benefits of the proposed method. Evaluations of three optimization problems on different Riemannian manifolds show that our method achieves state-of-the-art performance in terms of the convergence speed and the quality of optima. Xiaomeng Fan, Yuwei Wu 0001, Zhi Gao 0002, Yunde Jia, Mehrtash Harandi |
AAAI | 5 |
| 2022 | Adaptive Poincaré Point to Set Distance for Few-Shot ClassificationabstractLearning and generalizing from limited examples, i.e., few-shot learning, is of core importance to many real-world vision applications. A principal way of achieving few-shot learning is to realize an embedding where samples from different classes are distinctive. Recent studies suggest that embedding via hyperbolic geometry enjoys low distortion for hierarchical and structured data, making it suitable for few-shot learning. In this paper, we propose to learn a context-aware hyperbolic metric to characterize the distance between a point and a set associated with a learned set to set distance. To this end, we formulate the metric as a weighted sum on the tangent bundle of the hyperbolic space and develop a mechanism to obtain the weights adaptively, based on the constellation of the points. This not only makes the metric local but also dependent on the task in hand, meaning that the metric will adapt depending on the samples that it compares. We empirically show that such metric yields robustness in the presence of outliers and achieves a tangible improvement over baseline models. This includes the state-of-the-art results on five popular few-shot classification benchmarks, namely mini-ImageNet, tiered-ImageNet, Caltech-UCSD Birds-200-2011(CUB), CIFAR-FS, and FC100. Rongkai Ma, Pengfei Fang, Tom Drummond, Mehrtash Harandi |
AAAI | 4 |
| 2022 | A Differentiable Distance Approximation for Fairer Image Classification
Nicholas Rosa, Tom Drummond, Mehrtash Harandi |
ACCV (6) | 3 |
| 2022 | Implicit Motion Handling for Video Camouflaged Object DetectionabstractWe propose a new video camouflaged object detection (VCOD) framework that can exploit both short-term dynamics and long-term temporal consistency to detect camouflaged objects from video frames. An essential property of camouflaged objects is that they usually exhibit patterns similar to the background and thus make them hard to identify from still images. Therefore, effectively handling temporal dynamics in videos becomes the key for the VCOD task as the camouflaged objects will be noticeable when they move. However, current VCOD methods often leverage homography or optical flows to represent motions, where the detection error may accumulate from both the motion estimation error and the segmentation error. On the other hand, our method unifies motion estimation and object segmentation within a single optimization framework. Specifically, we build a dense correlation volume to implicitly capture motions between neighbouring frames and utilize the final segmentation supervision to optimize the implicit motion estimation and segmentation jointly. Furthermore, to enforce temporal consistency within a video sequence, we jointly utilize a spatio-temporal transformer to refine the short-term predictions. Extensive experiments on VCOD benchmarks demonstrate the architectural effectiveness of our approach. We also provide a large-scale VCOD dataset named MoCA-Mask with pixel-level handcrafted ground-truth masks and construct a comprehensive VCOD bench-mark with previous methods to facilitate research in this direction. Dataset Link: https://xueliancheng.github.io/SLT-Net-project. Xuelian Cheng, Huan Xiong, Deng-Ping Fan, Yiran Zhong, Mehrtash Harandi, Tom Drummond, ZongYuan Ge |
CVPR | 5 |
| 2022 | On Generalizing Beyond Domains in Cross-Domain Continual LearningabstractHumans have the ability to accumulate knowledge of new tasks in varying conditions, but deep neural networks of-ten suffer from catastrophic forgetting of previously learned knowledge after learning a new task. Many recent methods focus on preventing catastrophic forgetting under the assumption of train and test data following similar distributions. In this work, we consider a more realistic scenario of continual learning under domain shifts where the model must generalize its inference to an unseen domain. To this end, we encourage learning semantically meaningful features by equipping the classifier with class similarity metrics as learning parameters which are obtained through Mahalanobis similarity computations. Learning of the backbone representation along with these extra parameters is done seamlessly in an end-to-end manner. In addition, we propose an approach based on the exponential moving average of the parameters for better knowledge distillation. We demonstrate that, to a great extent, existing continual learning algorithms fail to handle the forgetting issue under multiple distributions, while our proposed approach learns new tasks under domain shift with accuracy boosts up to 10% on challenging datasets such as DomainNet and OfficeHome. Christian Simon, Masoud Faraki, Yi-Hsuan Tsai, Xiang Yu 0002, Samuel Schulter, Yumin Suh, Mehrtash Harandi, Manmohan Krishna Chandraker |
CVPR | 7 |
| 2022 | Learning Instance and Task-Aware Dynamic Kernels for Few-Shot Learning
Rongkai Ma, Pengfei Fang, Gil Avraham, Tianyu Zhu 0001, Tom Drummond, Mehrtash Harandi |
ECCV (20) | 7 |
| 2022 | Deep Laparoscopic Stereo Matching with Transformers
Xuelian Cheng, Yiran Zhong, Mehrtash Harandi, Tom Drummond, Zhiyong Wang 0001, ZongYuan Ge |
MICCAI (8) | 3 |
| 2022 | A Robust Volumetric Transformer for Accurate 3D Tumor Segmentation
Himashi Peiris, Munawar Hayat, Zhaolin Chen, Gary F. Egan, Mehrtash Harandi |
MICCAI (5) | 5 |
| 2022 | Hyperbolic Feature Augmentation via Distribution Estimation and Infinite Sampling on ManifoldsabstractLearning in hyperbolic spaces has attracted growing attention recently, owing to their capabilities in capturing hierarchical structures of data. However, existing learning algorithms in the hyperbolic space tend to overfit when limited data is given. In this paper, we propose a hyperbolic feature augmentation method that generates diverse and discriminative features in the hyperbolic space to combat overfitting. We employ a wrapped hyperbolic normal distribution to model augmented features, and use a neural ordinary differential equation module that benefits from meta-learning to estimate the distribution. This is to reduce the bias of estimation caused by the scarcity of data. We also derive an upper bound of the augmentation loss, which enables us to train a hyperbolic model by using an infinite number of augmentations. Experiments on few-shot learning and continual learning tasks show that our method significantly improves the performance of hyperbolic algorithms in scarce data regimes. Zhi Gao 0002, Yuwei Wu 0001, Yunde Jia, Mehrtash Harandi |
NeurIPS | 4 |
| 2022 | On Enforcing Better Conditioned Meta-Learning for Rapid Few-Shot AdaptationabstractInspired by the concept of preconditioning, we propose a novel method to increase adaptation speed for gradient-based meta-learning methods without incurring extra parameters. We demonstrate that recasting the optimisation problem to a non-linear least-squares formulation provides a principled way to actively enforce a well-conditioned parameter space for meta-learning models based on the concepts of the condition number and local curvature. Our comprehensive evaluations show that the proposed method significantly outperforms its unconstrained counterpart especially during initial adaptation steps, while achieving comparable or better overall results on several few-shot classification tasks – creating the possibility of dynamically choosing the number of adaptation steps at inference time. Markus Hiller, Mehrtash Harandi, Tom Drummond |
NeurIPS | 2 |
| 2022 | Rethinking Generalization in Few-Shot ClassificationabstractSingle image-level annotations only correctly describe an often small subset of an image’s content, particularly when complex real-world scenes are depicted. While this might be acceptable in many classification scenarios, it poses a significant challenge for applications where the set of classes differs significantly between training and test time. In this paper, we take a closer look at the implications in the context of few-shot learning. Splitting the input samples into patches and encoding these via the help of Vision Transformers allows us to establish semantic correspondences between local regions across images and independent of their respective class. The most informative patch embeddings for the task at hand are then determined as a function of the support set via online optimization at inference time, additionally providing visual interpretability of ‘what matters most’ in the image. We build on recent advances in unsupervised training of networks via masked image modelling to overcome the lack of fine-grained labels and learn the more general statistical structure of the data while avoiding negative image-level annotation influence, aka supervision collapse. Experimental results show the competitiveness of our approach, achieving new state-of-the-art results on four popular few-shot classification benchmarks for 5-shot and 1-shot scenarios. Markus Hiller, Rongkai Ma, Mehrtash Harandi, Tom Drummond |
NeurIPS | 3 |
| 2022 | Meta-Learning for Multi-Label Few-Shot ClassificationabstractEven with the luxury of having abundant data, multi-label classification is widely known to be a challenging task to address. This work targets the problem of multi-label meta-learning, where a model learns to predict multiple labels within a query (e.g., an image) by just observing a few supporting examples. In doing so, we first propose a benchmark for Few-Shot Learning (FSL) with multiple labels per sample. Next, we discuss and extend several solutions specifically designed to address the conventional and single-label FSL, to work in the multi-label regime. Lastly, we introduce a neural module to estimate the label count of a given sample by exploiting the relational inference. We will show empirically the benefit of the label count module, the label propagation algorithm, and the extensions of conventional FSL methods on three challenging datasets, namely MS-COCO, iMaterialist, and Open MIC. Overall, our thorough experiments suggest that the proposed label-propagation algorithm in conjunction with the neural label count module (NLC) shall be considered as the method of choice. Christian Simon, Piotr Koniusz, Mehrtash Harandi |
WACV | 3 |
| 2022 | Towards a Robust Differentiable Architecture Search under Label NoiseabstractNeural Architecture Search (NAS) is the game changer in designing robust neural architectures. Architectures designed by NAS outperform or compete with the best manual network designs in terms of accuracy, size, memory footprint and FLOPs. That said, previous studies focus on developing NAS algorithms for clean high quality data, a restrictive and somewhat unrealistic assumption. In this paper, focusing on the differentiable NAS algorithms, we show that vanilla NAS algorithms suffer from a performance loss if class labels are noisy. To combat this issue, we make use of the principle of information bottleneck as a regularizer. This leads us to develop a noise injecting operation that is included during the learning process, preventing the network from learning from noisy samples. Our empirical evaluations show that the noise injecting operation does not degrade the performance of the NAS algorithm if the data is indeed clean. In contrast, if the data is noisy, the architecture learned by our algorithm comfortably outperforms algorithms specifically equipped with sophisticated mechanisms to learn in the presence of label noise. In contrast to many algorithms designed to work in the presence of noisy labels, prior knowledge about the properties of the noise and its characteristics are not required for our algorithm. Christian Simon, Piotr Koniusz, Lars Petersson, Mehrtash Harandi |
WACV | 5 |
| 2022 | Learning Log-Determinant Divergences for Positive Definite MatricesabstractRepresentations in the form of Symmetric Positive Definite (SPD) matrices have been popularized in a variety of visual learning applications due to their demonstrated ability to capture rich second-order statistics of visual data. There exist several similarity measures for comparing SPD matrices with documented benefits. However, selecting an appropriate measure for a given problem remains a challenge and in most cases, is the result of a trial-and-error process. In this paper, we propose to learn similarity measures in a data-driven manner. To this end, we capitalize on the αβ-log-det divergence, which is a meta-divergence parametrized by scalars α and β, subsuming a wide family of popular information divergences on SPD matrices for distinct and discrete values of these parameters. Our key idea is to cast these parameters in a continuum and learn them from data. We systematically extend this idea to learn vector-valued parameters, thereby increasing the expressiveness of the underlying non-linear measure. We conjoin the divergence learning problem with several standard tasks in machine learning, including supervised discriminative dictionary learning and unsupervised SPD matrix clustering. We present Riemannian gradient descent schemes for optimizing our formulations efficiently, and show the usefulness of our method on eight standard computer vision tasks. Anoop Cherian, Panagiotis Stanitsas, Jue Wang 0010, Mehrtash Harandi, Vassilios Morellas, Nikolaos Papanikolopoulos |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | Attention in Attention Networks for Person RetrievalabstractThis paper generalizes the Attention in Attention (AiA) mechanism, in P. Fang et al., 2019 by employing explicit mapping in reproducing kernel Hilbert spaces to generate attention values of the input feature map. The AiA mechanism models the capacity of building inter-dependencies among the local and global features by the interaction of inner and outer attention modules. Besides a vanilla AiA module, termed linear attention with AiA, two non-linear counterparts, namely, second-order polynomial attention and Gaussian attention, are also proposed to utilize the non-linear properties of the input features explicitly, via the second-order polynomial kernel and Gaussian kernel approximation. The deep convolutional neural network, equipped with the proposed AiA blocks, is referred to as Attention in Attention Network (AiA-Net). The AiA-Net learns to extract a discriminative pedestrian representation, which combines complementary person appearance and corresponding part features. Extensive ablation studies verify the effectiveness of the AiA mechanism and the use of non-linear features hidden in the feature map for attention design. Furthermore, our approach outperforms current state-of-the-art by a considerable margin across a number of benchmarks. In addition, state-of-the-art performance is also achieved in the video person retrieval task with the assistance of the proposed AiA blocks. Pengfei Fang, Jieming Zhou, Soumava Kumar Roy, Pan Ji, Lars Petersson, Mehrtash Harandi |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2021 | Semi-Supervised Metric Learning: A Deep ResurrectionabstractDistance Metric Learning (DML) seeks to learn a discriminative embedding where similar examples are closer, and dissimilar examples are apart. In this paper, we address the problem of Semi-Supervised DML (SSDML) that tries to learn a metric using a few labeled examples, and abundantly available unlabeled examples. SSDML is important because it is infeasible to manually annotate all the examples present in a large dataset. Surprisingly, with the exception of a few classical approaches that learn a linear Mahalanobis metric, SSDML has not been studied in the recent years, and lacks approaches in the deep SSDML scenario. In this paper, we address this challenging problem, and revamp SSDML with respect to deep learning. In particular, we propose a stochastic, graph-based approach that first propagates the affinities between the pairs of examples from labeled data, to that of the unlabeled pairs. The propagated affinities are used to mine triplet based constraints for metric learning. We impose orthogonality constraint on the metric parameters, as it leads to a better performance by avoiding a model collapse. Ujjal Kr Dutta, Mehrtash Harandi, Chellu Chandra Sekhar |
AAAI | 2 |
| 2021 | Learning a Gradient-free Riemannian Optimizer on Tangent Spaces
Xiaomeng Fan, Zhi Gao 0002, Yuwei Wu 0001, Yunde Jia, Mehrtash Harandi |
AAAI | 5 |
| 2021 | Semantic-Aware Knowledge Distillation for Few-Shot Class-Incremental LearningabstractFew-shot class incremental learning (FSCIL) portrays the problem of learning new concepts gradually, where only a few examples per concept are available to the learner. Due to the limited number of examples for training, the techniques developed for standard incremental learning cannot be applied verbatim to FSCIL. In this work, we introduce a distillation algorithm to address the problem of FSCIL and propose to make use of semantic information during training. To this end, we make use of word embeddings as semantic information which is cheap to obtain and which facilitate the distillation process. Furthermore, we propose a method based on an attention mechanism on multiple parallel embeddings of visual data to align visual and semantic vectors, which reduces issues related to catastrophic forgetting. Via experiments on MiniImageNet, CUB200, and CIFAR100 dataset, we establish new state-of-the-art results by outperforming existing approaches. Ali Cheraghian, Shafin Rahman, Pengfei Fang, Soumava Kumar Roy, Lars Petersson, Mehrtash Harandi |
CVPR | 6 |
| 2021 | Reinforced Attention for Few-Shot Learning and BeyondabstractFew-shot learning aims to correctly recognize query samples from unseen classes given a limited number of support samples, often by relying on global embeddings of images. In this paper, we propose to equip the backbone network with an attention agent, which is trained by reinforcement learning. The policy gradient algorithm is employed to train the agent towards adaptively localizing the representative regions on feature maps over time. We further design a reward function based on the prediction of the held-out data, thus helping the attention mechanism to generalize better across the unseen classes. The extensive experiments show, with the help of the reinforced attention, that our embedding network has the capability to progressively generate a more discriminative representation in few-shot learning. Moreover, experiments on the task of image classification also show the effectiveness of the proposed design. Pengfei Fang, Weihao Li 0005, Tong Zhang 0023, Christian Simon, Mehrtash Harandi, Lars Petersson |
CVPR | 6 |
| 2021 | On Learning the Geodesic Path for Incremental LearningabstractNeural networks notoriously suffer from the problem of catastrophic forgetting, the phenomenon of forgetting the past knowledge when acquiring new knowledge. Overcoming catastrophic forgetting is of significant importance to emulate the process of "incremental learning", where the model is capable of learning from sequential experience in an efficient and robust way. State-of-the-art techniques for incremental learning make use of knowledge distillation towards preventing catastrophic forgetting. Therein, one updates the network while ensuring that the network’s responses to previously seen concepts remain stable throughout updates. This in practice is done by minimizing the dissimilarity between current and previous responses of the network one way or another. Our work contributes a novel method to the arsenal of distillation techniques. In contrast to the previous state of the art, we propose to firstly construct low-dimensional manifolds for previous and current responses and minimize the dissimilarity between the responses along the geodesic connecting the manifolds. This induces a more formidable knowledge distillation with smooth properties which preserves the past knowledge more efficiently as observed by our comprehensive empirical study.1 Christian Simon, Piotr Koniusz, Mehrtash Harandi |
CVPR | 3 |
| 2021 | Synthesized Feature based Few-Shot Class-Incremental Learning on a Mixture of SubspacesabstractFew-shot class incremental learning (FSCIL) aims to incrementally add sets of novel classes to a well-trained base model in multiple training sessions with the restriction that only a few novel instances are available per class. While learning novel classes, FSCIL methods gradually forget base (old) class training and overfit to a few novel class samples. Existing approaches have addressed this problem by computing the class prototypes from the visual or semantic word vector domain. In this paper, we propose addressing this problem using a mixture of subspaces. Subspaces define the cluster structure of the visual domain and help to describe the visual and semantic domain considering the overall distribution of the data. Additionally, we propose to employ a variational autoencoder (VAE) to generate synthesized visual samples for augmenting pseudo-feature while learning novel classes incrementally. The combined effect of the mixture of subspaces and synthesized features reduces the forgetting and overfitting problem of FSCIL. Extensive experiments on three image classification datasets show that our proposed method achieves competitive results compared to state-of-the-art methods. Ali Cheraghian, Shafin Rahman, Sameera Ramasinghe, Pengfei Fang, Christian Simon, Lars Petersson, Mehrtash Harandi |
ICCV | 7 |
| 2021 | Kernel Methods in Hyperbolic SpacesabstractEmbedding data in hyperbolic spaces has proven beneficial for many advanced machine learning applications such as image classification and word embeddings. However, working in hyperbolic spaces is not without difficulties as a result of its curved geometry (e.g., computing the Frechet mean of a set of points requires an iterative algorithm). Furthermore, in Euclidean spaces, one can resort to kernel machines that not only enjoy rich theoretical properties but that can also lead to superior representational power (e.g., infinite-width neural networks). In this paper, we introduce positive definite kernel functions for hyperbolic spaces. This brings in two major advantages, 1. kernelization will pave the way to seamlessly benefit from kernel machines in conjunction with hyperbolic embeddings, and 2. the rich structure of the Hilbert spaces associated with kernel machines enables us to simplify various operations involving hyperbolic data. That said, identifying valid kernel functions on curved spaces is not straightforward and is indeed considered an open problem in the learning community. Our work addresses this gap and develops several valid positive definite kernels in hyperbolic spaces, including the universal ones (e.g., RBF). We comprehensively study the proposed kernels on a variety of challenging tasks including few-shot learning, zero-shot learning, person reidentification and knowledge distillation, showing the superiority of the kernelization for hyperbolic representations. Pengfei Fang, Mehrtash Harandi, Lars Petersson |
ICCV | 2 |
| 2021 | Curvature Generation in Curved Spaces for Few-Shot LearningabstractFew-shot learning describes the challenging problem of recognizing samples from unseen classes given very few labeled examples. In many cases, few-shot learning is cast as learning an embedding space that assigns test samples to their corresponding class prototypes. Previous methods assume that data of all few-shot learning tasks comply with a fixed geometrical structure, mostly a Euclidean structure. Questioning this assumption that is clearly difficult to hold in real-world scenarios and incurs distortions to data, we propose to learn a task-aware curved embedding space by making use of the hyperbolic geometry. As a result, task-specific embedding spaces where suitable curvatures are generated to match the characteristics of data are constructed, leading to more generic embedding spaces. We then leverage on intra-class and inter-class context information in the embedding space to generate class prototypes for discriminative classification. We conduct a comprehensive set of experiments on inductive and transductive few-shot learning, demonstrating the benefits of our proposed method over existing embedding methods. Zhi Gao 0002, Yuwei Wu 0001, Yunde Jia, Mehrtash Harandi |
ICCV | 4 |
| 2021 | Learning Online for Unified Segmentation and Tracking ModelsabstractTracking requires building a discriminative model for the target in the inference stage. An effective way to achieve this is online learning, which can comfortably outperform models that are only trained offline. Recent research shows that visual tracking benefits significantly from the unification of visual tracking and segmentation due to its pixel-level discrimination. However, it imposes a great challenge to perform online learning for such a unified model. A segmentation model cannot easily learn from prior information given in the visual tracking scenario. In this paper, we propose TrackMLP: a novel meta-learning method optimized to learn from only partial information to resolve the imposed challenge. Our model is capable of extensively exploiting limited prior information hence possesses much stronger target-background discriminability than other online learning methods. Empirically, we show that our model achieves state-of-the-art performance and tangible improvement over competing models. Our model achieves improved average overlaps of 66.0%,67.1%, and 68.5% in VOT2019, VOT2018, and VOT2016 datasets, which are 6.4%, 7.3%, and 6.4% higher than our baseline. Code will be made publicly available. Tianyu Zhu 0001, Mehrtash Harandi, Rongkai Ma, Tom Drummond |
IJCNN | 2 |
| 2021 | Duo-SegNet: Adversarial Dual-Views for Semi-supervised Medical Image Segmentation
Himashi Peiris, Zhaolin Chen, Gary F. Egan, Mehrtash Harandi |
MICCAI (2) | 4 |
| 2021 | Set Augmented Triplet Loss for Video Person Re-IdentificationabstractModern video person re-identification (re-ID) machines are often trained using a metric learning approach, supervised by a triplet loss. The triplet loss used in video re-ID is usually based on so-called clip features, each aggregated from a few frame features. In this paper, we propose to model the video clip as a set and instead study the distance between sets in the corresponding triplet loss. In contrast to the distance between clip representations, the distance between clip sets considers the pair-wise similarity of each element (i.e., frame representation) between two sets. This allows the network to directly optimize the feature representation at a frame level. Apart from the commonly-used set distance metrics (e.g., ordinary distance and Hausdorff distance), we further propose a hybrid distance metric, tailored for the set-aware triplet loss. Also, we propose a hard positive set construction strategy using the learned class prototypes in a batch. Our proposed method achieves state-of-the-art results across several standard benchmarks, demonstrating the advantages of the proposed method. Pengfei Fang, Pan Ji, Lars Petersson, Mehrtash Harandi |
WACV | 4 |
| 2021 | Discrepant collaborative training by Sinkhorn divergences
Soumava Kumar Roy, Lars Petersson, Mehrtash Harandi |
Image Vis. Comput. | 4 |
| 2021 | Learning Saliency From Single Noisy Labelling: A Robust Model Fitting PerspectiveabstractThe advances made in predicting visual saliency using deep neural networks come at the expense of collecting large-scale annotated data. However, pixel-wise annotation is labor-intensive and overwhelming. In this paper, we propose to learn saliency prediction from a single noisy labelling, which is easy to obtain (e.g., from imperfect human annotation or from unsupervised saliency prediction methods). With this goal, we address a natural question: Can we learn saliency prediction while identifying clean labels in a unified framework? To answer this question, we call on the theory of robust model fitting and formulate deep saliency prediction from a single noisy labelling as robust network learning and exploit model consistency across iterations to identify inliers and outliers (i.e., noisy labels). Extensive experiments on different benchmark datasets demonstrate the superiority of our proposed framework, which can learn comparable saliency prediction with state-of-the-art fully supervised saliency methods. Furthermore, we show that simply by treating ground truth annotations as noisy labelling, our framework achieves tangible improvements over state-of-the-art methods. Jing Zhang 0052, Yuchao Dai, Tong Zhang 0023, Mehrtash Harandi, Nick Barnes, Richard I. Hartley |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Semi-Supervised Domain Adaptation via Asymmetric Joint Distribution MatchingabstractAn intrinsic problem in domain adaptation is the joint distribution mismatch between the source and target domains. Therefore, it is crucial to match the two joint distributions such that the source domain knowledge can be properly transferred to the target domain. Unfortunately, in semi-supervised domain adaptation (SSDA) this problem still remains unsolved. In this article, we therefore present an asymmetric joint distribution matching (AJDM) approach, which seeks a couple of asymmetric matrices to linearly match the source and target joint distributions under the relative chi-square divergence. Specifically, we introduce a least square method to estimate the divergence, which is free from estimating the two joint distributions. Furthermore, we show that our AJDM approach can be generalized to a kernel version, enabling it to handle nonlinearity in the data. From the perspective of Riemannian geometry, learning the linear and nonlinear mappings are both formulated as optimization problems defined on the product of Riemannian manifolds. Numerical experiments on synthetic and real-world data sets demonstrate the effectiveness of the proposed approach and testify its superiority over existing SSDA techniques. Sentao Chen, Mehrtash Harandi, Xiaona Jin, Xiaowei Yang 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Unsupervised Metric Learning with Synthetic ExamplesabstractDistance Metric Learning (DML) involves learning an embedding that brings similar examples closer while moving away dissimilar ones. Existing DML approaches make use of class labels to generate constraints for metric learning. In this paper, we address the less-studied problem of learning a metric in an unsupervised manner. We do not make use of class labels, but use unlabeled data to generate adversarial, synthetic constraints for learning a metric inducing embedding. Being a measure of uncertainty, we minimize the entropy of a conditional probability to learn the metric. Our stochastic formulation scales well to large datasets, and performs competitive to existing metric learning methods. Ujjal Kr Dutta, Mehrtash Harandi, Chellu Chandra Sekhar |
AAAI | 2 |
| 2020 | Revisiting Bilinear Pooling: A Coding PerspectiveabstractBilinear pooling has achieved state-of-the-art performance on fusing features in various machine learning tasks, owning to its ability to capture complex associations between features. Despite the success, bilinear pooling suffers from redundancy and burstiness issues, mainly due to the rank-one property of the resulting representation. In this paper, we prove that bilinear pooling is indeed a similarity-based coding-pooling formulation. This establishment then enables us to devise a new feature fusion algorithm, the factorized bilinear coding (FBC) method, to overcome the drawbacks of the bilinear pooling. We show that FBC can generate compact and discriminative representations with substantially fewer parameters. Experiments on two challenging tasks, namely image classification and visual question answering, demonstrate that our method surpasses the bilinear pooling technique by a large margin. Zhi Gao 0002, Yuwei Wu 0001, Xiaoxun Zhang, Jindou Dai, Yunde Jia, Mehrtash Harandi |
AAAI | 6 |
| 2020 | Channel Recurrent Attention Networks for Video Pedestrian Retrieval
Pengfei Fang, Pan Ji, Jieming Zhou, Lars Petersson, Mehrtash Harandi |
ACCV (6) | 5 |
| 2020 | Learning to Optimize on SPD ManifoldsabstractMany tasks in computer vision and machine learning are modeled as optimization problems with constraints in the form of Symmetric Positive Definite (SPD) matrices. Solving such optimization problems is challenging due to the non-linearity of the SPD manifold, making optimization with SPD constraints heavily relying on expert knowledge and human involvement. In this paper, we propose a meta-learning method to automatically learn an iterative optimizer on SPD manifolds. Specifically, we introduce a novel recurrent model that takes into account the structure of input gradients and identifies the updating scheme of optimization. We parameterize the optimizer by the recurrent model and utilize Riemannian operations to ensure that our method is faithful to the geometry of SPD manifolds. Compared with existing SPD optimizers, our optimizer effectively exploits the underlying data distribution and learns a better optimization trajectory in a data-driven manner. Extensive experiments on various computer vision tasks including metric nearness, clustering, and similarity learning demonstrate that our optimizer outperforms existing state-of-the-art methods consistently. Zhi Gao 0002, Yuwei Wu 0001, Yunde Jia, Mehrtash Harandi |
CVPR | 4 |
| 2020 | Adaptive Subspaces for Few-Shot LearningabstractObject recognition requires a generalization capability to avoid overfitting, especially when the samples are extremely few. Generalization from limited samples, usually studied under the umbrella of meta-learning, equips learning techniques with the ability to adapt quickly in dynamical environments and proves to be an essential aspect of life long learning. In this paper, we provide a framework for few-shot learning by introducing dynamic classifiers that are constructed from few samples. A subspace method is exploited as the central block of a dynamic classifier. We will empirically show that such modelling leads to robustness against perturbations (e.g., outliers) and yields competitive results on the task of supervised and semi-supervised few-shot classification. We also develop a discriminative form which can boost the accuracy even further. Our code is available at https://github.com/chrysts/dsn_fewshot Christian Simon, Piotr Koniusz, Richard Nock, Mehrtash Harandi |
CVPR | 4 |
| 2020 | On Modulating the Gradient for Meta-learning
Christian Simon, Piotr Koniusz, Richard Nock, Mehrtash Harandi |
ECCV (8) | 4 |
| 2020 | An Input Residual Connection for Simplifying Gated Recurrent Neural NetworksabstractGated Recurrent Neural Networks (GRNNs) are important models that continue to push the state-of-the-art solutions across different machine learning problems. However, they are composed of intricate components that are generally not well understood. We increase GRNN interpretability by linking the canonical Gated Recurrent Unit (GRU) design to the well-studied Hopfield network. This connection allowed us to identify network redundancies, which we simplified with an Input Residual Connection (IRC). We tested GRNNs against their IRC counterparts on language modelling. In addition, we proposed an Input Highway Connection (IHC) as an advance application of the IRC and then evaluated the most widely applied GRNN of the Long Short-Term Memory (LSTM) and IHC-LSTM on tasks of i) image generation and ii) learning to learn to update another learner-network. Despite parameter reductions, all IRC-GRNNs showed either comparative or superior generalisation than their baseline models. Furthermore, compared to LSTM, the IHC-LSTM removed 85.4% parameters on image generation. In conclusion, the IRC is applicable, but not limited, to the GRNN designs of GRUs and LSTMs but also to FastGRNNs, Simple Recurrent Units (SRUs), and Strongly-Typed Recurrent Neural Networks (T-RNNs). Nicholas I-Hsien Kuo, Mehrtash Harandi, Nicolas Fourrier, Christian Walder, Gabriela Ferraro, Hanna Suominen |
IJCNN | 2 |
| 2020 | Hierarchical Neural Architecture Search for Deep Stereo MatchingabstractTo reduce the human efforts in neural network design, Neural Architecture Search (NAS) has been applied with remarkable success to various high-level vision tasks such as classification and semantic segmentation. The underlying idea for the NAS algorithm is straightforward, namely, to allow the network the ability to choose among a set of operations (\eg convolution with different filter sizes), one is able to find an optimal architecture that is better adapted to the problem at hand. However, so far the success of NAS has not been enjoyed by low-level geometric vision tasks such as stereo matching. This is partly due to the fact that state-of-the-art deep stereo matching networks, designed by humans, are already sheer in size. Directly applying the NAS to such massive structures is computationally prohibitive based on the currently available mainstream computing resources. In this paper, we propose the first \emph{end-to-end} hierarchical NAS framework for deep stereo matching by incorporating task-specific human knowledge into the neural architecture search framework. Specifically, following the gold standard pipeline for deep stereo matching (\ie, feature extraction -- feature volume construction and dense matching), we optimize the architectures of the entire pipeline jointly. Extensive experiments show that our searched network outperforms all state-of-the-art deep stereo matching architectures and is ranked at the top 1 accuracy on KITTI stereo 2012, 2015, and Middlebury benchmarks, as well as the top 1 on SceneFlow dataset with a substantial improvement on the size of the network and the speed of inference. Code available at https://github.com/XuelianCheng/LEAStereo. Xuelian Cheng, Yiran Zhong, Mehrtash Harandi, Yuchao Dai, Xiaojun Chang, Hongdong Li, Tom Drummond, ZongYuan Ge |
NeurIPS | 3 |
| 2020 | Learning from Noisy Labels via Discrepant Collaborative TrainingabstractNoise is ubiquitous in the world around us. Difficulty in estimating the noise within a dataset makes learning from such a dataset a difficult and challenging task. In this paper, we propose a novel and effective learning framework in order to alleviate the adverse effects of noise within a dataset. Towards this aim, we modify a collaborative training framework to utilize discrepancy constraints between respective feature extractors enabling the learning of distinct, yet discriminative features, pacifying the adverse effects of noise. Empirical results of our proposed algorithm, Discrepant Collaborative Training (DCT), achieve competitive results against several current state-of-the-art algorithms across MNIST, CIFAR10 and CIFAR100, as well as large fine-grained image classification datasets such as CUBS-200-2011 and CARS196 for different levels of noise. Soumava Kumar Roy, Lars Petersson, Mehrtash Harandi |
WACV | 4 |
| 2020 | Devon: Deformable Volume Network for Learning Optical Flow
Jack Valmadre, Juho Kannala, Mehrtash Harandi, Philip Torr 0001 |
WACV | 5 |
| 2020 | Cross-Correlated Attention Networks for Person Re-Identification
Jieming Zhou, Soumava Kumar Roy, Pengfei Fang, Mehrtash Harandi, Lars Petersson |
Image Vis. Comput. | 4 |
| 2020 | Domain Adaptation by Joint Distribution Invariant ProjectionsabstractDomain adaptation addresses the learning problem where the training data are sampled from a source joint distribution (source domain), while the test data are sampled from a different target joint distribution (target domain). Because of this joint distribution mismatch, a discriminative classifier naively trained on the source domain often generalizes poorly to the target domain. In this paper, we therefore present a Joint Distribution Invariant Projections (JDIP) approach to solve this problem. The proposed approach exploits linear projections to directly match the source and target joint distributions under the L2-distance. Since the traditional kernel density estimators for distribution estimation tend to be less reliable as the dimensionality increases, we propose a least square method to estimate the L2-distance without the need to estimate the two joint distributions, leading to a quadratic problem with analytic solution. Furthermore, we introduce a kernel version of JDIP to account for inherent nonlinearity in the data. We show that the proposed learning problems can be naturally cast as optimization problems defined on the product of Riemannian manifolds. To be comprehensive, we also establish an error bound, theoretically explaining how our method works and contributes to reducing the target domain generalization error. Extensive empirical evidence demonstrates the benefits of our approach over state-of-the-art domain adaptation methods on several visual data sets. Sentao Chen, Mehrtash Harandi, Xiaona Jin, Xiaowei Yang 0003 |
IEEE Trans. Image Process. | 2 |
| 2020 | A Robust Distance Measure for Similarity-Based Classification on the SPD ManifoldabstractThe symmetric positive definite (SPD) matrices, forming a Riemannian manifold, are commonly used as visual representations. The non-Euclidean geometry of the manifold often makes developing learning algorithms (e.g., classifiers) difficult and complicated. The concept of similarity-based learning has been shown to be effective to address various problems on SPD manifolds. This is mainly because the similarity-based algorithms are agnostic to the geometry and purely work based on the notion of similarities/distances. However, existing similarity-based models on SPD manifolds opt for holistic representations, ignoring characteristics of information captured by SPD matrices. To circumvent this limitation, we propose a novel SPD distance measure for the similarity-based algorithm. Specifically, we introduce the concept of point-to-set transformation, which enables us to learn multiple lower dimensional and discriminative SPD manifolds from a higher dimensional one. For lower dimensional SPD manifolds obtained by the point-to-set transformation, we propose a tailored set-to-set distance measure by making use of the family of alpha-beta divergences. We further propose to learn the point-to-set transformation and the set-to-set distance measure jointly, yielding a powerful similarity-based algorithm on SPD manifolds. Our thorough evaluations on several visual recognition tasks (e.g., action classification and face recognition) suggest that our algorithm comfortably outperforms various state-of-the-art algorithms. Zhi Gao 0002, Yuwei Wu 0001, Mehrtash Harandi, Yunde Jia |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2019 | Min-Max Statistical Alignment for Transfer LearningabstractA profound idea in learning invariant features for transfer learning is to align statistical properties of the domains. In practice, this is achieved by minimizing the disparity between the domains, usually measured in terms of their statistical properties. We question the capability of this school of thought and propose to minimize the maximum disparity between domains. Furthermore, we develop an end-to-end learning scheme that enables us to benefit from the proposed min-max strategy in training deep models. We show that the min-max solution can outperform the existing statistical alignment solutions, and can compete with state-of-the-art solutions on two challenging learning tasks, namely, Unsupervised Domain Adaptation (UDA) and Zero-Shot Learning (ZSL). Samitha Herath, Mehrtash Harandi, Basura Fernando, Richard Nock |
CVPR | 2 |
| 2019 | Bilinear Attention Networks for Person RetrievalabstractThis paper investigates a novel Bilinear attention (Bi-attention) block, which discovers and uses second order statistical information in an input feature map, for the purpose of person retrieval. The Bi-attention block uses bilinear pooling to model the local pairwise feature interactions along each channel, while preserving the spatial structural information. We propose an Attention in Attention (AiA) mechanism to build inter-dependency among the second order local and global features with the intent to make better use of, or pay more attention to, such higher order statistical relationships. The proposed network, equipped with the proposed Bi-attention is referred to as Bilinear ATtention network (BAT-net). Our approach outperforms current state-of-the-art by a considerable margin across the standard benchmark datasets (e.g., CUHK03, Market-1501, DukeMTMC-reID and MSMT17). Pengfei Fang, Jieming Zhou, Soumava Kumar Roy, Lars Petersson, Mehrtash Harandi |
ICCV | 5 |
| 2019 | Siamese Networks: The Tale of Two ManifoldsabstractSiamese networks are non-linear deep models that have found their ways into a broad set of problems in learning theory, thanks to their embedding capabilities. In this paper, we study Siamese networks from a new perspective and question the validity of their training procedure. We show that in the majority of cases, the objective of a Siamese network is endowed with an invariance property. Neglecting the invariance property leads to a hindrance in training the Siamese networks. To alleviate this issue, we propose two Riemannian structures and generalize a well-established accelerated stochastic gradient descent method to take into account the proposed Riemannian structures. Our empirical evaluations suggest that by making use of the Riemannian geometry, we achieve state-of-the-art results against several algorithms for the challenging problem of fine-grained image classification. Soumava Kumar Roy, Mehrtash Harandi, Richard Nock, Richard I. Hartley |
ICCV | 2 |
| 2019 | Neural Collaborative Subspace ClusteringabstractWe introduce the Neural Collaborative Subspace Clustering, a neural model that discovers clusters of data points drawn from a union of low-dimensional subspaces. In contrast to previous attempts, our model runs without the aid of spectral clustering. This makes our algorithm one of the kinds that can gracefully scale to large datasets. At its heart, our neural model benefits from a classifier which determines whether a pair of points lies on the same subspace or not. Essential to our model is the construction of two affinity matrices, one from the classifier and the other from a notion of subspace self-expressiveness, to supervise training in a collaborative scheme. We thoroughly assess and contrast the performance of our model against various state-of-the-art clustering algorithms including deep subspace-based ones. Tong Zhang 0023, Pan Ji, Mehrtash Harandi, Wenbing Huang 0001, Hongdong Li |
ICML | 3 |
| 2019 | Using temporal information for recognizing actions from still images
Samitha Herath, Basura Fernando, Mehrtash Harandi |
Pattern Recognit. | 3 |
| 2019 | Toward Efficient Action Recognition: Principal Backpropagation for Training Two-Stream NetworksabstractIn this paper, we propose the novel principal backpropagation networks (PBNets) to revisit the backpropagation algorithms commonly used in training two-stream networks for video action recognition. We content that existing approaches always take all the frames/snippets for the backpropagation not optimal for video recognition since the desired actions only occur in a short period within a video. To remedy these drawbacks, we design a watch-and-choose mechanism. In particular, the watching stage exploits a dense snippet-wise temporal pooling strategy to discover the global characteristic for each input video, while the choosing phase only backpropagates a small number of representative snippets that are selected with two novel strategies, i.e., Max-rule and KL-rule. We prove that with the proposed selection strategies, performing the backpropagation on the selected subset is capable of decreasing the loss of the whole snippets as well. The proposed PBNets are evaluated on two standard video action recognition benchmarks UCF101 and HMDB51, where it surpasses the state of the arts consistently, but requiring less memory and computation to achieve high performance. Wenbing Huang 0001, Lijie Fan, Mehrtash Harandi, Lin Ma 0002, Huaping Liu 0001, Wei Liu 0005, Chuang Gan 0001 |
IEEE Trans. Image Process. | 3 |
| 2018 | Scalable Deep k-Subspace Clustering
Tong Zhang 0023, Pan Ji, Mehrtash Harandi, Richard I. Hartley, Ian D. Reid 0001 |
ACCV (5) | 3 |
| 2018 | Geometry Aware Constrained Optimization Techniques for Deep LearningabstractIn this paper, we generalize the Stochastic Gradient Descent (SGD) and RMSProp algorithms to the setting of Riemannian optimization. SGD is a popular method for large scale optimization. In particular, it is widely used to train the weights of Deep Neural Networks. However, gradients computed using standard SGD can have large variance, which is detrimental for the convergence rate of the algorithm. Other methods such as RMSProp and ADAM address this issue. Nevertheless, these methods cannot be directly applied to constrained optimization problems. In this paper, we extend some popular optimization algorithm to the Riemannian (constrained) setting. We substantiate our proposed extensions with a range of relevant problems in machine learning such as incremental Principal Component Analysis, computating the Riemannian centroids of SPD matrices, and Deep Metric Learning. We achieve competitive results against the state of the art for fine-grained object recognition datasets. Soumava Kumar Roy, Zakaria Mhammedi, Mehrtash Harandi |
CVPR | 3 |
| 2018 | Deep Unsupervised Saliency Detection: A Multiple Noisy Labeling PerspectiveabstractThe success of current deep saliency detection methods heavily depends on the availability of large-scale supervision in the form of per-pixel labeling. Such supervision, while labor-intensive and not always possible, tends to hinder the generalization ability of the learned models. By contrast, traditional handcrafted features based unsupervised saliency detection methods, even though have been surpassed by the deep supervised methods, are generally dataset-independent and could be applied in the wild. This raises a natural question that "Is it possible to learn saliency maps without using labeled data while improving the generalization ability?". To this end, we present a novel perspective to unsupervised saliency detection through learning from multiple noisy labeling generated by "weak" and "noisy" unsupervised handcrafted saliency methods. Our end-to-end deep learning framework for unsupervised saliency detection consists of a latent saliency prediction module and a noise modeling module that work collaboratively and are optimized jointly. Explicit noise modeling enables us to deal with noisy saliency maps in a probabilistic way. Extensive experimental results on various benchmarking datasets show that our model not only outperforms all the unsupervised saliency methods with a large margin but also achieves comparable performance with the recent state-of-the-art supervised deep saliency methods. Jing Zhang 0052, Tong Zhang 0023, Yuchao Dai, Mehrtash Harandi, Richard I. Hartley |
CVPR | 4 |
| 2018 | Museum Exhibit Identification Challenge for the Supervised Domain Adaptation and Beyond
Piotr Koniusz, Yusuf Tas, Mehrtash Harandi, Fatih Porikli |
ECCV (16) | 4 |
| 2018 | Dimensionality Reduction on SPD Manifolds: The Emergence of Geometry-Aware MethodsabstractRepresenting images and videos with Symmetric Positive Definite (SPD) matrices, and considering the Riemannian geometry of the resulting space, has been shown to yield high discriminative power in many visual recognition tasks. Unfortunately, computation on the Riemannian manifold of SPD matrices -especially of high-dimensional ones- comes at a high cost that limits the applicability of existing techniques. In this paper, we introduce algorithms able to handle high-dimensional SPD matrices by constructing a lower-dimensional SPD manifold. To this end, we propose to model the mapping from the high-dimensional SPD manifold to the low-dimensional one with an orthonormal projection. This lets us formulate dimensionality reduction as the problem of finding a projection that yields a low-dimensional manifold either with maximum discriminative power in the supervised scenario, or with maximum variance of the data in the unsupervised one. We show that learning can be expressed as an optimization problem on a Grassmann manifold and discuss fast solutions for special cases. Our evaluation on several classification tasks evidences that our approach leads to a significant accuracy gain over state-of-the-art methods. Mehrtash Harandi, Mathieu Salzmann, Richard I. Hartley |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2018 | Large-Scale Metric Learning: A Voyage From Shallow to DeepabstractDespite its attractive properties, the performance of the recently introduced Keep It Simple and Straightforward MEtric learning (KISSME) method is greatly dependent on principal component analysis as a preprocessing step. This dependence can lead to difficulties, e.g., when the dimensionality is not meticulously set. To address this issue, we devise a unified formulation for joint dimensionality reduction and metric learning based on the KISSME algorithm. Our joint formulation is expressed as an optimization problem on the Grassmann manifold, and hence enjoys the properties of Riemannian optimization techniques. Following the success of deep learning in recent years, we also devise end-to-end learning of a generic deep network for metric learning using our derivation. Masoud Faraki, Mehrtash Harandi, Fatih Porikli |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2018 | A Comprehensive Look at Coding Techniques on Riemannian ManifoldsabstractCore to many learning pipelines is visual recognition such as image and video classification. In such applications, having a compact yet rich and informative representation plays a pivotal role. An underlying assumption in traditional coding schemes [e.g., sparse coding (SC)] is that the data geometrically comply with the Euclidean space. In other words, the data are presented to the algorithm in vector form and Euclidean axioms are fulfilled. This is of course restrictive in machine learning, computer vision, and signal processing, as shown by a large number of recent studies. This paper takes a further step and provides a comprehensive mathematical framework to perform coding in curved and non-Euclidean spaces, i.e., Riemannian manifolds. To this end, we start by the simplest form of coding, namely, bag of words. Then, inspired by the success of vector of locally aggregated descriptors in addressing computer vision problems, we will introduce its Riemannian extensions. Finally, we study Riemannian form of SC, locality-constrained linear coding, and collaborative coding. Through rigorous tests, we demonstrate the superior performance of our Riemannian coding schemes against the state-of-the-art methods on several visual classification tasks, including head pose classification, video-based face recognition, and dynamic scene recognition. Masoud Faraki, Mehrtash Harandi, Fatih Porikli |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2017 | Generalized Rank Pooling for Activity RecognitionabstractMost popular deep models for action recognition split video sequences into short sub-sequences consisting of a few frames, frame-based features are then pooled for recognizing the activity. Usually, this pooling step discards the temporal order of the frames, which could otherwise be used for better recognition. Towards this end, we propose a novel pooling method, generalized rank pooling (GRP), that takes as input, features from the intermediate layers of a CNN that is trained on tiny sub-sequences, and produces as output the parameters of a subspace which (i) provides a low-rank approximation to the features and (ii) preserves their temporal order. We propose to use these parameters as a compact representation for the video sequence, which is then used in a classification setup. We formulate an objective for computing this subspace as a Riemannian optimization problem on the Grassmann manifold, and propose an efficient conjugate gradient scheme for solving it. Experiments on several activity recognition datasets show that our scheme leads to state-of-the-art performance. Anoop Cherian, Basura Fernando, Mehrtash Harandi, Stephen Gould |
CVPR | 3 |
| 2017 | Learning an Invariant Hilbert Space for Domain AdaptationabstractThis paper introduces a learning scheme to construct a Hilbert space (i.e., a vector space along its inner product) to address both unsupervised and semi-supervised domain adaptation problems. This is achieved by learning projections from each domain to a latent space along the Mahalanobis metric of the latent space to simultaneously minimizing a notion of domain variance while maximizing a measure of discriminatory power. In particular, we make use of the Riemannian optimization techniques to match statistical properties (e.g., first and second order statistics) between samples projected into the latent space from different domains. Upon availability of class labels, we further deem samples sharing the same label to form more compact clusters while pulling away samples coming from different classes. We extensively evaluate and contrast our proposal against state-of-the-art methods for the task of visual domain adaptation using both handcrafted and deep-net features. Our experiments show that even with a simple nearest neighbor classifier, the proposed method can outperform several state-of-the-art methods benefitting from more involved classification schemes. Samitha Herath, Mehrtash Harandi, Fatih Porikli |
CVPR | 2 |
| 2017 | Learning Discriminative αβ-Divergences for Positive Definite Matrices
Anoop Cherian, Panagiotis Stanitsas, Mehrtash Harandi, Vassilios Morellas, Nikolaos Papanikolopoulos |
ICCV | 3 |
| 2017 | Joint Dimensionality Reduction and Metric Learning: A Geometric TakeabstractTo be tractable and robust to data noise, existing metric learning algorithms commonly rely on PCA as a pre-processing step. How can we know, however, that PCA, or any other specific dimensionality reduction technique, is the method of choice for the problem at hand? The answer is simple: We cannot! To address this issue, in this paper, we develop a Riemannian framework to jointly learn a mapping performing dimensionality reduction and a metric in the induced space. Our experiments evidence that, while we directly work on high-dimensional features, our approach yields competitive runtimes with and higher accuracy than state-of-the-art metric learning algorithms. Mehrtash Harandi, Mathieu Salzmann, Richard I. Hartley |
ICML | 1 |
| 2017 | Efficient Optimization for Linear Dynamical Systems with Applications to Clustering and Sparse CodingabstractLinear Dynamical Systems (LDSs) are fundamental tools for modeling spatio-temporal data in various disciplines. Though rich in modeling, analyzing LDSs is not free of difficulty, mainly because LDSs do not comply with Euclidean geometry and hence conventional learning techniques can not be applied directly. In this paper, we propose an efficient projected gradient descent method to minimize a general form of a loss function and demonstrate how clustering and sparse coding with LDSs can be solved by the proposed method efficiently. To this end, we first derive a novel canonical form for representing the parameters of an LDS, and then show how gradient-descent updates through the projection on the space of LDSs can be achieved dexterously. In contrast to previous studies, our solution avoids any approximation in LDS modeling or during the optimization process. Extensive experiments reveal the superior performance of the proposed method in terms of the convergence and classification accuracy over state-of-the-art techniques. Wenbing Huang 0001, Mehrtash Harandi, Tong Zhang 0023, Lijie Fan, Fuchun Sun 0001, Junzhou Huang |
NIPS | 2 |
| 2017 | Going deeper into action recognition: A survey
Samitha Herath, Mehrtash Harandi, Fatih Porikli |
Image Vis. Comput. | 2 |
| 2017 | No fuss metric learning, a Hilbert space scenario
Masoud Faraki, Mehrtash Harandi, Fatih Porikli |
Pattern Recognit. Lett. | 2 |
| 2016 | When VLAD Met HilbertabstractIn many challenging visual recognition tasks where training data is limited, Vectors of Locally Aggregated Descriptors (VLAD) have emerged as powerful image/video representations that compete with or outperform state-of the-art approaches. In this paper, we address two fundamental limitations of VLAD: its requirement for the local descriptors to have vector form and its restriction to linear classifiers due to its high-dimensionality. To this end, we introduce a kernelized version of VLAD. This not only lets us inherently exploit more sophisticated classification schemes, but also enables us to efficiently aggregate nonvector descriptors (e.g., manifold-valued data) in the VLAD framework. Furthermore, we propose an approximate formulation that allows us to accelerate the coding process while still benefiting from the properties of kernel VLAD. Our experiments demonstrate the effectiveness of our approach at handling manifold-valued data, such as covariance descriptors, on several classification tasks. Our results also evidence the benefits of our nonlinear VLAD descriptors against the linear ones in Euclidean space using several standard benchmark datasets. Mehrtash Harandi, Mathieu Salzmann, Fatih Porikli |
CVPR | 1 |
| 2016 | Sparse Coding and Dictionary Learning with Linear Dynamical SystemsabstractLinear Dynamical Systems (LDSs) are the fundamental tools for encoding spatio-temporal data in various disciplines. To enhance the performance of LDSs, in this paper, we address the challenging issue of performing sparse coding on the space of LDSs, where both data and dictionary atoms are LDSs. Rather than approximate the extended observability with a finite-order matrix, we represent the space of LDSs by an infinite Grassmannian consisting of the orthonormalized extended observability subspaces. Via a homeomorphic mapping, such Grassmannian is embedded into the space of symmetric matrices, where a tractable objective function can be derived for sparse coding. Then, we propose an efficient method to learn the system parameters of the dictionary atoms explicitly, by imposing the symmetric constraint to the transition matrices of the data and dictionary systems. Moreover, we combine the state covariance into the algorithm formulation, thus further promoting the performance of the models with symmetric transition matrices. Comparative experimental evaluations reveal the superior performance of proposed methods on various tasks including video classification and tactile recognition. Wenbing Huang 0001, Fuchun Sun 0001, Le-le Cao, Deli Zhao, Huaping Liu 0001, Mehrtash Harandi |
CVPR | 6 |
| 2016 | Image set classification by symmetric positive semi-definite matricesabstractRepresenting images and videos by covariance descriptors and leveraging the inherent manifold structure of Symmetric Positive Definite (SPD) matrices leads to enhanced performances in various visual recognition tasks. However, when covariance descriptors are used to represent image sets, the result is often rank-deficient. Thus, most existing approaches adhere to blind perturbation with predefined regularizers just to be able to employ inference tools. To overcome this problem, we introduce novel similarity measures specifically designed for rank-deficient covariance descriptors, i.e., symmetric positive semi-definite matrices. In particular, we derive positive definite kernels that can be decomposed into the kernels on the cone of SPD matrices and kernels on the Grassmann manifolds. Our experiments evidence that, our method achieves superior results for image set classification on various recognition tasks including hand gesture classification, face recognition from video sequences, and dynamic scene categorization. Masoud Faraki, Mehrtash Harandi, Fatih Porikli |
WACV | 2 |
| 2016 | Distribution-Matching Embedding for Visual Domain AdaptationabstractDomain-invariant representations are key to addressing the domain shift problem where the training and test examples follow different distributions. Existing techniques that have attempted to match the distributions of the source and target domains typically compare these distributions in the original feature space. This space, however, may not be directly suitable for such a comparison, since some of the features may have been distorted by the domain shift, or may be domain specific. In this paper, we introduce a Distribution-Matching Embedding approach: An unsupervised domain adaptation method that overcomes this issue by mapping the data to a latent space where the distance between the empirical distributions of the source and target examples is minimized. In other words, we seek to extract the information that is invariant across the source and target data. In particular, we study two different distances to compare the source and target distributions: the Maximum Mean Discrepancy and the Hellinger distance. Furthermore, we show that our approach allows us to learn either a linear embedding, or a nonlinear one. We demonstrate the benefits of our approach on the tasks of visual object recognition, text categorization, and WiFi localization. Mahsa Baktash, Mehrtash Harandi, Mathieu Salzmann |
J. Mach. Learn. Res. | 2 |
| 2016 | Executable thematic special issue on pattern recognition techniques for indirect immunofluorescence images analysis
Mehrtash Harandi, Brian C. Lovell, Gennaro Percannella, Alessia Saggese, Mario Vento, Arnold Wiliem |
Pattern Recognit. Lett. | 1 |
| 2016 | Sparse Coding on Symmetric Positive Definite Manifolds Using Bregman DivergencesabstractThis paper introduces sparse coding and dictionary learning for symmetric positive definite (SPD) matrices, which are often used in machine learning, computer vision, and related areas. Unlike traditional sparse coding schemes that work in vector spaces, in this paper, we discuss how SPD matrices can be described by sparse combination of dictionary atoms, where the atoms are also SPD matrices. We propose to seek sparse coding by embedding the space of SPD matrices into the Hilbert spaces through two types of the Bregman matrix divergences. This not only leads to an efficient way of performing sparse coding but also an online and iterative scheme for dictionary learning. We apply the proposed methods to several computer vision tasks where images are represented by region covariance matrices. Our proposed algorithms outperform state-of-the-art methods on a wide range of classification tasks, including face recognition, action recognition, material classification, and texture categorization. Mehrtash Harandi, Richard I. Hartley, Brian C. Lovell, Conrad Sanderson |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2015 | Online Dictionary Learning on Symmetric Positive Definite Manifolds with Vision ApplicationsabstractSymmetric Positive Definite (SPD) matrices in the form of region covariances are considered rich descriptors for images and videos. Recent studies suggest that exploiting the Riemannian geometry of the SPD manifolds could lead to improved performances for vision applications. For tasks involving processing large-scale and dynamic data in computer vision, the underlying model is required to progressively and efficiently adapt itself to the new and unseen observations. Motivated by these requirements, this paper studies the problem of online dictionary learning on the SPD manifolds. We make use of the Stein divergence to recast the problem of online dictionary learning on the manifolds to a problem in Reproducing Kernel Hilbert Spaces, for which, we develop efficient algorithms by taking into account the geometric structure of the SPD manifolds. To our best knowledge, our work is the first study that provides a solution for online dictionary learning on the SPD manifolds. Empirical results on both large-scale image classification task and dynamic video processing tasks validate the superior performance of our approach as compared to several state-of-the-art algorithms. Shengping Zhang, Shiva Prasad Kasiviswanathan, Pong C. Yuen, Mehrtash Harandi |
AAAI | 4 |
| 2015 | More about VLAD: A leap from Euclidean to Riemannian manifoldsabstractThis paper takes a step forward in image and video coding by extending the well-known Vector of Locally Aggregated Descriptors (VLAD) onto an extensive space of curved Riemannian manifolds. We provide a comprehensive mathematical framework that formulates the aggregation problem of such manifold data into an elegant solution. In particular, we consider structured descriptors from visual data, namely Region Covariance Descriptors and linear subspaces that reside on the manifold of Symmetric Positive Definite matrices and the Grassmannian manifolds, respectively. Through rigorous experimental validation, we demonstrate the superior performance of this novel Riemannian VLAD descriptor on several visual classification tasks including video-based face recognition, dynamic scene recognition, and head pose classification. Masoud Faraki, Mehrtash Harandi, Fatih Porikli |
CVPR | 2 |
| 2015 | Riemannian coding and dictionary learning: Kernels to the rescueabstractWhile sparse coding on non-flat Riemannian manifolds has recently become increasingly popular, existing solutions either are dedicated to specific manifolds, or rely on optimization problems that are difficult to solve, especially when it comes to dictionary learning. In this paper, we propose to make use of kernels to perform coding and dictionary learning on Riemannian manifolds. To this end, we introduce a general Riemannian coding framework with its kernel-based counterpart. This lets us (i) generalize beyond the special case of sparse coding; (ii) introduce efficient solutions to two coding schemes; (iii) learn the kernel parameters; (iv) perform unsupervised and supervised dictionary learning in a much simpler manner than previous Riemannian coding methods. We demonstrate the effectiveness of our approach on three different types of non-flat manifolds, and illustrate its generality by applying it to Euclidean spaces, which also are Riemannian manifolds. Mehrtash Harandi, Mathieu Salzmann |
CVPR | 1 |
| 2015 | Approximate infinite-dimensional Region Covariance Descriptors for image classificationabstractWe introduce methods to estimate infinite-dimensional Region Covariance Descriptors (RCovDs) by exploiting two feature mappings, namely random Fourier features and the Nyström method. In general, infinite-dimensional RCovDs offer better discriminatory power over their low-dimensional counterparts. However, the underlying Riemannian structure, i.e., the manifold of Symmetric Positive Definite (SPD) matrices, is out of reach to great extent for infinite-dimensional RCovDs. To overcome this difficulty, we propose to approximate the infinite-dimensional RCovDs by making use of the aforementioned explicit mappings. We will empirically show that the proposed finite-dimensional approximations of infinite-dimensional RCovDs consistently outperform the low-dimensional RCovDs for image classification task, while enjoying the Riemannian structure of the SPD manifolds. Moreover, our methods achieve the state-of-the-art performance on three different image classification tasks. Masoud Faraki, Mehrtash Harandi, Fatih Porikli |
ICASSP | 2 |
| 2015 | Beyond Gauss: Image-Set Matching on the Riemannian Manifold of PDFsabstractState-of-the-art image-set matching techniques typically implicitly model each image-set with a Gaussian distribution. Here, we propose to go beyond these representations and model image-sets as probability distribution functions (PDFs) using kernel density estimators. To compare and match image-sets, we exploit Csiszar f-divergences, which bear strong connections to the geodesic distance defined on the space of PDFs, i.e., the statistical manifold. Furthermore, we introduce valid positive definite kernels on the statistical manifolds, which let us make use of more powerful classification schemes to match image-sets. Finally, we introduce a supervised dimensionality reduction technique that learns a latent space where f-divergences reflect the class labels of the data. Our experiments on diverse problems, such as video-based face recognition and dynamic texture classification, evidence the benefits of our approach over the state-of-the-art image-set matching methods. Mehrtash Harandi, Mathieu Salzmann, Mahsa Baktash |
ICCV | 1 |
| 2015 | Material Classification on Symmetric Positive Definite ManifoldsabstractThis paper tackles the problem of categorizing materials and textures by exploiting the second order statistics. To this end, we introduce the Extrinsic Vector of Locally Aggregated Descriptors (E-VLAD), a method to combine local and structured descriptors into a unified vector representation where each local descriptor is a Covariance Descriptor (CovD). In doing so, we make use of an accelerated method of obtaining a visual codebook where each atom is itself a CovD. We will then introduce an efficient way of aggregating local CovDs into a vector representation. Our method could be understood as an extrinsic extension of the highly acclaimed method of Vector of Locally Aggregated Descriptors [17] (or VLAD) to CovDs. We will show that the proposed method is extremely powerful in classifying materials/ textures and can outperform complex machineries even with simple classifiers. Masoud Faraki, Mehrtash Harandi, Fatih Porikli |
WACV | 2 |
| 2015 | Extrinsic Methods for Coding and Dictionary Learning on Grassmann Manifolds
Mehrtash Harandi, Richard I. Hartley, Chunhua Shen, Brian C. Lovell, Conrad Sanderson |
Int. J. Comput. Vis. | 1 |
| 2015 | Kernel Methods on Riemannian Manifolds with Gaussian RBF KernelsabstractIn this paper, we develop an approach to exploiting kernel methods with manifold-valued data. In many computer vision problems, the data can be naturally represented as points on a Riemannian manifold. Due to the non-Euclidean geometry of Riemannian manifolds, usual Euclidean computer vision and machine learning algorithms yield inferior results on such data. In this paper, we define Gaussian radial basis function (RBF)-based positive definite kernels on manifolds that permit us to embed a given manifold with a corresponding metric in a high dimensional reproducing kernel Hilbert space. These kernels make it possible to utilize algorithms developed for linear spaces on nonlinear manifold-valued data. Since the Gaussian RBF defined with any given metric is not always positive definite, we present a unified framework for analyzing the positive definiteness of the Gaussian RBF on a generic metric space. We then use the proposed framework to identify positive definite kernels on two specific manifolds commonly encountered in computer vision: the Riemannian manifold of symmetric positive definite matrices and the Grassmann manifold, i.e., the Riemannian manifold of linear subspaces of a Euclidean space. We show that many popular algorithms designed for Euclidean spaces, such as support vector machines, discriminant analysis and principal component analysis can be generalized to Riemannian manifolds with the help of such positive definite Gaussian kernels. Sadeep Jayasumana, Richard I. Hartley, Mathieu Salzmann, Hongdong Li, Mehrtash Harandi |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2015 | Novelty detection in human tracking based on spatiotemporal oriented energies
Ali Emami, Mehrtash Harandi, Farhad Dadgostar, Brian C. Lovell |
Pattern Recognit. | 2 |
| 2014 | Domain Adaptation on the Statistical ManifoldabstractIn this paper, we tackle the problem of unsupervised domain adaptation for classification. In the unsupervised scenario where no labeled samples from the target domain are provided, a popular approach consists in transforming the data such that the source and target distributions become similar. To compare the two distributions, existing approaches make use of the Maximum Mean Discrepancy (MMD). However, this does not exploit the fact that probability distributions lie on a Riemannian manifold. Here, we propose to make better use of the structure of this manifold and rely on the distance on the manifold to compare the source and target distributions. In this framework, we introduce a sample selection method and a subspace-based method for unsupervised domain adaptation, and show that both these manifold-based techniques outperform the corresponding approaches based on the MMD. Furthermore, we show that our subspace-based approach yields state-of-the-art results on a standard object recognition benchmark. Mahsa Baktash, Mehrtash Harandi, Brian C. Lovell, Mathieu Salzmann |
CVPR | 2 |
| 2014 | Bregman Divergences for Infinite Dimensional Covariance MatricesabstractWe introduce an approach to computing and comparing Covariance Descriptors (CovDs) in infinite-dimensional spaces. CovDs have become increasingly popular to address classification problems in computer vision. While CovDs offer some robustness to measurement variations, they also throw away part of the information contained in the original data by only retaining the second-order statistics over the measurements. Here, we propose to overcome this limitation by first mapping the original data to a high-dimensional Hilbert space, and only then compute the CovDs. We show that several Bregman divergences can be computed between the resulting CovDs in Hilbert space via the use of kernels. We then exploit these divergences for classification purpose. Our experiments demonstrate the benefits of our approach on several tasks, such as material and texture recognition, person re-identification, and action recognition from motion capture data. Mehrtash Harandi, Mathieu Salzmann, Fatih Porikli |
CVPR | 1 |
| 2014 | Optimizing over Radial Kernels on Compact ManifoldsabstractWe tackle the problem of optimizing over all possible positive definite radial kernels on Riemannian manifolds for classification. Kernel methods on Riemannian manifolds have recently become increasingly popular in computer vision. However, the number of known positive definite kernels on manifolds remain very limited. Furthermore, most kernels typically depend on at least one parameter that needs to be tuned for the problem at hand. A poor choice of kernel, or of parameter value, may yield significant performance drop-off. Here, we show that positive definite radial kernels on the unit n-sphere, the Grassmann manifold and Kendall's shape manifold can be expressed in a simple form whose parameters can be automatically optimized within a support vector machine framework. We demonstrate the benefits of our kernel learning algorithm on object, face, action and shape recognition. Sadeep Jayasumana, Richard I. Hartley, Mathieu Salzmann, Hongdong Li, Mehrtash Harandi |
CVPR | 5 |
| 2014 | From Manifold to Manifold: Geometry-Aware Dimensionality Reduction for SPD Matrices
Mehrtash Harandi, Mathieu Salzmann, Richard I. Hartley |
ECCV (2) | 1 |
| 2014 | Expanding the Family of Grassmannian Kernels: An Embedding Perspective
Mehrtash Harandi, Mathieu Salzmann, Sadeep Jayasumana, Richard I. Hartley, Hongdong Li |
ECCV (7) | 1 |
| 2014 | Object tracking via non-Euclidean geometry: A Grassmann approachabstractA robust visual tracking system requires an object appearance model that is able to handle occlusion, pose, and illumination variations in the video stream. This can be difficult to accomplish when the model is trained using only a single image. In this paper, we first propose a tracking approach based on affine subspaces (constructed from several images) which are able to accommodate the above-mentioned variations. We use affine subspaces not only to represent the object, but also the candidate areas that the object may occupy. We furthermore propose a novel approach to measure affine subspace-to-subspace distance via the use of non-Euclidean geometry of Grassmann manifolds. The tracking problem is then considered as an inference task in a Markov Chain Monte Carlo framework via particle filtering. Quantitative evaluation on challenging video sequences indicates that the proposed approach obtains considerably better performance than several recent state-of-the-art methods such as Tracking-Learning-Detection and MILtrack. Sareh Abolahrari Shirazi, Mehrtash Harandi, Brian C. Lovell, Conrad Sanderson |
WACV | 2 |
| 2014 | Discriminative Non-Linear Stationary Subspace Analysis for Video ClassificationabstractLow-dimensional representations are key to the success of many video classification algorithms. However, the commonly-used dimensionality reduction techniques fail to account for the fact that only part of the signal is shared across all the videos in one class. As a consequence, the resulting representations contain instance-specific information, which introduces noise in the classification process. In this paper, we introduce non-linear stationary subspace analysis: a method that overcomes this issue by explicitly separating the stationary parts of the video signal (i.e., the parts shared across all videos in one class), from its non-stationary parts (i.e., the parts specific to individual videos). Our method also encourages the new representation to be discriminative, thus accounting for the underlying classification problem. We demonstrate the effectiveness of our approach on dynamic texture recognition, scene classification and action recognition. Mahsa Baktash, Mehrtash Harandi, Brian C. Lovell, Mathieu Salzmann |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | Fisher tensors for classifying human epithelial cells
Masoud Faraki, Mehrtash Harandi, Arnold Wiliem, Brian C. Lovell |
Pattern Recognit. | 2 |
| 2013 | Improved Image Set Classification via Joint Sparse Approximated Nearest SubspacesabstractExisting multi-model approaches for image set classification extract local models by clustering each image set individually only once, with fixed clusters used for matching with other image sets. However, this may result in the two closest clusters to represent different characteristics of an object, due to different undesirable environmental conditions (such as variations in illumination and pose). To address this problem, we propose to constrain the clustering of each query image set by forcing the clusters to have resemblance to the clusters in the gallery image sets. We first define a Frobenius norm distance between subspaces over Grassmann manifolds based on reconstruction error. We then extract local linear subspaces from a gallery image set via sparse representation. For each local linear subspace, we adaptively construct the corresponding closest subspace from the samples of a probe image set by joint sparse representation. We show that by minimising the sparse representation reconstruction error, we approach the nearest point on a Grassmann manifold. Experiments on Honda, ETH-80 and Cambridge-Gesture datasets show that the proposed method consistently outperforms several other recent techniques, such as Affine Hull based Image Set Distance (AHISD), Sparse Approximated Nearest Points(SANP) and Manifold Discriminant Analysis (MDA). Shaokang Chen, Conrad Sanderson, Mehrtash Harandi, Brian C. Lovell |
CVPR | 3 |
| 2013 | Kernel Methods on the Riemannian Manifold of Symmetric Positive Definite MatricesabstractSymmetric Positive Definite (SPD) matrices have become popular to encode image information. Accounting for the geometry of the Riemannian manifold of SPD matrices has proven key to the success of many algorithms. However, most existing methods only approximate the true shape of the manifold locally by its tangent plane. In this paper, inspired by kernel methods, we propose to map SPD matrices to a high dimensional Hilbert space where Euclidean geometry applies. To encode the geometry of the manifold in the mapping, we introduce a family of provably positive definite kernels on the Riemannian manifold of SPD matrices. These kernels are derived from the Gaussian kernel, but exploit different metrics on the manifold. This lets us extend kernel-based algorithms developed for Euclidean spaces, such as SVM and kernel PCA, to the Riemannian manifold of SPD matrices. We demonstrate the benefits of our approach on the problems of pedestrian detection, object categorization, texture analysis, 2D motion segmentation and Diffusion Tensor Imaging (DTI) segmentation. Sadeep Jayasumana, Richard I. Hartley, Mathieu Salzmann, Hongdong Li, Mehrtash Harandi |
CVPR | 5 |
| 2013 | Unsupervised Domain Adaptation by Domain Invariant ProjectionabstractDomain-invariant representations are key to addressing the domain shift problem where the training and test examples follow different distributions. Existing techniques that have attempted to match the distributions of the source and target domains typically compare these distributions in the original feature space. This space, however, may not be directly suitable for such a comparison, since some of the features may have been distorted by the domain shift, or may be domain specific. In this paper, we introduce a Domain Invariant Projection approach: An unsupervised domain adaptation method that overcomes this issue by extracting the information that is invariant across the source and target domains. More specifically, we learn a projection of the data to a low-dimensional latent space where the distance between the empirical distributions of the source and target examples is minimized. We demonstrate the effectiveness of our approach on the task of visual object recognition and show that it outperforms state-of-the-art methods on a standard domain adaptation benchmark dataset. Mahsa Baktash, Mehrtash Harandi, Brian C. Lovell, Mathieu Salzmann |
ICCV | 2 |
| 2013 | Dictionary Learning and Sparse Coding on Grassmann Manifolds: An Extrinsic SolutionabstractRecent advances in computer vision and machine learning suggest that a wide range of problems can be addressed more appropriately by considering non-Euclidean geometry. In this paper we explore sparse dictionary learning over the space of linear subspaces, which form Riemannian structures known as Grassmann manifolds. To this end, we propose to embed Grassmann manifolds into the space of symmetric matrices by an isometric mapping, which enables us to devise a closed-form solution for updating a Grassmann dictionary, atom by atom. Furthermore, to handle non-linearity in data, we propose a kernelised version of the dictionary learning algorithm. Experiments on several classification tasks (face recognition, action recognition, dynamic texture classification) show that the proposed approach achieves considerable improvements in discrimination accuracy, in comparison to state-of-the-art methods such as kernelised Affine Hull Method and graph-embedding Grassmann discriminant analysis. Mehrtash Harandi, Conrad Sanderson, Chunhua Shen, Brian C. Lovell |
ICCV | 1 |
| 2013 | A Framework for Shape Analysis via Hilbert Space EmbeddingabstractWe propose a framework for 2D shape analysis using positive definite kernels defined on Kendall's shape manifold. Different representations of 2D shapes are known to generate different nonlinear spaces. Due to the nonlinearity of these spaces, most existing shape classification algorithms resort to nearest neighbor methods and to learning distances on shape spaces. Here, we propose to map shapes on Kendall's shape manifold to a high dimensional Hilbert space where Euclidean geometry applies. To this end, we introduce a kernel on this manifold that permits such a mapping, and prove its positive definiteness. This kernel lets us extend kernel-based algorithms developed for Euclidean spaces, such as SVM, MKL and kernel PCA, to the shape manifold. We demonstrate the benefits of our approach over the state-of-the-art methods on shape classification, clustering and retrieval. Sadeep Jayasumana, Mathieu Salzmann, Hongdong Li, Mehrtash Harandi |
ICCV | 4 |
| 2013 | Multi-shot person re-identification via relational Stein divergenceabstractPerson re-identification is particularly challenging due to significant appearance changes across separate camera views. In order to re-identify people, a representative human signature should effectively handle differences in illumination, pose and camera parameters. While general appearance-based methods are modelled in Euclidean spaces, it has been argued that some applications in image and video analysis are better modelled via non-Euclidean manifold geometry. To this end, recent approaches represent images as covariance matrices, and interpret such matrices as points on Riemannian manifolds. As direct classification on such manifolds can be difficult, in this paper we propose to represent each manifold point as a vector of similarities to class representers, via a recently introduced form of Bregman matrix divergence known as the Stein divergence. This is followed by using a discriminative mapping of similarity vectors for final classification. The use of similarity vectors is in contrast to the traditional approach of embedding manifolds into tangent spaces, which can suffer from representing the manifold structure inaccurately. Comparative evaluations on benchmark ETHZ and iLIDS datasets for the person re-identification task show that the proposed approach obtains better performance than recent techniques such as Histogram Plus Epitome, Partial Least Squares, and Symmetry-Driven Accumulation of Local Features. Azadeh Alavi, Mehrtash Harandi, Conrad Sanderson |
ICIP | 3 |
| 2013 | Non-Linear Stationary Subspace Analysis with Application to Video ClassificationabstractLow-dimensional representations are key to the success of many video classification algorithms. However, the commonly-used dimensionality reduction techniques fail to account for the fact that only part of the signal is shared across all the videos in one class. As a consequence, the resulting representations contain instance-specific information, which introduces noise in the classification process. In this paper, we introduce Non-Linear Stationary Subspace Analysis: A method that overcomes this issue by explicitly separating the stationary parts of the video signal (i.e., the parts shared across all videos in one class), from its non-stationary parts (i.e., specific to individual videos). We demonstrate the effectiveness of our approach on action recognition, dynamic texture classification and scene recognition. Mahsa Baktash, Mehrtash Harandi, Abbas Bigdeli, Brian C. Lovell, Mathieu Salzmann |
ICML (3) | 2 |
| 2013 | Relational divergence based classification on Riemannian manifoldsabstractA recent trend in computer vision is to represent images through covariance matrices, which can be treated as points on a special class of Riemannian manifolds. A popular way of analysing such manifolds is to embed them in Euclidean spaces, a process which can be interpreted as warping the feature space. Embedding manifolds is not without problems, as the manifold structure may not be accurately preserved. In this paper, we propose a new method for analysing Riemannian manifolds, where embedding into Euclidean spaces is not explicitly required. To this end, we propose to represent Riemannian points through their similarities to a set of reference points on the manifold, with the aid of the recently proposed Stein divergence, which is a symmetrised version of Bregman matrix divergence. Classification problems on manifolds are then effectively converted into the problem of finding appropriate machinery over the space of similarities, which can be tackled by conventional Euclidean learning methods such as linear discriminant analysis. Experiments on face recognition, person re-identification and texture classification show that the proposed method outperforms state-of-the-art approaches, such as Tensor Sparse Coding, Histogram Plus Epitome and the recent Riemannian Locality Preserving Projection. Azadeh Alavi, Mehrtash Harandi, Conrad Sanderson |
WACV | 2 |
| 2013 | Spatio-temporal covariance descriptors for action and gesture recognitionabstractWe propose a new action and gesture recognition method based on spatio-temporal covariance descriptors and a weighted Riemannian locality preserving projection approach that takes into account the curved space formed by the descriptors. The weighted projection is then exploited during boosting to create a final multiclass classification algorithm that employs the most useful spatio-temporal regions. We also show how the descriptors can be computed quickly through the use of integral video representations. Experiments on the UCF sport, CK+ facial expression and Cambridge hand gesture datasets indicate superior performance of the proposed method compared to several recent state-of-the-art techniques. The proposed method is robust and does not require additional processing of the videos, such as foreground detection, interest-point detection or tracking. Andres Sanin, Conrad Sanderson, Mehrtash Harandi, Brian C. Lovell |
WACV | 3 |
| 2013 | Kernel analysis on Grassmann manifolds for action recognition
Mehrtash Harandi, Conrad Sanderson, Sareh Abolahrari Shirazi, Brian C. Lovell |
Pattern Recognit. Lett. | 1 |
| 2012 | Combined Learning of Salient Local Descriptors and Distance Metrics for Image Set Face VerificationabstractIn contrast to comparing faces via single exemplars, matching sets of face images increases robustness and discrimination performance. Recent image set matching approaches typically measure similarities between subspaces or manifolds, while representing faces in a rigid and holistic manner. Such representations are easily affected by variations in terms of alignment, illumination, pose and expression. While local feature based representations are considerably more robust to such variations, they have received little attention within the image set matching area. We propose a novel image set matching technique, comprised of three aspects: (i) robust descriptors of face regions based on local features, partly inspired by the hierarchy in the human visual system, (ii) use of several subspace and exemplar metrics to compare corresponding face regions, (iii) jointly learning which regions are the most discriminative while finding the optimal mixing weights for combining metrics. Experiments on LFW, PIE and MOBIO face datasets show that the proposed algorithm obtains considerably better performance than several recent state of-the-art techniques, such as Local Principal Angle and the Kernel Affine Hull Method. Conrad Sanderson, Mehrtash Harandi, Yongkang Wong, Brian C. Lovell |
AVSS | 2 |
| 2012 | Sparse Coding and Dictionary Learning for Symmetric Positive Definite Matrices: A Kernel Approach
Mehrtash Harandi, Conrad Sanderson, Richard I. Hartley, Brian C. Lovell |
ECCV (2) | 1 |
| 2012 | Directional Space-Time Oriented Gradients for 3D Visual Pattern Analysis
Ehsan Norouznezhad, Mehrtash Harandi, Abbas Bigdeli, Mahsa Baktash, Adam Postula, Brian C. Lovell |
ECCV (3) | 2 |
| 2012 | K-tangent spaces on Riemannian manifolds for improved pedestrian detectionabstractFor covariance-based image descriptors, taking into account the curvature of the corresponding feature space has been shown to improve discrimination performance. This is often done through representing the descriptors as points on Riemannian manifolds, with the discrimination accomplished on a tangent space. However, such treatment is restrictive as distances between arbitrary points on the tangent space do not represent true geodesic distances, and hence do not represent the manifold structure accurately. In this paper we propose a general discriminative model based on the combination of several tangent spaces, in order to preserve more details of the structure. The model can be used as a weak learner in a boosting-based pedestrian detection framework. Experiments on the challenging INRIA and DaimlerChrysler datasets show that the proposed model leads to considerably higher performance than methods based on histograms of oriented gradients as well as previous Riemannian-based techniques. Andres Sanin, Conrad Sanderson, Mehrtash Harandi, Brian C. Lovell |
ICIP | 3 |
| 2012 | Clustering on Grassmann manifolds via kernel embedding with application to action analysisabstractWith the aim of improving the clustering of data (such as image sequences) lying on Grassmann manifolds, we propose to embed the manifolds into Reproducing Kernel Hilbert Spaces. To this end, we define a measure of cluster distortion and embed the manifolds such that the distortion is minimised. We show that the optimal solution is a generalised eigenvalue problem that can be solved very efficiently. Experiments on several clustering tasks (including human action clustering) show that in comparison to the recent intrinsic Grassmann k-means algorithm, the proposed approach obtains notable improvements in clustering accuracy, while also being several orders of magnitude faster. Sareh Abolahrari Shirazi, Mehrtash Harandi, Conrad Sanderson, Azadeh Alavi, Brian C. Lovell |
ICIP | 2 |
| 2012 | On robust biometric identity verification via sparse encoding of faces: Holistic vs local approachesabstractIn the field of face recognition, Sparse Representation (SR) has received considerable attention during the past few years. Most of the related literature focuses on holistic descriptors in closed-set identification applications. The underlying assumption in identification is that the gallery always has sufficient samples per subject to linearly reconstruct a query image. Unfortunately, such assumption is easily violated in the more challenging and realistic face verification scenario. A verification algorithm is required to determine if two faces (where one or both have not been seen before) belong to the same person, while explicitly taking into account the possibility of impostor attacks. In this paper, we first discuss why most of the SR literature is not applicable to verification problems. Motivated by the success of bag-of-words methods in the field of object recognition, which describe an image as a set of local patches or interest points, we then propose to tackle the verification problem by encoding each local face patch through SR. The locally encoded sparse vectors are pooled to form regional descriptors, where each descriptor covers a relatively large portion of the face. Experiments in various challenging conditions show that the proposed method achieves high and robust verification performance. Yongkang Wong, Mehrtash Harandi, Conrad Sanderson, Brian C. Lovell |
IJCNN | 2 |
| 2012 | Kernel analysis over Riemannian manifolds for visual recognition of actions, pedestrians and texturesabstractA convenient way of analysing Riemannian manifolds is to embed them in Euclidean spaces, with the embedding typically obtained by flattening the manifold via tangent spaces. This general approach is not free of drawbacks. For example, only distances between points to the tangent pole are equal to true geodesic distances. This is restrictive and may lead to inaccurate modelling. Instead of using tangent spaces, we propose embedding into the Reproducing Kernel Hilbert Space by introducing a Riemannian pseudo kernel. We furthermore propose to recast a locality preserving projection technique from Euclidean spaces to Riemannian manifolds, in order to demonstrate the benefits of the embedding. Experiments on several visual classification tasks (gesture recognition, person re-identification and texture classification) show that in comparison to tangent-based processing and state-of-the-art methods (such as tensor canonical correlation analysis), the proposed approach obtains considerable improvements in discrimination accuracy. Mehrtash Harandi, Conrad Sanderson, Arnold Wiliem, Brian C. Lovell |
WACV | 1 |
| 2011 | Graph embedding discriminant analysis on Grassmannian manifolds for improved image set matchingabstractA convenient way of dealing with image sets is to represent them as points on Grassmannian manifolds. While several recent studies explored the applicability of discriminant analysis on such manifolds, the conventional formalism of discriminant analysis suffers from not considering the local structure of the data. We propose a discriminant analysis approach on Grassmannian manifolds, based on a graph-embedding framework. We show that by introducing within-class and between-class similarity graphs to characterise intra-class compactness and inter-class separability, the geometrical structure of data can be exploited. Experiments on several image datasets (PIE, BANCA, MoBo, ETH-80) show that the proposed algorithm obtains considerable improvements in discrimination accuracy, in comparison to three recent methods: Grassmann Discriminant Analysis (GDA), Kernel GDA, and the kernel version of Affine Hull Image Set Distance. We further propose a Grassmannian kernel, based on canonical correlation between subspaces, which can increase discrimination accuracy when used in combination with previous Grassmannian kernels. Mehrtash Harandi, Conrad Sanderson, Sareh Abolahrari Shirazi, Brian C. Lovell |
CVPR | 1 |
| 2011 | Ensemble of furthest subspace pairs for enhanced image set matchingabstractRecently it has been shown that the performance of image set matching methods can be improved by clustering set samples into smaller and more coherent groups. Typically, set samples are treated independently during clustering, ie., clustering criteria have not been defined to exploit set characteristics. In this paper we introduce a novel approach to image set clustering by considering the similarities between subspaces instead of similarities between samples. We exploit an ensemble learning technique to create an ensemble of subspace pairs. Each pair has the property that its members are located at the furthest distance in the sense of distances between subspaces. Object recognition experiments on the CMU-MoBO and ETH-80 datasets show that the proposed method obtains higher discrimination accuracy in comparison to several benchmark methods as well as the recently proposed Kernel Affine Hull Method. Mehrtash Harandi, Conrad Sanderson, Abbas Bigdeli, Brian C. Lovell |
ICIP | 1 |
| 2010 | Image-set face recognition based on transductive learningabstractIn this paper we consider the problem of face recognition in a scenario when the query consists of a set of images and the gallery contains a single still image per subject. This is a more challenging problem compared to image-set to image-set matching and has wider applications in advanced surveillance, smart access control and human-computer interaction. Unfortunately most of the previous matching strategies in literature fail to work or deteriorate drastically if they are provided with one sample per class as the gallery data. In this paper we demonstrate how transductive learning can be utilized to map the image-set to single image matching problem into the recently-studied framework of set matching using canonical correlations. Experimental results on different challenging datasets reveal the efficiency of the proposed method against existing approaches. Mehrtash Harandi, Abbas Bigdeli, Brian C. Lovell |
ICIP | 1 |
| 2010 | Directed Random Subspace Method for Face RecognitionabstractWith growing attention to ensemble learning, in recent years various ensemble methods for face recognition have been proposed that show promising results. Among diverse ensemble construction approaches, random subspace method has received considerable attention in face recognition. Although random feature selection in random subspace method improves accuracy in general, it is not free of serious difficulties and drawbacks. In this paper we present a learning scheme to overcome some of the drawbacks of random feature selection in the random subspace method. The proposed learning method derives a feature discrimination map based on a measure of accuracy and uses it in a probabilistic recall mode to construct an ensemble of subspaces. Experiments on different face databases revealed that the proposed method gives superior performance over the well-known benchmarks and state of the art ensemble methods. Mehrtash Harandi, Majid Nili Ahmadabadi, Babak Nadjar Araabi, Abbas Bigdeli, Brian C. Lovell |
ICPR | 1 |
| 2009 | Optimal Local Basis: A Reinforcement Learning Approach for Face Recognition
Mehrtash Harandi, Majid Nili Ahmadabadi, Babak Nadjar Araabi |
Int. J. Comput. Vis. | 1 |
| 2007 | A Hierarchical Face Identification System Based on Facial ComponentsabstractIt is generally agreed that faces are not recognized only by utilizing some holistic search among all learned faces, but also through a feature analysis that aimed to specify more important features of each specific face. This paper addresses a novel decision strategy that efficiently uses both holistic and facial component (left eye, right eye, nose and mouth) feature analysis to recognize faces. The proposed algorithm uses the whole face features in the first step of recognition task. If the decision machine fails to assign a class (with high confidence) then the individual facial components are processed and the resulting information are combined with those obtained from the whole face to assign the output. Simulation studies justify the superior performance of the proposed method as compared to that of Eigenface method. Experimental results also show that the proposed system is robust against small errors in facial component extractor. Mehrtash Harandi, Majid Nili Ahmadabadi, Babak Nadjar Araabi |
AICCSA | 1 |
| 2004 | Face recognition using reinforcement learning
Mehrtash Harandi, Majid Nili Ahmadabadi, Babak Nadjar Araabi |
ICIP | 1 |
| 2004 | A SVM-based method for face recognition using a wavelet PCA representation of facesabstractThis paper proposes a new method of face representation which is used for face recognition by SVM. For face representation we have used a two-step method, first two-dimensional discrete wavelet transform (DWT) is used to transform the faces to a more discriminated space and then principal component analysis (PCA) is applied. The proposed method produced a significant improvement which includes a substantial reduction in error rate and in time of processing during the obtaining PCA orthonormal basis. Majid Safari, Mehrtash Harandi, Babak Nadjar Araabi |
ICIP | 2 |
| 2003 | Low bitrate image compression using self-organized Kohonen mapsabstractIn this paper, we propose a new image compression algorithm based on Kohonen self-organized maps. The compression is based on vector quantization (VQ) of the DCT coefficients of image blocks, where the VQ is implemented by a Kohonen network. At low bitrates, our proposed method performs better than an earlier compression scheme developed by Amerijckx et al. (1998) and shows better subjective results in comparison to JPEG. Mehrtash Harandi, Mohammad Gharavi-Alkhansari |
ICIP (2) | 1 |