Mete Ozay

dblp:26/5515 · also Mete Özay · DBLP profile ↗
← Back
73ranked-venue papers
10as first author
50since 2021 · last 2026
0000-0002-7189-7260ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 45 · 6 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 44 · 6 first-author · 29 since 2021Systems, architecture and hardware · 6 · 1 first-author · 4 since 2021Computer networks · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021
YearPublicationVenuePosition
2026 K-Merge: Online Continual Merging of Adapters for On-device Large Language Models
abstract
Donald Shenaj, Ondrej Bohdal, Taha Ceritli, Mete Ozay, Pietro Zanuttigh, Umberto Michieli. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Donald Shenaj, Ondrej Bohdal, Taha Ceritli, Mete Ozay, Pietro Zanuttigh, Umberto Michieli
ACL (1)4
2026 DP-DyLoRA: Fine-Tuning Transformer-Based Models On-Device under Differentially Private Federated Learning using Dynamic Low-Rank Adaptation
abstract
Federated learning (FL) allows clients to collaboratively train a global model without sharing their local data with a server. However, clients' contributions to the server can still leak sensitive information. Differential privacy (DP) addresses such leakage by providing formal privacy guarantees, with mechanisms that add randomness to the clients' contributions. The randomness makes it infeasible to train large transformer-based models, common in modern federated learning systems. In this work, we empirically evaluate the practicality of fine-tuning large scale on-device transformer-based models with differential privacy in a federated learning system. We conduct comprehensive experiments on various system properties for tasks spanning a multitude of domains: speech recognition, computer vision (CV) and natural language understanding (NLU). Our results show that full fine-tuning under differentially private federated learning (DP-FL) generally leads to huge performance degradation which can be alleviated by reducing the dimensionality of contributions through parameter-efficient fine-tuning (PEFT). Our benchmarks of existing DP-PEFT methods show that DP-Low-Rank Adaptation (DP-LoRA) consistently outperforms other methods. An even more promising approach, DyLoRA, which makes the low rank variable, when naively combined with FL would straightforwardly break differential privacy. We therefore propose an adaptation method that can be combined with differential privacy and call it DP-DyLoRA. Finally, we are able to reduce the accuracy degradation and word error rate (WER) increase due to DP to less than 2% and 7% respectively with 1 million clients and a stringent privacy budget of $ε=2$.
Karthikeyan Saravanan, Rogier C. van Dalen, Haaris Mehmood, David Tuckey, Mete Ozay
ICC6
2026 Mem-MLP: Real-Time 3D Human Motion Generation from Sparse Inputs
abstract
Realistic and smooth full-body tracking is crucial for immersive AR/VR applications. Existing systems primarily track head and hands via Head Mounted Devices (HMDs) and controllers, making the 3D full-body reconstruction incomplete. One potential approach is to generate the full-body motions from sparse inputs collected from limited sensors using a Neural Network (NN) model. In this paper, we propose a novel method based on a multi-layer perceptron (MLP) backbone that is enhanced with residual connections and a novel NN-component called Memory-Block. In particular, Memory-Block represents missing sensor data with trainable code-vectors, which are combined with the sparse signals from previous time instances to improve the temporal consistency. Furthermore, we formulate our solution as a multi-task learning problem, allowing our MLP-backbone to learn robust representations that boost accuracy. Our experiments show that our method outperforms state-of-the-art baselines by substantially reducing prediction errors. Moreover, it achieves 72 FPS on mobile HMDs that ultimately improves the accuracy-running time tradeoff.
Sinan Mutlu, Georgios-Fotios Angelis, Savas Özkan, Paul Wisbey, Anastasios Drosou, Mete Ozay
WACV6
2026 Guided Model Merging for Hybrid Data Learning: Leveraging Centralized Data to Refine Decentralized Models
abstract
Current network training paradigms primarily focus on either centralized or decentralized data regimes. However, in practice, data availability often exhibits a hybrid nature, where both regimes coexist. This hybrid setting presents new opportunities for model training, as the two regimes offer complementary trade-offs: decentralized data is abundant but subject to heterogeneity and communication constraints, while centralized data—though limited in volume and potentially unrepresentative—enables better curation and high-throughput access. Despite its potential, effectively combining these paradigms remains challenging, and few frameworks are tailored to hybrid data regimes. To address this, we propose a novel framework that constructs a model atlas from decentralized models and leverages centralized data to refine a global model within this structured space. The refined model is then used to reinitialize the decentralized models. Our method synergizes federated learning (to exploit decentralized data) and model merging (to utilize centralized data), enabling effective training under hybrid data availability. Theoretically, we show that our approach achieves faster convergence than methods relying solely on decentralized data, due to variance reduction in the merging process. Extensive experiments demonstrate that our framework consistently outperforms purely centralized, purely decentralized, and existing hybrid-adaptable methods. Notably, our method remains robust even when the centralized and decentralized data domains differ or when decentralized data contains noise, significantly broadening its applicability.
Junyi Zhu 0002, Ruicong Yao, Taha Ceritli, Savas Özkan, Matthew B. Blaschko, Eunchung Noh, Jeongwon Min, Cho Jung Min, Mete Ozay
WACV9
2025 DreamCache: Finetuning-Free Lightweight Personalized Image Generation via Feature Caching
abstract
Personalized image generation requires text-to-image generative models that capture the core features of a reference subject to allow for controlled generation across different contexts. Existing methods face challenges due to complex training requirements, high inference costs, limited flexibility, or a combination of these issues. In this paper, we introduce DreamCache, a scalable approach for efficient and high-quality personalized image generation. By caching a small number of reference image features from a subset of layers and a single timestep of the pretrained diffusion denoiser, DreamCache enables dynamic modulation of the generated image features through lightweight, trained conditioning adapters. DreamCache achieves state-of-the-art image and text alignment, utilizing an order of magnitude fewer extra parameters, and is both more computationally effective and versatile than existing models.1
Emanuele Aiello, Umberto Michieli, Diego Valsesia, Mete Ozay, Enrico Magli
CVPR4
2025 Accurate Scene Text Recognition with Efficient Model Scaling and Cloze Self-Distillation
abstract
Scaling architectures have been proven effective for improving Scene Text Recognition (STR), but the individual contribution of vision encoder and text decoder scaling remain under-explored. In this work, we present an in-depth empirical analysis and demonstrate that, contrary to previous observations, scaling the decoder yields significant performance gains, always exceeding those achieved by encoder scaling alone. We also identify label noise as a key challenge in STR, particularly in real-world data, which can limit the effectiveness of STR models. To address this, we propose Cloze Self-Distillation (CSD), a method that mitigates label noise by distilling a student model from context-aware soft predictions and pseudolabels generated by a teacher model. Additionally, we enhance the decoder architecture by introducing differential cross-attention for STR. Our methodology achieves state-of-the-art performance on 10 out of 11 benchmarks using only real data, while significantly reducing the parameter size and computational costs.
Andrea Maracani, Savas Özkan, Sijun Cho, Hyowon Kim, Eunchung Noh, Jeongwon Min, Cho Jung Min, Dookun Park, Mete Ozay
CVPR9
2025 Efficient Compositional Multi-tasking for On-device Large Language Models
abstract
Adapter parameters provide a mechanism to modify the behavior of machine learning models and have gained significant popularity in the context of large language models (LLMs) and generative AI.These parameters can be merged to support multiple tasks via a process known as task merging.However, prior work on merging in LLMs, particularly in natural language processing, has been limited to scenarios where each test example addresses only a single task.In this paper, we focus on on-device settings and study the problem of text-based compositional multi-tasking, where each test example involves the simultaneous execution of multiple tasks.For instance, generating a translated summary of a long text requires solving both translation and summarization tasks concurrently.To facilitate research in this setting, we propose a benchmark comprising four practically relevant compositional tasks.We also present an efficient method (Learnable Calibration) tailored for on-device applications, where computational resources are limited, emphasizing the need for solutions that are both resourceefficient and high-performing.Our contributions lay the groundwork for advancing the capabilities of LLMs in real-world multi-tasking scenarios, expanding their applicability to complex, resource-constrained use cases.Project page:
Ondrej Bohdal, Mete Ozay, Ji Joong Moon, Kyeng-Hun Lee, Hyeonmok Ko, Umberto Michieli
EMNLP2
2025 HydraOpt: Navigating the Efficiency-Performance Trade-off of Adapter Merging
abstract
Large language models (LLMs) often leverage adapters, such as low-rank-based adapters, to achieve strong performance on downstream tasks.However, storing a separate adapter for each task significantly increases memory requirements, posing a challenge for resourceconstrained environments such as mobile devices.Although model merging techniques can reduce storage costs, they typically result in substantial performance degradation.In this work, we introduce HydraOpt, a new model merging technique that capitalizes on the inherent similarities between the matrices of low-rank adapters.Unlike existing methods that produce a fixed trade-off between storage size and performance, HydraOpt allows us to navigate this spectrum of efficiency and performance.Our experiments show that Hy-draOpt significantly reduces storage size (48% reduction) compared to storing all adapters, while achieving competitive performance (0.2-1.8% drop).Furthermore, it outperforms existing merging techniques in terms of performance at the same or slightly worse storage efficiency.
Taha Ceritli, Ondrej Bohdal, Mete Ozay, Ji Joong Moon, Kyeng-Hun Lee, Hyeonmok Ko, Umberto Michieli
EMNLP3
2025 ValSub: Subsampling Validation Data to Mitigate Forgetting during ASR Personalization
abstract
Automatic Speech Recognition (ASR) is widely used within consumer devices such as mobile phones. Recently, personalization or on-device model fine-tuning has shown that adaptation of ASR models towards target user speech improves their performance over rare words or accented speech. Despite these gains, fine-tuning on user data (target domain) risks the personalized model to forget knowledge about its original training distribution (source domain) i.e. catastrophic forgetting, leading to subpar general ASR performance. A simple and efficient approach to combat catastrophic forgetting is to measure forgetting via a validation set that represents the source domain distribution. However, such validation sets are large and impractical for mobile devices. Towards this, we propose a novel method to subsample a substantially large validation set into a smaller one while maintaining the ability to estimate forgetting. We demonstrate the efficacy of such a dataset in mitigating forgetting by utilizing it to dynamically determine the number of ideal fine-tuning epochs. When measuring the deviations in per user fine-tuning epochs against a 50x larger validation set (oracle), our method achieves a lower mean-absolute-error (3.39) compared to randomly selected subsets of the same size (3.78-8.65). Unlike random baselines, our method consistently tracks the oracle’s behaviour across three different forgetting thresholds.
Haaris Mehmood, Karthikeyan Saravanan, Pablo Peso Parada, David Tuckey, Mete Ozay, Gil Ho Lee, Jungin Lee, Seokyeong Jung
ICASSP5
2025 A Study of Improving The Privacy-Utility Trade-off of Task-specific Models with Learnable Privacy
abstract
In recent years, machine learning (ML) models have been integrated into various applications and products to improve user experience. However, this approach raises significant concerns about the protection of private user data utilized for training the models. One limitation of vanilla privacy methods is that they can improve the robustness of the models against privacy attacks at the cost of accuracy while performing ML tasks. We propose a framework for implementing privacy models that learn privacy budgets to improve the trade-off between privacy of user data, task models, and their utility (task accuracy). The experimental results show that our framework dramatically enhances the task accuracy of ML models in image classification tasks while providing better privacy protection compared to the state-of-the-art methods.
Savas Özkan, Taha Ceritli, Jeongwon Min, Eunchung Noh, Jung Min Cho, Dookun Park, Mete Ozay
ICASSP7
2025 Hyper-Refinement for Low-Rank Adaptation
abstract
Parameter-efficient fine-tuning (PEFT) is utilized to adapt large pre-trained machine learning (ML) models to new tasks using a small number of trainable parameters. In particular, Low-Rank Adaptation (LoRA) is one of the prominent PEFT methods. To this end, we introduce a novel method that exploits data to improve the accuracy of models fine-tuned by LoRA. Our method implements a hypernetwork that generates refinement parameters using data to update the low-rank parameters of LoRA. The experimental results validate that our method improves the accuracy of large language models (LLMs) on language understanding tasks by 3% on average compared to LoRA and its variants.
Savas Özkan, Taha Ceritli, Jeongwon Min, Eunchung Noh, Jung Min Cho, Dookun Park, Mete Ozay
ICASSP7
2025 persoDA: Personalized Data Augmentation for Personalized ASR
abstract
Data augmentation (DA) is ubiquitously used in training of Automatic Speech Recognition (ASR) models. DA offers increased data variability, robustness and generalization against different acoustic distortions. Recently, personalization of ASR models on mobile devices has been shown to improve Word Error Rate (WER). This paper evaluates data augmentation in this context and proposes persoDA; a DA method driven by user’s data utilized to personalize ASR [1] –[3]. persoDA aims to augment training with data specifically tuned towards acoustic characteristics of the end-user, as opposed to standard augmentation based on Multi-Condition Training (MCT) that applies random reverberation and noises. Our evaluation with an ASR conformer-based baseline trained on Librispeech and per-sonalized for VOICES [4] shows that persoDA achieves a 13.9% relative WER reduction over using standard data augmentation (using random noise & reverberation). Furthermore, persoDA shows 16% to 20% faster convergence over MCT.
Pablo Peso Parada, Spyros Fontalis, Md Asif Jalal, Karthikeyan Saravanan, Anastasios Drosou, Mete Ozay, Gil Ho Lee, Jungin Lee, Seokyeong Jung
ICASSP6
2025 Controllable Forgetting Mechanism for Few-Shot Class-Incremental Learning
abstract
Class-incremental learning in the context of limited personal labeled samples (few-shot) is critical for numerous real-world applications, such as smart home devices. A key challenge in these scenarios is balancing the trade-off between adapting to new, personalized classes and maintaining the performance of the model on the original, base classes. Fine-tuning the model on novel classes often leads to the phenomenon of catastrophic forgetting, where the accuracy of base classes declines unpredictably and significantly. In this paper, we propose a simple yet effective mechanism to address this challenge by controlling the trade-off between novel and base class accuracy. We specifically target the ultra-low-shot scenario, where only a single example is available per novel class. Our approach introduces a Novel Class Detection (NCD) rule, which adjusts the degree of forgetting a priori while simultaneously enhancing performance on novel classes. We demonstrate the versatility of our solution by applying it to state-of-the-art Few-Shot Class-Incremental Learning (FSCIL) methods, showing consistent improvements across different settings. To better quantify the trade-off between novel and base class performance, we introduce new metrics: NCR@2FOR and NCR@5FOR. Our approach achieves up to a 30% improvement in novel class accuracy on the CIFAR100 dataset (1-shot, 1 novel class) while maintaining a controlled base class forgetting rate of 2%.
Kirill Paramonov, Mete Ozay, Eunju Yang, Ji Joong Moon, Umberto Michieli
ICASSP2
2025 LoRA.rar: Learning to Merge LoRAs via Hypernetworks for Subject-Style Conditioned Image Generation
abstract
Recent advancements in image generation models have enabled personalized image creation with both user-defined subjects (content) and styles. Prior works achieved personalization by merging corresponding low-rank adapters (LoRAs) through optimization-based methods, which are computationally demanding and unsuitable for real-time use on resource-constrained devices like smartphones. To address this, we introduce LoRA$.$rar, a method that not only improves image quality but also achieves a remarkable speedup of over $4000\times$ in the merging process. We collect a dataset of style and subject LoRAs and pre-train a hypernetwork on a diverse set of content-style LoRA pairs, learning an efficient merging strategy that generalizes to new, unseen content-style pairs, enabling fast, high-quality personalization. Moreover, we identify limitations in existing evaluation metrics for content-style quality and propose a new protocol using multimodal large language models (MLLMs) for more accurate assessment. Our method significantly outperforms the current state of the art in both content and style fidelity, as validated by MLLM assessments and human evaluations.
Donald Shenaj, Ondrej Bohdal, Mete Ozay, Pietro Zanuttigh, Umberto Michieli
ICCV3
2025 Efficient and Accurate Scene Text Recognition with Cascaded-Transformers
abstract
In recent years, vision transformers with text decoder have demonstrated remarkable performance on Scene Text Recognition (STR) due to their ability to capture long-range dependencies and contextual relationships with high learning capacity. However, the computational and memory demands of these models are significant, limiting their deployment in resource-constrained applications. To address this challenge, we propose an efficient and accurate STR system. Specifically, we focus on improving the efficiency of encoder models by introducing a cascaded-transformers structure. This structure progressively reduces the vision token size during the encoding step, effectively eliminating redundant tokens and reducing computational cost. Our experimental results confirm that our STR system achieves comparable performance to state-of-the-art baselines while substantially decreasing computational requirements. In particular, for large-models, the accuracy remains same, 92.77 → 92.68, while computational complexity is almost halved with our structure.
Savas Özkan, Andrea Maracani, Mete Ozay, Hyowon Kim, Sijun Cho, Eunchung Noh, Jeongwon Min, Jung Min Cho
MMSys3
2025 Continual Error Correction on Low-Resource Devices
abstract
The proliferation of AI models in everyday devices has highlighted a critical challenge: prediction errors that degrade user experience. While existing solutions focus on error detection, they rarely provide efficient correction mechanisms, especially for resource-constrained devices. We present a novel system enabling users to correct AI misclassifications through few-shot learning, requiring minimal computational resources and storage. Our approach combines server-side foundation model training with on-device prototype-based classification, enabling efficient error correction through prototype updates rather than model retraining. The system consists of two key components: (1) a server-side pipeline that leverages knowledge distillation to transfer robust feature representations from foundation models to device-compatible architectures, and (2) a device-side mechanism that enables ultra-efficient error correction through prototype adaptation. We demonstrate our system's effectiveness on both image classification and object detection tasks, achieving over 50% error correction in one-shot scenarios on Food-101 and Flowers-102 datasets while maintaining minimal forgetting (less than 0.02%) and negligible computational overhead. Our implementation, validated through an Android demonstration app, proves the system's practicality in real-world scenarios.
Kirill Paramonov, Mete Ozay, Aristeidis Mystakidis, Nikolaos Tsalikidis, Dimitrios Sotos, Anastasios Drosou, Dimitrios Tzovaras, Kiseok Chang, Sangdok Mo, Namwoong Kim, Woojong Yoo, Ji Joong Moon, Umberto Michieli
MMSys2
2025 Robust Depth Estimation Under Sensor Degradations: A Multi-Sensor Fusion Perspective
abstract
The significance of depth estimation has spurred recent endeavors to enhance it through Multi-Sensor Fusion (MSF). However, prevailing MSF methods exhibit limitations concerning accuracy and resilience when confronted with sensor degradations. While certain forms of degradation, such as suboptimal lighting and adverse weather conditions, can be mitigated by collecting pertinent data in data-driven learning, this approach proves ineffective for Out-of-Distribution (OOD) sensor degradations. In this paper, we propose a novel approach termed Combinable and Separable Multi-Sensor Fusion (CSMSF) designed to bolster depth estimation robustness against multiple sensor degradations. CSMSF hinges on four core principles: i) improved performance is achieved with an increased number of valid sensors, ii) a single valid sensor can independently enable its own depth estimation, iii) maintaining a judicious equilibrium between accuracy and model complexity, and iv) autonomous diagnosis of sensor observation failure. Leveraging these advantages, CSMSF identifies and rejects degraded sensors, allowing autonomous selection of valid sensors for scene depth estimation. The experimental results demonstrate the superior robustness of the proposed CSMSF, underscoring its efficacy in addressing challenges associated with sensor degradations across diverse environmental conditions.
Junjie Hu 0003, Chenyou Fan, Mete Ozay, Qing Gao 0002, Yulan Guo, Tin Lun Lam
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Unlocking Drone Perception in Low AGL Heights: Progressive Semi-Supervised Learning for Ground-to-Aerial Perception Knowledge Transfer
abstract
We explore the novel challenge of drone perception across varying low AGL (above ground level) heights, a task essential for dynamic tasks, unlike the fixed ground viewpoint in autonomous driving. Supervised learning for this incurs high annotation costs, and current semi-supervised methods struggle with viewpoint differences. In this paper, we introduce ground-to-aerial perception knowledge transfer and propose a progressive semi-supervised learning framework for drone perception using only labeled data from the ground viewpoint and unlabeled data from flying viewpoints. The framework hinges on four key components: 1) a dense viewpoint sampling strategy, segmenting the vertical flight height range into evenly distributed intervals; 2) nearest neighbor pseudo-labeling, inferring labels of the nearest neighbor viewpoint using a model learned on the preceding viewpoint; 3) MixView, generating augmented images among different viewpoints to mitigate viewpoint differences; and 4) a progressive distillation strategy, gradually learning until reaching the maximum flying height. To validate our approach, we create both synthesized and real-world datasets. Extensive experimental analyses reveal a remarkable relative accuracy improvement of 25.7% and 16.9% for the synthesized dataset and the real world, respectively. Code and datasets are available on https://github.com/FreeformRobotics/Progressive-Self-Distillation-for-Ground-to-Aerial-Perception-Knowledge-Transfer.
Junjie Hu 0003, Chenyou Fan, Mete Ozay, Yuan Gao 0024, Tin Lun Lam
IEEE Trans. Intell. Transp. Syst.3
2024 HOP to the Next Tasks and Domains for Continual Learning in NLP
abstract
Continual Learning (CL) aims to learn a sequence of problems (i.e., tasks and domains) by transferring knowledge acquired on previous problems, whilst avoiding forgetting of past ones. Different from previous approaches which focused on CL for one NLP task or domain in a specific use-case, in this paper, we address a more general CL setting to learn from a sequence of problems in a unique framework. Our method, HOP, permits to hop across tasks and domains by addressing the CL problem along three directions: (i) we employ a set of adapters to generalize a large pre-trained model to unseen problems, (ii) we compute high-order moments over the distribution of embedded representations to distinguish independent and correlated statistics across different tasks and domains, (iii) we process this enriched information with auxiliary heads specialized for each end problem. Extensive experimental campaign on 4 NLP applications, 5 benchmarks and 2 CL setups demonstrates the effectiveness of our HOP.
Umberto Michieli, Mete Ozay
AAAI2
2024 LoRA-Guard: Parameter-Efficient Guardrail Adaptation for Content Moderation of Large Language Models
abstract
Guardrails have emerged as comprehensive method of content moderation for large language models (LLMs), complementing safety alignment from fine-tuning.However, existing model-based guardrails are too memory intensive for use on resource-constrained computational devices such as mobile phones, an increasing number of which are running LLM-based applications locally.We introduce LoRA-Guard, a parameter-efficient guardrail adaptation method that relies on knowledge sharing between LLMs and guardrail models.LoRA-Guard extracts language features from the LLMs and adapts them for the content moderation task using low-rank adapters in a dual-path design which prevents any performance degradation on the generative task.We show that LoRA-Guard outperforms existing guardrail approaches while using 100-1000x fewer guardrail parameters, enabling on-device content moderation.
Hayder Elesedy, Pedro M. Esperança, Silviu Oprea, Mete Ozay
EMNLP4
2024 Deep Neural Network Models Trained with a Fixed Random Classifier Transfer Better Across Domains
abstract
The recently discovered Neural collapse (NC) phenomenon states that the last-layer weights of Deep Neural Networks (DNN), converge to the so-called Equiangular Tight Frame (ETF) simplex, at the terminal phase of their training. This ETF geometry is equivalent to vanishing within-class variability of the last layer activations. Inspired by NC properties, we explore in this paper the transferability of DNN models trained with their last layer weight fixed according to ETF. This enforces class separation by eliminating class covariance information, effectively providing implicit regularization. We show that DNN models trained with such a fixed classifier significantly improve transfer performance, particularly on out-of-domain datasets. On a broad range of fine-grained image classification datasets, our approach outperforms i) baseline methods that do not perform any covariance regularization (up to 22%), as well as ii) methods that explicitly whiten covariance of activations throughout training (up to 19%). Our findings suggest that DNNs trained with fixed ETF classifiers offer a powerful mechanism for improving transfer learning across domains.
Hafiz Tiomoko Ali, Umberto Michieli, Ji Joong Moon, Mete Ozay
ICASSP5
2024 FFT-Based Selection and Optimization of Statistics for Robust Recognition of Severely Corrupted Images
abstract
Improving model robustness in case of corrupted images is among the key challenges to enable robust vision systems on smart devices, such as robotic agents. Particularly, robust test-time performance is imperative for most of the applications. This paper presents a novel approach to improve robustness of any classification model, especially on severely corrupted images. Our method (FROST) employs high-frequency features to detect input image corruption type, and select layer-wise feature normalization statistics. FROST provides the state-of-the-art results for different models and datasets, outperforming competitors on ImageNet-C by up to 37.1% relative gain, improving baseline of 40.9% mCE on severe corruptions.
Elena Camuffo, Umberto Michieli, Ji Joong Moon, Mete Ozay
ICASSP5
2024 Object-Conditioned Bag of Instances for Few-Shot Personalized Instance Recognition
abstract
Nowadays, users demand for increased personalization of vision systems to localize and identify personal instances of objects (e.g., my dog rather than dog) from a few-shot dataset only. Despite outstanding results of deep networks on classical label-abundant benchmarks (e.g., those of the latest YOLOv8 model for standard object detection), they struggle to maintain within-class variability to represent different instances rather than object categories only. We construct an Object-conditioned Bag of Instances (OBoI) based on multiorder statistics of extracted features, where generic object detection models are extended to search and identify personal instances from the OBoI’s metric space, without need for backpropagation. By relying on multi-order statistics, OBoI achieves consistent superior accuracy in distinguishing different instances. In the results, we achieve 77.1% personal object recognition accuracy in case of 18 personal instances, showing about 12% relative gain over the state of the art.
Umberto Michieli, Ji Joong Moon, Mete Ozay
ICASSP4
2024 Joint Embedding Learning and Latent Subspace Probing for Cross-Domain Few-Shot Keyword Spotting
abstract
Probing classifiers (PCs) have been employed as one of the notable approaches for exploring properties of deep neural network (DNN) models in various tasks such as natural language processing (NLP) and computer vision (CV). In this approach, a PC is trained to predict some property (e.g. semantics in NLP) from features learned by pre-trained models. If the PC performs well, then it is concluded that the model learned information relevant for the property. Recent studies have demonstrated various methodological limitations of this approach, such as classifier selection, and focused on designing PCs. Instead, we are studying improvement of limitations of feature spaces used to train PCs for cross-domain few-shot keyword spotting (FKWS). To this end, we propose a Latent Subspace Probing (LSP) framework for jointly learning embedding of features onto latent subspaces of models together with PCs employed on the subspaces. We apply our LSP for FKWS aiming to identify and improve the keyword discrimination property of models. We quantitatively and qualitatively explore discrimination properties of learned embeddings of keywords with PCs. The results show that our proposed methods outperform the conventional and state-of-the-art methods up to 8% and 10% in 1 and 5 shot recognition tasks.
Mete Ozay
ICASSP1
2024 Texture and Normal Map Estimation for 3D Face Reconstruction
abstract
3D Morphable Models (3DMMs) can effectively capture human face shape and texture by exploiting face statistics computed on 3D scanned faces. However, the capacity of these models imposes limitations on shape details and texture expressiveness. Consequently, they provide low-detailed texture and normal map estimations. We propose a novel framework designed to estimate high-quality texture and normal maps in the UV space of a single input image. This framework comprises encoder, normalization and decoder models. To this end, we aim to learn rich and robust representations that remain unaffected under changing pose, expression and illumination by enhancing estimation details. Experimental results demonstrate that our framework outperforms the baseline significantly in terms of realism and identity preservation.
Savas Özkan, Mete Ozay, Tom Robinson
ICASSP2
2024 Towards Better Control Of Latent Spaces For Face Editing
abstract
Generative models can synthesize diverse and photo-realistic images that have demonstrated remarkable success in computer vision. Notably, Generative Adversarial Networks (GANs) trained for faces (i.e., StyleGAN2) can be considered as a powerful image generation pipeline. However, the entangled space of features learned by GANs restricts precise control of modifying the content of generated images. This paper introduces a framework designed to enhance the control for face editing in image generation by disentanglement of feature spaces of GANs (a.k.a. GAN spaces). Our framework aims to enable better control for modification of face concepts such as pose, expression, and illumination. For this purpose, the framework first learns multiple latent spaces for parameterization of face concepts learned by 3D Morphable face models. Then, it employs a Identity-Conditioned Attention Mechanism (ICAM) to decouple face identity representations from the parameterized face concepts. Moreover, we adapt variational autoencoders to model the hierarchical structure of GAN features by incorporating transformer networks for end-to-end optimization of model parameters. Our results show that our ICAM achieves state-of-the-art identity preservation and editing precision accuracy on benchmark datasets with improved training time and memory usage.
Savas Özkan, Mete Ozay
ICIP2
2024 Cross-Architecture Auxiliary Feature Space Translation for Efficient Few-Shot Personalized Object Detection
abstract
Recent years have seen object detection robotic systems deployed in several personal devices (e.g., home robots and appliances). This has highlighted a challenge in their design, i.e., they cannot efficiently update their knowledge to distinguish between general classes and user-specific instances (e.g., a dog vs. user’s dog). We refer to this challenging task as Instance-level Personalized Object Detection (IPOD). The personalization task requires many samples for model tuning and optimization in a centralized server, raising privacy concerns. An alternative is provided by approaches based on recent large-scale Foundation Models, but their compute costs preclude on-device applications. In our work we tackle both problems at the same time, designing a Few-Shot IPOD strategy called AuXFT. We introduce a conditional coarse-to-fine few-shot learner to refine the coarse predictions made by an efficient object detector, showing that using an off-the-shelf model leads to poor personalization due to neural collapse. Therefore, we introduce a Translator block that generates an auxiliary feature space where features generated by a self-supervised model (e.g., DINOv2) are distilled without impacting the performance of the detector. We validate AuXFT on three publicly available datasets and one in-house benchmark designed for the IPOD task, achieving remarkable gains in all considered scenarios with excellent time-complexity trade-off: AuXFT reaches a performance of 80% its upper bound at just 32% of the inference time, 13% of VRAM and 19% of the model size.
Francesco Barbato, Umberto Michieli, Ji Joong Moon, Pietro Zanuttigh, Mete Ozay
IROS5
2024 Enhanced Model Robustness to Input Corruptions by Per-corruption Adaptation of Normalization Statistics
abstract
Developing a reliable vision system is a fundamental challenge for robotic technologies (e.g., indoor service robots and outdoor autonomous robots) which can ensure reliable navigation even in challenging environments such as adverse weather conditions (e.g., fog, rain), poor lighting conditions (e.g., over/under exposure), or sensor degradation (e.g., blurring, noise), and can guarantee high performance in safety-critical functions. Current solutions proposed to improve model robustness usually rely on generic data augmentation techniques or employ costly test-time adaptation methods. In addition, most approaches focus on addressing a single vision task (typically, image recognition) utilising synthetic data. In this paper, we introduce Per-corruption Adaptation of Normalization statistics (PAN) to enhance the model robustness of vision systems. Our approach entails three key components: (i) a corruption type identification module, (ii) dynamic adjustment of normalization layer statistics based on identified corruption type, and (iii) real-time update of these statistics according to input data. PAN can integrate seamlessly with any convolutional model for enhanced accuracy in several robot vision tasks. In our experiments, PAN obtains robust performance improvement on challenging real-world corrupted image datasets (e.g., OpenLoris, ExDark, ACDC), where most of the current solutions tend to fail. Moreover, PAN outperforms the baseline models by 20-30% on synthetic benchmarks in object recognition tasks.
Elena Camuffo, Umberto Michieli, Simone Milani, Ji Joong Moon, Mete Ozay
IROS5
2024 Swiss DINO: Efficient and Versatile Vision Framework for On-device Personal Object Search
abstract
In this paper, we address a recent trend in robotic home appliances to include vision systems on personal devices, capable of personalizing the appliances on the fly. In particular, we formulate and address an important technical task of personal object search, which involves localization and identification of personal items of interest on images captured by robotic appliances, with each item referenced only by a few annotated images. The task is crucial for robotic home appliances and mobile systems, which need to process personal visual scenes or to operate with particular personal objects (e.g., for grasping or navigation). In practice, personal object search presents two main technical challenges. First, a robot vision system needs to be able to distinguish between many fine-grained classes, in the presence of occlusions and clutter. Second, the strict resource requirements for the on-device system restrict the usage of most state-of-the-art methods for few-shot learning and often prevent on-device adaptation. In this work, we propose Swiss DINO: a simple yet effective framework for one-shot personal object search based on the recent DINOv2 transformer model, which was shown to have strong zero-shot generalization properties. Swiss DINO handles challenging on-device personalized scene understanding requirements and does not require any adaptation training. We show significant improvement (up to 55%) in segmentation and recognition accuracy compared to the common lightweight solutions, and significant footprint reduction of backbone inference time (up to 100×) and GPU consumption (up to 10×) compared to the heavy transformer-based solutions1.
Kirill Paramonov, Jia-Xing Zhong, Umberto Michieli, Ji Joong Moon, Mete Ozay
IROS5
2024 A Modular System for Enhanced Robustness of Multimedia Understanding Networks via Deep Parametric Estimation
abstract
Performance degradation caused by corrupted multimedia samples is a critical challenge for machine learning models. Previously, three groups of approaches have been proposed to tackle this issue: i) enhancer and denoiser modules to improve the quality of the noisy data, ii) data augmentation approaches, and iii) domain adaptation strategies. All have drawbacks limiting applicability; the first requires paired clean-corrupted data for training and has an high computational cost, while the others can only be used on the same task they were trained on. In this paper, we propose SyMPIE to solve these shortcomings, designing a small, modular, and efficient system to enhance input data for robust downstream multimedia understanding with minimal computational cost. Our SyMPIE is pre-trained on an upstream task/network that should not match the downstream ones and does not need paired clean-corrupted samples. Our key insight is that most input corruptions found in real-world tasks can be modeled through global operations on color channels of images or spatial filters with small kernels. We validate our approach on multiple datasets and tasks, such as image classification (on ImageNetC, ImageNetC-Bar, VizWiz, and a newly proposed mixed corruption benchmark named ImageNetC-mixed) and semantic segmentation (on Cityscapes, ACDC, and DarkZurich) with consistent improvements of about 5% relative accuracy gain across the board1.
Francesco Barbato, Umberto Michieli, Mehmet Kerim Yucel, Pietro Zanuttigh, Mete Ozay
MMSys5
2024 Dense depth distillation with out-of-distribution simulated images
Junjie Hu 0003, Chenyou Fan, Mete Ozay, Hualie Jiang, Tin Lun Lam
Knowl. Based Syst.3
2023 Conceptual and Hierarchical Latent Space Decomposition for Face Editing
abstract
Generative Adversarial Networks (GANs) can produce photo-realistic results using an unconditional image-generation pipeline. However, the images generated by GANs (e.g., StyleGAN) are entangled in feature spaces, which makes it difficult to interpret and control the contents of images. In this paper, we present an encoder-decoder model that decomposes the entangled GAN space into a conceptual and hierarchical latent space in a self-supervised manner. The outputs of 3D morphable face models are leveraged to independently control image synthesis parameters like pose, expression, and illumination. For this purpose, a novel latent space decomposition pipeline is introduced using transformer networks and generative models. Later, this new space is used to optimize a transformer-based GAN space controller for face editing. In this work, a StyleGAN2 model for faces is utilized. Since our method manipulates only GAN features, the photo-realism of Style-GAN2 is fully preserved. The results demonstrate that our method qualitatively and quantitatively outperforms baselines in terms of identity preservation and editing precision.
Savas Özkan, Mete Ozay, Tom Robinson
ICCV2
2023 A Model for Every User and Budget: Label-Free and Personalized Mixed-Precision Quantization
Edward Fish, Umberto Michieli, Mete Ozay
INTERSPEECH3
2023 On-Device Speaker Anonymization of Acoustic Embeddings for ASR based on Flexible Location Gradient Reversal Layer
Md Asif Jalal, Pablo Peso Parada, Jisi Zhang, Mete Ozay, Karthikeyan Saravanan, Myoungji Han, Jungin Lee, Seokyeong Jung
INTERSPEECH4
2023 Online Continual Learning in Keyword Spotting for Low-Resource Devices via Pooling High-Order Temporal Statistics
Umberto Michieli, Pablo Peso Parada, Mete Ozay
INTERSPEECH3
2023 Online Continual Learning for Robust Indoor Object Recognition
abstract
Vision systems mounted on home robots need to interact with unseen classes in changing environments. Robots have limited computational resources, labelled data and storage capability. These requirements pose some unique challenges: models should adapt without forgetting past knowledge in a data- and parameter-efficient way. We characterize the problem as few-shot (FS) online continual learning (OCL), where robotic agents learn from a non-repeated stream of few-shot data updating only a few model parameters. Additionally, such models experience variable conditions at test time, where objects may appear in different poses (e.g., horizontal or vertical) and environments (e.g., day or night). To improve robustness of CL agents, we propose RobOCLe, which; 1) constructs an enriched feature space computing high order statistical moments from the embedded features of samples; and 2) computes similarity between high order statistics of the samples on the enriched feature space, and predicts their class labels. We evaluate robustness of CL models to train/test augmentations in various cases. We show that different moments allow RobOCLe to capture different properties of deformations, providing higher robustness with no decrease of inference speed.
Umberto Michieli, Mete Ozay
IROS2
2023 LRA&LDRA: Rethinking Residual Predictions for Efficient Shadow Detection and Removal
abstract
The majority of the state-of-the-art shadow removal models (SRMs) reconstruct whole input images, where their capacity is needlessly spent on reconstructing non-shadow regions. SRMs that predict residuals remedy this up to a degree, but fall short of providing an accurate and flexible solution. In this paper, we rethink residual predictions and propose Learnable Residual Attention (LRA) and Learnable Dense Reconstruction Attention (LDRA) modules, which operate over the input and the output of SRMs. These modules guide an SRM to concentrate on shadow region reconstruction, and limit reconstruction of non-shadow regions. The modules improve shadow removal (up to 20%) and detection accuracy across various backbones, and even improve the accuracy of other removal methods (up to 10%). In addition, the modules have minimal overhead (+<1MB memory) and are implemented in a few lines of code. Furthermore, to combat the challenge of training SRMs with small datasets, we present a synthetic dataset generation pipeline. Using our pipeline, we create a dataset called PITSA, which has 10 times more unique shadow-free images than the largest benchmark dataset. Pre-training models on the PITSA significantly improves shadow removal (+2 MAE on shadow regions) and detection accuracy of multiple methods. Our results show that LRA&LDRA, when plugged into a lightweight architecture pre-trained on the PITSA, outperform state-of-the-art shadow removal (+0.7 all-region MAE) and detection (+0.1 BER) methods on the benchmark ISTD and SRD datasets, despite running faster (+5%) and consuming less memory (×150).
Mehmet Kerim Yucel, Valia Dimaridou, Bruno Manganelli, Mete Ozay, Anastasios Drosou, Albert Saà-Garriga
WACV4
2023 Federated Learning via Attentive Margin of Semantic Feature Representations
abstract
Federated learning (FL) in Internet of Things (IoT) systems enables distributed model training using a large corpus of decentralized training data dispersed among multiple IoT clients. In this distributed setting, system and statistical heterogeneity, in the form of highly imbalanced, and nonindependent and identically distributed (non-i.i.d.) data stored on multiple devices, are likely to hinder model training. Existing methods aggregate models disregarding the internal representations being learned, which yet play an essential role to solve the pursued task, especially in the case of deep learning modules. To leverage feature representations in an FL framework, we introduce a method, called FedMargin, which computes client deviations using margins over feature representations learned on distributed data, and applies them to drive federated optimization via an attention mechanism. Local and aggregated margins are jointly exploited, taking into account local representation shift and representation discrepancy with the global model. In addition, we propose three methods to analyse statistical properties of feature representations learned in FL, in order to elucidate the relationship between accuracy, margins, and feature discrepancy of FL models. In experimental analyses, FedMargin demonstrates state-of-the-art accuracy and convergence rate across image classification and semantic segmentation benchmarks by enabling maximum margin training of FL models. Moreover, FedMargin reduces the uncertainty of predictions of FL models compared to the baseline. In this work, we also evaluate FL models on dense prediction tasks, such as semantic segmentation, proving the versatility of the proposed approach.
Umberto Michieli, Marco Toldo, Mete Ozay
IEEE Internet Things J.3
2023 Task guided representation learning using compositional models for zero-shot domain adaptation
Shuang Liu 0002, Mete Ozay
Neural Networks2
2023 Deep Depth Completion From Extremely Sparse Data: A Survey
abstract
Depth completion aims at predicting dense pixel-wise depth from an extremely sparse map captured from a depth sensor, e.g., LiDARs. It plays an essential role in various applications such as autonomous driving, 3D reconstruction, augmented reality, and robot navigation. Recent successes on the task have been demonstrated and dominated by deep learning based solutions. In this article, for the first time, we provide a comprehensive literature review that helps readers better grasp the research trends and clearly understand the current advances. We investigate the related studies from the design aspects of network architectures, loss functions, benchmark datasets, and learning strategies with a proposal of a novel taxonomy that categorizes existing methods. Besides, we present a quantitative comparison of model performance on three widely used benchmarks, including indoor and outdoor datasets. Finally, we discuss the challenges of prior works and provide readers with some insights for future research directions.
Junjie Hu 0003, Chenyu Bao, Mete Ozay, Chenyou Fan, Qing Gao 0002, Honghai Liu 0001, Tin Lun Lam
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 Bring Evanescent Representations to Life in Lifelong Class Incremental Learning
abstract
In Class Incremental Learning (CIL), a classification model is progressively trained at each incremental step on an evolving dataset of new classes, while at the same time, it is required to preserve knowledge of all the classes ob-served so far. Prototypical representations can be lever-aged to model feature distribution for the past data and in-ject information of former classes in later incremental steps without resorting to stored exemplars. However, if not up-dated, those representations become increasingly outdated as the incremental learning progresses with new classes. To address the aforementioned problems, we propose a frame-work which aims to (i) model the semantic drift by learning the relationship between representations of past and novel classes among incremental steps, and (ii) estimate the feature drift, defined as the evolution of the represen-tations learned by models at each incremental step. Se-mantic and feature drifts are then jointly exploited to infer up-to-date representations of past classes (evanescent rep-resentations), and thereby infuse past knowledge into incre-mental training. We experimentally evaluate our framework achieving exemplar-free SotA results on multiple bench-marks. In the ablation study, we investigate nontrivial relationships between evanescent representations and models.
Marco Toldo, Mete Ozay
CVPR2
2022 Feature Kernel Distillation
Bobby He, Mete Ozay
ICLR2
2022 Exploring the Gap between Collapsed & Whitened Features in Self-Supervised Learning
abstract
Avoiding feature collapse, when a Neural Network (NN) encoder maps all inputs to a constant vector, is a shared implicit desideratum of various methodological advances in self-supervised learning (SSL). To that end, whitened features have been proposed as an explicit objective to ensure uncollapsed features \cite{zbontar2021barlow,ermolov2021whitening,hua2021feature,bardes2022vicreg}. We identify power law behaviour in eigenvalue decay, parameterised by exponent $\beta{\geq}0$, as a spectrum that bridges between the collapsed & whitened feature extremes. We provide theoretical & empirical evidence highlighting the factors in SSL, like projection layers & regularisation strength, that influence eigenvalue decay rate, & demonstrate that the degree of feature whitening affects generalisation, particularly in label scarce regimes. We use our insights to motivate a novel method, PMP (PostMan-Pat), which efficiently post-processes a pretrained encoder to enforce eigenvalue decay rate with power law exponent $\beta$, & find that PostMan-Pat delivers improved label efficiency and transferability across a range of SSL methods and encoder architectures.
Bobby He, Mete Ozay
ICML2
2022 FedNST: Federated Noisy Student Training for Automatic Speech Recognition
Haaris Mehmood, Agnieszka Dobrowolska, Karthikeyan Saravanan, Mete Ozay
INTERSPEECH4
2022 pMCT: Patched Multi-Condition Training for Robust Speech Recognition
abstract
We propose a novel Patched Multi-Condition Training (pMCT) method for robust Automatic Speech Recognition (ASR).pMCT employs Multi-condition Audio Modification and Patching (MAMP) via mixing patches of the same utterance extracted from clean and distorted speech.Training using patchmodified signals improves robustness of models in noisy reverberant scenarios.Our proposed pMCT is evaluated on the LibriSpeech dataset showing improvement over using vanilla Multi-Condition Training (MCT).For analyses on robust ASR, we employed pMCT on the VOiCES dataset which is a noisy reverberant dataset created using utterances from LibriSpeech.In the analyses, pMCT achieves 23.1% relative WER reduction compared to the MCT.
Pablo Peso Parada, Agnieszka Dobrowolska, Karthikeyan Saravanan, Mete Ozay
INTERSPEECH4
2022 Exploring Targeted and Stealthy False Data Injection Attacks via Adversarial Machine Learning
abstract
State estimation methods used in cyber–physical systems (CPSs), such as smart grid, are vulnerable to false data injection attacks (FDIAs). Although substantial deep learning methods have been proposed to detect such attacks, deep neural networks (DNNs) are highly susceptible to adversarial attacks, which modify input of DNNs with unnoticeable but malicious perturbations. This article proposes a method to explore targeted and stealthy FDIAs via adversarial machine learning. We pose FDIAs as sparse optimization problems to achieve initial attack objectives and remain stealthy during attacks. We propose a parallel optimization algorithm to efficiently solve the problems and explore additional sparse-state attacks. The experimental results show that for IEEE 14-bus and 118-bus systems, the success rate of two-state sparse attacks with small-scale targets is as high as 80%. In addition, the attack success rate can continue to increase as the number of attack states increases. The proposed attacks demonstrate that attackers can implement attacks that can bypass both bad data detectors and neural network detectors while keeping the initial attack objectives unchanged, which is a critical and urgent security threat in CPS.
Jiwei Tian, Buhong Wang, Zhen Wang 0020, Mete Ozay
IEEE Internet Things J.6
2022 Joint Adversarial Example and False Data Injection Attacks for State Estimation in Power Systems
abstract
Although state estimation using a bad data detector (BDD) is a key procedure employed in power systems, the detector is vulnerable to false data injection attacks (FDIAs). Substantial deep learning methods have been proposed to detect such attacks. However, deep neural networks are susceptible to adversarial attacks or adversarial examples, where slight changes in inputs may lead to sharp changes in the corresponding outputs in even well-trained networks. This article introduces the joint adversarial example and FDIAs (AFDIAs) to explore various attack scenarios for state estimation in power systems. Considering that perturbations added directly to measurements are likely to be detected by BDDs, our proposed method of adding perturbations to state variables can guarantee that the attack is stealthy to BDDs. Then, malicious data that are stealthy to both BDDs and deep learning-based detectors can be generated. Theoretical and experimental results show that our proposed state-perturbation-based AFDIA method (S-AFDIA) can carry out attacks stealthy to both conventional BDDs and deep learning-based detectors, while our proposed measurement-perturbation-based adversarial FDIA method (M-AFDIA) succeeds if only deep learning-based detectors are used. The comparative experiments show that our proposed methods provide better performance than state-of-the-art methods. Besides, the ultimate effect of attacks can also be optimized using the proposed joint attack methods.
Jiwei Tian, Buhong Wang, Zhen Wang 0020, Kunrui Cao, Mete Ozay
IEEE Trans. Cybern.6
2021 A Mixed Quantization Network for Efficient Mobile Inverse Tone Mapping
Juan Borrego-Carazo, Mete Ozay, Frederik Laboyrie, Paul Wisbey
BMVC2
2021 A New Approach to Design Symmetry Invariant Neural Networks
abstract
We investigate a new method to design$G$-invariant neural networks that approximate functions invariant to the action of a given permutation subgroup$G$of the symmetric group on input data. The key element of the new network architecture is a$G$-invariant transformation module, which produces a$G$-invariant latent representation of the input data. This latent representation is then processed with a multi-layer perceptron in the network. We prove the universality of the new architecture, discuss its properties and highlight its computational and memory efficiency. Theoretical considerations are supported by numerical experiments involving different network configurations, which demonstrate the efficiency and strong generalization properties of the new approach to design symmetry invariant neural networks, in comparison to other$G$-invariant neural architectures.
Piotr Kicki, Piotr Skrzypczynski, Mete Ozay
IJCNN3
2021 Learning from experience for rapid generation of local car maneuvers
Piotr Kicki, Tomasz Gawron, Krzysztof Cwian, Mete Ozay, Piotr Skrzypczynski
Eng. Appl. Artif. Intell.4
2019 Improving Head Pose Estimation with a Combined Loss and Bounding Box Margin Adjustment
abstract
We address a problem of estimating pose of a person's head from its RGB image. The employment of CNNs for the problem has contributed to significant improvement in accuracy in recent works. However, we show that the following two methods, despite their simplicity, can attain further improvement: (i) proper adjustment of the margin of bounding box of a detected face, and (ii) choice of loss functions. We show that the integration of these two methods achieve the new state-of-the-art on standard benchmark datasets for in-the-wild head pose estimation. The Tensorflow implementation of our work is available at https://github.com/MingzhenShao/HeadPose.
Mingzhen Shao, Zhun Sun, Mete Ozay, Takayuki Okatani
FG3
2019 A Generative Model of Underwater Images for Active Landmark Detection and Docking
abstract
Underwater active landmarks (UALs) are widely used for short-range underwater navigation in underwater robotics tasks. Detection of UALs is challenging due to large variance of underwater illumination, water quality and change of camera viewpoint. Moreover, improvement of detection accuracy relies upon statistical diversity of images used to train detection models. We propose a generative adversarial network, called Tank-to-field GAN (T2FGAN), to learn generative models of underwater images, and use the learned models for data augmentation to improve detection accuracy. To this end, first a T2FGAN is trained using images of UALs captured in a tank. Then, the learned model of the T2FGAN is used to generate images of UALs according to different water quality, illumination, pose and landmark configurations (WIPCs). In experimental analyses, we first explore statistical properties of images of UALs generated by T2FGAN under various WIPCs for active landmark detection. Then, we use the generated images for training detection algorithms. Experimental results show that training detection algorithms using the generated images can improve detection accuracy. In field experiments, underwater docking tasks are successfully performed in a lake by employing detection models trained on datasets generated by T2FGAN.
Shuang Liu 0002, Mete Ozay, Takayuki Okatani
IROS2
2019 Fine-grained Optimization of Deep Neural Networks
abstract
In recent studies, several asymptotic upper bounds on generalization errors on deep neural networks (DNNs) are theoretically derived. These bounds are functions of several norms of weights of the DNNs, such as the Frobenius and spectral norms, and they are computed for weights grouped according to either input and output channels of the DNNs. In this work, we conjecture that if we can impose multiple constraints on weights of DNNs to upper bound the norms of the weights, and train the DNNs with these weights, then we can attain empirical generalization errors closer to the derived theoretical bounds, and improve accuracy of the DNNs. To this end, we pose two problems. First, we aim to obtain weights whose different norms are all upper bounded by a constant number. To achieve these bounds, we propose a two-stage renormalization procedure; (i) normalization of weights according to different norms used in the bounds, and (ii) reparameterization of the normalized weights to set a constant and finite upper bound of their norms. In the second problem, we consider training DNNs with these renormalized weights. To this end, we first propose a strategy to construct joint spaces (manifolds) of weights according to different constraints in DNNs. Next, we propose a fine-grained SGD algorithm (FG-SGD) for optimization on the weight manifolds to train DNNs with assurance of convergence to minima. Experimental analyses show that image classification accuracy of baseline DNNs can be boosted using FG-SGD on collections of manifolds identified by multiple constraints.
Mete Ozay
NeurIPS1
2019 Revisiting Single Image Depth Estimation: Toward Higher Resolution Maps With Accurate Object Boundaries
abstract
This paper considers the problem of single image depth estimation. The employment of convolutional neural networks (CNNs) has recently brought about significant advancements in the research of this problem. However, most existing methods suffer from loss of spatial resolution in the estimated depth maps; a typical symptom is distorted and blurry reconstruction of object boundaries. In this paper, toward more accurate estimation with a focus on depth maps with higher spatial resolution, we propose two improvements to existing approaches. One is about the strategy of fusing features extracted at different scales, for which we propose an improved network architecture consisting of four modules: an encoder, decoder, multi-scale feature fusion module, and refinement module. The other is about loss functions for measuring inference errors used in training. We show that three loss terms, which measure errors in depth, gradients and surface normals, respectively, contribute to improvement of accuracy in an complementary fashion. Experimental results show that these two improvements enable to attain higher accuracy than the current state-of-the-arts, which is given by finer resolution reconstruction, for example, with small objects and object boundaries.
Junjie Hu 0003, Mete Ozay, Yan Zhang 0055, Takayuki Okatani
WACV2
2018 Training CNNs With Normalized Kernels
abstract
Several methods of normalizing convolution kernels have been proposed in the literature to train convolutional neural networks (CNNs), and have shown some success. However, our understanding of these methods has lagged behind their success in application; there are a lot of open questions, such as why a certain type of kernel normalization is effective and what type of normalization should be employed for each (e.g., higher or lower) layer of a CNN. As the first step towards answering these questions, we propose a framework that enables us to use a variety of kernel normalization methods at any layer of a CNN. A naive integration of kernel normalization with a general optimization method, such as SGD, often entails instability while updating parameters. Thus, existing methods employ ad-hoc procedures to empirically assure convergence. In this study, we pose estimation of convolution kernels under normalization constraints as constraint-free optimization on kernel submanifolds that are identified by the employed constraints. Note that naive application of the established optimization methods for matrix manifolds to the aforementioned problems is not feasible because of the hierarchical nature of CNNs. To this end, we propose an algorithm for optimization on kernel manifolds in CNNs by appropriate scaling of the space of kernels based on structure of CNNs and statistics of data. We theoretically prove that the proposed algorithm has assurance of almost sure convergence to a solution at single minimum. Our experimental results show that the proposed method can successfully train popular CNN models using several different types of kernel normalization methods. Moreover, they show that the proposed method improves classification performance of baseline CNNs, and provides state-of-the-art performance for major image classification benchmarks.
Mete Ozay, Takayuki Okatani
AAAI1
2018 Feature Quantization for Defending Against Distortion of Images
abstract
In this work, we address the problem of improving robustness of convolutional neural networks (CNNs) to image distortion. We argue that higher moment statistics of feature distributions can be shifted due to image distortion, and the shift leads to performance decrease and cannot be reduced by ordinary normalization methods as observed in our experimental analyses. In order to mitigate this effect, we propose an approach base on feature quantization. To be specific, we propose to employ three different types of additional non-linearity in CNNs: i) a floor function with scalable resolution, ii) a power function with learnable exponents, and iii) a power function with data-dependent exponents. In the experiments, we observe that CNNs which employ the proposed methods obtain better performance in both generalization performance and robustness for various distortion types for large scale benchmark datasets. For instance, a ResNet-50 model equipped with the proposed method (+HPOW) obtains 6.95%, 5.26% and 5.61% better accuracy on the ILSVRC-12 classification tasks using images distorted with motion blur, salt and pepper and mixed distortions.
Zhun Sun, Mete Ozay, Yan Zhang 0055, Xing Liu 0010, Takayuki Okatani
CVPR2
2018 Exploiting the Potential of Standard Convolutional Autoencoders for Image Restoration by Evolutionary Search
abstract
Researchers have applied deep neural networks to image restoration tasks, in which they proposed various network architectures, loss functions, and training methods. In particular, adversarial training, which is employed in recent studies, seems to be a key ingredient to success. In this paper, we show that simple convolutional autoencoders (CAEs) built upon only standard network components, i.e., convolutional layers and skip connections, can outperform the state-of-the-art methods which employ adversarial training and sophisticated loss functions. The secret is to search for good architectures using an evolutionary algorithm. All we did was to train the optimized CAEs by minimizing the l2 loss between reconstructed images and their ground truths using the ADAM optimizer. Our experimental results show that this approach achieves 27.8 dB peak signal to noise ratio (PSNR) on the CelebA dataset and 33.3 dB on the SVHN dataset, compared to 22.8 dB and 19.0 dB provided by the former state-of-the-art methods, respectively.
Masanori Suganuma, Mete Ozay, Takayuki Okatani
ICML2
2018 Deep Structured Energy-Based Image Inpainting
abstract
In this paper, we propose a structured image inpainting method employing an energy based model. In order to learn structural relationship between patterns observed in images and missing regions of the images, we employ an energy-based structured prediction method. The structural relationship is learned by minimizing an energy function which is defined by a simple convolutional neural network. The experimental results on various benchmark datasets show that our proposed method significantly outperforms the state-of-the-art methods which use Generative Adversarial Networks (GANs). We obtained 497.35 mean squared error (MSE) on the Olivetti face dataset compared to 833.0 MSE provided by the state-of-the-art method. Moreover, we obtained 28.4 dB peak signal to noise ratio (PSNR) on the SVHN dataset and 23.53 dB on the CelebA dataset, compared to 22.3 dB and 21.3 dB, provided by the state-of-the-art methods, respectively. The code is publicly available1.
Fazil Altinel, Mete Ozay, Takayuki Okatani
ICPR2
2017 Truncating Wide Networks Using Binary Tree Architectures
abstract
In this paper, we propose a binary tree architecture to truncate architecture of wide networks by reducing the width of the networks. More precisely, in the proposed architecture, the width is incrementally reduced from lower layers to higher layers in order to increase the expressive capacity of networks with a less increase on parameter size. Also, in order to ease the gradient vanishing problem, features obtained at different layers are concatenated to form the output of our architecture. By employing the proposed architecture on a baseline wide network, we can construct and train a new network with same depth but considerably less number of parameters. In our experimental analyses, we observe that the proposed architecture enables us to obtain better parameter size and accuracy trade-off compared to baseline networks using various benchmark image classification datasets. The results show that our model can decrease the classification error of a baseline from 20:43% to 19:22% on Cifar-100 using only 28% of parameters that the baseline has. Code is available at https://github.com/ZhangVision/bitnet.
Yan Zhang 0055, Mete Ozay, Shuohao Li, Takayuki Okatani
ICCV2
2016 Design of Kernels in Convolutional Neural Networks for Image Classification
Zhun Sun, Mete Ozay, Takayuki Okatani
ECCV (7)2
2016 Integrating deep features for material recognition
abstract
This paper considers the problem of material recognition. Motivated by observation of close interconnections between material and object recognition, we study how to select and integrate multiple features obtained by different models of Convolutional Neural Networks (CNNs) trained in a transfer learning setting. To be specific, we first compute activations of features using representations on images to select a set of samples which are best represented by the features. Then, we measure uncertainty of the features by computing entropy of class distributions for each sample set. Finally, we compute contribution of each feature to representation of classes for feature selection and integration. Experimental results show that the proposed method achieves state-of-the-art performance on two benchmark datasets for material recognition. Additionally, we introduce a new material dataset, named EFMD, which extends Flickr Material Database (FMD). By the employment of the EFMD for transfer learning, we achieve 84.0% ± 1.8% accuracy on the FMD dataset, which is close to the reported human performance 84.9%.
Yan Zhang 0055, Mete Ozay, Xing Liu 0010, Takayuki Okatani
ICPR2
2016 Machine Learning Methods for Attack Detection in the Smart Grid
abstract
Attack detection problems in the smart grid are posed as statistical learning problems for different attack scenarios in which the measurements are observed in batch or online settings. In this approach, machine learning algorithms are used to classify measurements as being either secure or attacked. An attack detection framework is provided to exploit any available prior knowledge about the system and surmount constraints arising from the sparse structure of the problem in the proposed approach. Well-known batch and online learning algorithms (supervised and semisupervised) are employed with decision- and feature-level fusion to model the attack detection problem. The relationships between statistical and geometric properties of attack vectors employed in the attack scenarios and learning algorithms are analyzed to detect unobservable attacks using statistical learning methods. The proposed algorithms are examined on various IEEE test systems. Experimental analyses show that machine learning algorithms can detect attacks with performances higher than attack detection algorithms that employ state vector estimation methods in the proposed attack detection framework.
Mete Ozay, Inaki Esnaola, Fatos T. Yarman-Vural, Sanjeev R. Kulkarni, H. Vincent Poor
IEEE Trans. Neural Networks Learn. Syst.1
2015 Compositional Hierarchical Representation of Shape Manifolds for Classification of Non-manifold Shapes
abstract
We address the problem of statistical learning of shape models which are invariant to translation, rotation and scale in compositional hierarchies when data spaces of measurements and shape spaces are not topological manifolds. In practice, this problem is observed while modeling shapes having multiple disconnected components, e.g. partially occluded shapes in cluttered scenes. We resolve the aforementioned problem by first reformulating the relationship between data and shape spaces considering the interaction between Receptive Fields (RFs) and Shape Manifolds (SMs) in a compositional hierarchical shape vocabulary. Then, we suggest a method to model the topological structure of the SMs for statistical learning of the geometric transformations of the shapes that are defined by group actions on the SMs. For this purpose, we design a disjoint union topology using an indexing mechanism for the formation of shape models on SMs in the vocabulary, recursively. We represent the topological relationship between shape components using graphs, which are aggregated to construct a hierarchical graph structure for the shape vocabulary. To this end, we introduce a framework to implement the indexing mechanisms for the employment of the vocabulary for structural shape classification. The proposed approach is used to construct invariant shape representations. Results on benchmark shape classification outperform state-of-the-art methods.
Mete Ozay, Ümit Rusen Aktas, Jeremy L. Wyatt, Ales Leonardis
ICCV1
2014 A Graph Theoretic Approach for Object Shape Representation in Compositional Hierarchies Using a Hybrid Generative-Descriptive Model
Ümit Rusen Aktas, Mete Ozay, Ales Leonardis, Jeremy L. Wyatt
ECCV (3)2
2014 Modeling the Brain Connectivity for Pattern Analysis
abstract
An information theoretic approach is proposed to estimate the degree of connectivity for each voxel with its neighboring voxels. The neighborhood system is defined by spatial and functional connectivity metrics. Then, a local mesh of variable size is formed around each voxel using spatial or functional neighborhood. The mesh arc weights, called Mesh Arc Descriptors (MAD), are estimated by a linear regression model fitted to the voxel intensity values of the functional Magnetic Resonance Images (fMRI). Finally, the error term of the linear regression equation is used to estimate the mesh size for a voxel by optimizing Akaike's information Criterion, Bayesian Information Criterion and Rissanen's Minimum Description Length. fMRI measurements are obtained during a memory encoding and retrieval experiment performed on a subject who is exposed to the stimuli from 10 semantic categories. For each sample, a k-NN classifier is trained using the Mesh Arc Descriptors (MAD) having the variable mesh sizes. The classification performances reflect that the suggested variable-size Mesh Arc Descriptors represents the mental states better than the classical multi-voxel pattern representation. Moreover, we observe that the degree of connectivities in the brain greatly varies for each voxel.
Itir Önal, Emre Aksan, Burak Velioglu, Orhan Firat, Mete Ozay, Ilke Öztekin, Fatos T. Yarman-Vural
ICPR5
2014 Semi-supervised Segmentation Fusion of Multi-spectral and Aerial Images
abstract
A Semi-supervised Segmentation Fusion algorithm is proposed using consensus and distributed learning. The aim of Unsupervised Segmentation Fusion (USF) is to achieve a consensus among different segmentation outputs obtained from different segmentation algorithms by computing an approximate solution to the NP problem with less computational complexity. Semi-supervision is incorporated in USF using a new algorithm called Semi-supervised Segmentation Fusion (SSSF). In SSSF, side information about the co-occurrence of pixels in the same or different segments is formulated as the constraints of a convex optimization problem. The results of the experiments employed on artificial and real-world benchmark multi-spectral and aerial images show that the proposed algorithms perform better than the individual state-of-the art segmentation algorithms.
Mete Ozay
ICPR1
2014 A hierarchical approach for joint multi-view object pose estimation and categorization
abstract
We propose a joint object pose estimation and categorization approach which extracts information about object poses and categories from the object parts and compositions constructed at different layers of a hierarchical object representation algorithm, namely Learned Hierarchy of Parts (LHOP) [7]. In the proposed approach, we first employ the LHOP to learn hierarchical part libraries which represent entity parts and compositions across different object categories and views. Then, we extract statistical and geometric features from the part realizations of the objects in the images in order to represent the information about object pose and category at each different layer of the hierarchy. Unlike the traditional approaches which consider specific layers of the hierarchies in order to extract information to perform specific tasks, we combine the information extracted at different layers to solve a joint object pose estimation and categorization problem using distributed optimization algorithms. We examine the proposed generative-discriminative learning approach and the algorithms on two benchmark 2-D multi-view image datasets. The proposed approach and the algorithms outperform state-of-the-art classification, regression and feature extraction algorithms. In addition, the experimental results shed light on the relationship between object categorization, pose estimation and the part realizations observed at different layers of the hierarchy.
Mete Ozay, Krzysztof Walas, Ales Leonardis
ICRA1
2013 An information theoretic approach to classify cognitive states using fMRI
abstract
In this study, an information theoretic approach is proposed to model brain connectivity during a cognitive processing task, measured by functional Magnetic Resonance Imaging (fMRI). For this purpose, a local mesh of varying size is formed around each voxel. The arc weights of each mesh are estimated using a linear regression model by minimizing the squared error. Then, the optimal mesh size for each sample, that represents the information distribution in the brain, is estimated by minimizing various information criteria which employ the mean square error of linear regression model. The estimated mesh size shows the degree of locality or degree of connectivity of the voxels for the underlying cognitive process. The samples are generated during an fMRI experiment employing item recognition (IR) and judgment of recency (JOR) tasks. For each sample, estimated arc weights of the local mesh with optimal size are used to classify whether it belongs to IR or JOR tasks. Results indicate that the suggested connectivity model with optimal mesh size for each sample represent the information distribution in the brain better than the state-of-the art methods.
Itir Önal, Mete Ozay, Orhan Firat, Ilke Öztekin, Fatos T. Yarman-Vural
BIBE2
2013 Mesh learning for object classification using fMRI measurements
abstract
Machine learning algorithms have been widely used as reliable methods for modeling and classifying cognitive processes using functional Magnetic Resonance Imaging (fMRI) data. In this study, we aim to classify fMRI measurements recorded during an object recognition experiment. Previous studies focus on Multi Voxel Pattern Analysis (MVPA) which feeds a set of active voxels in a concatenated vector form to a machine learning algorithm to train and classify the cognitive processes. In most of the MVPA methods, after an image preprocessing step, the voxel intensity values are fed to a classifier to train and recognize the underlying cognitive process. Sometimes, the fMRI data is further processed for de-noising or feature selection where techniques, such as Generalized Linear Model (GLM), Independent Component Analysis (ICA) or Principal Component Analysis are employed. Although these techniques are proved to be useful in MVPA, they do not model the spatial connectivity among the voxels. In this study, we attempt to represent the local relations among the voxel intensity values by forming a mesh network around each voxel to model the relationship of a voxel and its surroundings. The degree of connectivity of a voxel to its surroundings is represented by the arc weights of each mesh. The arc weights, which are estimated by a linear regression model, are fed to a classifier to discriminate the brain states during an object recognition task. This approach, called Mesh Learning, provides a powerful tool to analyze various cognitive states using fMRI data. Compared to traditional studies which focus either merely on multi-voxel pattern vectors or their reduced-dimension versions, the suggested Mesh Learning provides a better representation of object recognition task. Various machine learning algorithms are tested to compare the suggested Mesh Learning to the state-of-the art MVPA techniques. The performance of the Mesh Learning is shown to be higher than that of the available MVPA techniques.
Omer Ekmekci, Orhan Firat, Mete Ozay, Ilke Öztekin, Fatos T. Yarman-Vural, Uygar Öztekin
ICIP3
2013 Fusion of image segmentation algorithms using consensus clustering
abstract
A new segmentation fusion method is proposed that ensembles the output of several segmentation algorithms applied on a remotely sensed image. The candidate segmentation sets are processed to achieve a consensus segmentation using a stochastic optimization algorithm based on the Filtered Stochastic BOEM (Best One Element Move) method. For this purpose, Filtered Stochastic BOEM is reformulated as a segmentation fusion problem by designing a new distance learning approach. The proposed algorithm also embeds the computation of the optimum number of clusters into the segmentation fusion problem.
Mete Ozay, Fatos T. Yarman-Vural, Sanjeev R. Kulkarni, H. Vincent Poor
ICIP1
2013 Sparse Attack Construction and State Estimation in the Smart Grid: Centralized and Distributed Models
abstract
New methods that exploit sparse structures arising in smart grid networks are proposed for the state estimation problem when data injection attacks are present. First, construction strategies for unobservable sparse data injection attacks on power grids are proposed for an attacker with access to all network information and nodes. Specifically, novel formulations for the optimization problem that provide a flexible design of the trade-off between performance and false alarm are proposed. In addition, the centralized case is extended to a distributed framework for both the estimation and attack problems. Different distributed scenarios are proposed depending on assumptions that lead to the spreading of the resources, network nodes and players. Consequently, for each of the presented frameworks a corresponding optimization problem is introduced jointly with an algorithm to solve it. The validity of the presented procedures in real settings is studied through extensive simulations in the IEEE test systems.
Mete Ozay, Inaki Esnaola, Fatos T. Yarman-Vural, Sanjeev R. Kulkarni, H. Vincent Poor
IEEE J. Sel. Areas Commun.1
2012 Automatic building detection with feature space fusion using ensemble learning
abstract
This paper proposes a novel approach to building detection problem in satellite images. The proposed method employs a two layer hierarchical classification mechanism for ensemble learning. After an initial segmentation, each segment is classified by N different classifiers using different features at the first layer. The class membership values of the segments, which are obtained from different base layer classifiers, are ensembled to form a new fusion space, which forms a linearly separable simplex. Then, this simplex is partitioned by a linear classifier at the meta layer. The paper presents the performance results of the proposed model and comparisons with the state of the art classifiers.
Çaglar Senaras, Baris Yüksel, Mete Ozay, Fatos T. Yarman-Vural
IGARSS3
2009 A new decision fusion technique for image classification
abstract
In this study, we introduce a new image classification technique using decision fusion. The proposed technique, called Meta-Fuzzified Yield Value (Meta-FYV), is based on two-layer Stacked Generalization (SG) architecture. At the base-layer, the system, receives a set of feature vectors of various dimensions and dynamical ranges and outputs hypotheses through fuzzy transformations. Then, the hypotheses created by the base layer transformations are concatenated for building a regression equation at meta-layer. Experimental evidence indicates that the Meta-FYV is superior compared to one of the most successful Fuzzy SG methods, introduced by Akbas.
Mete Ozay, Fatos T. Yarman-Vural
ICIP1