VLDB 2026 Research / reviewers in the wild / expert
Adrian G. Bors
dblp:94/1481
· DBLP profile ↗
141ranked-venue papers
23as first author
58since 2021 · last 2026
0000-0001-7838-0021ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 99 · 18 first-author · 36 since 2021Artificial intelligence and machine learning · 65 · 6 first-author · 38 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 3 · 1 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSecurity and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Adaptive and Expandable Mixture Model for Continual LearningabstractContinuous learning constitutes a fundamental capability of artificial intelligence systems, enabling them to incrementally assimilate novel information without succumbing to catastrophic forgetting. Recent research has leveraged Pre-Trained Models (PTMs) to enhance continual learning efficacy. Nevertheless, prevailing methodologies typically depend on a singular pre-trained backbone and freeze all pre-trained parameters to mitigate network forgetting, thereby constraining adaptability to emerging tasks. In this study, we introduce an innovative PTM-based framework featuring a Dual-Representation Backbone Architecture (DRBA), which integrates both invariant and evolved representation networks to concurrently capture static and dynamic features. Building upon DRBA, we propose an Adaptive and Expandable Mixture Model (AEMM) that incrementally incorporates new expert modules with minimal parameter overhead to accommodate the learning of each novel task. To further augment adaptability, we develop a Dynamic Adaptive Representation Fusion Mechanism (DARFM) that processes outputs from both representation networks and autonomously generates data-driven adaptive weights, optimizing the contribution of each representation. This mechanism yields an adaptive, semantically enriched composite representation, thereby maximizing positive knowledge transfer. Additionally, we propose a Dynamic Knowledge Calibration Mechanism (DKCM), comprising prediction and representation calibration processes, to ensure consistency in both predictions and feature representations. This approach achieves a balance between stability and plasticity, even when learning complex datasets. Empirical evaluations substantiate that the proposed approach attains state-of-the-art performance. Fei Ye 0004, YongCheng Zhong, Qihe Liu, Adrian G. Bors, Jingling Sun, Jinyu Guo, Shijie Zhou 0002 |
AAAI | 4 |
| 2026 | Continual Learning across multiple domains via a Dynamic Expandable and Mergeable Model
Fei Ye 0004, Ruilong Yu, Qihe Liu, Adrian G. Bors, Jingling Sun, Rongyao Hu, Shijie Zhou 0002 |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | Task-free continual generative modelling via dynamic teacher-student framework
Fei Ye 0004, Adrian G. Bors |
Expert Syst. Appl. | 2 |
| 2026 | Online task-free continual learning via Expansible Vision Transformer
Fei Ye 0004, Adrian G. Bors |
Pattern Recognit. | 2 |
| 2025 | Lifelong Scalable Generative System via Online Maximum Mean DiscrepancyabstractDiffusion-based models have been recently shown to be high-quality data generators. However, their performance severely degrades when training on non-stationary changing data distributions in an online manner, due to the catastrophic forgetting. In this paper, we propose enabling the diffusion model with a novel Dynamic Expansion Memory Unit (DEMU) methodology that adaptively creates new memory buffers, to be added to a memory system, in order to preserve information deemed critical for training the model. Having a selective memory unit is essential for training diffusion networks, which are expensive to train, especially when deployed in resource-constrained environments. A Maximum Mean Discrepancy (MMD) based expansion mechanism, that evaluates probabilistic distances between each of the previously defined memory buffers and the newly given data, and uses them as expansion signals, is employed for ensuring the diversity of information learning. We propose a new model expansion mechanism to automatically add new diffusion models as experts in a mixture system, which enhances the multi-domain image generation performance. Also a novel memory compaction approach is proposed to automatically remove statistically overlapping memory units, through a graph relationship evaluation, preventing the limitless expansion of DEMU. Comprehensive results show that the proposed approach performs better than the state-of-the-art. Fei Ye 0004, Adrian G. Bors |
AAAI | 2 |
| 2025 | Continual Unsupervised Generative Modelling via Online Optimal TransportabstractLately, deep generative models have achieved excellent results after learning pre-defined and static data distribution. Meanwhile, their performance on continual learning suffers from degeneration, caused by catastrophic forgetting. In this paper, we study the unsupervised generative modelling in a more realistic continual learning scenario, where class and task information are absent during both training and inference learning phases. To implement this goal, the proposed memory approach consists of a temporary memory system, which stores data examples while a dynamic expansion memory system would gradually preserve those samples that are crucial for long-term memorization. A novel memory expansion mechanism is then proposed, by employing optimal transport distances between the statistics of memorized samples and each newly seen datum. This paper proposes the Sinkhorn-based Dual Dynamic Memory (SDDM) method, by considering Sinkhorn distance as an optimal transport measure, for evaluating the significance of the data to be stored in the memory buffer. The Sinkhorn transport algorithm leads to preserving a diversity of samples within a compact memory capacity. The memory buffering approach does not interact with the model's training process and can be optimized independently in both supervised and unsupervised learning without any modifications. Moreover, we also propose a novel dynamic model expansion mechanism to automatically increase the model's capacity whenever necessary, which can deal with infinite data streams and further improve the model's performance. Experimental results show that the proposed approach achieves state-of-the-art performance in both supervised and unsupervised learning. Fei Ye 0004, Adrian G. Bors |
AAAI | 2 |
| 2025 | Dynamic Expansion Diffusion Learning for Lifelong Generative ModellingabstractThe diffusion model has lately been shown to achieve remarkable performances through its ability of generating high quality images. However, current diffusion model studies consider only learning from a single data distribution, resulting in catastrophic forgetting when attempting to learn new data. In this paper, we explore a more realistic learning scenario where training data is continuously acquired. We propose the Dynamic Expansion Diffusion Model (DEDM) for addressing catastrophic forgetting and data distribution shifts under Online Task-Free Continual Learning (OTFCL) paradigm. New diffusion components are added to a mixture model following the evaluation of a criterion which compares the probabilistic representation of the new data with the existing knowledge of the DEDM model. In addition, to maintain an optimal architecture, we propose a component discovery approach that ensures the diversity of knowledge while minimizing the total number of parameters in the DEDM. Furthermore, we show how the proposed DEDM can be implemented as a teacher module in a unified framework for representation learning. In this approach, knowledge distillation is proposed for training a student module aiming to compress the teacher's knowledge into the latent space of the student. Fei Ye 0004, Adrian G. Bors |
AAAI | 2 |
| 2025 | Online Task-Free Continual Learning via Dynamic Expansionable Memory DistributionabstractRecent continuous learning (CL) research primarily addresses catastrophic forgetting within a straightforward learning framework where class and task information are predefined. However, in the Task-Free Continual Learning (TFCL), representing a more realistic and challenging CL scenarios, such information is typically absent. In this paper, we address the online TFCL by introducing an innovative memory management approach, by incorporating a dynamic memory system for storing selected data representatives from evolving distributions while a dynamically expandable memory system enables the retention of essential long-term knowledge. The proposed dynamic expandable memory system manages a series of memory distributions, each designed to represent the information from a distinct data category. A new memory expansion mechanism that assesses the proximity between incoming samples and existing memory distributions is proposed for evaluating when to add new memory distributions into the system. Additionally, a novel memory distribution augmentation technique is proposed for selectively gathering suitable samples for each memory distribution, enhancing the statistical robustness over time. To prevent memory saturation before the training phase, we introduce a memory distribution reduction strategy that automatically eliminates overlapping memory distributions, ensuring adequate capacity for accommodating new information in subsequent learning episodes. We conduct a series of experiments demonstrating that our proposed approach attains state-of-the-art performance in both supervised and unsupervised learning contexts. The source code is available at https://github.com/dtuzi123/DEMD. Fei Ye 0004, Adrian G. Bors |
CVPR | 2 |
| 2025 | Localised Frequency Latent Domain Watermarking of DDIM Generated ImagesabstractStable Diffusion models, relying on iterative generative latent diffusion processes, have recently achieved remarkable results in producing realistic and diverse images. Meanwhile, the widespread application of generative models raised significant concerns about the origins of image content or the infringement of intellectual property rights. Consequently, a method for identifying AI generated images and/or other information about their origins is imperatively necessary. To address these requirements we propose to embed watermarks during one of the diffusion iterative steps of the DDIM. Such watermarks are required to be recoverable while also robust to possible changes to the generated watermarked images. The watermarks are embedded in the localized regions of the latent space frequencies. The binary watermarks are detected from the generated watermarked images by means of a CNN watermark detector. The robustness of the CNN watermark detector is improved through training by considering various distortions to the watermarked images. Qiran Lai, Adrian G. Bors |
ICASSP | 2 |
| 2025 | Online Continual Learning via Dynamic Expandable Recursive ModelabstractThe continual learning (CL) of novel concepts from new environments represents a popular and important topic aiming to manage catastrophic forgetting. Research studies have developed dynamic expansion models to deal with network forgetting in CL. Existing CL models usually explore the full capacity of activating parameters and representations while ignoring the previously learned representations when learning new tasks. In this paper, we propose a novel dynamic expansion model that incrementally accumulates and incorporates all previously learned representations into defining new experts to add to a mixture of experts in a recursive manner, aiming to reuse previously learned parameters and features to promote future task learning. We define a graph structure having each expert as a component node. We then propose a novel expandable expert graph attention mechanism that dynamically optimizes the graph when learning new tasks, maximizing the positive knowledge transfer. In addition, we propose a novel expert cooperation mechanism to promote the cooperation between all previous experts and with the currently updated expert. Furthermore, we propose a novel memory optimization approach, which encourages each expert to capture and learn completely different information, further improving performance. We provide the results of a series of experiments demonstrating that the proposed approach outperforms the state-of-the-art performance in CL. Fei Ye 0004, Adrian G. Bors |
ACM Multimedia | 2 |
| 2025 | Learning Multi-Source and Robust Representations for Continual LearningabstractPlasticity and stability denote the ability to assimilate new tasks while preserving previously acquired knowledge, representing two important concepts in continual learning. Recent research addresses stability by leveraging pre-trained models to provide informative representations, yet the efficacy of these methods is highly reliant on the choice of the pre-trained backbone, which may not yield optimal plasticity. This paper addresses this limitation by introducing a streamlined and potent framework that orchestrates multiple different pre-trained backbones to derive semantically rich multi-source representations. We propose an innovative Multi-Scale Interaction and Dynamic Fusion (MSIDF) technique to process and selectively capture the most relevant parts of multi-source features through a series of learnable attention modules, thereby helping to learn better decision boundaries to boost performance. Furthermore, we introduce a novel Multi-Level Representation Optimization (MLRO) strategy to adaptively refine the representation networks, offering adaptive representations that enhance plasticity. To mitigate over-regularization issues, we propose a novel Adaptive Regularization Optimization (ARO) method to manage and optimize a switch vector that selectively governs the updating process of each representation layer, which promotes the new task learning. The proposed MLRO and ARO approaches are collectively optimized within a unified optimization framework to achieve an optimal trade-off between plasticity and stability. Our extensive experimental evaluations reveal that the proposed framework attains state-of-the-art performance. The source code of our algorithm is available at https://github.com/CL-Coder236/LMSRR. Fei Ye 0004, YongCheng Zhong, Qihe Liu, Adrian G. Bors, Jingling Sun, Rongyao Hu, Shijie Zhou 0002 |
NeurIPS | 4 |
| 2025 | Dynamic Siamese Expansion Framework for Improving Robustness in Online Continual LearningabstractContinual learning requires the model to continually capture novel information without forgetting prior knowledge. Nonetheless, existing studies predominantly address the catastrophic forgetting, often neglecting enhancements in model robustness. Consequently, these methodologies fall short in real-time applications, such as autonomous driving, where data samples frequently exhibit noise due to environmental and lighting variations, thereby impairing model efficacy and causing safety issues. In this paper, we address robustness in continual learning systems by introducing an innovative approach, the Dynamic Siamese Expansion Framework (DSEF) that employs a Siamese backbone architecture, comprising static and dynamic components, to facilitate the learning of both global and local representations over time. Specifically, the proposed framework dynamically generates a lightweight expert for each novel task, leveraging the Siamese backbone to enable rapid adaptation. A novel Robust Dynamic Representation Optimization (RDRO) approach is proposed to incrementally update the dynamic backbone by maintaining all previously acquired representations and prediction patterns of historical experts, thereby fostering new task learning without inducing detrimental knowledge transfer. Additionally, we propose a novel Robust Feature Fusion (RFF) approach to incrementally amalgamate robust representations from all historical experts into the expert construction process. A novel mutual information-based technique is employed to derive adaptive weights for feature fusion by assessing the knowledge relevance between historical experts and the new task, thus maximizing positive knowledge transfer effects. A comprehensive experimental evaluation, benchmarking our approach against established baselines, demonstrates that our method achieves state-of-the-art performance even under adversarial attacks. Fei Ye 0004, Qihe Liu, Junlin Chen, Adrian G. Bors, Jingling Sun, Rongyao Hu, Shijie Zhou 0002 |
NeurIPS | 5 |
| 2025 | Learning Expandable and Adaptable Representations for Continual LearningabstractExtant studies predominantly address catastrophic forgetting within a simplified continual learning paradigm, typically confined to a singular data domain. Conversely, real-world applications frequently encompass multiple, evolving data domains, wherein models often struggle to retain many critical past information, thereby leading to performance degradation. This paper addresses this complex scenario by introducing a novel dynamic expansion approach called Learning Expandable and Adaptable Representations (LEAR). This framework orchestrates a collaborative backbone structure, comprising global and local backbones, designed to capture both general and task-specific representations. Leveraging this collaborative backbone, the proposed framework dynamically create a lightweight expert to delineate decision boundaries for each novel task, thereby facilitating the prediction process. To enhance new task learning, we introduce a novel Mutual Information-Based Prediction Alignment approach, which incrementally optimizes the global backbone via a mutual information metric, ensuring consistency in the prediction patterns of historical experts throughout the optimization phase. To mitigate network forgetting, we propose a Kullback–Leibler (KL) Divergence-Based Feature Alignment approach, which employs a probabilistic distance measure to prevent significant shifts in critical local representations. Furthermore, we introduce a novel Hilbert-Schmidt Independence Criterion (HSIC)-Based Collaborative Optimization approach, which encourages the local and global backbones to capture distinct semantic information in a collaborative manner, thereby mitigating information redundancy and enhancing model performance. Moreover, to accelerate new task learning, we propose a novel Expert Selection Mechanism that automatically identifies the most relevant expert based on data characteristics. This selected expert is then utilized to initialize a new expert, thereby fostering positive knowledge transfer. This approach also enables expert selection during the testing phase without requring any task information. Empirical results demonstrate that the proposed framework achieves state-of-the-art performance. Ruilong Yu, Mingyan Liu, Fei Ye 0004, Adrian G. Bors, Rongyao Hu, Jingling Sun, Shijie Zhou 0002 |
NeurIPS | 4 |
| 2025 | Evolving Ensemble Model based on Hilbert Schmidt Independence Criterion for task-free continual learning
Fei Ye 0004, Adrian G. Bors |
Neurocomputing | 2 |
| 2025 | Online task-free continual learning via discrepancy mechanism
Fei Ye 0004, Adrian G. Bors |
Knowl. Based Syst. | 2 |
| 2025 | Continual Unsupervised Generative ModelingabstractVariational Autoencoders (VAEs), can achieve remarkable results in single tasks, by learning data representations, image generation, or image-to-image translation among others. However, VAEs suffer from loss of information when aiming to continuously learn a sequence of different data domains. This is caused by the catastrophic forgetting, which affects all machine learning methods. This paper addresses the problem of catastrophic forgetting by developing a new theoretical framework which derives an upper bound to the negative sample log-likelihood when continuously learning sequences of tasks. These theoretical derivations provide new insights into the forgetting behavior of learning models, showing that their optimal performance is achieved when a dynamic mixture expansion model adds new components whenever learning new tasks. In our approach we optimize the model size by introducing the Dynamic Expansion Graph Model (DEGM) that dynamically builds a graph structure promoting the positive knowledge transfer when learning new tasks. In addition, we propose a Dynamic Expansion Graph Adaptive Mechanism (DEGAM) that generates adaptive weights to regulate the graph structure, further improving the positive knowledge transfer effectiveness. Experimental results show that the proposed methodology performs better than other baselines in continual learning. Fei Ye 0004, Adrian G. Bors |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Training a Dynamic Growing Mixture Model for Lifelong LearningabstractLifelong learning (LLL) defines a training paradigm that aims to continuously acquire and capture new concepts from a sequence of tasks without forgetting. Recently, dynamic expansion models (DEMs) have been proposed to address catastrophic forgetting under the LLL paradigm. However, the efficiency of DEMs lacks a thorough explanation based on theoretical analysis. In this article, we develop a new theoretical framework that interprets the forgetting process of the DEM as increasing the statistical discrepancy distance between the distribution of the probabilistic representation of the new data and the previously learned knowledge. The theoretical analysis shows that adding new components to a mixture model represents a trade-off between model complexity and its performance. Inspired by the theoretical analysis, we introduce a new DEM, called the growing mixture model (GMM), where generative data components are added according to the novelty of the incoming task information compared to what is already known. A new component selection mechanism considering the model's already acquired knowledge is employed for updating new DEM's components, promoting efficient future task learning. We also train a compact student model with samples drawn through the generative mechanisms of the GMM, aiming to accumulate cross-domain representations over time. By employing the student model, we can significantly reduce the number of parameters and make quick inferences during the testing phase. Fei Ye 0004, Adrian G. Bors |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | BotCF: Improving the Social Bot Detection Performance By Focusing on the Community FeaturesabstractVarious malicious activities performed by social bots have brought a crisis of trust to online social networks. Existing social bot detection methods often overlook the significance of community structure features and effective fusion strategies for multimodal features. To counter these limitations, we propose BotCF, a novel social bot detection method that incorporates community features and utilizes cross-attention fusion for multimodal features. In BotCF, we extract community features using a community division algorithm based on deep autoencoder-like non-negative matrix factorization. These features capture the social interactions and relationships within the network, providing valuable insights for bot detection. Furthermore, we employ cross-attention fusion to integrate the features of the account’s semantic content, properties, and community structure. This fusion strategy allows the model to learn the interdependencies between different modalities, leading to a more comprehensive representation of each account. Extensive experiments conducted on three publicly available benchmark datasets (Twibot20, Twibot22, and Cresci-2015) demonstrate the effectiveness of BotCF. Compared to state-of-the-art social bot detection models, BotCF achieves significant improvements in accuracy, with an average increase of 1.86%, 1.67%, and 0.47% on the respective datasets. The detection accuracy is boosted to 86.53%, 81.33%, and 98.21%, respectively. Feng Liu 0045, Zhenyu Li 0004, Chunfang Yang, Daofu Gong, Fenlin Liu, Rui Ma 0011, Adrian G. Bors |
IEEE Trans. Netw. Serv. Manag. | 7 |
| 2024 | Task-Free Dynamic Sparse Vision Transformer for Continual LearningabstractVision Transformers (ViTs) represent self-attention-based network backbones shown to be efficient in many individual tasks, but which have not been explored in Task-Free Continual Learning (TFCL) so far. Most existing ViT-based approaches for Continual Learning (CL) are relying on task information. In this study, we explore the advantages of the ViT in a more challenging CL scenario where the task boundaries are unavailable during training. To address this learning paradigm, we propose the Task-Free Dynamic Sparse Vision Transformer (TFDSViT), which can dynamically build new sparse experts, where each expert leverages sparsity to allocate the model's capacity for capturing different information categories over time. To avoid forgetting and ensure efficiency in reusing the previously learned knowledge in subsequent learning, we propose a new dynamic dual attention mechanism consisting of the Sparse Attention (SA') and Knowledge Transfer Attention (KTA) modules. The SA' refrains from updating some previously learned attention blocks for preserving prior knowledge. The KTA uses and regulates the information flow of all previously learned experts for learning new patterns. The proposed dual attention mechanism can simultaneously relieve forgetting and promote knowledge transfer for a dynamic expansion model in a task-free manner. We also propose an energy-based dynamic expansion mechanism using the energy as a measure of novelty for the incoming samples which provides appropriate expansion signals leading to a compact network architecture for TFDSViT. Extensive empirical studies demonstrate the effectiveness of TFDSViT. The code and supplementary material (SM) are available at https://github.com/dtuzi123/TFDSViT. Fei Ye 0004, Adrian G. Bors |
AAAI | 2 |
| 2024 | Task-Free Continual Generation and Representation Learning via Dynamic Expansionable Memory ClusterabstractHuman brains can continually acquire and learn new skills and knowledge over time from a dynamically changing environment without forgetting previously learnt information. Such a capacity can selectively transfer some important and recently seen information to the persistent knowledge regions of the brain. Inspired by this intuition, we propose a new memory-based approach for image reconstruction and generation in continual learning, consisting of a temporary and evolving memory, with two different storage strategies, corresponding to the temporary and permanent memorisation. The temporary memory aims to preserve up-to-date information while the evolving memory can dynamically increase its capacity in order to preserve permanent knowledge information. This is achieved by the proposed memory expansion mechanism that selectively transfers those data samples deemed as important from the temporary memory to new clusters defined within the evolved memory according to an information novelty criterion. Such a mechanism promotes the knowledge diversity among clusters in the evolved memory, resulting in capturing more diverse information by using a compact memory capacity. Furthermore, we propose a two-step optimization strategy for training a Variational Autoencoder (VAE) to implement generation and representation learning tasks, which updates the generator and inference models separately using two optimisation paths. This approach leads to a better trade-off between generation and reconstruction performance. We show empirically and theoretically that the proposed approach can learn meaningful latent representations while generating diverse images from different domains. The source code and supplementary material (SM) are available at https://github.com/dtuzi123/DEMC. Fei Ye 0004, Adrian G. Bors |
AAAI | 2 |
| 2024 | Online Task-Free Continual Generative and Discriminative Learning via Dynamic Cluster MemoryabstractOnline Task-Free Continual Learning (OTFCL) aims to learn novel concepts from streaming data without accessing task information. Most memory-based approaches used in OTFCL are not suitable for unsupervised learning because they require accessing supervised signals to implement their sample selection mechanisms. In this study, we address this issue by proposing a novel memory management approach, namely the Dynamic Cluster Memory (DCM), which builds new memory clusters to capture distribution shifts over time without accessing any supervised signals. DCM introduces a novel memory expansion mechanism based on the knowledge discrepancy criterion, which evaluates the novelty of the incoming data as the signal for the memory expansion, ensuring a compact memory capacity. We also propose a new sample selection approach that automatically stores incoming data samples with similar semantic information in the same memory cluster, while also facilitating the knowledge diversity among memory clusters. Further-more, a novel memory pruning approach is proposed to automatically remove overlapping memory clusters through a graph relation evaluation, ensuring a fixed memory capacity while maintaining the diversity among the samples stored in the memory. The proposed DCM is model-free, plug-and-play, and can be used in both supervised and unsupervised learning without modifications. Empirical results on OTFCL experiments show that the proposed DCM outperforms the state-of-the-art while requiring fewer data samples to be stored. The source code is available at https://github.com/dtuzi123/DCM. Fei Ye 0004, Adrian G. Bors |
CVPR | 2 |
| 2024 | Temporal Transformer Encoder for Video Class Incremental LearningabstractCurrent video classification approaches suffer from catastrophic forgetting when they are retrained on new databases. Continual learning aims to enable a classification system with learning from a succession of tasks without forgetting. In this paper we propose to use a transformer-based video class incremental learning model. During a succession of learning steps, at each training time, the transformer is used to extract characteristic spatio-temporal features from videos corresponding to a set of classes. When new video classification tasks become available, we train new classifier modules with the transformer-extracted features, gradually building a mixture model. The proposed methodology enables continual class learning in videos without being required to consider the learning of an initial set of classes, leading to low computation and memory requirements. The proposed model is evaluated on standard action recognition datasets including UCF101 and HMDB51, which are split into sets of classes, to be learnt sequentially. Our proposed method significantly outperforms the baselines on all datasets. Nattapong Kurpukdee, Adrian G. Bors |
ICIP | 2 |
| 2024 | Self-Supervised Adversarial Variational LearningabstractA natural approach for representation learning is to combine the inference mechanisms of VAEs and the generation capabilities of GANs, within a new model, namely VAEGAN. Most existing VAEGAN models would jointly train the generator and inference modules, which has limitations when learning representations generated by a pre-trained GAN model without data. In this paper, we develop a novel hybrid model, called the Self-Supervised Adversarial Variational Learning (SS-AVL) which introduces a two-step optimization procedure training separately the generator and the inference model. The primary advantage of SS-AVL over existing VAEGAN models is that SS-AVL optimizes the inference models in a self-supervised learning manner where the samples used for training the inference models are drawn from the generator distribution instead of using real samples. This can allow SS-AVL to learn representations from arbitrary GAN models without using real data. Additionally, we employ information maximization into the context of increasing the maximum likelihood, which encourages SS-AVL to learn meaningful latent representations. We perform extensive experiments to demonstrate the effectiveness of the proposed SS-AVL model. Fei Ye 0004, Adrian G. Bors |
Pattern Recognit. | 2 |
| 2024 | Lifelong Dual Generative Adversarial Nets Learning in TandemabstractContinually capturing novel concepts without forgetting is one of the most critical functions sought for in artificial intelligence systems. However, even the most advanced deep learning networks are prone to quickly forgetting previously learned knowledge after training with new data. The proposed lifelong dual generative adversarial networks (LD-GANs) consist of two generative adversarial networks (GANs), namely, a Teacher and an Assistant teaching each other in tandem while successively learning a series of tasks. A single discriminator is used to decide the realism of generated images by the dual GANs. A new training algorithm, called the lifelong self knowledge distillation (LSKD) is proposed for training the LD-GAN while learning each new task during lifelong learning (LLL). LSKD enables the transfer of knowledge from one more knowledgeable player to the other jointly with learning the information from a newly given dataset, within an adversarial playing game setting. In contrast to other LLL models, LD-GANs are memory efficient and does not require freezing any parameters after learning each given task. Furthermore, we extend the LD-GANs to being the Teacher module in a Teacher-Student network for assimilating data representations across several domains during LLL. Experimental results indicate a better performance for the proposed framework in unsupervised lifelong representation learning when compared to other methods. Fei Ye 0004, Adrian G. Bors |
IEEE Trans. Cybern. | 2 |
| 2024 | Lifelong Generative Adversarial AutoencoderabstractLifelong learning describes an ability that enables humans to continually acquire and learn new information without forgetting. This capability, common to humans and animals, has lately been identified as an essential function for an artificial intelligence system aiming to learn continuously from a stream of data during a certain period of time. However, modern neural networks suffer from degenerated performance when learning multiple domains sequentially and fail to recognize past learned tasks after being retrained. This corresponds to catastrophic forgetting and is ultimately induced by replacing the parameters associated with previously learned tasks with new values. One approach in lifelong learning is the generative replay mechanism (GRM) that trains a powerful generator as the generative replay network, implemented by a variational autoencoder (VAE) or a generative adversarial network (GAN). In this article, we study the forgetting behavior of GRM-based learning systems by developing a new theoretical framework in which the forgetting process is expressed as an increase in the model's risk during the training. Although many recent attempts have provided high-quality generative replay samples by using GANs, they are limited to mainly downstream tasks due to the lack of inference. Inspired by the theoretical analysis while aiming to address the drawbacks of existing approaches, we propose the lifelong generative adversarial autoencoder (LGAA). LGAA consists of a generative replay network and three inference models, each addressing the inference of a different type of latent variable. The experimental results show that LGAA learns novel visual concepts without forgetting and can be applied to a wide range of downstream tasks. Fei Ye 0004, Adrian G. Bors |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | Learning Dynamic Latent Spaces for Lifelong Generative ModellingabstractTask Free Continual Learning (TFCL) aims to capture novel concepts from non-stationary data streams without forgetting previously learned knowledge. Mixture models, which add new components when certain conditions are met, have shown promising results in TFCL tasks. However, such approaches do not make use of the knowledge already accumulated for positive knowledge transfer. In this paper, we develop a new model, namely the Online Recursive Variational Autoencoder (ORVAE). ORVAE utilizes the prior knowledge by selectively incorporating the newly learnt information, by adding new components, according to the knowledge already known from the past learnt data. We introduce a new attention mechanism to regularize the structural latent space in which the most important information is reused while the information that interferes with novel samples is inactivated. The proposed attention mechanism can maximize the benefit from the forward transfer for learning novel information without forgetting previously learnt knowledge. We perform several experiments which show that ORVAE achieves state-of-the-art results under TFCL. Fei Ye 0004, Adrian G. Bors |
AAAI | 2 |
| 2023 | Lifelong Compression Mixture Model via Knowledge Relationship GraphabstractTask-Free Continual Learning (TFCL) represents a challenging scenario for lifelong learning because the model, under this paradigm, does not access any task information. The Dynamic Expansion Model (DEM) has shown promising results in this scenario due to its scalability and generalisation power. However, DEM focuses only on addressing forgetting and ignores minimizing the model size, which limits its deployment in practical systems. In this work, we aim to simultaneously address network forgetting and model size optimization by developing the Lifelong Compression Mixture Model (LGMM) equipped with the Maximum Mean Discrepancy (MMD) based expansion criterion for model expansion. A diversity-aware sample selection approach is proposed to selectively store a variety of samples to promote information diversity among the components of the LGMM, which allows more knowledge to be captured with an appropriate model size. In order to avoid having multiple components with similar knowledge in the LGMM, we propose a data-free component discarding mechanism that evaluates a knowledge relation graph matrix describing the relevance between each pair of components. A greedy selection procedure is proposed to identify and remove the redundant components from the LGMM. The proposed discarding mechanism can be performed during or after the training. Experiments on different datasets show that LGMM achieves the best performance for TFCL. Fei Ye 0004, Adrian G. Bors |
AAAI | 2 |
| 2023 | Lifelong Variational Autoencoder via Online Adversarial Expansion StrategyabstractThe Variational Autoencoder (VAE) suffers from a significant loss of information when trained on a non-stationary data distribution. This loss in VAE models, called catastrophic forgetting, has not been studied theoretically before. We analyse the forgetting behaviour of a VAE in continual generative modelling by developing a new lower bound on the data likelihood, which interprets the forgetting process as an increase in the probability distance between the generator's distribution and the evolved data distribution. The proposed bound shows that a VAE-based dynamic expansion model can achieve better performance if its capacity increases appropriately considering the shift in the data distribution. Based on this analysis, we propose a novel expansion criterion that aims to preserve the information diversity among the VAE components, while ensuring that it acquires more knowledge with fewer parameters. Specifically, we implement this expansion criterion from the perspective of a multi-player game and propose the Online Adversarial Expansion Strategy (OAES), which considers all previously learned components as well as the currently updated component as multiple players in a game, while an adversary model evaluates their performance. The proposed OAES can dynamically estimate the discrepancy between each player and the adversary without accessing task information. This leads to the gradual addition of new components while ensuring the knowledge diversity among all of them. We show theoretically and empirically that the proposed extension strategy can enable a VAE model to achieve the best performance given an appropriate model size. Fei Ye 0004, Adrian G. Bors |
AAAI | 2 |
| 2023 | Continual Variational Autoencoder via Continual Generative Knowledge DistillationabstractHumans and other living beings have the ability of short and long-term memorization during their entire lifespan. However, most existing Continual Learning (CL) methods can only account for short-term information when training on infinite streams of data. In this paper, we develop a new unsupervised continual learning framework consisting of two memory systems using Variational Autoencoders (VAEs). We develop a Short-Term Memory (STM), and a parameterised scalable memory implemented by a Teacher model aiming to preserve the long-term information. To incrementally enrich the Teacher's knowledge during training, we propose the Knowledge Incremental Assimilation Mechanism (KIAM), which evaluates the knowledge similarity between the STM and the already accumulated information as signals to expand the Teacher's capacity. Then we train a VAE as a Student module and propose a new Knowledge Distillation (KD) approach that gradually transfers generative knowledge from the Teacher to the Student module. To ensure the quality and diversity of knowledge in KD, we propose a new expert pruning approach that selectively removes the Teacher's redundant parameters, associated with unnecessary experts which have learnt overlapping information with other experts. This mechanism further reduces the complexity of the Teacher's module while ensuring the diversity of knowledge for the KD procedure. We show theoretically and empirically that the proposed framework can train a statistically diversified Teacher module for continual VAE learning which is applicable to learning infinite data streams. Fei Ye 0004, Adrian G. Bors |
AAAI | 2 |
| 2023 | Few but Informative Local Hash Code Matching for Image RetrievalabstractContent-based image retrieval (CBIR) aims to search for the most similar images from an extensive database to a given query content. Existing CBIR works either represent each image with a compact global feature vector or extract a large number of highly compressed low-dimensional local features, where each contains limited information. In this research study, we propose an expressive local feature extraction pipeline and a many-to-many local feature matching method for large-scale CBIR. Unlike existing local feature methods, which tend to extract large amounts of low-dimensional local features from each image, the proposed method models characteristic feature representations for each image, aiming to employ fewer but more expressive local features. For further improving the results, an end-to-end trainable hash encoding layer is used for extracting compact but informative codes from images. The proposed many-to-many local feature matching is then directly performed on the hash feature vectors from input images, leading to new state-of-the-art performance on several benchmark datasets. Zechao Hu 0001, Adrian G. Bors |
ICASSP | 2 |
| 2023 | Enabling Large-Scale Image Search with Co-Attention MechanismabstractContent-based image retrieval (CBIR) consists of searching the most similar images to a given query. Most existing attention mechanisms for CBIR are query non-sensitive and are only based on single candidate image’s feature regardless of the actual query content. This can result in incorrect regions especially when the target object is not salient or surrounded by distractors. This paper proposes an efficient and effective query sensitive co-attention mechanism for large scale CBIR tasks. Local feature selection and clustering are employed to reduce the computation cost caused by the query sensitivity. Experimental results indicate that the proposed co-attention method can generate good co-attention maps even under challenging situations leading to a new state of the art performance on several benchmark datasets. Zechao Hu 0001, Adrian G. Bors |
ICASSP | 2 |
| 2023 | Dynamic Scalable Self-Attention Ensemble for Task-Free Continual LearningabstractContinual learning represents a challenging task for modern deep neural networks due to the catastrophic forgetting following the adaptation of network parameters to new tasks. In this paper, we address a more challenging learning paradigm called Task-Free Continual Learning (TFCL), in which the task information is missing during the training. To deal with this problem, we introduce the Dynamic Scalable Self-Attention Ensemble (DSSAE) model, which dynamically adds new Vision Transformer (ViT) based-experts to deal with the data distribution shift during the training. To avoid frequent expansions and ensure an appropriate number of experts for the model, we propose a new dynamic expansion mechanism that evaluates the novelty of incoming samples as expansion signals. Furthermore, the proposed expansion mechanism does not require knowing the task information or the class label, which can be used in a realistic learning environment. Empirical results demonstrate that the proposed DSSAE achieves state-of-the-art performance in a series of TFCL experiments. Fei Ye 0004, Adrian G. Bors |
ICASSP | 2 |
| 2023 | Compressing Cross-Domain Representation via Lifelong Knowledge DistillationabstractMost Knowledge Distillation (KD) approaches focus on the discriminative information transfer and assume that the data is provided in batches during training stages. In this paper, we address a more challenging scenario in which different tasks are presented sequentially, at different times, and the learning goal is to transfer the generative factors of visual concepts learned by a Teacher module to a compact latent space represented by a Student module. In order to achieve this, we develop a new Lifelong Knowledge Distillation (LKD) framework where we train an infinite mixture model as the Teacher which automatically increases its capacity to deal with a growing number of tasks. In order to ensure a compact architecture and to avoid forgetting, we propose to measure the relevance of the knowledge from a new task for a set of experts making up the Teacher module, guiding each expert to capture the probabilistic characteristics of several similar domains. The network architecture is expanded only when learning an entirely different task. The Student is implemented as a lightweight probabilistic generative model. The experiments show that LKD can train a compressed Student module that achieves the state of the art results with fewer parameters. Fei Ye 0004, Adrian G. Bors |
ICASSP | 2 |
| 2023 | Wasserstein Expansible Variational Autoencoder for Discriminative and Generative Continual LearningabstractTask-Free Continual Learning (TFCL) represents a challenging learning paradigm where a model is trained on the non-stationary data distributions without any knowledge of the task information, thus representing a more practical approach. Despite promising achievements by the Variational Autoencoder (VAE) mixtures in continual learning, such methods ignore the redundancy among the probabilistic representations of their components when performing model expansion, leading to mixture components learning similar tasks. This paper proposes the Wasserstein Expansible Variational Autoencoder (WEVAE), which evaluates the statistical similarity between the probabilistic representation of new data and that represented by each mixture component and then uses it for deciding when to expand the model. Such a mechanism can avoid unnecessary model expansion while ensuring the knowledge diversity among the trained components. In addition, we propose an energy-based sample selection approach that assigns high energies to novel samples and low energies to the samples which are similar to the model’s knowledge. Extensive empirical studies on both supervised and unsupervised benchmark tasks demonstrate that our model outperforms all competing methods. The code is available at https://github.com/dtuzi123/WEVAE/. Fei Ye 0004, Adrian G. Bors |
ICCV | 2 |
| 2023 | Self-Evolved Dynamic Expansion Model for Task-Free Continual LearningabstractTask-Free Continual Learning (TFCL) aims to learn new concepts from a stream of data without any task information. The Dynamic Expansion Model (DEM) has shown promising results in TFCL by dynamically expanding the model’s capacity to deal with shifts in the data distribution. However, existing approaches only consider the recognition of the input shift as the expansion signal and ignore the correlation between the newly incoming data and previously learned knowledge, resulting in adding and training unnecessary parameters. In this paper, we propose a novel and effective framework for TFCL, which dynamically expands the architecture of a DEM model through a self-assessment mechanism evaluating the diversity of knowledge among existing experts as expansion signals. This mechanism ensures learning additional underlying data distributions with a compact model structure. A novelty-aware sample selection approach is proposed to manage the memory buffer that forces the newly added expert to learn novel information from a data stream, which further promotes the diversity among experts. Moreover, we also propose to reuse previously learned representation information for learning new incoming data by using knowledge transfer in TFCL, which has not been explored before. The DEM expansion and training are regularized through a gradient updating mechanism to gradually explore the positive forward transfer, further improving the performance. Empirical results on TFCL benchmarks show that the proposed framework outperforms the state-of-the-art while using a reasonable number of parameters. The code is available at https://github.com/dtuzi123/SEDEM/. Fei Ye 0004, Adrian G. Bors |
ICCV | 2 |
| 2023 | Simultaneous Watermarking and Draco 3D Object Compression MethodabstractIn our present society, 3D objects play an increasingly important role in many different domains, especially with the development of meta environments. Many applications demand large 3D objects, which makes their compression a requirement. Meanwhile, 3D objects are also exposed to various security problems during online storage and transmission. Consequently, security aspects such as copyright and authentication are essential for 3D objects. In this paper, we propose a simultaneous 3D object watermarking and compression method based on Draco. A watermarking step is integrated within Draco, Google’s 3D compression method, which is rapidly being universally adopted as standard. The proposed method enables a large bit embedding rate, which can be used for copyright information. The proposed method is also, to the best of our knowledge, the first watermarking method for 3D objects compressed with Draco. Bianca Jansen Van Rensburg, Adrian G. Bors, William Puech, Jean-Pierre Pedeboy |
ICIP | 2 |
| 2023 | Enabling the Encoder-Empowered GAN-based Video Generators for Long Video GenerationabstractDespite the remarkable progress in the video generation field, generating videos of longer-term remains challenging due to the challenge of sustaining the temporal consistency and continuity in the resulting synthesized movement while ensuring realism. In this paper, we propose a recall mechanism for enabling an encoder-empowered short-term video generator to produce long-term videos. This mechanism connects smoothly short video clips by modeling their temporal connections. We propose the Recall Encoder-GAN3 (REncGAN3), which enables an Encoder-based Generative Adversarial Network (GAN) to connect short generated video clips into longer sequences of hundreds of frames. The recall mechanism, defined through a loss function, enables an appropriate plasticity-continuity balance in the resulting long video stream. The proposed long-term video generation method ensures the generation of several hundred frames displaying consistent movement, which is non-repetitive while the computational memory costs are similar to those of short video generation models. Adrian G. Bors |
ICIP | 2 |
| 2023 | Masked Image Residual Learning for Scaling Deeper Vision TransformersabstractDeeper Vision Transformers (ViTs) are more challenging to train. We expose a degradation problem in deeper layers of ViT when using masked image modeling (MIM) for pre-training.
To ease the training of deeper ViTs, we introduce a self-supervised learning framework called $\textbf{M}$asked $\textbf{I}$mage $\textbf{R}$esidual $\textbf{L}$earning ($\textbf{MIRL}$), which significantly alleviates the degradation problem, making scaling ViT along depth a promising direction for performance upgrade. We reformulate the pre-training objective for deeper layers of ViT as learning to recover the residual of the masked image.
We provide extensive empirical evidence showing that deeper ViTs can be effectively optimized using MIRL and easily gain accuracy from increased depth.
With the same level of computational complexity as ViT-Base and ViT-Large, we instantiate $4.5{\times}$ and $2{\times}$ deeper ViTs, dubbed ViT-S-54 and ViT-B-48.
The deeper ViT-S-54, costing $3{\times}$ less than ViT-Large, achieves performance on par with ViT-Large.
ViT-B-48 achieves 86.2\% top-1 accuracy on ImageNet.
On one hand, deeper ViTs pre-trained with MIRL exhibit excellent generalization capabilities on downstream tasks, such as object detection and semantic segmentation. On the other hand, MIRL demonstrates high pre-training efficiency. With less pre-training time, MIRL yields competitive performance compared to other approaches. Guoxi Huang, Hongtao Fu, Adrian G. Bors |
NeurIPS | 3 |
| 2023 | Co-attention enabled content-based image retrievalabstractContent-based image retrieval (CBIR) aims to provide the most similar images to a given query. Feature extraction plays an essential role in retrieval performance within a CBIR pipeline. Current CBIR studies would either uniformly extract feature information from the input image and use it directly or employ some trainable spatial weighting module which is then used for similarity comparison between pairs of query and candidate matching images. These spatial weighting modules are normally query non-sensitive and only based on the knowledge learned during the training stage. They may focus towards incorrect regions, especially when the target image is not salient or is surrounded by distractors. This paper proposes an efficient query sensitive co-attention1 mechanism for large-scale CBIR tasks. In order to reduce the extra computation cost required by the query sensitivity to the co-attention mechanism, the proposed method employs clustering of the selected local features. Experimental results indicate that the co-attention maps can provide the best retrieval results on benchmark datasets under challenging situations, such as having completely different image acquisition conditions between the query and its match image. Zechao Hu 0001, Adrian G. Bors |
Neural Networks | 2 |
| 2023 | Dynamic Self-Supervised Teacher-Student Network LearningabstractLifelong learning (LLL) represents the ability of an artificial intelligence system to learn successively a sequence of different databases. In this paper we introduce the Dynamic Self-Supervised Teacher-Student Network (D-TS), representing a more general LLL framework, where the Teacher is implemented as a dynamically expanding mixture model which automatically increases its capacity to deal with a growing number of tasks. We propose the Knowledge Discrepancy Score (KDS) criterion for measuring the relevance of the incoming information characterizing a new task when compared to the existing knowledge accumulated by the Teacher module from its previous training. The KDS ensures a light Teacher architecture while also enabling to reuse the learned knowledge whenever appropriate, accelerating the learning of given tasks. The Student module is implemented as a lightweight probabilistic generative model. We introduce a novel self-supervised learning procedure for the Student that allows to capture cross-domain latent representations from the entire knowledge accumulated by the Teacher as well as from novel data. We perform several experiments which show that D-TS can achieve the state of the art results in LLL while requiring fewer parameters than other methods. Fei Ye 0004, Adrian G. Bors |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Lifelong Mixture of Variational AutoencodersabstractIn this article, we propose an end-to-end lifelong learning mixture of experts. Each expert is implemented by a variational autoencoder (VAE). The experts in the mixture system are jointly trained by maximizing a mixture of individual component evidence lower bounds (MELBO) on the log-likelihood of the given training samples. The mixing coefficients in the mixture model control the contributions of each expert in the global representation. These are sampled from a Dirichlet distribution whose parameters are determined through nonparametric estimation during lifelong learning. The model can learn new tasks fast when these are similar to those previously learned. The proposed lifelong mixture of VAE (L-MVAE) expands its architecture with new components when learning a completely new task. After the training, our model can automatically determine the relevant expert to be used when fed with new data samples. This mechanism benefits both the memory efficiency and the required computational cost as only one expert is used during the inference. The L-MVAE inference model is able to perform interpolations in the joint latent space across the data domains associated with different tasks and is shown to be efficient for disentangled learning representation. Fei Ye 0004, Adrian G. Bors |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2023 | SSR-Net: A Spatial Structural Relation Network for Vehicle Re-identificationabstractVehicle re-identification (Re-ID) represents the task aiming to identify the same vehicle from images captured by different cameras. Recent years have seen various feature learning-based approaches merely focusing on feature representations including global features or local features to obtain more subtle details to identify highly similar vehicles. However, few such methods consider the spatial geometrical structure relationship among local regions or between the global and local regions. By contrast, in this study, we propose a Spatial Structural Relation Network (SSR-Net) that explores the above-mentioned two kinds of relations simultaneously to learn more discriminative features by modeling the spatial structure information and global context information. In this article, we propose to adopt a Graph Convolution Network (GCN), for modeling spatial structural relationships among characteristic features. The GCN model aggregating the local and global features is shown to be more discriminative and robust to several car image transformations. To improve the performance of our proposed network, we jointly combine the classification loss with metric learning loss. Extensive experiments conducted on the public VehicleID and VeRi-776 datasets validate the effectiveness of our approach in comparison with recent works. Zheming Xu, Congyan Lang, Songhe Feng, Tao Wang 0011, Adrian G. Bors, Hongzhe Liu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2022 | Lifelong Generative Modelling Using Dynamic Expansion Graph ModelabstractVariational Autoencoders (VAEs) suffer from degenerated performance, when learning several successive tasks. This is caused by catastrophic forgetting. In order to address the knowledge loss, VAEs are using either Generative Replay (GR) mechanisms or Expanding Network Architectures (ENA). In this paper we study the forgetting behaviour of VAEs using a joint GR and ENA methodology, by deriving an upper bound on the negative marginal log-likelihood. This theoretical analysis provides new insights into how VAEs forget the previously learnt knowledge during lifelong learning. The analysis indicates the best performance achieved when considering model mixtures, under the ENA framework, where there are no restrictions on the number of components. However, an ENA-based approach may require an excessive number of parameters. This motivates us to propose a novel Dynamic Expansion Graph Model (DEGM). DEGM expands its architecture, according to the novelty associated with each new database, when compared to the information already learnt by the network from previous tasks. DEGM training optimizes knowledge structuring, characterizing the joint probabilistic representations corresponding to the past and more recently learned tasks. We demonstrate that DEGM guarantees optimal performance for each task while also minimizing the required number of parameters. Fei Ye 0004, Adrian G. Bors |
AAAI | 2 |
| 2022 | Continual Variational Autoencoder Learning via Online Cooperative Memorization
Fei Ye 0004, Adrian G. Bors |
ECCV (23) | 2 |
| 2022 | Learning an Evolved Mixture Model for Task-Free Continual LearningabstractRecently, continual learning (CL) has gained significant interest because it enables deep learning models to acquire new knowledge without forgetting previously learnt information. However, most existing works require knowing the task identities and boundaries, which is not realistic in a real context. In this paper, we address a more challenging and realistic setting in CL, namely the Task-Free Continual Learning (TFCL) in which a model is trained on non-stationary data streams with no explicit task information. To address TFCL, we introduce an evolved mixture model whose network architecture is dynamically expanded to adapt to the data distribution shift. We implement this expansion mechanism by evaluating the probability distance between the knowledge stored in each mixture model component and the current memory buffer using the Hilbert Schmidt Independence Criterion (HSIC). We further introduce two simple dropout mechanisms to selectively remove stored examples in order to avoid memory overload while preserving memory diversity. Empirical results demonstrate that the proposed approach achieves excellent performance. Fei Ye 0004, Adrian G. Bors |
ICIP | 2 |
| 2022 | Dot-Product Based Global and Local Feature Fusion for Image SearchabstractContent-based image retrieval (CBIR) consists in searching the most similar images to the query content from a given pool of images or database. Existing works' success relies on taking advantage of both local and global feature information leading to better retrieval performance than when using either of these. Lately, CBIR area has been dominated by the two-stage image retrieval framework which utilizes global features to get initial retrieval results, while using local features for reranking in a second stage. In this study, instead of utilizing local and global features separately during two stages, we propose to use a dot-product based local and global (DPLG) feature fusion module leading to a comprehensive global feature descriptor. The proposed fusion module is jointly end-to-end trained within the convolution backbone structure. According to the experimental results, the proposed module achieves new state-of-the-art results on some benchmark datasets. Zechao Hu 0001, Adrian G. Bors |
ICIP | 2 |
| 2022 | Predicting Human Perception of Scene ComplexityabstractIt is apparent that humans are intrinsically capable of determining the degree of complexity present in an image; but it is unclear which regions in that image lead humans towards evaluating an image as complex or simple. Here, we develop a novel deep learning model for predicting human perception of the complexity of natural scene images in order to address these problems. For a given image, our approach, ComplexityNet, can generate both single-score complexity ratings and two-dimensional per-pixel complexity maps. These complexity maps indicate the regions of scenes that humans find to be complex, or simple. Drawing on work in the cognitive sciences we integrate metrics for scene clutter and scene symmetry, and conclude that the proposed metrics do indeed boost neural network performance when predicting complexity. Cameron P. Kyle-Davidson, Adrian G. Bors, Karla K. Evans |
ICIP | 2 |
| 2022 | Encoder Enabled Gan-based Video GeneratorsabstractThis research study proposes a compatible encoder-enabled video generating method. The encoder-enabled method adds an inference mechanism for enhancing the ability of Generative Adversarial Networks (GAN) based video generators. The proposed video generating method is called Encoding GAN3 (EncGAN3) and decomposes the video into two streams representing content and movement, respectively. The proposed model consists of three processing modules, representing Encoder, Generator and Discriminator, each trained separately, by considering its own loss function. Enc-GAN3 is shown to generate videos of high quality, according to both visual and numerical results. Adrian G. Bors |
ICIP | 2 |
| 2022 | Expressive Local Feature Match for Image SearchabstractContent-based image retrieval (CBIR) aims to search the most similar images to a given query content, from a large pool of images. Existing state of the art works would extract a compact global feature vector for each image and then evaluate their similarity. Although some CBIR works utilize local features to get better retrieval results, they either would require extra codebook training or use re-ranking for improving the retrieved results. In this work, we propose a many-to-many local feature matching for large scale CBIR tasks. Unlike existing local feature based algorithms which tend to extract large amounts of short-dimensional local features from each image, the characteristic feature representation in the proposed approach is modeled for each image aiming to employ fewer but more expressive local features. Characteristic latent features are selected using k-means clustering and then fed into a similarity measure, without using complex matching kernels or codebook references. Despite the straightforwardness of the proposed CBIR method, experimental results indicate state of art results on several benchmark datasets. Zechao Hu 0001, Adrian G. Bors |
ICPR | 2 |
| 2022 | Task-Free Continual Learning via Online Discrepancy Distance LearningabstractLearning from non-stationary data streams, also called Task-Free Continual Learning (TFCL) remains challenging due to the absence of explicit task information in most applications. Even though recently some algorithms have been proposed for TFCL, these methods lack theoretical guarantees. Moreover, there are no theoretical studies about forgetting during TFCL. This paper develops a new theoretical analysis framework that derives generalization bounds based on the discrepancy distance between the visited samples and the entire information made available for training the model. This analysis provides new insights into the forgetting behaviour in classification tasks. Inspired by this theoretical model, we propose a new approach enabled with the dynamic component expansion mechanism for a mixture model, namely Online Discrepancy Distance Learning (ODDL). ODDL estimates the discrepancy between the current memory and the already accumulated knowledge as an expansion signal aiming to ensure a compact network architecture with optimal performance. We then propose a new sample selection approach that selectively stores the samples into the memory buffer through the discrepancy-based measure, further improving the performance. We perform several TFCL experiments with the proposed methodology, which demonstrate that the proposed approach achieves the state of the art performance. Fei Ye 0004, Adrian G. Bors |
NeurIPS | 2 |
| 2022 | Busy-Quiet Video Disentangling for Video ClassificationabstractIn video data, busy motion details from moving regions are conveyed within a specific frequency bandwidth in the frequency domain. Meanwhile, the rest of the frequencies of video data are encoded with quiet information with substantial redundancy, which causes low processing efficiency in existing video models that take as input raw RGB frames. In this paper, we consider allocating intenser computation for the processing of the important busy information and less computation for that of the quiet information. We design a trainable Motion Band-Pass Module (MBPM) for separating busy information from quiet information in raw video data. By embedding the MBPM into a two-pathway CNN architecture, we define a Busy-Quiet Net (BQN). The efficiency of BQN is determined by avoiding redundancy in the feature space processed by the two pathways: one operating on Quiet features of low-resolution, while the other processes Busy features. The proposed BQN outperforms many recent video processing models on Something-Something V1, Kinetics400, UCF101 and HMDB51 datasets. The code is available at: https://github.com/guoxih/busy-quiet-net. Guoxi Huang, Adrian G. Bors |
WACV | 2 |
| 2022 | Lifelong Teacher-Student Network LearningabstractA unique cognitive capability of humans consists in their ability to acquire new knowledge and skills from a sequence of experiences. Meanwhile, artificial intelligence systems are good at learning only the last given task without being able to remember the databases learnt in the past. We propose a novel lifelong learning methodology by employing a Teacher-Student network framework. While the Student module is trained with a new given database, the Teacher module would remind the Student about the information learnt in the past. The Teacher, implemented by a Generative Adversarial Network (GAN), is trained to preserve and replay past knowledge corresponding to the probabilistic representations of previously learnt databases. Meanwhile, the Student module is implemented by a Variational Autoencoder (VAE) which infers its latent variable representation from both the output of the Teacher module as well as from the newly available database. Moreover, the Student module is trained to capture both continuous and discrete underlying data representations across different domains. The proposed lifelong learning framework is applied in supervised, semi-supervised and unsupervised training. Fei Ye 0004, Adrian G. Bors |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | BQN: Busy-Quiet Net Enabled by Motion Band-Pass Module for Action RecognitionabstractA rich video data representation can be realized by means of spatio-temporal frequency analysis. In this research study we show that a video can be disentangled, following the learning of video characteristics according to their spatio-temporal properties, into two complementary information components, dubbed Busy and Quiet. The Busy information characterizes the boundaries of moving regions, moving objects, or regions of change in movement. Meanwhile, the Quiet information encodes global smooth spatio-temporal structures defined by substantial redundancy. We design a trainable Motion Band-Pass Module (MBPM) for separating Busy and Quiet-defined information, in raw video data. We model a Busy-Quiet Net (BQN) by embedding the MBPM into a two-pathway CNN architecture. The efficiency of BQN is determined by avoiding redundancy in the feature spaces defined by the two pathways. While one pathway processes the Busy features, the other processes Quiet features at lower spatio-temporal resolutions reducing both memory and computational costs. Through experiments we show that the proposed MBPM can be used as a plug-in module in various CNN backbone architectures, significantly boosting their performance. The proposed BQN is shown to outperform many recent video models on Something-Something V1, Kinetics400, UCF101 and HMDB51 datasets. Guoxi Huang, Adrian G. Bors |
IEEE Trans. Image Process. | 2 |
| 2022 | Deep Mixture Generative AutoencodersabstractVariational autoencoders (VAEs) are one of the most popular unsupervised generative models that rely on learning latent representations of data. In this article, we extend the classical concept of Gaussian mixtures into the deep variational framework by proposing a mixture of VAEs (MVAE). Each component in the MVAE model is implemented by a variational encoder and has an associated subdecoder. The separation between the latent spaces modeled by different encoders is enforced using the d -variable Hilbert-Schmidt independence criterion (dHSIC). Each component would capture different data variational features. We also propose a mechanism for finding the appropriate number of VAE components for a given task, leading to an optimal architecture. The differentiable categorical Gumbel-softmax distribution is used in order to generate dropout masking parameters within the end-to-end backpropagation training framework. Extensive experiments show that the proposed MVAE model can learn a rich latent data representation and is able to discover additional underlying data representation factors. Fei Ye 0004, Adrian G. Bors |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Lifelong Infinite Mixture Model Based on Knowledge-Driven Dirichlet ProcessabstractRecent research efforts in lifelong learning propose to grow a mixture of models to adapt to an increasing number of tasks. The proposed methodology shows promising results in overcoming catastrophic forgetting. However, the theory behind these successful models is still not well understood. In this paper, we perform the theoretical analysis for lifelong learning models by deriving the risk bounds based on the discrepancy distance between the probabilistic representation of data generated by the model and that corresponding to the target dataset. Inspired by the theoretical analysis, we introduce a new lifelong learning approach, namely the Lifelong Infinite Mixture (LIMix) model, which can automatically expand its network architectures or choose an appropriate component to adapt its parameters for learning a new task, while preserving its previously learnt information. We propose to incorporate the knowledge by means of Dirichlet processes by using a gating mechanism which computes the dependence between the knowledge learnt previously and stored in each component, and a new set of data. Besides, we train a compact Student model which can accumulate cross-domain representations over time and make quick inferences. The code is available at https://github.com/dtuzi123/Lifelong-infinite-mixture-model. Fei Ye 0004, Adrian G. Bors |
ICCV | 2 |
| 2021 | InfoVAEGAN: Learning Joint Interpretable Representations by Information Maximization and Maximum LikelihoodabstractLearning disentangled and interpretable representations is an important step towards accomplishing comprehensive data representations on the manifold. In this paper, we propose a novel representation learning algorithm which combines the inference abilities of Variational Autoencoders (VAE) with the generalization capability of Generative Adversarial Networks (GAN). The proposed model, called InfoVAEGAN, consists of three networks: Encoder, Generator and Discriminator. InfoVAEGAN aims to jointly learn discrete and continuous interpretable representations in an unsupervised manner by using two different data-free log-likelihood functions onto the variables sampled from the generator’s distribution. We propose a two-stage algorithm for optimizing the inference network separately from the generator training. Moreover, we enforce the learning of interpretable representations through the maximization of the mutual information between the existing latent variables and those created through generative and inference processes. Fei Ye 0004, Adrian G. Bors |
ICIP | 2 |
| 2021 | Lifelong Twin Generative Adversarial NetworksabstractIn this paper, we propose a new continuously learning generative model, called the Lifelong Twin Generative Adversarial Networks (LT-GANs). LT-GANs learns a sequence of tasks from several databases and its architecture consists of three components: two identical generators, namely the Teacher and Assistant, and one Discriminator. In order to allow for the LT-GANs to learn new concepts without forgetting, we introduce a new lifelong training approach, namely Lifelong Adversarial Knowledge Distillation (LAKD), which encourages the Teacher and Assistant to alternately teach each other, while learning a new database. This training approach favours transferring knowledge from a more knowledgeable player to another player which knows less information about a previously given task. Fei Ye 0004, Adrian G. Bors |
ICIP | 2 |
| 2021 | Learning joint latent representations based on information maximization
Fei Ye 0004, Adrian G. Bors |
Inf. Sci. | 2 |
| 2020 | Conditional Attention for Content-based Image Retrieval
Zechao Hu 0001, Adrian G. Bors |
BMVC | 2 |
| 2020 | Learning Latent Representations Across Multiple Data Domains Using Lifelong VAEGAN
Fei Ye 0004, Adrian G. Bors |
ECCV (20) | 2 |
| 2020 | Learning Spatio-Temporal Representations With Temporal Squeeze PoolingabstractIn this paper, we propose a new video representation learn¬ing method, named Temporal Squeeze (TS) pooling, which can extract the essential movement information from a long sequence of video frames and map it into a set of few im¬ages, named Squeezed Images. By embedding the Tempo¬ral Squeeze pooling as a layer into off-the-shelf Convolution Neural Networks (CNN), we design anew video classification model, named Temporal Squeeze Network (TeSNet). The re¬sulting Squeezed Images contain the essential movement in¬formation from the video frames, corresponding to the op¬timization of the video classification task. We evaluate our architecture on two video classification benchmarks, and the results achieved are compared to the state-of-the-art. Guoxi Huang, Adrian G. Bors |
ICASSP | 2 |
| 2020 | Region-based Non-local Operation for Video ClassificationabstractConvolutional Neural Networks (CNNs) model long-range dependencies by deeply stacking convolution operations with small window sizes, which makes the optimizations difficult. This paper presents region-based non-local (RNL) operations as a family of self-attention mechanisms, which can directly capture long-range dependencies without using a deep stack of local operations. Given an intermediate feature map, our method recalibrates the feature at a position by aggregating the information from the neighboring regions of all positions. By combining a channel attention module with the proposed RNL, we design an attention chain, which can be integrated into the off-the-shelf CNNs for end-to-end training. We evaluate our method on two video classification benchmarks. The experimental results of our method outperform other attention mechanisms, and we achieve state-of-the-art performance on the Something-Something V1 dataset. The code is available at: https://github.com/guoxih/region-based-non-local-network. Guoxi Huang, Adrian G. Bors |
ICPR | 2 |
| 2020 | Steganalysis of meshes based on 3D wavelet multiresolution analysis
Zhenyu Li 0004, Adrian G. Bors |
Inf. Sci. | 2 |
| 2020 | Defining Image Memorability Using the Visual Memory SchemaabstractMemorability of an image is a characteristic determined by the human observers' ability to remember images they have seen. Yet recent work on image memorability defines it as an intrinsic property that can be obtained independent of the observer. The current study aims to enhance our understanding and prediction of image memorability, improving upon existing approaches by incorporating the properties of cumulative human annotations. We propose a new concept called the Visual Memory Schema (VMS) referring to an organization of image components human observers share when encoding and recognizing images. The concept of VMS is operationalised by asking human observers to define memorable regions of images they were asked to remember during an episodic memory test. We then statistically assess the consistency of VMSs across observers for either correctly or incorrectly recognised images. The associations of the VMSs with eye fixations and saliency are analysed separately as well. Lastly, we adapt various deep learning architectures for the reconstruction and prediction of memorable regions in images and analyse the results when using transfer learning at the outputs of different convolutional network layers. Erdem Akagündüz, Adrian G. Bors, Karla K. Evans |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Selection of Robust and Relevant Features for 3-D SteganalysisabstractWhile 3-D steganography and digital watermarking represent methods for embedding information into 3-D objects, 3-D steganalysis aims to find the hidden information. Previous research studies have shown that by estimating the parameters modeling the statistics of 3-D features and feeding them into a classifier we can identify whether a 3-D object carries secret information. For training the steganalyzer, such features are extracted from cover and stego pairs, representing the original 3-D objects and those carrying hidden information. However, in practical applications, the steganalyzer would have to distinguish stego-objects from cover-objects, which most likely have not been used during the training. This represents a significant challenge for existing steganalyzers, raising a challenge known as the cover source mismatch (CSM) problem, which is due to the significant limitation of their generalization ability. This paper proposes a novel feature selection algorithm taking into account both feature robustness and relevance in order to mitigate the CSM problem in 3-D steganalysis. In the context of the proposed methodology, new shapes are generated by distorting those used in the training. Then a subset of features is selected from a larger given set, by assessing their effectiveness in separating cover-objects from stego-objects among the generated sets of objects. Two different measures are used for selecting the appropriate features: 1) the Pearson correlation coefficient and 2) the mutual information criterion. Zhenyu Li 0004, Adrian G. Bors |
IEEE Trans. Cybern. | 2 |
| 2019 | Predicting Visual Memory Schemas with Variational Autoencoders
Cameron P. Kyle-Davidson, Adrian G. Bors, Karla K. Evans |
BMVC | 2 |
| 2018 | 3D Steganalysis Using the Extended Local Feature Setabstract3D steganalysis aims to find the changes embedded through steganographic or information hiding algorithms into 3D models. This research study proposes to use new 3D features, such as the edge vectors, represented in both Cartesian and Laplacian coordinate systems, together with other steganalytic features, for improving the results of 3D steganalysers. In this way the local feature vector used by the steganalyzer is extended to 124 dimensions. We test the performance of the extended local feature set, and compare it to four other steganalytic features, when detecting the stego-objects watermarked by six information hiding algorithms. Zhenyu Li 0004, Daofu Gong, Fenlin Liu, Adrian G. Bors |
ICIP | 4 |
| 2018 | Editorial: Special Issue on Machine Vision
Edwin R. Hancock, Richard C. Wilson 0001, William A. P. Smith, Adrian G. Bors, Nick E. Pears |
Int. J. Comput. Vis. | 4 |
| 2017 | Recognizing Interactions Between People from Video Sequences
Kyle Stephens, Adrian G. Bors |
CAIP (1) | 2 |
| 2017 | Rethinking the high capacity 3D steganography: Increasing its resistance to steganalysisabstract3D steganography is used in order to embed or hide information into 3D objects without causing visible or machine detectable modifications. In this paper we rethink about a high capacity 3D steganography based on the Hamiltonian path quantization, and increase its resistance to steganalysis. We analyze the parameters that may influence the distortion of a 3D shape as well as the resistance of the steganography to 3D steganalysis. According to the experimental results, the proposed high capacity 3D steganographic method has an increased resistance to steganalysis. Zhenyu Li 0004, Sebastien Beugnon, William Puech, Adrian G. Bors |
ICIP | 4 |
| 2017 | Steganalysis of 3D objects using statistics of local feature sets
Zhenyu Li 0004, Adrian G. Bors |
Inf. Sci. | 2 |
| 2016 | Group activity recognition on outdoor scenesabstractIn this research study, we propose an automatic group activity recognition approach by modelling the interdependencies of group activity features over time. Unlike in simple human activity recognition approaches, the distinguishing characteristics of group activities are often determined by how the movement of people are influenced by one another. We propose to model the group interdependences in both motion and location spaces. These spaces are extended to time-space and time-movement spaces and modelled using Kernel Density Estimation (KDE). Such representations are then fed into a machine learning classifier which identifies the group activity. Unlike other approaches to group activity recognition, we do not rely on the manual annotation of pedestrian tracks from the video sequence. Kyle Stephens, Adrian G. Bors |
AVSS | 2 |
| 2016 | 3D mesh steganalysis using local shape featuresabstractSteganalysis aims to identify those changes performed in a specific media with the intention to hide information. In this paper we assess the efficiency, in finding hidden information, of several local feature detectors. In the proposed 3D steanalysis approach we first smooth the cover object and its corresponding stego-object obtained after embedding a given message. We use various operators in order to extract local features from both the cover and stego-objects, and their smoothed versions. Machine learning algorithms are then used for learning to discriminate between those 3D objects which are used as carriers of hidden information and those are not used. The proposed 3D steganalysis methodology is shown to provide superior performance to other approaches in a well known database of 3D objects. Zhenyu Li 0004, Adrian G. Bors |
ICASSP | 2 |
| 2016 | Selection of robust features for the Cover Source Mismatch problem in 3D steganalysisabstractThis paper introduces a novel method for extracting sets of feature from 3D objects characterising a robust steganalyzer. Specifically, the proposed steganalyzer should mitigate the Cover Source Mismatch (CSM) paradigm. A steganalyzer is considered as a classifier aiming to identify separately cover and stego objects. A steganalyzer behaves as a classifier by considering a set of features extracted from cover stego pairs of 3D objects as inputs during the training stage. However, during the testing stage, the steganalyzer would have to identify whether specific information was hidden in a set of 3D objects which can be different from those used during the training. Addressing the CSM paradigm corresponds to testing the generalization ability of the steganalyzer when introducing distortions in the cover objects before hiding information through steganography. Our method aims to select those 3D features that model best the changes introduced in objects by steganography or information hiding and moreover they are able to generalize for different objects, not present in the training set. The proposed robust steganalysis approach is tested when considering changes in 3D objects such as those produced by mesh simplification and additive noise. The results obtained from this study show that the steganalyzers trained with the selected set of robust features achieve better detection accuracy of the changes embedded in the objects, when compared to other sets of features. Zhenyu Li 0004, Adrian G. Bors |
ICPR | 2 |
| 2016 | Human group activity recognition based on modelling moving regions interdependenciesabstractIn this research study, we model the interdependency of actions performed by people in a group in order to identify their activity. Unlike single human activity recognition, in interacting groups the local movement activity is usually influenced by the other persons in the group. We propose a model to describe the discriminative characteristics of group activity by considering the relations between motion flows and the locations of moving regions. The inputs of the proposed model are jointly represented in time-space and time-movement spaces. These spaces are modelled using Kernel Density Estimation (KDE) which is then fed into a machine learning classifier. Unlike in other group-based human activity recognition algorithms, the proposed methodology is automatic and does not rely on any pedestrian detection or on the manual annotation of tracks. Kyle Stephens, Adrian G. Bors |
ICPR | 2 |
| 2015 | Observing human activities using movement modellingabstractIn this paper we propose an unsupervised method based on a scene observation approach to human activity detection from video sequences. The approach adopted has two stages: modelling the categories of movements in the scene while the second stage consists of detecting new activities. Human activity is decided based on analysing movement in the scene. Both optical flow, estimated from pairs or frames, and medium term tracking are considered for representing the movement. Statistical modelling using mixtures of Gaussians, with each component representing a specific movement is considered. The tracking approach considers several consecutive frames and uses streaklines for modelling the corresponding movement. During the first stage, a dictionary of activities, characteristic to the observed scene, is formed. New activities are identified using Kullback-Leibler (KL) divergence between any new distribution of motion vectors and the existing ones from the dictionary. The proposed methodology is applied on various video sequences representing both indoor and outdoor scenes. Kyle Stephens, Adrian G. Bors |
AVSS | 2 |
| 2015 | Robust Learning from Ortho-Diffusion Decompositions
Sravan Gudivada, Adrian G. Bors |
CAIP (1) | 2 |
| 2015 | Content Based Image Retrieval Based on Modelling Human Visual Attention
Alex Papushoy, Adrian G. Bors |
CAIP (1) | 2 |
| 2015 | Ortho-diffusion decompositions for face recognition from low quality imagesabstractWe propose a new approach for recognizing human from images of low quality. An ortho-diffusion decomposition is used on graph representations of images. This is implemented by a recursive algorithm in three steps on either the covariance matrix or on the correlation of the training set. The first stage consists of an orthonormal decomposition implemented through the modified Gram-Schmidt with pivoting the columns. The other two stages consists of the data reduction and diffusion on graph representations. The data reduction ensures that the most significant features are preserved and together with the diffusion step ensures robustness to a variety of data corruption factors. The proposed methodology produces a set of ortho-diffusion bases representing the quintessential information from the training data set. The resulting orhto-diffusion bases are used to model face images when considering low resolution and corruption by various noise distributions. Sravan Gudivada, Adrian G. Bors |
ICIP | 2 |
| 2015 | Visual attention for content based image retrievalabstractA new image retrieval method, based on human visual attention models, called query by saliency content retrieval (QSCR) is presented in this paper. Each image is segmented and a set of characteristic features is evaluated for each region. The saliency for each image region, as it would be perceived by a human observer, is estimated for each region and then used for image retrieval. Images displaying similar features and characterized by similar saliency are then retrieved from the database. Both local and global saliency are considered in the retrieval process. The proposed method ranks the similarity between the query and the a set of given images using the Earth Mover Distance algorithm. Alex Papushoy, Adrian G. Bors |
ICIP | 2 |
| 2015 | Ortho-diffusion decompositions of graph-based representation of images
Sravan Gudivada, Adrian G. Bors |
Pattern Recognit. | 2 |
| 2015 | Editorial
Edwin R. Hancock, Richard C. Wilson 0001, Adrian G. Bors, William A. P. Smith |
Pattern Recognit. | 3 |
| 2014 | Correcting 3D scenes estimated from sets of multi-view images using shape-from-contoursabstractThis paper proposes enforcing the consistency with segmented contours when modelling scenes with multiple objects from multi-view images. A certain rough initialization of the 3D scene is assumed to be available and in the case of multiple objects inconsistencies are expected. In the proposed shape-from-contours approach images are segmented and back-projections of segmented contours are used for enforcing the consistency of the segmented contours with 3D objects from the scene. We provide a study for the physical requirements for detecting occlusions when reconstructing 3-D scenes with multiple objects. Matthew Grum, Adrian G. Bors |
ICIP | 2 |
| 2014 | Cryptanalysis aspects in 3-D watermarkingabstract3-D object security is increasingly brought to the attention of the public by the expansion of new multimedia technologies such as the 3-D printing. In the development of crypto-security systems of 3-D objects, we can identify two major directions represented by the cryptography and digital watermarking. A good security system has to be format compliant, has to preserve the original bit rate and, whenever possible, it should be reversible. Watermarking methodology has the advantage of ensuring that the embedded hidden message can be verified at any processing stage such as the transmission, storage and when visualizing the embedding media. In this paper, we review the previous work in 3-D security and analyze the crypto-security of a 3-D watermarking method which embeds information by mesh surface distortion minimization. Then, we discuss future avenues of research by presenting emerging applications. Vincent Itier, William Puech, Adrian G. Bors |
ICIP | 3 |
| 2014 | 3D modeling of multiple-object scenes from sets of images
Matthew Grum, Adrian G. Bors |
Pattern Recognit. | 2 |
| 2013 | Watermark Optimization of 3D Shapes for Minimal Distortion and High Robustness
Adrian G. Bors, Ming Luo 0002 |
CAIP (2) | 1 |
| 2013 | Enforcing Consistency of 3D Scenes with Multiple Objects Using Shape-from-Contours
Matthew Grum, Adrian G. Bors |
CAIP (1) | 2 |
| 2013 | Orthonormal Diffusion Decompositions of Images for Optical Flow Estimation
Sravan Gudivada, Adrian G. Bors |
CAIP (2) | 2 |
| 2013 | Optimized 3D Watermarking for Minimal Surface DistortionabstractThis paper proposes a new approach to 3D watermarking by ensuring the optimal preservation of mesh surfaces. A new 3D surface preservation function metric is defined consisting of the distance of a vertex displaced by watermarking to the original surface, to the watermarked object surface as well as the actual vertex displacement. The proposed method is statistical, blind, and robust. Minimal surface distortion according to the proposed function metric is enforced during the statistical watermark embedding stage using Levenberg-Marquardt optimization method. A study of the watermark code crypto-security is provided for the proposed methodology. According to the experimental results, the proposed methodology has high robustness against the common mesh attacks while preserving the original object surface during watermarking. Adrian G. Bors, Ming Luo 0002 |
IEEE Trans. Image Process. | 1 |
| 2011 | Surface-Preserving Robust Watermarking of 3-D ShapesabstractThis paper describes a new statistical approach for watermarking mesh representations of 3-D graphical objects. A robust digital watermarking method has to mitigate among the requirements of watermark invisibility, robustness, embedding capacity and key security. The proposed method employs a mesh propagation distance metric procedure called the fast marching method (FMM), which defines regions of equal geodesic distance width calculated with respect to a reference location on the mesh. Each of these regions is used for embedding a single bit. The embedding is performed by changing the normalized distribution of local geodesic distances from within each region. Two different embedding methods are used by changing the mean or the variance of geodesic distance distributions. Geodesic distances are slightly modified statistically by displacing the vertices in their existing triangle planes. The vertex displacements, performed according to the FMM, ensure a minimal surface distortion while embedding the watermark code. Robustness to a variety of attacks is shown according to experimental results. Ming Luo 0002, Adrian G. Bors |
IEEE Trans. Image Process. | 2 |
| 2010 | Detecting Vorticity in Optical Flow of FluidsabstractIn this paper we apply the diffusion framework to dense optical flow estimation. Local image information is represented by matrices of gradients between paired locations. Diffusion distances are modelled as sums of eigenvectors weighted by their eigenvalues extracted following the eigen decomposion of these matrices. Local optical flow is estimated by correlating diffusion distances characterizing features from different frames. A feature confidence factor is defined based on the local correlation efficiency when compared to that of its neighbourhood. High confidence optical flow estimates are propagated to areas of lower confidence. Ashish Doshi, Adrian G. Bors |
ICPR | 2 |
| 2010 | Optical Flow Estimation Using Diffusion DistancesabstractIn this paper we apply the diffusion framework to dense optical flow estimation. Local image information is represented by matrices of gradients between paired locations. Diffusion distances are modelled as sums of eigenvectors weighted by their eigenvalues extracted following the eigen decomposion of these matrices. Local optical flow is estimated by correlating diffusion distances characterizing features from different frames. A feature confidence factor is defined based on the local correlation efficiency when compared to that of its neighbourhood. High confidence optical flow estimates are propagated to areas of lower confidence. Szymon Wartak, Adrian G. Bors |
ICPR | 2 |
| 2010 | Smoothing of optical flow using robustified diffusion kernels
Ashish Doshi, Adrian G. Bors |
Image Vis. Comput. | 2 |
| 2010 | Robust Processing of Optical Flow of FluidsabstractThis paper proposes a new approach, coupling physical models and image estimation techniques, for modelling the movement of fluids. The fluid flow is characterized by turbulent movement and dynamically changing patterns which poses challenges to existing optical flow estimation methods. The proposed methodology, which relies on Navier-Stokes equations, is used for processing fluid optical flow by using a succession of stages such as advection, diffusion and mass conservation. A robust diffusion step jointly considering the local data geometry and its statistics is embedded in the proposed framework. The diffusion kernel is Gaussian with the covariance matrix defined by the local second derivatives. Such an anisotropic kernel is able to implicitly detect changes in the vector field orientation and to diffuse accordingly. A new approach is developed for detecting fluid flow structures such as vortices. The proposed methodology is applied on artificially generated vector fields as well as on various image sequences. Ashish Doshi, Adrian G. Bors |
IEEE Trans. Image Process. | 2 |
| 2009 | Bayesian Estimation of Kernel Bandwidth for Nonparametric Modelling
Adrian G. Bors, Nikolaos Nasios |
ICANN (2) | 1 |
| 2009 | Blind and robust mesh watermarking using manifold harmonicsabstractIn this paper, we present a new blind and robust 3-D mesh watermarking scheme that makes use of the recently proposed manifold harmonics analysis. The mesh spectrum coefficient amplitudes obtained by using this analysis are quite robust against various attacks, including connectivity changes. A blind 16-bit watermark is embedded through an iterative scalar Costa quantization of the low frequency coefficient amplitudes. The imperceptibility of the watermark is ensured since the human visual system has been proved insensitive to the mesh low frequency components modification. The embedded watermark is experimentally robust against both geometry and connectivity attacks. Comparison results with two state-of-the-art methods are provided. Kai Wang 0002, Ming Luo 0002, Adrian G. Bors, Florence Denis |
ICIP | 3 |
| 2009 | Local Patch Blind Spectral Watermarking Method for 3D Graphics
Ming Luo 0002, Kai Wang 0002, Adrian G. Bors, Guillaume Lavoué |
IWDW | 3 |
| 2009 | Shape watermarking based on minimizing the quadric error metricabstractBlind and robust watermarking of 3D object aims to embed codes into a 3D object such that the object is not visually distorted from the original shape. An essential condition is that the message should be securely extracted even after the graphical object was processed. In this paper, we propose a novel blind and robust mesh watermarking method based on the quadric error metric. The vertices are firstly grouped into bins using a secret key according to their distances to the object center. The statistics of the distances in each bin is modified when embedding the message. A novel quadric selective vertex placement scheme is proposed for finding the best location of each vertex, following watermark embedding, such that the resulting shape distortion is minimal. Experimental results show that the proposed method reduces the distortion to a minimum in the 3D shape. Ming Luo 0002, Adrian G. Bors |
Shape Modeling International | 2 |
| 2009 | Kernel Bandwidth Estimation for Nonparametric ModelingabstractKernel density estimation is a nonparametric procedure for probability density modeling, which has found several applications in various fields. The smoothness and modeling ability of the functional approximation are controlled by the kernel bandwidth. In this paper, we describe a Bayesian estimation method for finding the bandwidth from a given data set. The proposed bandwidth estimation method is applied in three different computational-intelligence methods that rely on kernel density estimation: 1) scale space; 2) mean shift; and 3) quantum clustering. The third method is a novel approach that relies on the principles of quantum mechanics. This method is based on the analogy between data samples and quantum particles and uses the SchrOdinger potential as a cost function. The proposed methodology is used for blind-source separation of modulated signals and for terrain segmentation based on topography information. Adrian G. Bors, Nikolaos Nasios |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2008 | Multiple image disparity correction for 3-D scene representationabstractThis paper analyses the reconstruction of a 3D scene from multiple images. The proposed approach uses a voxel representation for initializing an implicit radial basis function (RBF) model of the 3D scene. The voxel representation is produced by the space carving algorithm. Disparity errors are identified as inconsistencies between the 3D patches of the scene surface and their pixel block correspondents in the images. A matching algorithm is used for estimating the disparity errors after aligning all these pixel block regions along epipolar lines, each corresponding to a 3D patch of scene surface. An algorithm is derived for updating the RBF center locations using the disparity errors. Experiments are performed on two different sets of images and numerical results are provided when comparing the resulting 3D scene with the ground truth given by a laser scanner. Matthew Grum, Adrian G. Bors |
ICIP | 2 |
| 2008 | Principal Component Analysis of spectral coefficients for mesh watermarkingabstractThis paper proposes a new robust 3-D object blind watermarking method using constraints in the spectral domain. Mesh watermarking in spectral domain has the property of spreading the information in unpredictable ways, thus increasing the security of the watermark. In the proposed method, firstly, the Laplacian matrix of the graphical object mesh is eigen-decomposed. The coefficients corresponding to the higher spectra are split into sets and each set is used for embedding one bit. A bit of 1 is embedded by introducing an asymmetry in the 3-D distribution of the spectral coefficients from the given set, while the distribution symmetry is enforced in the case when embedding a bit of 0. The Principal Component Analysis (PCA) is used for embedding the constraints in the spectral domain by ensuring a minimal distortion. Comparison results are provided for various attacks. Ming Luo 0002, Adrian G. Bors |
ICIP | 2 |
| 2008 | Kernel bandwidth estimation in methods based on probability density function modellingabstractIn kernel density estimation methods, an approximation of the data probability density function is achieved by locating a kernel function at each data location. The smoothness of the functional approximation and the modelling ability are controlled by the kernel bandwidth. In this paper we propose a Bayesian estimation method for finding the kernel bandwidth. The distribution corresponding to the bandwidth is estimated from distributions characterizing the second order statistics estimates calculated from local neighbourhoods. The proposed bandwidth estimation method is applied in three different kernel density estimation based approaches: scale space, mean shift and quantum clustering. The third method is a novel pattern recognition approach using the principles of quantum mechanics. Adrian G. Bors, Nikolaos Nasios |
ICPR | 1 |
| 2008 | Enforcing image consistency in multiple 3-D object modellingabstractIn this paper we present a new approach for modelling multiple object scenes using images taken from various viewpoints. The voxel representation produced by the space carving is used for estimating the parameters of radial basis function functions (RBFs). A new methodology is proposed for correcting the modelling errors within the RBF framework, by enforcing consistency with the input images by matching along epipolar lines, and enforcing silhouette constraints using the visual hull. The proposed methodology is applied on two datasets consisting of several objects while providing quantitative and qualitative assessment. Matthew Grum, Adrian G. Bors |
ICPR | 2 |
| 2007 | Navier-Stokes Formulation for Modelling Turbulent Optical FlowabstractThis paper proposes a physics-based methodology for the analysis of optical flows displaying complex patterns. Turbulent motion, such as that exhibited by fluid substances, can be modelled using fluid dynamics principles. Together with supplemental equations, such as the conservation of mass, and well formulated boundary conditions, the Navier-Stokes equations can be used to model complex fluid motion estimated from image sequences. In this paper, we propose to use a robust kernel which adapts to the local data geometry in the diffusion stage of the Navier-Stokes formulation. The proposed kernel is Gaussian and embeds the Hessian of the local data as its covariance matrix. The local Hessian models the variation of the flow in a certain neighbourhood. Moreover, we use a robust statistics mechanism in order to eliminate the outliers from the estimation process. The proposed methodology is applied on artificial vector fields and in image sequences showing atmospheric and solar phenomena. 1 Ashish Doshi, Adrian G. Bors |
BMVC | 2 |
| 2007 | Refining Implicit Function Representations of 3-D ScenesabstractThis paper considers the problem of modelling a 3-D scene from calibrated images taken from multiple viewpoints. The initial 3-D information is acquired using probabilistic space carving which provides a voxel representation consistent with the given set of images. The scene is afterwards modelled as an implicit surface using radial basis functions (RBF). The mixture of multiorder basis functions models a smoothed 3-D scene representation while providing compactness. We use correspondences between pairs of image patches in order to update the RBF centres for improving the 3-D scene representation. The RBF centre updating leads to improving the consistency between the 3-D model and the given set of images. The proposed method is applied on a complex 3-D scene displaying various objects. 1 Matthew Grum, Adrian G. Bors |
BMVC | 2 |
| 2007 | Kernel-based classification using quantum mechanics
Nikolaos Nasios, Adrian G. Bors |
Pattern Recognit. | 2 |
| 2006 | Robust Diffusion of Structural Flows for Volumetric Image InterpolationabstractIn this paper we propose a set of algorithms that combine the anisotropic smoothing using the heat kernel with the outlier rejection capability of robust statistics. The proposed algorithms are applied on structural vector flows that model the internal shape variation in volumetric images. The 3D shapes are represented by sparse cross-sections along the main axis of the object. The dual directional block matching algorithm is used to initially extract the structural flows. This algorithm uses block matching between pixel blocks from consecutive images representing sparse cross-sections through a volume. Two flows are produced using forward and reverse matching along the main axis of the 3D object. After smoothing, the structural flows are used for slice interpolation. Experimental results provide a comparison among the given algorithms when used for digital 3D reconstruction of an incisor and of two human bones. Ashish Doshi, Adrian G. Bors |
ICIP | 2 |
| 2006 | Selective Encryption of Human Skin in JPEG ImagesabstractIn this study we propose a new approach for selective encryption in the Huffman coding of the Discrete Cosine Transform (DCT) coefficients using the Advanced Encryption Standard (AES). The objective is to partially encrypt the human face in an image or video sequence. This approach is based on the AES stream ciphering using Variable Length Coding (VLC) of the Huffman's vector. The proposed scheme allows the decryption of a specific region of the image and results in a significant reduction in encrypting and decrypting processing time. It also provides a constant bit rate while maintaining the JPEG and MPEG bitstream compliance. José M. Rodrigues, William Puech, Adrian G. Bors |
ICIP | 3 |
| 2006 | Watermarking mesh-based representations of 3-D objects using local momentsabstractA new methodology for fingerprinting and watermarking three-dimensional (3-D) graphical objects is proposed in this paper. The 3-D graphical objects are described by means of polygonal meshes. The information to be embedded is provided as a binary code. A watermarking methodology has two stages: embedding and detecting the information that has been embedded in the given media. The information is embedded by means of local geometrical perturbations while maintaining the local connectivity. A neighborhood localized measure is used for selecting appropriate vertices for watermarking. A study is undertaken in order to verify the suitability of this measure for selecting vertices from regions where geometrical perturbations are less perceptible. Two different watermarking algorithms, that do not require the original 3-D graphical object in the detection stage, are proposed. The two algorithms differ with respect to the type of constraint to be embedded in the local structure: by using parallel planes and bounding ellipsoids, respectively. The information capacity of various 3-D meshes is analyzed when using the proposed 3-D watermarking algorithms. The robustness of the 3-D watermarking algorithms is tested to noise perturbation and to object cropping. Adrian G. Bors |
IEEE Trans. Image Process. | 1 |
| 2006 | Variational learning for Gaussian mixture modelsabstractThis paper proposes a joint maximum likelihood and Bayesian methodology for estimating Gaussian mixture models. In Bayesian inference, the distributions of parameters are modeled, characterized by hyperparameters. In the case of Gaussian mixtures, the distributions of parameters are considered as Gaussian for the mean, Wishart for the covariance, and Dirichlet for the mixing probability. The learning task consists of estimating the hyperparameters characterizing these distributions. The integration in the parameter space is decoupled using an unsupervised variational methodology entitled variational expectation-maximization (VEM). This paper introduces a hyperparameter initialization procedure for the training algorithm. In the first stage, distributions of parameters resulting from successive runs of the expectation-maximization algorithm are formed. Afterward, maximum-likelihood estimators are applied to find appropriate initial values for the hyperparameters. The proposed initialization provides faster convergence, more accurate hyperparameter estimates, and better generalization for the VEM training algorithm. The proposed methodology is applied in blind signal detection and in color image segmentation. Nikolaos Nasios, Adrian G. Bors |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2005 | Optical Flow Diffusion with Robustified Kernels
Ashish Doshi, Adrian G. Bors |
CAIP | 2 |
| 2005 | Finding the Number of Clusters for Nonparametric Segmentation
Nikolaos Nasios, Adrian G. Bors |
CAIP | 2 |
| 2005 | Variational segmentation of color imagesabstractA variational Bayesian framework is employed in the paper for image segmentation using color clustering. A Gaussian mixture model is used to represent color distributions. Variational expectation-maximization (VEM) algorithm takes into account the uncertainty in the parameter estimation ensuring a lower bound on the approximation error. In the variational Bayesian approach we integrate over distributions of parameters. The processing task in this case consists of estimating the hyperparameters of these distributions. We propose a maximum log-likelihood initialization approach for the variational expectation-maximization (VEM) algorithm. The proposed algorithm is applied to image segmentation using color clustering when representing the images in the L*u*v color coordinate system. Nikolaos Nasios, Adrian G. Bors |
ICIP (2) | 2 |
| 2005 | Nonparametric clustering using quantum mechanicsabstractThis paper introduces a new nonparametric estimation approach that can be used for data that is not necessarily Gaussian distributed. The proposed approach employs the Shrodinger partial differential equation. We assume that each data sample is associated with a quantum physics particle that has a radial field around its value. We consider a statistical estimation approach for finding the size of the influence field around each data sample. By implementing the Shrodinger equation we obtain a potential field that is assimilated with the data density. The regions of minima in the potential are determined by calculating the local Hessian on the potential hypersurface. The quantum clustering approach is applied for blind separation of signals and for segmenting SAR images of terrain based on surface normal orientation. Nikolaos Nasios, Adrian G. Bors |
ICIP (3) | 2 |
| 2004 | Segmentation of colour images using variational expectation-maximization algorithmabstractThe approach proposed in this paper takes into account the uncertainty in colour modelling by employing variational Bayesian estimation. Mixtures of Gaussians are considered for modelling colour images. Distributions of parameters characterising colour regions are inferred from data statistics. The Variational Expectation-Maximization (VEM) algorithm is used for estimating the hyperparameters corresponding to distributions of parameters. A maximum a posteriori approach employing a dual expectation-maximization (EM) algorithm is considered for the hyperparameter initialisation of the VEM algorithm. In the first stage, the EM algorithm is applied on the given colour image, while the second EM algorithm is used on distributions of parameters resulted from several runs of the first stage EM. The VEM algorithm is used for segmenting several colour images. Nikolaos Nasios, Adrian G. Bors |
BMVC | 2 |
| 2004 | Watermarking 3D shapes using local momentsabstractIn this paper a method using local moments is proposed for watermarking 3D shapes modelled by mesh surfaces. The moments are used for separating two regions in selected areas. Two different approaches are adopted for 3D shape watermarking, according to the boundary surface between the two regions: by using parallel planes and bounding ellipsoids, respectively. The proposed algorithms change locations of 3D vertices by placing them in one of the two regions, according to a given code. Both algorithms use a vertex selection procedure followed by the embedding. The paper provides a statistical analysis of distortions caused in shapes by geometrical perturbations. Experimental results are provided for various 3D graphical objects. Adrian G. Bors |
ICIP | 1 |
| 2003 | Blind Source Separation USing Variational Expectation-Maximization Algorithm
Nikolaos Nasios, Adrian G. Bors |
CAIP | 2 |
| 2003 | Surface acquisition from single gray-scale imagesabstractIn this paper we show how a system for performing automatic surface model acquisition from single object views can be designed. The surface acquisition process is a two step one. Firstly, the surface normals are computed using a shape-from-shading algorithm. Secondly, the field of surface normals is integrated into a 3D surface. For the surface integration step, we have performed experiments with two alternatives. The first of these is a geometric surface integration algorithm. The second alternative comprises a graph-spectral surface integration algorithm. We present results on images of classical statues and provide a preliminary quantitative study. Antonio Robles-Kelly, Adrian G. Bors, Edwin R. Hancock |
ICIP (3) | 2 |
| 2003 | Variational Gaussian mixtures for blind source detectionabstractBayesian algorithms have lately been used in a large variety of applications. This paper proposes a new methodology for hyperparameter initialization in the Variational Bayes (VB) algorithm. We employ a dual expectation-maximization (EM) algorithm as the initialization stage in the VB-based learning. In the first stage, the EM algorithm is used on the given data set while the second EM algorithm is applied on distributions of parameters resulted from several runs of the first stage EM. The graphical model case study considered in this paper consists of a mixture of Gaussians. Appropriate conjugate prior distributions are considered for modelling the parameters. The proposed methodology is applied on blind source separation of modulated signals. Nikolaos Nasios, Adrian G. Bors |
SMC | 2 |
| 2003 | Terrain Analysis Using Radar Shape-from-ShadingabstractThis paper develops a maximum a posteriori (MAP) probability estimation framework for shape-from-shading (SFS) from synthetic aperture radar (SAR) images. The aim is to use this method to reconstruct surface topography from a single radar image of relatively complex terrain. Our MAP framework makes explicit how the recovery of local surface orientation depends on the whereabouts of terrain edge features and the available radar reflectance information. To apply the resulting process to real world radar data, we require probabilistic models for the appearance of terrain features and the relationship between the orientation of surface normals and the radar reflectance. We show that the SAR data can be modeled using a Rayleigh-Bessel distribution and use this distribution to develop a maximum likelihood algorithm for detecting and labeling terrain edge features. Moreover, we show how robust statistics can be used to estimate the characteristic parameters of this distribution. We also develop an empirical model for the SAR reflectance function. Using the reflectance model, we perform Lambertian correction so that a conventional SFS algorithm can be applied to the radar data. The initial surface normal direction is constrained to point in the direction of the nearest ridge or ravine feature. Each surface normal must fall within a conical envelope whose axis is in the direction of the radar illuminant. The extent of the envelope depends on the corrected radar reflectance and the variance of the radar signal statistics. We explore various ways of smoothing the field of surface normals using robust statistics. Finally, we show how to reconstruct the terrain surface from the smoothed field of surface normal vectors. The proposed algorithm is applied to various SAR data sets containing relatively complex terrain structure. Adrian G. Bors, Edwin R. Hancock, Richard C. Wilson 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2002 | Recovering height information from SAR images of terrainabstractWe suggest a new approach for recovering 3-D depth from a. field of surface normals. The surface normals are extracted from synthetic aperture radar (SAR) images using the radar reflectivity function and local statistical estimations. Reconstructing the 3-D surface from the flow of surface normals has an inexact solution due to missing information. A two stage algorithm is proposed for calculating the height information. In the first stage a site of unknown height and one of known height are chosen according to the orientation of the surface normal component in the horizontal plane. In the second stage a gradient updating algorithm is used for calculating the unknown height. We apply the proposed algorithm on SAR images of the terrain. Adrian G. Bors, Edwin R. Hancock |
ICIP (2) | 1 |
| 2002 | Watermarking 3D modelsabstractCopyright protection of graphical objects and models is important for protecting author rights in animation, multimedia, computer-aided design (CAD), virtual reality, medical imaging, etc. We suggest a blind watermarking algorithm for 3D models and objects. A string of bits, generated according to a key, is embedded in the geometrical structure of the graphical object by changing the locations of certain vertices. The criterion to choose these vertices ensures a minimal visibility of the distortions in the watermarked object. A bit encoding 1 is associated with positioning a vertex inside a volume modelled by the geometry of its neighbourhood, while a bit encoding 0 positions the vertex outside such a volume. The proposed watermarking algorithm is applied on various 3D models. Thomas Harte, Adrian G. Bors |
ICIP (3) | 2 |
| 2002 | Binary Morphological Shape-Based Interpolation Applied to 3D Tooth ReconstructionabstractIn this paper, we propose an interpolation algorithm using a mathematical morphology morphing approach. The aim of this algorithm is to reconstruct the n-dimensional object from a group of (n - 1)-dimensional sets representing sections of that object. The morphing transformation modifies pairs of consecutive sets such that they approach in shape and size. The interpolated set is achieved when the two consecutive sets are made idempotent by the morphing transformation. We prove the convergence of the morphological morphing. The entire object is modeled by successively interpolating a certain number of intermediary sets between each two consecutive given sets. We apply the interpolation algorithm for three-dimensional tooth reconstruction. Adrian G. Bors, Lefteris Kechagias, Ioannis Pitas |
IEEE Trans. Medical Imaging | 1 |
| 2001 | Shape-based interpolation using morphological morphingabstractWe propose an interpolation algorithm using a morphological morphing approach. The aim of this algorithm is to reconstruct an n-dimensional object from a group of (n-1)-dimensional sets representing object sections. The morphing transformation modifies consecutive sets so that they approach in shape and size. When the two morphed sets become idempotent we generate a new set. The entire object is modeled by successively interpolating a certain number of intermediary sets between each two consecutive initial sets. The interpolation algorithm is used for 3D tooth reconstruction. Adrian G. Bors, Lefteris Kechagias, Ioannis Pitas |
ICIP (2) | 1 |
| 2001 | Hierarchical watermarking depending on local constraintsabstractWe propose a new watermarking technique where the watermark is embedded according to two keys. The first key is used to embed a code bit in a block of pixels. The second is used to generate the whole sequence of code bits The watermark is embedded in spatial domain by adding or subtracting a random digital pattern to the given image signal. The embedding depth level depends on the spectral density distribution of the DCT coefficients and on the JPEG quantization table. The label consists of a set of bits which are embedded locally in a rectangular set of blocks and it is repeated over the entire image. After detecting individual bits, the retrieved label is verified by performing a XOR operation to the watermark code. Christopher D. Coltman, Adrian G. Bors |
ICIP (3) | 2 |
| 2001 | Hierarchical iterative eigendecomposition for motion segmentationabstractThis paper applies a new clustering approach for identifying and segmenting motion in image sequences. We estimate a matrix whose entries represent similarity probabilities between local motion estimates. We adopt a two step iterative algorithm which consists of a variant of the expectation maximization algorithm for segmenting regions with similar motion. The proposed algorithm updates cluster memberships in one step while it maximizes the expected log-likelihood in the second step. The performance of the algorithm is improved greatly by the use of modal sharpening. Antonio Robles-Kelly, Adrian G. Bors, Edwin R. Hancock |
ICIP (2) | 2 |
| 2001 | Projection distortion analysis for flattened image mosaicing from straight uniform generalized cylinders
William Puech, Adrian G. Bors, Ioannis Pitas, Jean-Marc Chassery |
Pattern Recognit. | 2 |
| 2000 | A Bayesian Framework for Radar Shape-from-ShadingabstractThis paper introduces a Bayesian approach to shape-from-shading ($F$) which is applied to terrain recovery in Synthetic Aperture Radar ($AR) images. The Bayesian model relates the recovery of 3-D shape information to the original 2-D radar intensity and to edges separating different topographic regions. First, we model the image amplitude distribution and the reflection function in $AR images. Using a maximum log-likelihood feature detector derived from the image statistics we identify the ridges and ravines in the terrain image. These topographic features are used to constrain the recovery of suace normals in the shapefrom -shading process. Finally, the suace normals are smoothed using rvbust statistics operators. Adrian G. Bors, Edwin R. Hancock, Richard C. Wilson 0001 |
CVPR | 1 |
| 2000 | Terrain Feature Identification by Modeling Radar Image StatisticsabstractWe propose a new statistical model for SAR images. According to this model, the SAR image amplitude follows a product of Rayleigh and Bessel functions. We derive the maximum likelihood feature detector for extracting terrain features from synthetic aperture radar (SAR) images. The terrain features are classified as ridges and ravines according to their statistical properties and surrounding neighborhood. These salient features are used as constraints for estimating the SAR terrain surface. Adrian G. Bors, Edwin R. Hancock, Richard C. Wilson 0001 |
ICIP | 1 |
| 2000 | Terrain Modeling in Synthetic Aperture Radar Images Using Shape-from-ShadingabstractWe introduce a new approach for recovering shape-from-shading (SFS) from synthetic aperture radar (SAR) images of the terrain. Three contributions are proposed: 1) we show how the direction of surface normals is constrained by the geometry of the radar reflectivity cone; 2) we show how topographic features can be used as boundary constraints on the recovered surface normals; and 3) the resulting field of surface normals is smoothed using robust statistics. Adrian G. Bors, Edwin R. Hancock, Richard C. Wilson 0001 |
ICPR | 1 |
| 2000 | Prediction and tracking of moving objects in image sequencesabstractWe employ a prediction model for moving object velocity and location estimation derived from Bayesian theory. The optical flow of a certain moving object depends on the history of its previous values. A joint optical flow estimation and moving object segmentation algorithm is used for the initialization of the tracking algorithm. The segmentation of the moving objects is determined by appropriately classifying the unlabeled and the occluding regions. Segmentation and optical flow tracking is used for predicting future frames. Adrian G. Bors, Ioannis Pitas |
IEEE Trans. Image Process. | 1 |
| 1999 | Object classification in 3-D images using alpha-trimmed mean radial basis function networkabstractWe propose a pattern classification based approach for simultaneous three-dimensional (3-D) object modeling and segmentation in image volumes. The 3-D objects are described as a set of overlapping ellipsoids. The segmentation relies on the geometrical model and graylevel statistics. The characteristic parameters of the ellipsoids and of the graylevel statistics are embedded in a radial basis function (RBF) network and they are found by means of unsupervised training. A new robust training algorithm for RBF networks based on alpha-trimmed mean statistics is employed in this study. The extension of the Hough transform algorithm in the 3-D space by employing a spherical coordinate system is used for ellipsoidal center estimation. We study the performance of the proposed algorithm and we present results when segmenting a stack of microscopy images. Adrian G. Bors, Ioannis Pitas |
IEEE Trans. Image Process. | 1 |
| 1999 | Multimodal decision-level fusion for person authenticationabstractThe use of clustering algorithms for decision-level data fusion is proposed. Person authentication results coming from several modalities (e.g., still image, speech), are combined by using fuzzy k-means (FKM) and fuzzy vector quantization (FVQ) algorithms, and a median radial basis function (MRBF) network. The quality measure of the modalities data is used for fuzzification. Two modifications of the FKM and FVQ algorithms, based on a fuzzy vector distance definition, are proposed to handle the fuzzy data and utilize the quality measure. Simulations show that fuzzy clustering algorithms have better performance compared to the classical clustering algorithms and other known fusion algorithms. MRBF has better performance especially when two modalities are combined. Moreover, the use of the quality via the proposed modified algorithms increases the performance of the fusion system. Vassilios Chatzis, Adrian G. Bors, Ioannis Pitas |
IEEE Trans. Syst. Man Cybern. Part A | 2 |
| 1998 | Motion and Segmentation Prediction in Image Sequences based on Moving Object Tracking
Adrian G. Bors, Ioannis Pitas |
ICIP (3) | 1 |
| 1998 | Optical flow estimation and moving object segmentation based on median radial basis function networkabstractVarious approaches have been proposed for simultaneous optical flow estimation and segmentation in image sequences. In this study, the moving scene is decomposed into different regions with respect to their motion, by means of a pattern recognition scheme. The inputs of the proposed scheme are the feature vectors representing still image and motion information. Each class corresponds to a moving object. The classifier employed is the median radial basis function (MRBF) neural network. An error criterion function derived from the probability estimation theory and expressed as a function of the moving scene model is used as the cost function. Each basis function is activated by a certain image region. Marginal median and median of the absolute deviations from the median (MAD) estimators are employed for estimating the basis function parameters. The image regions associated with the basis functions are merged by the output units in order to identify moving objects. Adrian G. Bors, Ioannis Pitas |
IEEE Trans. Image Process. | 1 |
| 1997 | Mosaicing of Flattened Images from Straight Homogeneous Generalized Cylinders
Adrian G. Bors, William Puech, Ioannis Pitas, Jean-Marc Chassery |
CAIP | 1 |
| 1997 | Perspective distortion analysis for mosaicing images painted on cylindrical surfacesabstractA set of monocular images of a curved painting is taken from different viewpoints around its curved surface. After deriving the surface localization in the camera coordinate system we backproject the image on the curved surface and we flatten it. We analyze the perspective distortions of the scene in the case when it is mapped on a cylindrical surface. Based on the result of this analysis we derive the necessary number of views in order to represent the entire scene depicted on a cylindrical surface. We employ a matching-based mosaicing method for reconstructing the scene from the curved surface. The proposed method is appropriate to be used for painting reconstruction. Adrian G. Bors, William Puech, Ioannis Pitas, Jean-Marc Chassery |
ICASSP | 1 |
| 1996 | Image watermarking using DCT domain constraintsabstractWatermarking algorithms are used for image copyright protection. The algorithms proposed select certain blocks in the image based on a Gaussian network classifier. The pixel values of the selected blocks are modified such that their discrete cosine transform (DCT) coefficients fulfil a constraint imposed by the watermark code. Two different constraints are considered. The first approach consists of embedding a linear constraint among selected DCT coefficients and the second one defines circular detection regions in the DCT domain. A rule for generating the DCT parameters of distinct watermarks is provided. The watermarks embedded by the proposed algorithms are resistant to JPEG compression. Adrian G. Bors, Ioannis Pitas |
ICIP (3) | 1 |
| 1996 | Mosaicing of paintings on curved surfacesabstractThe paper presents an approach for reconstructing images painted on curved surfaces. A set of monocular images is taken from different viewpoints in order to mosaic and represent the entire scene. By using a priori knowledge about the support surface of the picture, we derive the surface localization in the camera coordinate system. An automatic mosaicing method is applied on the patterned images in order to obtain the complete scene. The mosaiced scene is visualized on a new synthetic surface by a mapping procedure. William Puech, Jean-Marc Chassery, Adrian G. Bors, Ioannis Pitas |
WACV | 3 |
| 1996 | Median radial basis function neural networkabstractRadial basis functions (RBFs) consist of a two-layer neural network, where each hidden unit implements a kernel function. Each kernel is associated with an activation region from the input space and its output is fed to an output unit. In order to find the parameters of a neural network which embeds this structure we take into consideration two different statistical approaches. The first approach uses classical estimation in the learning stage and it is based on the learning vector quantization algorithm and its second-order statistics extension. After the presentation of this approach, we introduce the median radial basis function (MRBF) algorithm based on robust estimation of the hidden unit parameters. The proposed algorithm employs the marginal median for kernel location estimation and the median of the absolute deviations for the scale parameter estimation. A histogram-based fast implementation is provided for the MRBF algorithm. The theoretical performance of the two training algorithms is comparatively evaluated when estimating the network weights. The network is applied in pattern classification problems and in optical flow segmentation. Adrian G. Bors, Ioannis Pitas |
IEEE Trans. Neural Networks | 1 |
| 1995 | Segmentation and Estimation of the Optical Flow
Adrian G. Bors, Ioannis Pitas |
CAIP | 1 |