EDBT 2026 Demo / reviewers in the wild / expert
Gianfranco Doretto
dblp:61/3857
· DBLP profile ↗
53ranked-venue papers
7as first author
23since 2021 · last 2026
0000-0002-8921-6646ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 36 · 7 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 5 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | iStructTab: Structured Feature Sequencing for Multimodal Learning of Image and Tabular Data
Al-Zadid Sultan Bin Habib, Md Younus Ahamed, Prashnna Gyawali, Gianfranco Doretto, Donald A. Adjeroh |
ICPR (8) | 4 |
| 2025 | Improving Accuracy and Generalization for Efficient Visual TrackingabstractEfficient visual trackers overfit to their training distributions and lack generalization abilities, resulting in them performing well on their respective in-distribution (ID) test sets and not as well on out-of-distribution (OOD) sequences, imposing limitations to their deployment in-the-wild under constrained resources. We introduce Siam-ABC, a highly efficient Siamese tracker that significantly improves tracking performance, even on OOD sequences. SiamABC takes advantage of new architectural designs in the way it bridges the dynamic variability of the target, and of new losses for training. Also, it directly addresses OOD tracking generalization by including a fast backward-free dynamic test-time adaptation method that continuously adapts the model according to the dynamic visual changes of the target. Our extensive experiments suggest that Siam-ABC shows remarkable performance gains in OOD sets while maintaining accurate performance on the ID benchmarks. SiamABC outperforms MixFormerV2-S by 7.6% on the OOD AVisT benchmark while being 3x faster (100 FPS) on a CPU. Our code and models are available at https://wvuvl.github.io/SiamABC/. Ram J. Zaveri, Shivang Patel, Gianfranco Doretto |
WACV | 4 |
| 2025 | ItpCtrl-AI: End-to-end interpretable and controllable artificial intelligence by modeling radiologists' intentions
Trong-Thang Pham, Jacob Brecheisen, Carol C. Wu, Hien Van Nguyen, Zhigang Deng 0001, Donald A. Adjeroh, Gianfranco Doretto, Arabinda Choudhary, T. Hoang Ngan Le |
Artif. Intell. Medicine | 7 |
| 2024 | FG-CXR: A Radiologist-Aligned Gaze Dataset for Enhancing Interpretability in Chest X-Ray Report Generation
Trong-Thang Pham, Ngoc-Vuong Ho, Nhat-Tan Bui, Thinh Phan, Brijesh Patel 0001, Donald A. Adjeroh, Gianfranco Doretto, Anh Nguyen 0003, Carol C. Wu, T. Hoang Ngan Le |
ACCV (6) | 7 |
| 2024 | A Framework for Evaluating Model Trustworthiness in Classification of Very High Resolution Histopathology ImagesabstractIn computer vision, one approach to explaining a deep learning model’s decision is to show regions of visual evidence upon which the model makes a decision. Typically, this evidence is represented in the form of a saliency map which conveys how much an image region is contributing to the model’s decision. For a model to be trustworthy, it is expected that this saliency region should provide relevant information. In this work, we use model "trustworthiness" or "rationale" to describe how much relevant information the model is using to determine the image class. For medical images, this information connects to biological relevance. For very high resolution histopathology image applications, such as gigapixel whole-slide image classification, where patch-based multiple-instance based learning approach is taken to determine the patch label, this biological relevance has to be determined both at the patch and the image level. In this work, we present a novel patch-based model trustworthiness evaluation framework for very high resolution histopathology images. Our trustworthiness framework takes two approaches: spatial overlap based and feature based evaluation. For the overlap based approach, we check overlap with the annotation provided with the database to see if they have biological relevance, since for tumor positive patches only high probability regions from within the annotated regions are likely to be relevant. For feature based approach, we train an interpretability model using the sub-patches of the training set, extract features and cluster them. Then based on the distance from these clusters we determine if there is any biological rationale behind the prediction. Finally, we propose four patch-level and four image-level rationale metrics that evaluate the biological relevance of the information used by the classifier to decide on the patch class. Our experiment using the CAMELYON16 dataset shows the efficacy of this approach for model trustworthiness evaluation and explainability. Mohammad Iqbal Nouyed, Gianfranco Doretto, Donald A. Adjeroh |
BIBM | 2 |
| 2024 | Multi-label Classification using Self-Supervised Learning: Addressing Class Inter-Dependency and Data ImbalanceabstractDeveloping multi-label classification models under significant class imbalance, and when annotating data requires expert-level knowledge remains a major challenge. Additionally, interdependency and correlation among labels are common in multi-label problems. In this work, we introduce a novel framework to address these challenges. Our approach extracts robust discriminative features from unlabeled data through self-supervised contrastive learning and uses an adaptive data augmentation mechanism (ACBA) to balance the dataset. Independent binary classifiers are trained for each class, using a new custom Focal Weighted Cross-Entropy (FWCE) loss function to focus on hard-to-classify examples. A correlation learning module then refines predictions by integrating statistical and domain-specific knowledge. Finally, a meta-learner, employing a Gated Recurrent Unit (GRU) and multi-head attention, identifies complex relationships between classes, even for those that rarely occur together. We used the detection of thoracic diseases using chest X-rays, a domain with a major class imbalance and highly associated labels, to validate our approach. Our findings demonstrate the potential of our method to apply to other medical and non-medical imaging scenarios with similar multi-label classification problems. Ghazaleh Mirzaee, Gianfranco Doretto, Donald A. Adjeroh |
ICMLA | 2 |
| 2024 | TabSeq: A Framework for Deep Learning on Tabular Data via Sequential Ordering
Al-Zadid Sultan Bin Habib, Kesheng Wang, Mary-Anne Hartley, Gianfranco Doretto, Donald A. Adjeroh |
ICPR (4) | 4 |
| 2024 | Efficient Classification of Histopathology Images Using Highly Imbalanced Data
Mohammad Iqbal Nouyed, Mary-Anne Hartley, Gianfranco Doretto, Donald A. Adjeroh |
ICPR (2) | 3 |
| 2024 | Open-Fusion: Real-time Open-Vocabulary 3D Mapping and Queryable Scene RepresentationabstractPrecise 3D environmental mapping with semantics is essential in robotics. Existing methods often rely on pre-defined concepts during training or are time-intensive when generating semantic maps. This paper presents Open-Fusion, an approach for real-time open-vocabulary 3D mapping and queryable scene representation using RGB-D data. Open-Fusion harnesses the power of a pretrained vision-language foundation model (VLFM) for open-set semantic comprehension and employs the Truncated Signed Distance Function (TSDF) for swift 3D scene reconstruction. By leveraging the VLFM, we extract region-based embeddings and their associated confidence maps. These are then integrated with the 3D knowledge from TSDF using an enhanced Hungarian-based feature-matching mechanism. In particular, Open-Fusion delivers outstanding annotation-free 3D segmentation for open vocabulary query without the need for additional 3D training. Benchmark tests on the ScanNet dataset against leading zero-shot methods highlight Open-Fusion’s superiority. Furthermore, it seamlessly combines the strengths of region-based VLFM and TSDF, facilitating real-time 3D scene comprehension that includes object concepts and open-world semantics. We encourage the readers to view the demos on our project page: https://uark-aicv.github.io/OpenFusion Kashu Yamazaki, Taisei Hanyu, Viet-Khoa Vo-Ho, Thang Pham, Gianfranco Doretto, Anh Nguyen 0003, T. Hoang Ngan Le |
ICRA | 6 |
| 2024 | ZEETAD: Adapting Pretrained Vision-Language Model for Zero-Shot End-to-End Temporal Action DetectionabstractTemporal action detection (TAD) involves the localization and classification of action instances within untrimmed videos. While standard TAD follows fully supervised learning with closed-set setting on large training data, recent zero-shot TAD methods showcase the promising open-set setting by leveraging large-scale contrastive visual-language (ViL) pretrained models. However, existing zero-shot TAD methods have limitations on how to properly construct the strong relationship between two interdependent tasks of localization and classification and adapt ViL model to video understanding. In this work, we present ZEE-TAD, featuring two modules: dual-localization and zero-shot proposal classification. The former is a Transformer-based module that detects action events while selectively collecting crucial semantic embeddings for later recognition. The latter one, CLIP-based module, generates semantic embeddings from text and frame inputs for each temporal unit. Additionally, we enhance discriminative capability on unseen classes by minimally updating the frozen CLIP encoder with lightweight adapters. Extensive experiments on THUMOS14 and ActivityNet-1.3 datasets demonstrate our approach’s superior performance in zero-shot TAD and effective knowledge transfer from ViL models to unseen action categories. Code is available at https: //github.com/UARK-AICV/ZEETAD. Thinh Phan, Viet-Khoa Vo-Ho, Duy Le 0004, Gianfranco Doretto, Donald A. Adjeroh, T. Hoang Ngan Le |
WACV | 4 |
| 2024 | Online continual decoding of streaming EEG signal with a balanced and informative memory buffer
Tiehang Duan, Zhenyi Wang 0001, Fang Li 0011, Gianfranco Doretto, Donald A. Adjeroh, Yiyi Yin, Cui Tao |
Neural Networks | 4 |
| 2023 | Replay with Stochastic Neural Transformation for Online Continual EEG ClassificationabstractBrain computer interface (BCI) systems used for clinical assistance purposes such as wheelchair control require decoding of streaming brain signals i.e. electroencephalography (EEG) signals over a long period of time with subject shift in the middle. Numerous challenges arise during this online continual brain signal decoding process: 1) the EEG decoder needs to deal with streaming EEG signals from sequentially arriving subjects, with no data available beforehand for large-scale pretraining; 2) the EEG decoder should avoid catastrophic forgetting on previous subjects after learning on a new subject; 3) the EEG decoder should perform well on noisy signals with high variance across subjects. We proposed a principled replay-based approach for this general decoding scenario, forming a bi-level optimization framework with stochastic neural transformation for dynamic memory evolution, making them representative in feature space and encouraging the model to generalize well. The evolved signal segments are stored and replayed during later decoding stages to achieve optimal model performance on all previous subjects. The stochastic neural transformation performed in inner sup of bi-level optimization significantly enhances the diversity of stored signal segments and improves model robustness during online continual decoding. We perform detailed theoretical analysis on model’s generalization ability in addition to the empirical evaluations. We construct multiple new benchmarks to mimic real-world online sequential EEG decoding scenarios with underlying subject shifts. The extensive evaluation of the proposed approach shows it outperforms related strong baselines by a large margin. Tiehang Duan, Zhenyi Wang 0001, Gianfranco Doretto, Fang Li 0011, Cui Tao, Donald A. Adjeroh |
BIBM | 3 |
| 2023 | Distributionally Robust Cross Subject EEG DecodingabstractRecently, deep learning has shown to be effective for Electroencephalography (EEG) decoding tasks. Yet, its performance can be negatively influenced by two key factors: 1) the high variance and different types of corruption that are inherent in the signal, 2) the EEG datasets are usually relatively small given the acquisition cost, annotation cost and amount of effort needed. Data augmentation approaches for alleviation of this problem have been empirically studied, with augmentation operations on spatial domain, time domain or frequency domain handcrafted based on expertise of domain knowledge. In this work, we propose a principled approach to perform dynamic evolution on the data for improvement of decoding robustness. The approach is based on distributionally robust optimization and achieves robustness by optimizing on a family of evolved data distributions instead of the single training data distribution. We derived a general data evolution framework based on Wasserstein gradient flow (WGF) and provides two different forms of evolution within the framework. Intuitively, the evolution process helps the EEG decoder to learn more robust and diverse features. It is worth mentioning that the proposed approach can be readily integrated with other data augmentation approaches for further improvements. We performed extensive experiments on the proposed approach and tested its performance on different types of corrupted EEG signals. The model significantly outperforms competitive baselines on challenging decoding scenarios. Tiehang Duan, Zhenyi Wang 0001, Gianfranco Doretto, Fang Li 0011, Cui Tao, Donald A. Adjeroh |
ECAI | 3 |
| 2023 | AG-ReID 2023: Aerial-Ground Person Re-identification Challenge ResultsabstractPerson re-identification (Re-ID) on aerial-ground platforms has emerged as an intriguing topic within computer vision, presenting a plethora of unique challenges. Highflying altitudes of aerial cameras make persons appear differently in terms of viewpoints, poses, and resolution compared to the images of the same person viewed from ground cameras. Despite its potential, few algorithms have been developed for person re-identification on aerial-ground data, mainly due to the absence of comprehensive datasets. In response, we have collected a large-scale dataset and organized the Aerial-Ground person Re-IDentification Challenge (AG-ReID2023) to foster advancements in the field. The dataset comprises 100,502 images with 1,615 unique identities, including 51,530 training images featuring 807 identities. The test set is divided into two subsets: Aerial to Ground (808 ids, 4,348 query images, 19,259 gallery images) and Ground to Aerial (808 ids, 4,151 query images, 21,214 gallery images). In addition, we manually annotate individuals with their matching IDs across cameras and provide 15 soft attribute labels. The AG-ReID2023 Challenge in conjunction with the 7thIEEE International Joint Conference on Biometrics (IJCB) has garnered interest from numerous institutes, resulting in the submission of five distinct algorithms. We provide an in-depth examination of the evaluation outcomes and present our findings from the contest. For additional details, kindly refer to the official website1.1https://agreid23.github.io. Kien Nguyen Thanh, Clinton Fookes, Sridha Sridharan, Feng Liu 0037, Xiaoming Liu 0002, Arun Ross, Dana Michalski, Debayan Deb, Mahak Kothari, Manisha Saini, Dawei Du, Scott McCloskey, Gabriel Bertocco, Fernanda A. Andaló, Terrance E. Boult, Anderson Rocha 0001, Haidong Zhu, Zhaoheng Zheng, Ramakant Nevatia, Zaigham A. Randhawa, Sinan Sabri, Gianfranco Doretto |
IJCB | 23 |
| 2023 | More Synergy, Less Redundancy: Exploiting Joint Mutual Information for Self-Supervised LearningabstractSelf-supervised learning (SSL) is now a serious competitor for supervised learning, even though it does not require data annotation. Several baselines have attempted to make SSL models exploit information about data distribution, and less dependent on the augmentation effect. However, there is no clear consensus on whether maximizing or minimizing the mutual information between representations of augmentation views practically contribute to improvement or degradation in performance of SSL models. This paper is a fundamental work where, we investigate the role of mutual information in SSL, and reformulate the problem of SSL in the context of a new perspective on mutual information. To this end, we consider joint mutual information from the perspective of partial information decomposition (PID) as a key step in reliable multivariate information measurement. PID enables us to decompose joint mutual information into three important components, namely, unique information, redundant information and synergistic information. Our framework aims for minimizing the redundant information between views and the desired target representation while maximizing the synergistic information at the same time. Our experiments lead to a re-calibration of two redundancy reduction baselines, and a proposal for a new SSL training protocol. Experimental results on multiple datasets and two downstream tasks show the effectiveness of this framework. Salman Mohamadi, Gianfranco Doretto, Donald A. Adjeroh |
ICIP | 2 |
| 2023 | A Framework for Token-Based Scene-Classification of Remote Sensing ImagesabstractRemote sensing images often come in very high resolutions. For instance, large images of resolution 1000×1000 are quite common in remote sensing applications. Thus, one key challenge is how to effectively analysis such very high resolution images using low computational resources, while maintain a reasonable performance as appropriate to the specific problem domain, such as scene classification, object segmentation, or object detection. In this work, we propose a framework for addressing this challenge, with focus on the specific problem of scene-classification for remote-sensed aerial imagery. Various approaches have been proposed for scene-based classification of remote sensing imagery [1] , [2] , [3] . For a survey of approaches and challenges, see [4] . See also [5] . Mohammad Iqbal Nouyed, Gianfranco Doretto, Donald A. Adjeroh |
IGARSS | 2 |
| 2023 | CellTranspose: Few-shot Domain Adaptation for Cellular Instance SegmentationabstractAutomated cellular instance segmentation is a process utilized for accelerating biological research for the past two decades, and recent advancements have produced higher quality results with less effort from the biologist. Most current endeavors focus on completely cutting the researcher out of the picture by generating highly generalized models. However, these models invariably fail when faced with novel data, distributed differently than the ones used for training. Rather than approaching the problem with methods that presume the availability of large amounts of target data and computing power for retraining, in this work we address the even greater challenge of designing an approach that requires minimal amounts of new annotated data as well as training time. We do so by designing specialized contrastive losses that leverage the few annotated samples very efficiently. A large set of results show that 3 to 5 annotations lead to models with accuracy that: 1) significantly mitigate the covariate shift effects; 2) matches or surpasses other adaptation methods; 3) even approaches methods that have been fully retrained on the target distribution. The adaptation training is only a few minutes, paving a path towards a balance between model performance, computing requirements and expert-level annotation needs. Matthew R. Keaton, Ram J. Zaveri, Gianfranco Doretto |
WACV | 3 |
| 2023 | FUSSL: Fuzzy Uncertain Self Supervised LearningabstractSelf supervised learning (SSL) has become a very successful technique to harness the power of unlabeled data, with no annotation effort. A number of developed approaches are evolving with the goal of outperforming supervised alternatives, which have been relatively successful. Similar to some other disciplines in deep representation learning, one main issue in SSL is robustness of the approaches under different settings. In this paper, for the first time, we recognise the fundamental limits of SSL coming from the use of a single-supervisory signal. To address this limitation, we leverage the power of uncertainty representation to devise a robust and general standard hierarchical learning/training protocol for any SSL baseline, regardless of their assumptions and approaches. Essentially, using the information bottleneck principle, we decompose feature learning into a two-stage training procedure, each with a distinct supervision signal. This double supervision approach is captured in two key steps: 1) invariance enforcement to data augmentation, and 2) fuzzy pseudo labeling (both hard and soft annotation). This simple, yet, effective protocol which enables cross-class/cluster feature learning, is instantiated via an initial training of an ensemble of models through invariance enforcement to data augmentation as first training phase, and then assigning fuzzy labels to the original samples for the second training phase. We consider multiple alternative scenarios with double supervision and evaluate the effectiveness of our approach on recent baselines, covering four different SSL paradigms, including geometrical, contrastive, non-contrastive, and hard/soft whitening (redundancy reduction) baselines. We performed extensive experiments under multiple settings to show that the proposed training protocol consistently improves the performance of the former baselines, independent of their respective underlying principles. Salman Mohamadi, Gianfranco Doretto, Donald A. Adjeroh |
WACV | 2 |
| 2022 | Deep Active Ensemble Sampling for Image Classification
Salman Mohamadi, Gianfranco Doretto, Donald A. Adjeroh |
ACCV (7) | 2 |
| 2022 | Efficient Classification of Very High Resolution Histopathological ImagesabstractOver the years, deep learning approaches have shown significant improvement in various image understanding tasks. However, analysis of high resolution images still remains a major challenge. Apart from the huge computational resources required for such images, the large image sizes make it difficult to extract effective contextual information needed for important tasks, such as classification, segmentation, or clustering of such images. In this work, we address the challenge of high resolution image classification u sing a new discriminative patch selection approach. We embed our patch selection approach inside a novel classification framework, supporting potential use of different pre-trained learning models. We show results on a high resolution image dataset, namely, gigapixel whole slide tissue images for cancer tumors. We demonstrate the performance of the proposed approaches using comparative analysis with state-of-the art methods on this dataset. Mohammad Iqbal Nouyed, Gianfranco Doretto, Donald A. Adjeroh |
BIBM | 2 |
| 2021 | Detecting Drug-Drug Interactions using Protein Sequence-Structure Similarity NetworksabstractAdverse drug events represent a key challenge in public health, especially with respect to drug safety profiling and drug surveillance. Drug-drug interactions represent one of the most popular types of adverse drug events. Most computational approaches to this problem have used different types of data, such as drug chemical structure, information about protein targets, side effects, pathways, etc to predict potential interactions between drugs. In this work, we study the question of whether using just genetic information about the drugs can provide significant information about the potential safety profile for a given drug. We propose a novel neural network model to predict adverse drug events using only data about the protein sequence and protein structure associated with the drug targets. We compare the results with those from the state-of-the-art methods on this problem. Our results show that the proposed method is quite competitive, at times outperforming the state-of-the-art. Saminur Islam, Ahmed Abbasi, Nitin Agarwal 0001, Wanhong Zheng, Gianfranco Doretto, Donald A. Adjeroh |
BIBM | 5 |
| 2021 | Human Age Estimation from Gene Expression Data using Artificial Neural NetworksabstractThe study of signatures of aging in terms of genomic biomarkers can be uniquely helpful in understanding the mechanisms of aging and developing models to accurately predict the age. Prior studies have employed gene expression and DNA methylation data aiming at accurate prediction of age. In this line, we propose a new framework for human age estimation using information from human dermal fibroblast gene expression data. First, we propose a new spatial representation as well as a data augmentation approach for gene expression data. Next in order to predict the age, we design an architecture of neural network and apply it to this new representation of the original and augmented data, as an ensemble classification approach. Our experimental results suggest the superiority of the proposed framework over state-of-the-art age estimation methods using DNA methylation and gene expression data. Salman Mohamadi, Nasser M. Nasrabadi, Gianfranco Doretto, Donald A. Adjeroh |
BIBM | 3 |
| 2021 | Deep learning for biological age estimationabstractModern machine learning techniques (such as deep learning) offer immense opportunities in the field of human biological aging research. Aging is a complex process, experienced by all living organisms. While traditional machine learning and data mining approaches are still popular in aging research, they typically need feature engineering or feature extraction for robust performance. Explicit feature engineering represents a major challenge, as it requires significant domain knowledge. The latest advances in deep learning provide a paradigm shift in eliciting meaningful knowledge from complex data without performing explicit feature engineering. In this article, we review the recent literature on applying deep learning in biological age estimation. We consider the current data modalities that have been used to study aging and the deep learning architectures that have been applied. We identify four broad classes of measures to quantify the performance of algorithms for biological age estimation and based on these evaluate the current approaches. The paper concludes with a brief discussion on possible future directions in biological aging research using deep learning. This study has significant potentials for improving our understanding of the health status of individuals, for instance, based on their physical activities, blood samples and body shapes. Thus, the results of the study could have implications in different health care settings, from palliative care to public health. Syed Ashiqur Rahman, Peter Giacobbi, Lee Pyles, Charles J. Mullett, Gianfranco Doretto, Donald A. Adjeroh |
Briefings Bioinform. | 5 |
| 2020 | Adversarial Latent AutoencodersabstractAutoencoder networks are unsupervised approaches aiming at combining generative and representational properties by learning simultaneously an encoder-generator map. Although studied extensively, the issues of whether they have the same generative power of GANs, or learn disentangled representations, have not been fully addressed. We introduce an autoencoder that tackles these issues jointly, which we call Adversarial Latent Autoencoder (ALAE). It is a general architecture that can leverage recent improvements on GAN training procedures. We designed two autoencoders: one based on a MLP encoder, and another based on a StyleGAN generator, which we call StyleALAE. We verify the disentanglement properties of both architectures. We show that StyleALAE can not only generate 1024x1024 face images with comparable quality of StyleGAN, but at the same resolution can also produce face reconstructions and manipulations based on real images. This makes ALAE the first autoencoder able to compare with, and go beyond the capabilities of a generator-only type of architecture. Stanislav Pidhorskyi, Donald A. Adjeroh, Gianfranco Doretto |
CVPR | 3 |
| 2018 | Deep Supervised Hashing with Spherical Embedding
Stanislav Pidhorskyi, Quinn Jones, Saeid Motiian, Donald A. Adjeroh, Gianfranco Doretto |
ACCV (4) | 5 |
| 2018 | Generative Probabilistic Novelty Detection with Adversarial AutoencodersabstractNovelty detection is the problem of identifying whether a new data point is considered to be an inlier or an outlier. We assume that training data is available to describe only the inlier distribution. Recent approaches primarily leverage deep encoder-decoder network architectures to compute a reconstruction error that is used to either compute a novelty score or to train a one-class classifier. While we too leverage a novel network of that kind, we take a probabilistic approach and effectively compute how likely it is that a sample was generated by the inlier distribution. We achieve this with two main contributions. First, we make the computation of the novelty probability feasible because we linearize the parameterized manifold capturing the underlying structure of the inlier distribution, and show how the probability factorizes and can be computed with respect to local coordinates of the manifold tangent space. Second, we improve the training of the autoencoder network. An extensive set of results show that the approach achieves state-of-the-art performance on several benchmark datasets. Stanislav Pidhorskyi, Ranya Almohsen, Gianfranco Doretto |
NeurIPS | 3 |
| 2017 | Unified Deep Supervised Domain Adaptation and GeneralizationabstractThis work provides a unified framework for addressing the problem of visual supervised domain adaptation and generalization with deep models. The main idea is to exploit the Siamese architecture to learn an embedding subspace that is discriminative, and where mapped visual domains are semantically aligned and yet maximally separated. The supervised setting becomes attractive especially when only few target data samples need to be labeled. In this scenario, alignment and separation of semantic probability distributions is difficult because of the lack of data. We found that by reverting to point-wise surrogates of distribution distances and similarities provides an effective solution. In addition, the approach has a high “speed” of adaptation, which requires an extremely low number of labeled target training samples, even one per category can be effective. The approach is extended to domain generalization. For both applications the experiments show very promising results. Saeid Motiian, Marco Piccirilli, Donald A. Adjeroh, Gianfranco Doretto |
ICCV | 4 |
| 2017 | Few-Shot Adversarial Domain AdaptationabstractThis work provides a framework for addressing the problem of supervised domain adaptation with deep models. The main idea is to exploit adversarial learning to learn an embedded subspace that simultaneously maximizes the confusion between two domains while semantically aligning their embedding. The supervised setting becomes attractive especially when there are only a few target data samples that need to be labeled. In this few-shot learning scenario, alignment and separation of semantic probability distributions is difficult because of the lack of data. We found that by carefully designing a training scheme whereby the typical binary adversarial discriminator is augmented to distinguish between four different classes, it is possible to effectively address the supervised adaptation problem. In addition, the approach has a high “speed” of adaptation, i.e. it requires an extremely low number of labeled target training samples, even one per category can be effective. We then extensively compare this approach to the state of the art in domain adaptation in two experiments: one using datasets for handwritten digit recognition, and one using datasets for visual object recognition. Saeid Motiian, Quinn Jones, Seyed Mehdi Iranmanesh, Gianfranco Doretto |
NIPS | 4 |
| 2017 | Online Human Interaction Detection and Recognition With Multiple CamerasabstractWe address the problem of detecting and recognizing online the occurrence of human interactions as seen by a network of multiple cameras. We represent interactions by forming temporal trajectories, coupling together the body motion of each individual and their proximity relationships with others, and also sound whenever available. Such trajectories are modeled with kernel state-space (KSS) models. Their advantage is being suitable for the online interaction detection, recognition, and also for fusing information from multiple cameras, while enabling a fast implementation based on online recursive updates. For recognition, in order to compare interaction trajectories in the space of KSS models, we design so-called pairwise kernels with a special symmetry. For detection, we exploit the geometry of linear operators in Hilbert space, and extend to KSS models the concept of parity space, originally defined for linear models. For fusion, we combine KSS models with kernel construction and multiview learning techniques. We extensively evaluate the approach on four single view publicly available data sets, and we also introduce, and will make public, a new challenging human interactions data set that we have collected using a network of three cameras. The results show that the approach holds promise to become an effective building block for the analysis of real-time human behavior from multiple cameras. Saeid Motiian, Farzad Siyahjani, Ranya Almohsen, Gianfranco Doretto |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2016 | Information Bottleneck Learning Using Privileged Information for Visual RecognitionabstractWe explore the visual recognition problem from a main data view when an auxiliary data view is available during training. This is important because it allows improving the training of visual classifiers when paired additional data is cheaply available, and it improves the recognition from multi-view data when there is a missing view at testing time. The problem is challenging because of the intrinsic asymmetry caused by the missing auxiliary view during testing. We account for such view during training by extending the information bottleneck method, and by combining it with risk minimization. In this way, we establish an information theoretic principle for leaning any type of visual classifier under this particular setting. We use this principle to design a large-margin classifier with an efficient optimization in the primal space. We extensively compare our method with the state-of-the-art on different visual recognition datasets, and with different types of auxiliary data, and show that the proposed framework has a very promising potential. Saeid Motiian, Marco Piccirilli, Donald A. Adjeroh, Gianfranco Doretto |
CVPR | 4 |
| 2016 | Information Bottleneck Domain Adaptation with Privileged Information for Visual Recognition
Saeid Motiian, Gianfranco Doretto |
ECCV (7) | 2 |
| 2015 | A Supervised Low-Rank Method for Learning Invariant SubspacesabstractSparse representation and low-rank matrix decomposition approaches have been successfully applied to several computer vision problems. They build a generative representation of the data, which often requires complex training as well as testing to be robust against data variations induced by nuisance factors. We introduce the invariant components, a discriminative representation invariant to nuisance factors, because it spans subspaces orthogonal to the space where nuisance factors are defined. This allows developing a framework based on geometry that ensures a uniform inter-class separation, and a very efficient and robust classification based on simple nearest neighbor. In addition, we show how the approach is equivalent to a local metric learning, where the local metrics (one for each class) are learned jointly, rather than independently, thus avoiding the risk of overfitting without the need for additional regularization. We evaluated the approach for face recognition with highly corrupted training and testing data, obtaining very promising results. Farzad Siyahjani, Ranya Almohsen, Sinan Sabri, Gianfranco Doretto |
ICCV | 4 |
| 2014 | Online geometric human interaction segmentation and recognitionabstractWe address the problem of online temporal segmentation and recognition of human interactions in video sequences. The complexity of the high-dimensional data variability representing interactions is handled by combining kernel methods with linear models, giving rise to kernel regression and kernel state space models. By exploiting the geometry of linear operators in Hilbert space, we show how the concept of parity space, defined for linear models, generalizes to the kernellized extensions. This provides a powerful and flexible framework for online temporal segmentation and recognition. We extensively evaluate the approach on a publicly available dataset, and on a new challenging human interactions dataset that we have collected. The results show that the approach holds the promise to become an effective building block for the analysis in real-time of human behavior. Farzad Siyahjani, Saeid Motiian, Harika Bharthavarapu, Sajid Sharlemin, Gianfranco Doretto |
ICME | 5 |
| 2012 | Learning a Context Aware Dictionary for Sparse Representation
Farzad Siyahjani, Gianfranco Doretto |
ACCV (2) | 2 |
| 2012 | M-VIVIE: A multi-thread video indexer via identity extraction
Maria De Marsico, Gianfranco Doretto, Daniel Riccio |
Pattern Recognit. Lett. | 2 |
| 2010 | Region moments: Fast invariant descriptors for detecting small image structuresabstractThis paper presents region moments, a class of appearance descriptors based on image moments applied to a pool of image features. A careful design of the moments and the image features, makes the descriptors scale and rotation invariant, and therefore suitable for vehicle detection from aerial video, where targets appear at different scales and orientations. Region moments are linearly related to the image features. Thus, comparing descriptors by computing costly geodesic distances and non-linear classifiers can be avoided, because Euclidean geometry and linear classifiers are still effective. The descriptor computation is made efficient by designing a fast procedure based on the integral representation. An extensive comparison between region moments and the region covariance descriptors, reports theoretical, qualitative, and quantitative differences among them, with a clear advantage of the region moments, when used for detecting small image structures, such as vehicles in aerial video. The proposed descriptors hold the promise to become an effective building block in other applications. Gianfranco Doretto |
CVPR | 1 |
| 2010 | Boosting for transfer learning with multiple sourcesabstractTransfer learning allows leveraging the knowledge of source domains, available a priori, to help training a classifier for a target domain, where the available data is scarce. The effectiveness of the transfer is affected by the relationship between source and target. Rather than improving the learning, brute force leveraging of a source poorly related to the target may decrease the classifier performance. One strategy to reduce this negative transfer is to import knowledge from multiple sources to increase the chance of finding one source closely related to the target. This work extends the boosting framework for transferring knowledge from multiple sources. Two new algorithms, MultiSource-TrAdaBoost, and TaskTrAdaBoost, are introduced, analyzed, and applied for object category recognition and specific object detection. The experiments demonstrate their improved performance by greatly reducing the negative transfer as the number of sources increases. TaskTrAdaBoost is a fast algorithm enabling rapid retraining over new targets. Gianfranco Doretto |
CVPR | 2 |
| 2009 | A model change detection approach to dynamic scene modelingabstractIn this work we propose a dynamic scene model to provide information about the presence of salient motion in the scene, and that could be used for focusing the attention of a pan/tilt/zoom camera, or for background modeling purposes. Rather than proposing a set of saliency detectors, we define what we mean by salient motion, and propose a precise model for it. Detecting salient motion becomes equivalent to detecting a model change. We derive optimal online procedures to solve this problem, which enable a very fast implementation. Promising results show that our model can effectively detect salient motion even in severely cluttered scenes, and while a camera is panning and tilting. Seon Joo Kim, Gianfranco Doretto, Jens Rittscher, Peter H. Tu, Nils Krahnstoever, Marc Pollefeys |
AVSS | 2 |
| 2009 | Intelligent Video for Protecting Crowded Sports VenuesabstractIntelligent video in urban settings can be challenging due the presence of crowds, clutter, poor camera placement and continuously changing light conditions. The surveillance of sports venues is particularly difficult, because thousands of people can enter or exit a venue in short periods of time. This paper presents a case study of successfully monitoring a sports venue using a multi-camera multi-target tracking system. The system performed site-wide tracking throughout a network of calibrated cameras and was able to accurately track thousands of people in real-time under challenging conditions. The extracted tracking information was used to detect a range of real-time events such as crowd formation, left luggage, and loitering. In addition all video,track and event information was indexed and stored to allow operators to perform playback and forensic search. This paper will present an overview of the deployed system and discuss the challenges that were encountered during the deployment. Nils Krahnstoever, Peter H. Tu, Ting Yu 0003, Kedar A. Patwardhan, Donald Hamilton, C. Greco, Gianfranco Doretto |
AVSS | 8 |
| 2008 | Face alignment via boosted ranking modelabstractFace alignment seeks to deform a face model to match it with the features of the image of a face by optimizing an appropriate cost function. We propose a new face model that is aligned by maximizing a score function, which we learn from training data, and that we impose to be concave. We show that this problem can be reduced to learning a classifier that is able to say whether or not by switching from one alignment to a new one, the model is approaching the correct fitting. This relates to the ranking problem where a number of instances need to be ordered. For training the model, we propose to extend GentleBoost [23] to rank-learning. Extensive experimentation shows the superiority of this approach to other learning paradigms, and demonstrates that this model exceeds the alignment performance of the state-of-the-art. Xiaoming Liu 0002, Gianfranco Doretto |
CVPR | 3 |
| 2008 | Unified Crowd Segmentation
Peter H. Tu, Thomas Sebastian, Gianfranco Doretto, Nils Krahnstoever, Jens Rittscher, Ting Yu 0003 |
ECCV (4) | 3 |
| 2007 | Shape and Appearance Context ModelingabstractIn this work we develop appearance models for computing the similarity between image regions containing deformable objects of a given class in realtime. We introduce the concept of shape and appearance context. The main idea is to model the spatial distribution of the appearance relative to each of the object parts. Estimating the model entails computing occurrence matrices. We introduce a generalization of the integral image and integral histogram frameworks, and prove that it can be used to dramatically speed up occurrence computation. We demonstrate the ability of this framework to recognize an individual walking across a network of cameras. Finally, we show that the proposed approach outperforms several other methods. Gianfranco Doretto, Thomas Sebastian, Jens Rittscher, Peter H. Tu |
ICCV | 2 |
| 2006 | Joint Recognition of Complex Events and Track MatchingabstractWe present a novel method for jointly performing recognition of complex events and linking fragmented tracks into coherent, long-duration tracks. Many event recognition methods require highly accurate tracking, and may fail when tracks corresponding to event actors are fragmented or partially missing. However, these conditions occur frequently from occlusions, traffic and tracking errors. Recently, methods have been proposed for linking track fragments from multiple objects under these difficult conditions. Here, we develop a method for solving these two problems jointly. A hypothesized event model, represented as a Dynamic Bayes Net, supplies data-driven constraints on the likelihood of proposed track fragment matches. These event-guided constraints are combined with appearance and kinematic constraints used in the previous track linking formulation. The result is the most likely track linking solution given the event model, and the highest event score given all of the track fragments. The event model with the highest score is determined to have occurred, if the score exceeds a threshold. Results demonstrated on a busy scene of airplane servicing activities, where many non-event movers and long fragmented tracks are present, show the promise of the approach to solving the joint problem. Michael T. Chan, Anthony Hoogs, Rahul Bhotika, A. G. Amitha Perera, John Schmiederer, Gianfranco Doretto |
CVPR (2) | 6 |
| 2006 | Dynamic Shape and Appearance ModelsabstractWe propose a model of the joint variation of shape and appearance of portions of an image sequence. The model is conditionally linear, and can be thought of as an extension of active appearance models to exploit the temporal correlation of adjacent image frames. Inference of the model parameters can be performed efficiently using established numerical optimization techniques borrowed from finite-element analysis and system identification techniques. Gianfranco Doretto, Stefano Soatto |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2005 | Modeling Dynamic Scenes with Active AppearanceabstractIn this work, we propose a model for video scenes that contains temporal variability in shape and appearance. We propose a conditionally linear model akin to a dynamic extension of active appearance models. We formulate the problem variationally, and propose a framework where a model complexity cost dictates the "modeling responsibility" of each of the factors: appearance, shape and motion. We render the learning problem well-posed by reverting to a physical and a dynamic prior, and use the finite element method to compute a numerical solution. We illustrate our model to learn and simulate the shape, appearance, and motion of scenes that exhibit some form of temporal regularity, intended in a statistical sense. Gianfranco Doretto |
CVPR (1) | 1 |
| 2004 | Spatially Homogeneous Dynamic Textures
Gianfranco Doretto, Eagle Jones, Stefano Soatto |
ECCV (2) | 1 |
| 2003 | Editable Dynamic TexturesabstractWe present a simple and efficient algorithm for modifying the temporal behavior of "dynamic textures," i.e. sequences of images that exhibit some form of temporal regularity, such as flowing water, steam, smoke, flames, foliage of trees in wind. The main goal is to design algorithms for synthesizing and editing realistic sequences of images of dynamic scenes that exhibit some form of temporal stationarity. This is an image-based rendering task, and in particular we are interested in synthesizing the temporal behavior of the scene. Gianfranco Doretto, Stefano Soatto |
CVPR (2) | 1 |
| 2003 | Dynamic Texture SegmentationabstractWe address the problem of segmenting a sequence of images of natural scenes into disjoint regions that are characterized by constant spatio-temporal statistics. We model the spatio-temporal dynamics in each region by Gauss-Markov models, and infer the model parameters as well as the boundary of the regions in a variational optimization framework. Numerical results demonstrate that - in contrast to purely texture-based segmentation schemes - our method is effective in segmenting regions that differ in their dynamics even when spatial statistics are identical. Gianfranco Doretto, Daniel Cremers, Paolo Favaro, Stefano Soatto |
ICCV | 1 |
| 2003 | Dynamic Textures
Gianfranco Doretto, Alessandro Chiuso, Ying Nian Wu, Stefano Soatto |
Int. J. Comput. Vis. | 1 |
| 2002 | A Frequency Domain Technique for Range Data RegistrationabstractThis work introduces an original method for registering pairs of 3D views consisting of range data sets which operates in the frequency domain. The Fourier transform allows the decoupling of the estimate of the rotation parameters from the estimate of the translation parameters, our algorithm exploits this well-known property by suggesting a three-step procedure. The rotation parameters are estimated by the first two steps through convenient representations and projections of the Fourier transforms' magnitudes and the translational displacement is recovered by the third step by means of a standard phase correlation technique after compensating one of the two views for rotation. The performance of the algorithm, which is well-suited for unsupervised registration, is clearly assessed through extensive testing with several objects and shows that good and robust estimates of 3D rigid motion are achievable. Our algorithm can be used as a prealignment tool for more accurate space-domain registration techniques, like the ICP algorithm. Luca Lucchese, Gianfranco Doretto, Guido M. Cortelazzo |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2001 | Dynamic Texture RecognitionabstractDynamic textures are sequences of images that exhibit some form of temporal stationarity, such as waves, steam, and foliage. We pose the problem of recognizing and classifying dynamic textures in the space of dynamical systems where each dynamic texture is uniquely represented. Since the space is non-linear, a distance between models must be defined We examine three different distances in the space of autoregressive models and assess their power. Payam Saisan, Gianfranco Doretto, Ying Nian Wu, Stefano Soatto |
CVPR (2) | 2 |
| 2001 | Dynamic TexturesabstractDynamic textures are sequences of images of moving scenes that exhibit certain stationarity properties in time; these include sea-waves, smoke, foliage, whirlwind but also talking faces, traffic scenes etc. We present a novel characterization of dynamic textures that poses the problems of modelling, learning, recognizing and synthesizing dynamic textures on a firm analytical footing. We borrow tools from system identification to capture the "essence" of dynamic textures; we do so by learning (i.e. identifying) models that are optimal in the sense of maximum likelihood or minimum prediction error variance. For the special case of second-order stationary processes we identify the model in closed form. Once learned, a model has predictive power and can be used for extrapolating synthetic sequences to infinite length with negligible computational cost. We present experimental evidence that, within our framework, even low dimensional models can capture very complex visual phenomena. Stefano Soatto, Gianfranco Doretto, Ying Nian Wu |
ICCV | 2 |
| 1998 | Free-form Textured Surfaces Registration by a Frequency Domain TechniqueabstractFree-form 3-D surfaces registration is a fundamental problem in 3-D imaging, typically approached by extensions or variations of the ICP algorithm. This work presents a new frequency domain technique for 3-D view registration, totally different form any other techniques for 3-D motion estimation also, based on the Fourier transform. The proposed method can give a non-feature-based method for unsupervised registration of 3-D views. The obtained results are useful "per se" in applications targeted to visual quality or can serve as good starting point for the ICP algorithm when a higher precision is needed. Guido M. Cortelazzo, Gianfranco Doretto, Luca Lucchese |
ICIP (1) | 2 |