VLDB 2026 Research / reviewers in the wild / expert
Basura Fernando
dblp:01/9558
· DBLP profile ↗
81ranked-venue papers
18as first author
39since 2021 · last 2026
0000-0002-6920-9916ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 63 · 16 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 51 · 11 first-author · 25 since 2021Applied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PKR-QA: A Benchmark for Procedural Knowledge Reasoning with Knowledge Module LearningabstractWe introduce PKR-QA (Procedural Knowledge Reasoning Question Answering), a new benchmark for question answering over procedural tasks that require structured reasoning. PKR-QA is constructed semi-automatically using a procedural knowledge graph (PKG), which encodes task-specific knowledge across diverse domains. The PKG is built by curating and linking information from the COIN instructional video dataset and the ontology, enriched with commonsense knowledge from ConceptNet and structured outputs from Large Language Models (LLMs), followed by manual verification. To generate question-answer pairs, we design graph traversal templates where each template is applied systematically over PKG. To enable interpretable reasoning, we propose a neurosymbolic approach called Knowledge Module Learning (KML), which learns procedural relations via neural modules and composes them for structured reasoning with LLMs. Experiments demonstrate that this paradigm improves reasoning performance on PKR-QA and enables step-by-step reasoning traces that facilitate interpretability. Thanh-Son Nguyen 0001, Tzeh Yuan Neoh, Hao Zhang 0047, Ee Yeo Keat, Basura Fernando |
AAAI | 6 |
| 2026 | Improving Temporal Action Segmentation via Constraint-Aware Decoding
Ee Yeo Keat, Debaditya Roy, Hao Zhang 0047, Basura Fernando |
ICPR (11) | 5 |
| 2026 | PointTFA$^{m}$: Multi-Modal, Training-Free Adaptation for Point Cloud Understanding
Jinmeng Wu, Youxiang Hu, Hao Zhang 0047, Basura Fernando, Yanbin Hao, Hanyu Hong |
IEEE Trans. Multim. | 5 |
| 2025 | PhysReason: A Comprehensive Benchmark towards Physics-Based ReasoningabstractLarge language models demonstrate remarkable capabilities across various domains, especially mathematics and logic reasoning. However, current evaluations overlook physics-based reasoning - a complex task requiring physics theorems and constraints. We present PhysReason, a 1,200-problem benchmark comprising knowledge-based (25%) and reasoning-based (75%) problems, where the latter are divided into three difficulty levels (easy, medium, hard). Notably, problems require an average of 8.1 solution steps, with hard requiring 15.6, reflecting the complexity of physics-based reasoning. We propose the Physics Solution Auto Scoring Framework, incorporating efficient answer-level and comprehensive step-level evaluations. Top-performing models like Deepseek-R1, Gemini-2.0-Flash-Thinking, and o3-mini-high achieve less than 60% on answer-level evaluation, with performance dropping from knowledge questions (75.11%) to hard problems (31.95%). Through step-level evaluation, we identified four key bottlenecks: Physics Theorem Application, Physics Process Understanding, Calculation, and Physics Condition Analysis. These findings position PhysReason as a novel and comprehensive benchmark for evaluating physics-based reasoning capabilities in large language models. Xinyu Zhang 0021, Yanrui Wu, Chengyou Jia, Basura Fernando, Zheng Shou 0001, Lingling Zhang 0005, Jun Liu 0036 |
ACL (1) | 6 |
| 2025 | Diagram-Driven Course Questions GenerationabstractXinyu Zhang, Lingling Zhang, Yanrui Wu, Muye Huang, Wenjun Wu, Bo Li, Shaowei Wang, Basura Fernando, Jun Liu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Xinyu Zhang 0021, Lingling Zhang 0005, Yanrui Wu, Muye Huang, Basura Fernando, Jun Liu 0002 |
EMNLP | 8 |
| 2025 | Improving Open-vocabulary Video Visual Relation Detection with Decomposed Prompt Learning and Relation AdjustmentabstractOpen-vocabulary video visual relation detection (VidVRD) expands the scope of detecting object relations in videos to include unseen categories. It marks considerable advancement in recognizing novel relations solely by training on a base set, thus extending the frontiers of automated video understanding. However, the performance of current methods on novel predicates remains significantly inferior to that on base categories. We attribute this discrepancy to two primary factors: (1) A significant task misalignment between the Visual Relation Detection (VRD) task and the pre-trained models’ visual feature extractors, which are often designed for tasks like video-text retrieval and image-text retrieval, resulting in poor generalization to the novel set. (2) The relatively small size and limited vocabulary of open-vocabulary datasets, which create a substantial gap between base and novel predicates. Consequently, text prompts trained on the base set fail to generalize effectively to the novel set. To address these issues, we propose two improvement measures: (1) We decompose base and novel relations into actional and spatial patterns and introduce an innovative text prompt learning method that leverages the shared patterns between base and novel relations. (2) We develop a relation probability adjustment mechanism that utilizes reliable base relation predictions to adjust the probabilities of relations in novel classes by considering their overlaps in either actional or spatial contents. Experimental results on the benchmark dataset demonstrate significant performance improvements. Ming Pei, Yi Tan 0001, Yanbin Hao, Hao Zhang 0047, Jinmeng Wu, Basura Fernando, Xun Yang 0001 |
ICASSP | 6 |
| 2025 | Imore: Implicit Program-Guided Reasoning for Human Motion QA
Chinthani Sugandhika, Ee Yeo Keat, Eric P. Xing, Deepu Rajan, Basura Fernando |
ICCV | 8 |
| 2025 | You Only Communicate Once: One-shot Federated Low-Rank Adaptation of MLLMabstractMultimodal Large Language Models (MLLMs) with Federated Learning (FL) can quickly adapt to privacy-sensitive tasks, but face significant challenges such as high communication costs and increased attack risks, due to their reliance on multi-round communication. To address this, One-shot FL (OFL) has emerged, aiming to complete adaptation in a single client-server communication. However, existing adaptive ensemble OFL methods still need more than one round of communication, because correcting heterogeneity-induced local bias relies on aggregated global supervision, meaning they still do not achieve true one-shot communication. In this work, we make the first attempt to achieve true one-shot communication for MLLMs under OFL, by investigating whether implicit (i.e., initial rather than aggregated) global supervision alone can effectively correct local training bias. Our key finding from the empirical study is that imposing directional supervision on local training substantially mitigates client conflicts and local bias. Building on this insight, we propose YOCO, in which directional supervision with sign-regularized LoRA B enforces global consistency, while sparsely regularized LoRA A preserves client-specific adaptability. Experiments demonstrate that YOCO cuts communication to $\sim$0.03\% of multi-round FL while surpassing those methods in several multimodal scenarios and consistently outperforming all one-shot competitors. Binqian Xu, Haiyang Mei, Zechen Bai, Jinjin Gong, Rui Yan 0010, Guosen Xie, Yazhou Yao, Basura Fernando, Xiangbo Shu |
NeurIPS | 8 |
| 2025 | CoFFT: Chain of Foresight-Focus Thought for Visual Language ModelsabstractDespite significant advances in Vision Language Models (VLMs), they remain constrained by the complexity and redundancy of visual input.
When images contain large amounts of irrelevant information, VLMs are susceptible to interference, thus generating excessive task-irrelevant reasoning processes or even hallucinations.
This limitation stems from their inability to discover and process the required regions during reasoning precisely.
To address this limitation, we present the Chain of Foresight-Focus Thought (CoFFT), a novel training-free approach that enhances VLMs' visual reasoning by emulating human visual cognition.
Each Foresight-Focus Thought consists of three stages:
(1) Diverse Sample Generation: generates diverse reasoning samples to explore potential reasoning paths, where each sample contains several reasoning steps;
(2) Dual Foresight Decoding: rigorously evaluates these samples based on both visual focus and reasoning progression, adding the first step of optimal sample to the reasoning process;
(3) Visual Focus Adjustment: precisely adjust visual focus toward regions most beneficial for future reasoning, before returning to stage (1) to generate subsequent reasoning samples until reaching the final answer.
These stages function iteratively, creating an interdependent cycle where reasoning guides visual focus and visual focus informs subsequent reasoning.
Empirical results across multiple benchmarks using Qwen2.5-VL, InternVL-2.5, and Llava-Next demonstrate consistent performance improvements of 3.1-5.8\% with controllable increasing computational overhead. Xinyu Zhang 0021, Lingling Zhang 0005, Chengyou Jia, Zhuohang Dang, Basura Fernando, Jun Liu 0036, Zheng Shou 0001 |
NeurIPS | 6 |
| 2025 | Effective Scene Graph Generation by Statistical Relation DistillationabstractAnnotating scene graphs for images is a time-consuming task, resulting in many instances of missing relations within existing datasets. In this paper, we introduce the Statistical Relation Distillation (SRD) method to enhance scene graph datasets. SRD leverages human-annotated relations alongside object-to-object and predicate-to-predicate similarities to reinforce the existence likelihood of scene graph relations. Moreover, SRD can augment relational frequency using relations of non-selected object and predicate categories that are usually omitted by scene graph generation (SGG) task. The output from SRD derives the prior probability which is combined with model-predicted probabilities to annotate missing relations in training images and subsequently re-train SGG models on the augmented dataset. We evaluate our proposed method on Visual Genome and GQA-200 datasets. Experimental results show that training on the augmented dataset enhances the performance of prominent scene-graph generation models. The implementation code is at https://github.com/LUNAProject22/SRD. Thanh-Son Nguyen 0001, Basura Fernando |
WACV | 3 |
| 2025 | Deduce and Select Evidences with Language Models for Training-Free Video Goal InferenceabstractWe introduce ViDSE, a Video framework that Deduce and Selects visual Evidence for training-free video goal inference using language models. Unlike approaches that directly apply vision-language models (VLM) or combine VLM+LLM to process dense video visuals, ViDSE explicitly selects relevant visual evidence (e.g., frames) based on the hypothesis deduced by the LLM. This approach not only im-proves accuracy but also reveals the logical process behind the model's decisions, enhancing explainability. Our exper-iments demonstrate that this selection process significantly reduces ambiguity in the subsequent inference reasoning stage and outperforms VLM-only and VLM+LLM models on goal inference tasks such as CrossTask and COIN. We further validate ViDSE’ s generalizability and robustness on action recognition tasks, such as ActivityNet and UCF-101, under training-free and open-vocabulary conditions. We observe that ViDSE easily generalizes to other video tasks (e.g., action recognition) requiring filtering of redundant and irrelevant information. Ee Yeo Keat, Hao Zhang 0047, Alexander Matyasko, Basura Fernando |
WACV | 4 |
| 2025 | Learning to Visually Connect Actions and Their Effects
Paritosh Parmar, Eric Peh, Basura Fernando |
WACV | 3 |
| 2025 | Situational Scene Graph for Structured Human-Centric Situation Understanding
Chinthani Sugandhika, Deepu Rajan, Basura Fernando |
WACV | 4 |
| 2025 | Inferring Past Human Actions in Homes with Abductive ReasoningabstractAbductive reasoning aims to make the most likely inference for a given set of incomplete observations. In this paper, we introduce “Abductive Past Action Inference”, a novel research task aimed at identifying the past actions performed by individuals within homes to reach specific states captured in a single image, using abductive inference. The research explores three key abductive inference problems: past action set prediction, past action sequence prediction, and abductive past action verification. We introduce several models tailored for abductive past action inference, including a relational graph neural network, a relational bilinear pooling model, and a relational transformer model. Notably, the newly proposed object-relational bilinear graph encoder-decoder (BiGED) model emerges as the most effective among all methods evaluated, demonstrating good proficiency in handling the intricacies of the Action Genome dataset. The contributions of this research significantly advance the ability of deep learning models to reason about current scene evidence and make highly plausible inferences about past human actions. This advancement enables a deeper understanding of events and behaviors, which can enhance decision-making and improve system capabilities across various real-world applications such as Human-Robot Interaction and Elderly Care and Health Monitoring. Code and data available at https://github.com/LUNAProject22/AAR Clement Tan, Chai Kiat Yeo, Cheston Tan, Basura Fernando |
WACV | 4 |
| 2025 | Effectively Leveraging CLIP for Generating Situational Summaries of Images and Videos
Dhruv Verma, Debaditya Roy, Basura Fernando |
Int. J. Comput. Vis. | 3 |
| 2025 | Clothing Purification with Causality Meets Vision-Language Pretraining Models
Zhengwei Yang 0001, Huilin Zhu, Nan Lei, Basura Fernando, Zheng Wang 0007 |
Int. J. Comput. Vis. | 4 |
| 2024 | Dissecting Multimodality in VideoQA Transformer Models by Impairing Modality FusionabstractWhile VideoQA Transformer models demonstrate competitive performance on standard benchmarks, the reasons behind their success are not fully understood. Do these models capture the rich multimodal structures and dynamics from video and text jointly? Or are they achieving high scores by exploiting biases and spurious features? Hence, to provide insights, we design QUAG (QUadrant AveraGe), a lightweight and non-parametric probe, to conduct dataset-model combined representation analysis by impairing modality fusion. We find that the models achieve high performance on many datasets without leveraging multimodal representations. To validate QUAG further, we design QUAG-attention, a less-expressive replacement of self-attention with restricted token interactions. Models with QUAG-attention achieve similar performance with significantly fewer multiplication operations without any finetuning. Our findings raise doubts about the current models’ abilities to learn highly-coupled multimodal representations. Hence, we design the CLAVI (Complements in LAnguage and VIdeo) dataset, a stress-test dataset curated by augmenting real-world videos to have high modality coupling. Consistent with the findings of QUAG, we find that most of the models achieve near-trivial performance on CLAVI. This reasserts the limitations of current models for learning highly-coupled multimodal representations, that is not evaluated by the current datasets. Ishaan Singh Rawal, Alexander Matyasko, Shantanu Jaiswal, Basura Fernando, Cheston Tan |
ICML | 4 |
| 2024 | Predicting the Next Action by Modeling the Abstract Goal
Debaditya Roy, Basura Fernando |
ICPR (15) | 2 |
| 2024 | PointTFA: Training-Free Clustering Adaption for Large 3D Point Cloud Models
Jinmeng Wu, Hao Zhang 0047, Basura Fernando, Yanbin Hao, Hanyu Hong |
IJCAI | 4 |
| 2024 | RCA: Region Conditioned Adaptation for Visual Abductive Reasoning
Hao Zhang 0047, Ee Yeo Keat, Basura Fernando |
ACM Multimedia | 3 |
| 2024 | Learning to Reason Iteratively and Parallelly for Complex Visual Reasoning ScenariosabstractComplex visual reasoning and question answering (VQA) is a challenging task that requires compositional multi-step processing and higher-level reasoning capabilities beyond the immediate recognition and localization of objects and events. Here, we introduce a fully neural Iterative and Parallel Reasoning Mechanism (IPRM) that combines two distinct forms of computation -- iterative and parallel -- to better address complex VQA scenarios. Specifically, IPRM's "iterative" computation facilitates compositional step-by-step reasoning for scenarios wherein individual operations need to be computed, stored, and recalled dynamically (e.g. when computing the query “determine the color of pen to the left of the child in red t-shirt sitting at the white table”). Meanwhile, its "parallel'' computation allows for the simultaneous exploration of different reasoning paths and benefits more robust and efficient execution of operations that are mutually independent (e.g. when counting individual colors for the query: "determine the maximum occurring color amongst all t-shirts'"). We design IPRM as a lightweight and fully-differentiable neural module that can be conveniently applied to both transformer and non-transformer vision-language backbones. It notably outperforms prior task-specific methods and transformer-based attention modules across various image and video VQA benchmarks testing distinct complex reasoning capabilities such as compositional spatiotemporal reasoning (AGQA), situational reasoning (STAR), multi-hop reasoning generalization (CLEVR-Humans) and causal event linking (CLEVRER-Humans). Further, IPRM's internal computations can be visualized across reasoning steps, aiding interpretability and diagnosis of its errors. Shantanu Jaiswal, Debaditya Roy, Basura Fernando, Cheston Tan |
NeurIPS | 3 |
| 2024 | CausalChaos! Dataset for Comprehensive Causal Action Question Answering Over Longer Causal Chains Grounded in Dynamic Visual ScenesabstractCausal video question answering (QA) has garnered increasing interest, yet existing datasets often lack depth in causal reasoning. To address this gap, we capitalize on the unique properties of cartoons and construct CausalChaos!, a novel, challenging causal Why-QA dataset built upon the iconic "Tom and Jerry" cartoon series. Cartoons use the principles of animation that allow animators to create expressive, unambiguous causal relationships between events to form a coherent storyline. Utilizing these properties, along with thought-provoking questions and multi-level answers (answer and detailed causal explanation), our questions involve causal chains that interconnect multiple dynamic interactions between characters and visual scenes. These factors demand models to solve more challenging, yet well-defined causal relationships. We also introduce hard incorrect answer mining, including a causally confusing version that is even more challenging. While models perform well, there is much room for improvement, especially, on open-ended answers. We identify more advanced/explicit causal relationship modeling & joint modeling of vision and language as the immediate areas for future efforts to focus upon. Along with the other complementary datasets, our new challenging dataset will pave the way for these developments in the field. Dataset and Code: https://github.com/LUNAProject22/CausalChaos Paritosh Parmar, Eric Peh, Ruirui Chen 0002, Ting En Lam, Elston Tan, Basura Fernando |
NeurIPS | 7 |
| 2024 | DoFIT: Domain-aware Federated Instruction Tuning with Alleviated Catastrophic ForgettingabstractFederated Instruction Tuning (FIT) advances collaborative training on decentralized data, crucially enhancing model's capability and safeguarding data privacy. However, existing FIT methods are dedicated to handling data heterogeneity across different clients (i.e., client-aware data heterogeneity), while ignoring the variation between data from different domains (i.e., domain-aware data heterogeneity). When scarce data needs supplementation from related fields, these methods lack the ability to handle domain heterogeneity in cross-domain training. This leads to domain-information catastrophic forgetting in collaborative training and therefore makes model perform sub-optimally on the individual domain. To address this issue, we introduce DoFIT, a new Domain-aware FIT framework that alleviates catastrophic forgetting through two new designs. First, to reduce interference information from the other domain, DoFIT finely aggregates overlapping weights across domains on the inter-domain server side. Second, to retain more domain information, DoFIT initializes intra-domain weights by incorporating inter-domain information into a less-conflicted parameter space. Experimental results on diverse datasets consistently demonstrate that DoFIT excels in cross-domain collaborative training and exhibits significant advantages over conventional FIT methods in alleviating catastrophic forgetting. Code is available at [this link](https://github.com/1xbq1/DoFIT). Binqian Xu, Xiangbo Shu, Haiyang Mei, Zechen Bai, Basura Fernando, Zheng Shou 0001, Jinhui Tang 0001 |
NeurIPS | 5 |
| 2024 | Interaction Region Visual Transformer for Egocentric Action AnticipationabstractHuman-object interaction (HOI) and temporal dynamics along the motion paths are the most important visual cues for egocentric action anticipation. Especially, interaction regions covering objects and the human hand reveal significant visual cues to predict future human actions. However, how to incorporate and capture these important visual cues in modern video Transformer architecture remains a challenge. We leverage the effective MotionFormer that models motion dynamics to incorporate interaction regions using spatial cross-attention and further infuse contextual information using trajectory cross-attention to obtain an interaction-centric video representation for action anticipation. We term our model InAViT which achieves state-of-the-art action anticipation performance on large-scale egocentric datasets EPICKTICHENS100 (EK100) and EGTEA Gaze+. On the EK100 evaluation server, InAViT is on top of the public leader board (at the time of submission) where it outperforms the second-best model by 3.3% on mean-top5 recall. The code is available1. Debaditya Roy, Ramanathan Rajendiran, Basura Fernando |
WACV | 3 |
| 2023 | Memory-efficient Temporal Moment Localization in Long VideosabstractCristian Rodriguez-Opazo, Edison Marrese-Taylor, Basura Fernando, Hiroya Takamura, Qi Wu. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Cristian Rodriguez Opazo, Edison Marrese-Taylor, Basura Fernando, Hiroya Takamura, Qi Wu 0001 |
EACL | 3 |
| 2023 | Semi-supervised multimodal coreference resolution in image narrationsabstractIn this paper, we study multimodal coreference resolution, specifically where a longer descriptive text, i.e., a narration is paired with an image.This poses significant challenges due to fine-grained image-text alignment, inherent ambiguity present in narrative language, and unavailability of large annotated training sets.To tackle these challenges, we present a data efficient semi-supervised approach that utilizes image-narration pairs to resolve coreferences and narrative grounding in a multimodal context.Our approach incorporates losses for both labeled and unlabeled data within a crossmodal framework.Our evaluation shows that the proposed approach outperforms strong baselines both quantitatively and qualitatively, for the tasks of coreference resolution and narrative grounding. Arushi Goel, Basura Fernando, Frank Keller, Hakan Bilen |
EMNLP | 2 |
| 2023 | Who are you referring to? Coreference resolution in image narrationsabstractCoreference resolution aims to identify words and phrases which refer to the same entity in a text, a core task in natural language processing. In this paper, we extend this task to resolving coreferences in long-form narrations of visual scenes. First, we introduce a new dataset with annotated coreference chains and their bounding boxes, as most existing image-text datasets only contain short sentences without coreferring expressions or labeled chains. We propose a new technique that learns to identify coref-erence chains using weak supervision, only from image-text pairs and a regularization using prior linguistic knowledge. Our model yields large performance gains over several strong baselines in resolving coreferences. We also show that coreference resolution helps improve grounding narratives in images. Arushi Goel, Basura Fernando, Frank Keller, Hakan Bilen |
ICCV | 2 |
| 2023 | Energy-based Self-Training and Normalization for Unsupervised Domain AdaptationabstractWe propose an Unsupervised Domain Adaptation (UDA) method by making use of Energy-Based Learning (EBL) and demonstrate 1. EBL can be used to improve the instance selection for a self-training task on the unlabelled target domain, and 2. alignment and normalizing energy scores can learn domain-invariant representations. For the former, we show that an energy-based selection criterion can be used to model instance selections by mimicking the joint distribution between data and predictions in the target domain. As per learning domain invariant representations, we show that stable domain alignment can be achieved by a combined energy alignment and an energy normalization process. We implement our method in consistent with the vision-transformer (ViT) backbone and show that our proposed method can outperform state-of-the-art ViT based UDA methods on diverse benchmarks (DomainNet, Office-Home, and VISDA2017). Samitha Herath, Basura Fernando, Ehsan Abbasnejad, Munawar Hayat, Shahram Khadivi, Mehrtash Harandi, Seyed Hamid Rezatofighi, Gholamreza Haffari |
ICCV | 2 |
| 2022 | Not All Relations are Equal: Mining Informative Labels for Scene Graph GenerationabstractScene graph generation (SGG) aims to capture a wide variety of interactions between pairs of objects, which is essential for full scene understanding. Existing SGG methods trained on the entire set of relations fail to acquire complex reasoning about visual and textual correlations due to various biases in training data. Learning on trivial relations that indicate generic spatial configuration like ‘on’ instead of informative relations such as ‘parked on’ does not enforce this complex reasoning, harming generalization. To address this problem, we propose a novel framework for SGG training that exploits relation labels based on their informativeness. Our model-agnostic training procedure imputes missing informative relations for less informative samples in the training data and trains a SGG model on the imputed labels along with existing annotations. We show that this approach can successfully be used in conjunction with state-of-the-art SGG methods and improves their performance significantly in multiple metrics on the standard Visual Genome benchmark. Furthermore, we obtain considerable improvements for unseen triplets in a more challenging zero-shot setting. Arushi Goel, Basura Fernando, Frank Keller, Hakan Bilen |
CVPR | 2 |
| 2022 | 3D Equivariant Graph Implicit Functions
Yunlu Chen, Basura Fernando, Hakan Bilen, Matthias Nießner, Efstratios Gavves |
ECCV (3) | 2 |
| 2022 | TDAM: Top-Down Attention Module for Contextually Guided Feature Selection in CNNs
Shantanu Jaiswal, Basura Fernando, Cheston Tan |
ECCV (25) | 2 |
| 2022 | Action anticipation using latent goal learningabstractTo get something done, humans perform a sequence of actions dictated by a goal. So, predicting the next action in the sequence becomes easier once we know the goal that guides the entire activity. We present an action anticipation model that uses goal information in an effective manner. Specifically, we use a latent goal representation as a proxy for the "real goal" of the sequence and use this goal information when predicting the next action. We design a model to compute the latent goal representation from the observed video and use it to predict the next action. We also exploit two properties of goals to propose new losses for training the model. First, the effect of the next action should be closer to the latent goal than the observed action, termed as "goal closeness". Second, the latent goal should remain consistent before and after the execution of the next action which we coined as "goal consistency". Using this technique, we obtain state-of-the-art action anticipation performance on scripted datasets 50Salads and Breakfast that have predefined goals in all their videos. We also evaluate the latent goal-based model on EPIC-KITCHENS55 which is an unscripted dataset with multiple goals being pursued simultaneously. Even though this is not an ideal setup for using latent goals, our model is able to predict the next noun better than existing approaches on both seen and unseen kitchens in the test set.1 Debaditya Roy, Basura Fernando |
WACV | 2 |
| 2022 | Effective Multimodal Encoding for Image Paragraph CaptioningabstractIn this paper, we present a regularization-based image paragraph generation method. We propose a novel multimodal encoding generator (MEG) to generate effective multimodal encoding that captures not only an individual sentence but also visual and paragraph-sequential information. By utilizing the encoding generated by MEG, we regularize a paragraph generation model that allows us to improve the results of the captioning model in all the evaluation metrics. With the support of the proposed MEG model for regularization, our paragraph generation model obtains state-of-the-art results on the Stanford paragraph dataset once further optimized with reinforcement learning. Moreover, we perform extensive empirical analysis on the capabilities of MEG encoding. A qualitative visualization based on t-distributed stochastic neighbor embedding (t-SNE) illustrates that sentence encoding generated by MEG captures some level of semantic information. We also demonstrate that the MEG encoding captures meaningful textual and visual information by performing multimodal sentence retrieval tasks and image instance retrieval given a paragraph query. Thanh-Son Nguyen 0001, Basura Fernando |
IEEE Trans. Image Process. | 2 |
| 2021 | Anticipating Human Actions by Correlating Past With the Future With Jaccard Similarity MeasuresabstractWe propose a framework for early action recognition and anticipation by correlating past features with the future using three novel similarity measures called Jaccard vector similarity, Jaccard cross-correlation and Jaccard Frobenius inner product over covariances. Using these combinations of novel losses and using our framework, we obtain state-of-the-art results for early action recognition in UCF101 and JHMDB datasets by obtaining 91.7 % and 83.5 % accuracy respectively for an observation percentage of 20. Similarly, we obtain state-of-the-art results for Epic-Kitchen55 and Breakfast datasets for action anticipation by obtaining 20.35 and 41.8 top-1 accuracy respectively. Basura Fernando, Samitha Herath |
CVPR | 1 |
| 2021 | Neural Feature Matching in Implicit 3D RepresentationsabstractRecently, neural implicit functions have achieved impressive results for encoding 3D shapes. Conditioning on low-dimensional latent codes generalises a single implicit function to learn shared representation space for a variety of shapes, with the advantage of smooth interpolation. While the benefits from the global latent space do not correspond to explicit points at local level, we propose to track the continuous point trajectory by matching implicit features with the latent code interpolating between shapes, from which we corroborate the hierarchical functionality of the deep implicit functions, where early layers map the latent code to fitting the coarse shape structure, and deeper layers further refine the shape details. Furthermore, the structured representation space of implicit functions enables to apply feature matching for shape deformation, with the benefits to handle topology and semantics inconsistency, such as from an armchair to a chair with no arms, without explicit flow functions or manual annotations. Yunlu Chen, Basura Fernando, Hakan Bilen, Thomas Mensink, Efstratios Gavves |
ICML | 2 |
| 2021 | FlowCaps: Optical Flow Estimation with Capsule Networks For Action RecognitionabstractCapsule networks (CapsNets) have recently shown promise to excel in most computer vision tasks, especially pertaining to scene understanding. In this paper, we explore CapsNet's capabilities in optical flow estimation, a task at which convolutional neural networks (CNNs) have already outperformed other approaches. We propose a CapsNet-based architecture, termed FlowCaps, which attempts to a) achieve better correspondence matching via finer-grained, motion-specific, and more-interpretable encoding crucial for optical flow estimation, b) perform better-generalizable optical flow estimation, c) utilize lesser ground truth data, and d) significantly reduce the computational complexity in achieving good performance, in comparison to its CNN-counterparts. Vinoj Jayasundara 0001, Debaditya Roy, Basura Fernando |
WACV | 3 |
| 2021 | DORi: Discovering Object Relationships for Moment Localization of a Natural Language Query in a VideoabstractThis paper studies the task of temporal moment localization in long untrimmed videos using natural language queries. Given a query sentence, the goal is to determine the start and end of the relevant segment within the video. Our key innovation is to learn a video feature embedding through a language-conditioned message-passing algorithm suitable for temporal moment localization which captures the relationships between humans, objects and activities in the video. These relationships are obtained by a spatial sub-graph that contextualizes the scene representation using detected objects and human features conditioned in the language query. Moreover, a temporal sub-graph captures the activities within the video through time. Our method is evaluated on three standard benchmark datasets, and we also introduce YouCookII as a new benchmark for this task. Experiments show our method outperforms state-of-the-art methods on these datasets, confirming the effectiveness of our approach. Cristian Rodriguez Opazo, Edison Marrese-Taylor, Basura Fernando, Hongdong Li, Stephen Gould |
WACV | 3 |
| 2021 | Weakly supervised action segmentation with effective use of attention and self-attention
Yan Bin Ng, Basura Fernando |
Comput. Vis. Image Underst. | 2 |
| 2021 | Action Anticipation Using Pairwise Human-Object Interactions and TransformersabstractThe ability to anticipate future actions of humans is useful in application areas such as automated driving, robot-assisted manufacturing, and smart homes. These applications require representing and anticipating human actions involving the use of objects. Existing methods that use human-object interactions for anticipation require object affordance labels for every relevant object in the scene that match the ongoing action. Hence, we propose to represent every pairwise human-object (HO) interaction using only their visual features. Next, we use cross-correlation to capture the second-order statistics across human-object pairs in a frame. Cross-correlation produces a holistic representation of the frame that can also handle a variable number of human-object pairs in every frame of the observation period. We show that cross-correlation based frame representation is more suited for action anticipation than attention-based and other second-order approaches. Furthermore, we observe that using a transformer model for temporal aggregation of frame-wise HO representations results in better action anticipation than other temporal networks. So, we propose two approaches for constructing an end-to-end trainable multi-modal transformer (MM-Transformer; code at https://github.com/debadityaroy/MM-Transformer_ActAnt) model that combines the evidence across spatio-temporal, motion, and HO representations. We show the performance of MM-Transformer on procedural datasets like 50 Salads and Breakfast, and an unscripted dataset like EPIC-KITCHENS55. Finally, we demonstrate that the combination of human-object representation and MM-Transformers is effective even for long-term anticipation. Debaditya Roy, Basura Fernando |
IEEE Trans. Image Process. | 2 |
| 2020 | What do CNNs gain by imitating the visual development of primate infants?
Shantanu Jaiswal, Dongkyu Choi, Basura Fernando |
BMVC | 3 |
| 2020 | How does simulating aspects of primate infant visual development inform training of CNNs?
Shantanu Jaiswal, Dongkyu Choi, Basura Fernando |
CogSci | 3 |
| 2020 | Intention Inference in a Dynamic Multi-Goal Environment
Desmond C. Ong, Marie Therese Robles Quieta, Basura Fernando |
CogSci | 3 |
| 2020 | Weakly Supervised Gaussian Networks for Action DetectionabstractDetecting temporal extents of human actions in videos is a challenging computer vision problem that requires detailed manual supervision including frame-level labels. This expensive annotation process limits deploying action detectors to a limited number of categories. We propose a novel method, called WSGN, that learns to detect actions from weak supervision, using only video-level labels. WSGN learns to exploit both video-specific and dataset-wide statistics to predict relevance of each frame to an action category. This strategy leads to significant gains in action detection for two standard benchmarks THU-MOS14 and Charades. Our method obtains excellent results compared to state-of-the-art methods that uses similar features and loss functions on THUMOS14 dataset. Similarly, our weakly supervised method is only 0.3% mAP behind a state-of-the-art supervised method on challenging Charades dataset for action localization. Basura Fernando, Cheston Tan, Hakan Bilen |
WACV | 1 |
| 2020 | Hallucinating Unaligned Face Images by Multiscale Transformative Discriminative Networks
Xin Yu 0002, Fatih Porikli, Basura Fernando, Richard I. Hartley |
Int. J. Comput. Vis. | 3 |
| 2020 | Semantic Face Hallucination: Super-Resolving Very Low-Resolution Face Images with Supplementary AttributesabstractGiven a tiny face image, existing face hallucination methods aim at super-resolving its high-resolution (HR) counterpart by learning a mapping from an exemplary dataset. Since a low-resolution (LR) input patch may correspond to many HR candidate patches, this ambiguity may lead to distorted HR facial details and wrong attributes such as gender reversal and rejuvenation. An LR input contains low-frequency facial components of its HR version while its residual face image, defined as the difference between the HR ground-truth and interpolated LR images, contains the missing high-frequency facial details. We demonstrate that supplementing residual images or feature maps with additional facial attribute information can significantly reduce the ambiguity in face super-resolution. To explore this idea, we develop an attribute-embedded upsampling network, which consists of an upsampling network and a discriminative network. The upsampling network is composed of an autoencoder with skip-connections, which incorporates facial attribute vectors into the residual features of LR inputs at the bottleneck of the autoencoder, and deconvolutional layers used for upsampling. The discriminative network is designed to examine whether super-resolved faces contain the desired attributes or not and then its loss is used for updating the upsampling network. In this manner, we can super-resolve tiny (16×16 pixels) unaligned face images with a large upscaling factor of 8× while reducing the uncertainty of one-to-many mappings remarkably. By conducting extensive evaluations on a large-scale dataset, we demonstrate that our method achieves superior face hallucination results and outperforms the state-of-the-art. Xin Yu 0002, Basura Fernando, Richard I. Hartley, Fatih Porikli |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Forecasting Future Action Sequences With Attention: A New Approach to Weakly Supervised Action ForecastingabstractFuture human action forecasting from partial observations of activities is an important problem in many practical applications such as assistive robotics, video surveillance and security. We present a method to forecast actions for the unseen future of the video using a neural machine translation technique that uses encoder-decoder architecture. The input to this model is the observed RGB video, and the objective is to forecast the correct future symbolic action sequence. Unlike prior methods that make action predictions for some unseen percentage of video one for each frame, we predict the complete action sequence that is required to accomplish the activity. We coin this task action sequence forecasting. To cater for two types of uncertainty in the future predictions, we propose a novel loss function. We show a combination of optimal transport and future uncertainty losses help to improve results. We evaluate our model in three challenging video datasets (Charades, MPII cooking and Breakfast). We extend our action sequence forecasting model to perform weakly supervised action forecasting on two challenging datasets, the Breakfast and the 50Salads. Specifically, we propose a model to predict actions of future unseen frames without using frame level annotations during training. Using Fisher vector features, our supervised model outperforms the state-of-the-art action forecasting model by 0.83% and 7.09% on the Breakfast and the 50Salads datasets respectively. Our weakly supervised model is only 0.6% behind the most recent state-of-the-art supervised model and obtains comparable results to other published fully supervised methods, and sometimes even outperforms them on the Breakfast dataset. Most interestingly, our weakly supervised model outperforms prior models by 1.04% leveraging on proposed weakly supervised architecture, and effective use of attention mechanism and loss functions. Yan Bin Ng, Basura Fernando |
IEEE Trans. Image Process. | 2 |
| 2019 | Min-Max Statistical Alignment for Transfer LearningabstractA profound idea in learning invariant features for transfer learning is to align statistical properties of the domains. In practice, this is achieved by minimizing the disparity between the domains, usually measured in terms of their statistical properties. We question the capability of this school of thought and propose to minimize the maximum disparity between domains. Furthermore, we develop an end-to-end learning scheme that enables us to benefit from the proposed min-max strategy in training deep models. We show that the min-max solution can outperform the existing statistical alignment solutions, and can compete with state-of-the-art solutions on two challenging learning tasks, namely, Unsupervised Domain Adaptation (UDA) and Zero-Shot Learning (ZSL). Samitha Herath, Mehrtash Harandi, Basura Fernando, Richard Nock |
CVPR | 3 |
| 2019 | Visual Permutation LearningabstractWe present a principled approach to uncover the structure of visual data by solving a deep learning task coined visual permutation learning. The goal of this task is to find the permutation that recovers the structure of data from shuffled versions of it. In the case of natural images, this task boils down to recovering the original image from patches shuffled by an unknown permutation matrix. Permutation matrices are discrete, thereby posing difficulties for gradient-based optimization methods. To this end, we resort to a continuous approximation using doubly-stochastic matrices and formulate a novel bi-level optimization problem on such matrices that learns to recover the permutation. Unfortunately, such a scheme leads to expensive gradient computations. We circumvent this issue by further proposing a computationally cheap scheme for generating doubly stochastic matrices based on Sinkhorn iterations. To implement our approach we propose DeepPermNet, an end-to-end CNN model for this task. The utility of DeepPermNet is demonstrated on three challenging computer vision problems, namely, relative attributes learning, supervised learning-to-rank, and self-supervised representation learning. Our results show state-of-the-art performance on the Public Figures and OSR benchmarks for relative attributes learning, chronological and interestingness image ranking for supervised learning-to-rank, and competitive results in the classification and segmentation tasks of the PASCAL VOC dataset for self-supervised representation learning. Rodrigo Santa Cruz, Basura Fernando, Anoop Cherian, Stephen Gould |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Using temporal information for recognizing actions from still images
Samitha Herath, Basura Fernando, Mehrtash Harandi |
Pattern Recognit. | 2 |
| 2018 | VIENA ^2 : A Driving Anticipation Dataset
Mohammad Sadegh Ali Akbarian, Fatemehsadat Saleh, Mathieu Salzmann, Basura Fernando, Lars Petersson, Lars Andersson |
ACCV (1) | 4 |
| 2018 | Super-Resolving Very Low-Resolution Face Images With Supplementary AttributesabstractGiven a tiny face image, existing face hallucination methods aim at super-resolving its high-resolution (HR) counterpart by learning a mapping from an exemplar dataset. Since a low-resolution (LR) input patch may correspond to many HR candidate patches, this ambiguity may lead to distorted HR facial details and wrong attributes such as gender reversal. An LR input contains low-frequency facial components of its HR version while its residual face image, defined as the difference between the HR ground-truth and interpolated LR images, contains the missing high-frequency facial details. We demonstrate that supplementing residual images or feature maps with additional facial attribute information can significantly reduce the ambiguity in face super-resolution. To explore this idea, we develop an attribute-embedded upsampling network, which consists of an upsampling network and a discriminative network. The upsampling network is composed of an autoencoder with skip-connections, which incorporates facial attribute vectors into the residual features of LR inputs at the bottleneck of the autoencoder and deconvolutional layers used for upsampling. The discriminative network is designed to examine whether super-resolved faces contain the desired attributes or not and then its loss is used for updating the upsampling network. In this manner, we can super-resolve tiny (16×16 pixels) unaligned face images with a large upscaling factor of 8× while reducing the uncertainty of one-to-many mappings remarkably. By conducting extensive evaluations on a large-scale dataset, we demonstrate that our method achieves superior face hallucination results and outperforms the state-of-the-art. Xin Yu 0002, Basura Fernando, Richard I. Hartley, Fatih Porikli |
CVPR | 2 |
| 2018 | Action Anticipation with RBF Kernelized Feature Mapping RNN
Yuge Shi, Basura Fernando, Richard I. Hartley |
ECCV (10) | 2 |
| 2018 | Face Super-Resolution Guided by Facial Component Heatmaps
Xin Yu 0002, Basura Fernando, Bernard Ghanem, Fatih Porikli, Richard I. Hartley |
ECCV (9) | 2 |
| 2018 | Neural Algebra of ClassifiersabstractThe world is fundamentally compositional, so it is natural to think of visual recognition as the recognition of basic visually primitives that are composed according to well-defined rules. This strategy allows us to recognize unseen complex concepts from simple visual primitives. However, the current trend in visual recognition follows a data greedy approach where huge amounts of data are required to learn models for any desired visual concept. In this paper, we build on the compositionality principle and develop an "algebra" to compose classifiers for complex visual concepts. To this end, we learn neural network modules to perform boolean algebra operations on simple visual classifiers. Since these modules form a complete functional set, a classifier for any complex visual concept defined as a boolean expression of primitives can be obtained by recursively applying the learned modules, even if we do not have a single training sample. As our experiments show, using such a framework, we can compose classifiers for complex visual concepts outperforming standard baselines on two well-known visual recognition benchmarks. Finally, we present a qualitative analysis of our method and its properties. Rodrigo Santa Cruz, Basura Fernando, Anoop Cherian, Stephen Gould |
WACV | 2 |
| 2018 | Action Recognition with Dynamic Image NetworksabstractWe introduce the concept of dynamic image, a novel compact representation of videos useful for video analysis, particularly in combination with convolutional neural networks (CNNs). A dynamic image encodes temporal data such as RGB or optical flow videos by using the concept of 'rank pooling'. The idea is to learn a ranking machine that captures the temporal evolution of the data and to use the parameters of the latter as a representation. We call the resulting representation dynamic image because it summarizes the video dynamics in addition to appearance. This powerful idea allows to convert any video to an image so that existing CNN models pre-trained with still images can be immediately extended to videos. We also present an efficient approximate rank pooling operator that runs two orders of magnitude faster than the standard ones with any loss in ranking performance and can be formulated as a CNN layer. To demonstrate the power of the representation, we introduce a novel four stream CNN architecture which can learn from RGB and optical flow frames as well as from their dynamic image representations. We show that the proposed network achieves state-of-the-art performance, 95.5 and 72.5 percent accuracy, in the UCF101 and HMDB51, respectively. Hakan Bilen, Basura Fernando, Efstratios Gavves, Andrea Vedaldi |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2017 | Generalized Rank Pooling for Activity RecognitionabstractMost popular deep models for action recognition split video sequences into short sub-sequences consisting of a few frames, frame-based features are then pooled for recognizing the activity. Usually, this pooling step discards the temporal order of the frames, which could otherwise be used for better recognition. Towards this end, we propose a novel pooling method, generalized rank pooling (GRP), that takes as input, features from the intermediate layers of a CNN that is trained on tiny sub-sequences, and produces as output the parameters of a subspace which (i) provides a low-rank approximation to the features and (ii) preserves their temporal order. We propose to use these parameters as a compact representation for the video sequence, which is then used in a classification setup. We formulate an objective for computing this subspace as a Riemannian optimization problem on the Grassmann manifold, and propose an efficient conjugate gradient scheme for solving it. Experiments on several activity recognition datasets show that our scheme leads to state-of-the-art performance. Anoop Cherian, Basura Fernando, Mehrtash Harandi, Stephen Gould |
CVPR | 2 |
| 2017 | DeepPermNet: Visual Permutation LearningabstractWe present a principled approach to uncover the structure of visual data by solving a novel deep learning task coined visual permutation learning. The goal of this task is to find the permutation that recovers the structure of data from shuffled versions of it. In the case of natural images, this task boils down to recovering the original image from patches shuffled by an unknown permutation matrix. Unfortunately, permutation matrices are discrete, thereby posing difficulties for gradient-based methods. To this end, we resort to a continuous approximation of these matrices using doubly-stochastic matrices which we generate from standard CNN predictions using Sinkhorn iterations. Unrolling these iterations in a Sinkhorn network layer, we propose DeepPermNet, an end-to-end CNN model for this task. The utility of DeepPermNet is demonstrated on two challenging computer vision problems, namely, (i) relative attributes learning and (ii) self-supervised representation learning. Our results show state-of-the-art performance on the Public Figures and OSR benchmarks for (i) and on the classification and segmentation tasks on the PASCAL VOC dataset for (ii). Rodrigo Santa Cruz, Basura Fernando, Anoop Cherian, Stephen Gould |
CVPR | 2 |
| 2017 | Self-Supervised Video Representation Learning with Odd-One-Out NetworksabstractWe propose a new self-supervised CNN pre-training technique based on a novel auxiliary task called odd-one-out learning. In this task, the machine is asked to identify the unrelated or odd element from a set of otherwise related elements. We apply this technique to self-supervised video representation learning where we sample subsequences from videos and ask the network to learn to predict the odd video subsequence. The odd video subsequence is sampled such that it has wrong temporal order of frames while the even ones have the correct temporal order. Therefore, to generate a odd-one-out question no manual annotation is required. Our learning machine is implemented as multi-stream convolutional neural network, which is learned end-to-end. Using odd-one-out networks, we learn temporal representations for videos that generalizes to other related tasks such as action recognition. On action classification, our method obtains 60.3% on the UCF101 dataset using only UCF101 data for training which is approximately 10% better than current state-of-the-art self-supervised learning methods. Similarly, on HMDB51 dataset we outperform self-supervised state-of-the art methods by 12.7% on action classification task. Basura Fernando, Hakan Bilen, Efstratios Gavves, Stephen Gould |
CVPR | 1 |
| 2017 | Guided Open Vocabulary Image Captioning with Constrained Beam SearchabstractExisting image captioning models do not generalize well to out-of-domain images containing novel scenes or objects.This limitation severely hinders the use of these models in real world applications dealing with images in the wild.We address this problem using a flexible approach that enables existing deep captioning architectures to take advantage of image taggers at test time, without re-training.Our method uses constrained beam search to force the inclusion of selected tag words in the output, and fixed, pretrained word embeddings to facilitate vocabulary expansion to previously unseen tag words.Using this approach we achieve state of the art results for out-of-domain captioning on MSCOCO (and improved results for in-domain captioning).Perhaps surprisingly, our results significantly outperform approaches that incorporate the same tag predictions into the learning algorithm.We also show that we can significantly improve the quality of generated ImageNet captions by leveraging ground-truth labels. Peter Anderson 0001, Basura Fernando, Mark Johnson 0001, Stephen Gould |
EMNLP | 2 |
| 2017 | Encouraging LSTMs to Anticipate Actions Very EarlyabstractIn contrast to the widely studied problem of recognizing an action given a complete sequence, action anticipation aims to identify the action from only partially available videos. As such, it is therefore key to the success of computer vision applications requiring to react as early as possible, such as autonomous navigation. In this paper, we propose a new action anticipation method that achieves high prediction accuracy even in the presence of a very small percentage of a video sequence. To this end, we develop a multi-stage LSTM architecture that leverages context-aware and action-aware features, and introduce a novel loss function that encourages the model to predict the correct class as early as possible. Our experiments on standard benchmark datasets evidence the benefits of our approach; We outperform the state-of-the-art action anticipation methods for early prediction by a relative increase in accuracy of 22.0% on JHMDB-21, 14.0% on UT-Interaction and 49.9% on UCF-101. Mohammad Sadegh Ali Akbarian, Fatemehsadat Saleh, Mathieu Salzmann, Basura Fernando, Lars Petersson, Lars Andersson |
ICCV | 4 |
| 2017 | Discriminatively Learned Hierarchical Rank Pooling Networks
Basura Fernando, Stephen Gould |
Int. J. Comput. Vis. | 1 |
| 2017 | Rank Pooling for Action RecognitionabstractWe propose a function-based temporal pooling method that captures the latent structure of the video sequence data - e.g., how frame-level features evolve over time in a video. We show how the parameters of a function that has been fit to the video data can serve as a robust new video representation. As a specific example, we learn a pooling function via ranking machines. By learning to rank the frame-level features of a video in chronological order, we obtain a new representation that captures the video-wide temporal dynamics of a video, suitable for action recognition. Other than ranking functions, we explore different parametric models that could also explain the temporal changes in videos. The proposed functional pooling methods, and rank pooling in particular, is easy to interpret and implement, fast to compute and effective in recognizing a wide variety of actions. We evaluate our method on various benchmarks for generic action, fine-grained action and gesture recognition. Results show that rank pooling brings an absolute improvement of 7-10 average pooling baseline. At the same time, rank pooling is compatible with and complementary to several appearance and local motion based methods and features, such as improved trajectories and deep learning features. Basura Fernando, Efstratios Gavves, José Oramas M., Amir Ghodrati, Tinne Tuytelaars |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2016 | Dynamic Image Networks for Action RecognitionabstractWe introduce the concept of dynamic image, a novel compact representation of videos useful for video analysis especially when convolutional neural networks (CNNs) are used. The dynamic image is based on the rank pooling concept and is obtained through the parameters of a ranking machine that encodes the temporal evolution of the frames of the video. Dynamic images are obtained by directly applying rank pooling on the raw image pixels of a video producing a single RGB image per video. This idea is simple but powerful as it enables the use of existing CNN models directly on video data with fine-tuning. We present an efficient and effective approximate rank pooling operator, speeding it up orders of magnitude compared to rank pooling. Our new approximate rank pooling CNN layer allows us to generalize dynamic images to dynamic feature maps and we demonstrate the power of our new representations on standard benchmarks in action recognition achieving state-of-the-art performance. Hakan Bilen, Basura Fernando, Efstratios Gavves, Andrea Vedaldi, Stephen Gould |
CVPR | 2 |
| 2016 | Discriminative Hierarchical Rank Pooling for Activity RecognitionabstractWe present hierarchical rank pooling, a video sequence encoding method for activity recognition. It consists of a network of rank pooling functions which captures the dynamics of rich convolutional neural network features within a video sequence. By stacking non-linear feature functions and rank pooling over one another, we obtain a high capacity dynamic encoding mechanism, which is used for action recognition. We present a method for jointly learning the video representation and activity classifier parameters. Our method obtains state-of-the art results on three important activity recognition benchmarks: 76.7% on Hollywood2, 66.9% on HMDB51 and, 91.4% on UCF101. Basura Fernando, Peter Anderson 0001, Marcus Hutter, Stephen Gould |
CVPR | 1 |
| 2016 | SPICE: Semantic Propositional Image Caption Evaluation
Peter Anderson 0001, Basura Fernando, Mark Johnson 0001, Stephen Gould |
ECCV (5) | 2 |
| 2016 | Learning End-to-end Video Classification with Rank-PoolingabstractWe introduce a new model for representation learning and classification of video sequences. Our model is based on a convolutional neural network coupled with a novel temporal pooling layer. The temporal pooling layer relies on an inner-optimization problem to efficiently encode temporal semantics over arbitrarily long video clips into a fixed-length vector representation. Importantly, the representation and classification parameters of our model can be estimated jointly in an end-to-end manner by formulating learning as a bilevel optimization problem. Furthermore, the model can make use of any existing convolutional neural network architecture (e.g., AlexNet or VGG) without modification or introduction of additional parameters. We demonstrate our approach on action and activity recognition tasks. Basura Fernando, Stephen Gould |
ICML | 1 |
| 2015 | Modeling video evolution for action recognitionabstractIn this paper we present a method to capture video-wide temporal information for action recognition. We postulate that a function capable of ordering the frames of a video temporally (based on the appearance) captures well the evolution of the appearance within the video. We learn such ranking functions per video via a ranking machine and use the parameters of these as a new video representation. The proposed method is easy to interpret and implement, fast to compute and effective in recognizing a wide variety of actions. We perform a large number of evaluations on datasets for generic action recognition (Hollywood2 and HMDB51), fine-grained actions (MPII- cooking activities) and gestures (Chalearn). Results show that the proposed method brings an absolute improvement of 7–10%, while being compatible with and complementary to further improvements in appearance and local motion based methods. Basura Fernando, Efstratios Gavves, José Oramas M., Amir Ghodrati, Tinne Tuytelaars |
CVPR | 1 |
| 2015 | Dataset fingerprints: Exploring image collections through data miningabstractAs the amount of visual data increases, so does the need for summarization tools that can be used to explore large image collections and to quickly get familiar with their content. In this paper, we propose dataset fingerprints, a new and powerful method based on data mining that extracts meaningful patterns from a set of images. The discovered patterns are compositions of discriminative mid-level features that co-occur in several images. Compared to earlier work, ours stands out because i) it's fully unsupervised, ii) discovered patterns cover large parts of the images, often corresponding to full objects or meaningful parts thereof, and iii) different patterns are connected based on co-occurrence, allowing a user to “browse” the images from one pattern to the next and to group patterns in a semantically meaningful manner. Konstantinos Rematas, Basura Fernando, Frank Dellaert, Tinne Tuytelaars |
CVPR | 2 |
| 2015 | Learning to Rank Based on SubsequencesabstractWe present a supervised learning to rank algorithm that effectively orders images by exploiting the structure in image sequences. Most often in the supervised learning to rank literature, ranking is approached either by analysing pairs of images or by optimizing a list-wise surrogate loss function on full sequences. In this work we propose MidRank, which learns from moderately sized sub-sequences instead. These sub-sequences contain useful structural ranking information that leads to better learnability during training and better generalization during testing. By exploiting sub-sequences, the proposed MidRank improves ranking accuracy considerably on an extensive array of image ranking applications and datasets. Basura Fernando, Efstratios Gavves, Damien Muselet, Tinne Tuytelaars |
ICCV | 1 |
| 2015 | Guiding the Long-Short Term Memory Model for Image Caption GenerationabstractIn this work we focus on the problem of image caption generation. We propose an extension of the long short term memory (LSTM) model, which we coin gLSTM for short. In particular, we add semantic information extracted from the image as extra input to each unit of the LSTM block, with the aim of guiding the model towards solutions that are more tightly coupled to the image content. Additionally, we explore different length normalization strategies for beam search to avoid bias towards short sentences. On various benchmark datasets such as Flickr8K, Flickr30K and MS COCO, we obtain results that are on par with or better than the current state-of-the-art. Xu Jia 0012, Efstratios Gavves, Basura Fernando, Tinne Tuytelaars |
ICCV | 3 |
| 2015 | Location recognition over large time lags
Basura Fernando, Tatiana Tommasi, Tinne Tuytelaars |
Comput. Vis. Image Underst. | 1 |
| 2015 | Local Alignments for Fine-Grained Categorization
Efstratios Gavves, Basura Fernando, Cees Snoek, Arnold W. M. Smeulders, Tinne Tuytelaars |
Int. J. Comput. Vis. | 2 |
| 2015 | Joint cross-domain classification and subspace learning for unsupervised adaptation
Basura Fernando, Tatiana Tommasi, Tinne Tuytelaars |
Pattern Recognit. Lett. | 1 |
| 2014 | Color features for dating historical color imagesabstractEstimating the age of historical photographs is a challenging task for human beings. Only recently this task has been addressed in computational image analysis perspective. The characteristics of the device used to acquire each photograph are discriminative features for this task. We aim at extracting such characteristics from a historical color photographs. The acquisition device mainly effects two properties of the colors: the distribution of their derivatives and the angles drawn by three consecutive pixels in the RGB space. We propose two color features that take advantage of these observations. We show that these two color descriptors (namely color derivatives and color angles) attain the state-of-the-art in the context of image dating. Basura Fernando, Damien Muselet, Rahat Khan, Tinne Tuytelaars |
ICIP | 1 |
| 2014 | Mining Mid-level Features for Image Classification
Basura Fernando, Élisa Fromont, Tinne Tuytelaars |
Int. J. Comput. Vis. | 1 |
| 2013 | Unsupervised Visual Domain Adaptation Using Subspace AlignmentabstractIn this paper, we introduce a new domain adaptation (DA) algorithm where the source and target domains are represented by subspaces described by eigenvectors. In this context, our method seeks a domain adaptation solution by learning a mapping function which aligns the source subspace with the target one. We show that the solution of the corresponding optimization problem can be obtained in a simple closed form, leading to an extremely fast algorithm. We use a theoretical result to tune the unique hyper parameter corresponding to the size of the subspaces. We run our method on various datasets and show that, despite its intrinsic simplicity, it outperforms state of the art DA methods. Basura Fernando, Amaury Habrard, Marc Sebban, Tinne Tuytelaars |
ICCV | 1 |
| 2013 | Mining Multiple Queries for Image Retrieval: On-the-Fly Learning of an Object-Specific Mid-level RepresentationabstractIn this paper we present a new method for object retrieval starting from multiple query images. The use of multiple queries allows for a more expressive formulation of the query object including, e.g., different viewpoints and/or viewing conditions. This, in turn, leads to more diverse and more accurate retrieval results. When no query images are available to the user, they can easily be retrieved from the internet using a standard image search engine. In particular, we propose a new method based on pattern mining. Using the minimal description length principle, we derive the most suitable set of patterns to describe the query object, with patterns corresponding to local feature configurations. This results in a powerful object-specific mid-level image representation. The archive can then be searched efficiently for similar images based on this representation, using a combination of two inverted file systems. Since the patterns already encode local spatial information, good results on several standard image retrieval datasets are obtained even without costly re-ranking based on geometric verification. Basura Fernando, Tinne Tuytelaars |
ICCV | 1 |
| 2013 | Fine-Grained Categorization by AlignmentsabstractThe aim of this paper is fine-grained categorization without human interaction. Different from prior work, which relies on detectors for specific object parts, we propose to localize distinctive details by roughly aligning the objects using just the overall shape, since implicit to fine-grained categorization is the existence of a super-class shape shared among all classes. The alignments are then used to transfer part annotations from training images to test images (supervised alignment), or to blindly yet consistently segment the object in a number of regions (unsupervised alignment). We furthermore argue that in the distinction of fine grained sub-categories, classification-oriented encodings like Fisher vectors are better suited for describing localized information than popular matching oriented features like HOG. We evaluate the method on the CU-2011 Birds and Stanford Dogs fine-grained datasets, outperforming the state-of-the-art. Efstratios Gavves, Basura Fernando, Cees Snoek, Arnold W. M. Smeulders, Tinne Tuytelaars |
ICCV | 2 |
| 2012 | Discriminative feature fusion for image classificationabstractBag-of-words-based image classification approaches mostly rely on low level local shape features. However, it has been shown that combining multiple cues such as color, texture, or shape is a challenging and promising task which can improve the classification accuracy. Most of the state-of-the-art feature fusion methods usually aim to weight the cues without considering their statistical dependence in the application at hand. In this paper, we present a new logistic regression-based fusion method, called LRFF, which takes advantage of the different cues without being tied to any of them. We also design a new marginalized kernel by making use of the output of the regression model. We show that such kernels, surprisingly ignored so far by the computer vision community, are particularly well suited to achieve image classification tasks. We compare our approach with existing methods that combine color and shape on three datasets. The proposed learning-based feature fusion process clearly outperforms the state-of-the art fusion methods for image classification. Basura Fernando, Élisa Fromont, Damien Muselet, Marc Sebban |
CVPR | 1 |
| 2012 | Effective Use of Frequent Itemset Mining for Image Classification
Basura Fernando, Élisa Fromont, Tinne Tuytelaars |
ECCV (1) | 1 |
| 2012 | Supervised learning of Gaussian mixture models for visual vocabulary generation
Basura Fernando, Élisa Fromont, Damien Muselet, Marc Sebban |
Pattern Recognit. | 1 |