Cornelia Fermüller

dblp:f/CorneliaFermuller · DBLP profile ↗
← Back
135ranked-venue papers
26as first author
35since 2021 · last 2025
0000-0003-2044-2386ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 120 · 23 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 57 · 15 first-author · 7 since 2021Systems, architecture and hardware · 38 · 17 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-authorDatabases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Repurposing Pre-trained Video Diffusion Models for Event-based Video Interpolation
abstract
Video Frame Interpolation aims to recover realistic missing frames between observed frames, generating a high-frame-rate video from a low-frame-rate video. However, without additional guidance, the large motion between frames makes this problem ill-posed. Event-based Video Frame Interpolation (EVFI) addresses this challenge by using sparse, high-temporal-resolution event measurements as motion guidance. This guidance allows EVFI methods to significantly outperform frame-only methods. However, to date, EVFI methods have relied on a limited set of paired event-frame training data, severely limiting their performance and generalization capabilities. In this work, we overcome the limited data challenge by adapting pre-trained video diffusion models trained on internet-scale datasets to EVFI. We experimentally validate our approach on real-world EVFI datasets, including a new one that we introduce. Our method outperforms existing methods and generalizes across cameras far better than existing approaches.
Jingxi Chen, Brandon Yushan Feng, Haoming Cai, Tianfu Wang 0007, Levi Burner, Dehao Yuan, Cornelia Fermüller, Christopher A. Metzler, Yiannis Aloimonos
CVPR7
2025 NatSGLD: A Dataset with Speech, Gesture, Logic, and Demonstration for Robot Learning in Natural Human-Robot Interaction
abstract
Recent advances in multimodal Human-Robot Interaction (HRI) datasets emphasize the integration of speech and gestures, allowing robots to absorb explicit knowledge and tacit understanding. However, existing datasets primarily focus on elementary tasks like object pointing and pushing, limiting their applicability to complex domains. They prioritize simpler human command data but place less emphasis on training robots to correctly interpret tasks and respond appropriately. To address these gaps, we present the NatSGLD dataset, which was collected using a Wizard of Oz (WoZ) method, where participants interacted with a robot they believed to be autonomous. NatSGLD records humans' multimodal commands (speech and gestures), each paired with a demonstration trajectory and a Linear Temporal Logic (LTL) formula that provides a ground-truth interpretation of the commanded tasks. This dataset serves as a foundational resource for research at the intersection of HRI and machine learning. By providing multimodal inputs and detailed annotations, NatSGLD enables exploration in areas such as multimodal instruction following, plan recognition, and human-advisable reinforcement learning from demonstrations. We release the dataset and code under the MIT License at https://www.snehesh.com/natsgld/to support future HRI research.
Snehesh Shrestha, Yantian Zha, Saketh Banagiri, Ge Gao 0001, Yiannis Aloimonos, Cornelia Fermüller
HRI6
2025 Learning Normal Flow Directly from Events
Dehao Yuan, Levi Burner, Jiayi Wu 0005, Jingxi Chen, Yiannis Aloimonos, Cornelia Fermüller
ICCV7
2025 Air-FAR: Fast and Adaptable Routing for Aerial Navigation in Large-Scale Complex Unknown Environments
abstract
This paper presents a novel approach for realtime 3D navigation in large-scale complex environments by introducing a hierarchical 3D visibility graph (V-graph) and an efficient path search method. The proposed algorithm addresses the computational challenges of V-graph construction and shortest path search on the graph simultaneously. By introducing hierarchical 3D V-graph construction with heuristic visibility update, the 3D V-graph is constructed in$O\left(K \cdot n^{2} \log n\right)$time, which guarantees real-time performance. The proposed iterative divide-and-conquer path search method can achieve near-optimal path solutions within the constraints of realtime operations. The algorithm ensures efficient 3D V-graph construction and path search. Extensive simulated and realworld environments validated that our algorithm reduces the travel time by 42%, achieves up to 24.8% higher trajectory efficiency, and runs faster than most benchmarks by orders of magnitude in complex environments. The code and developed simulator have been open-sourced to facilitate future research.
Botao He, Guofei Chen, Cornelia Fermüller, Yiannis Aloimonos, Ji Zhang 0003
ICRA3
2025 Search-Based Path Planning in Interactive Environments Among Movable Obstacles
abstract
This paper investigates Path planning Among Movable Obstacles (PAMO), which seeks a minimum cost collision-free path among static obstacles from start to goal while allowing the robot to push away movable obstacles (i.e., objects) along its path when needed. To develop planners that are complete and optimal for PAMO, the planner has to search a giant state space involving both the location of the robot as well as the locations of the objects, which grows exponentially with respect to the number of objects. This paper leverages a simple yet under-explored idea that, only a small fraction of this giant state space needs to be searched during planning as guided by a heuristic, and most of the objects far away from the robot are intact, which thus leads to runtime efficient algorithms. Based on this idea, this paper introduces two PAMO formulations, i.e., bi-objective and resource constrained problems in an occupancy grid, and develops PAMO*, a planning method with completeness and solution optimality guarantees, to solve the two problems. We then further extend PAMO* to hybrid-state PAMO* to plan in continuous spaces with high-fidelity interaction between the robot and the objects. Our results show that, PAMO* can often find optimal solutions within a second in cluttered maps with up to 400 objects.
Zhongqiang Ren, Bunyod Suvonov, Guofei Chen, Botao He, Yijie Liao, Cornelia Fermüller, Ji Zhang 0003
ICRA6
2025 FeelAnyForce: Estimating Contact Force Feedback from Tactile Sensation for Vision-Based Tactile Sensors
abstract
In this paper, we tackle the problem of estimating 3D contact forces using vision-based tactile sensors. In particular, our goal is to estimate contact forces over a large range (up to 15 N) on any objects while generalizing across different vision-based tactile sensors. Thus, we collected a dataset of over 200K indentations using a robotic arm that pressed various indenters onto a GelSight Mini sensor mounted on a force sensor and then used the data to train a multi-head transformer for force regression. Strong generalization is achieved via accurate data collection and multi-objective optimization that leverages depth contact images. Despite being trained only on primitive shapes and textures, the regressor achieves a mean absolute error of 4% on a dataset of unseen real-world objects. We further evaluate our approach's generalization capability to other GelSight mini and DIGIT sensors, and propose a reproducible calibration procedure for other sensors. Finally, the method was evaluated on real-world tasks, including weighing objects and controlling the deformation of delicate objects. Supplementary material and demo are available at http://prg.cs.umd.edu/FeelAnyForce.
Amir-Hossein Shahidzadeh, Gabriele M. Caddeo, Koushik Alapati, Lorenzo Natale, Cornelia Fermüller, Yiannis Aloimonos
ICRA5
2025 ViewActive: Active viewpoint optimization from a single image
abstract
When observing objects, humans benefit from their spatial visualization and mental rotation ability to envision potential optimal viewpoints based on the current observation. This capability is crucial for enabling robots to achieve efficient and robust scene perception during operation, as optimal viewpoints provide essential and informative features for accurately representing scenes in 2D images, thereby enhancing downstream tasks.To endow robots with this human-like active viewpoint optimization capability, we propose ViewActive, a modernized machine learning approach drawing inspiration from aspect graph, which provides viewpoint optimization guidance based solely on the current 2D image input. Specifically, we introduce the 3D Viewpoint Quality Field (VQF), a compact and consistent representation for viewpoint quality distribution similar to an aspect graph, composed of three general-purpose viewpoint quality metrics: self-occlusion ratio, occupancy-aware surface normal entropy, and visual entropy. We utilize pre-trained image encoders to extract robust visual and semantic features, which are then decoded into the 3D VQF, allowing our model to generalize effectively across diverse objects, including unseen categories. The lightweight ViewActive network (72 FPS on a single GPU) significantly enhances the performance of state-of-the-art object recognition pipelines and can be integrated into real-time motion planning for robotic applications. Our code and dataset are available here https://github.com/jiayi-wu-umd/ViewActive.
Jiayi Wu 0005, Xiaomin Lin 0002, Botao He, Cornelia Fermüller, Yiannis Aloimonos
IROS4
2025 VioPose: Violin Performance 4D Pose Estimation by Hierarchical Audiovisual Inference
abstract
Musicians delicately control their bodies to generate music. Sometimes, their motions are too subtle to be captured by the human eye. To analyze how they move to produce the music, we need to estimate precise 4D human pose (3D pose over time). However, current state-of-the-art (SoTA) visual pose estimation algorithms struggle to produce accurate monocular 4D poses because of occlusions, partial views, and human-object interactions. They are limited by the viewing angle, pixel density, and sampling rate of the cameras and fail to estimate fast and subtle movements, such as in the musical effect of vibrato. We leverage the direct causal relationship between the music produced and the human motions creating them to address these challenges. We propose VioPose: a novel multimodal network that hierarchically estimates dynamics. High-level features are cascaded to low-level features and integrated into Bayesian updates. Our architecture is shown to produce accurate pose sequences, facilitating precise motion analysis, and outperforms SoTA. As part of this work, we collected the largest and the most diverse calibrated violin-playing dataset, including video, sound, and 3D motion capture poses. Code and dataset can be found in our project page https://sj-yoo.info/viopose/.
Seong Jong Yoo, Snehesh Shrestha, Irina Muresanu, Cornelia Fermüller
WACV4
2024 Decodable and Sample Invariant Continuous Object Encoder
abstract
We propose Hyper-Dimensional Function Encoding (HDFE). Given samples of a continuous object (e.g. a function), HDFE produces an explicit vector representation of the given object, invariant to the sample distribution and density. Sample distribution and density invariance enables HDFE to consistently encode continuous objects regardless of their sampling, and therefore allows neural networks to receive continuous objects as inputs for machine learning tasks, such as classification and regression. Besides, HDFE does not require any training and is proved to map the object into an organized embedding space, which facilitates the training of the downstream tasks. In addition, the encoding is decodable, which enables neural networks to regress continuous objects by regressing their encodings. Therefore, HDFE serves as an interface for processing continuous objects. We apply HDFE to function-to-function mapping, where vanilla HDFE achieves competitive performance with the state-of-the-art algorithm. We apply HDFE to point cloud surface normal estimation, where a simple replacement from PointNet to HDFE leads to 12\% and 15\% error reductions in two benchmarks. In addition, by integrating HDFE into the PointNet-based SOTA network, we improve the SOTA baseline by 2.5\% and 1.7\% on the same benchmarks.
Dehao Yuan, Furong Huang, Cornelia Fermüller, Yiannis Aloimonos
ICLR3
2024 A Linear Time and Space Local Point Cloud Geometry Encoder via Vectorized Kernel Mixture (VecKM)
abstract
We propose VecKM, a local point cloud geometry encoder that is descriptive and efficient to compute. VecKM leverages a unique approach by vectorizing a kernel mixture to represent the local point cloud. Such representation’s descriptiveness is supported by two theorems that validate its ability to reconstruct and preserve the similarity of the local shape. Unlike existing encoders downsampling the local point cloud, VecKM constructs the local geometry encoding using all neighboring points, producing a more descriptive encoding. Moreover, VecKM is efficient to compute and scalable to large point cloud inputs: VecKM reduces the memory cost from $(n^2+nKd)$ to $(nd+np)$; and reduces the major runtime cost from computing $nK$ MLPs to $n$ MLPs, where $n$ is the size of the point cloud, $K$ is the neighborhood size, $d$ is the encoding dimension, and $p$ is a marginal factor. The efficiency is due to VecKM’s unique factorizable property that eliminates the need of explicitly grouping points into neighbors. In the normal estimation task, VecKM demonstrates not only 100x faster inference speed but also highest accuracy and strongest robustness. In classification and segmentation tasks, integrating VecKM as a preprocessing module achieves consistently better performance than the PointNet, PointNet++, and point transformer baselines, and runs consistently faster by up to 10 times.
Dehao Yuan, Cornelia Fermüller, Tahseen Rabbani, Furong Huang, Yiannis Aloimonos
ICML2
2024 AcTExplore: Active Tactile Exploration on Unknown Objects
abstract
Tactile exploration plays a crucial role in understanding object structures for fundamental robotics tasks such as grasping and manipulation. However, efficiently exploring such objects using tactile sensors is challenging, primarily due to the large-scale unknown environments and limited sensing coverage of these sensors. To this end, we present AcTExplore, an active tactile exploration method driven by reinforcement learning for object reconstruction at scales that automatically explores the object surfaces in a limited number of steps. Through sufficient exploration, our algorithm incrementally collects tactile data and reconstructs 3D shapes of the objects as well, which can serve as a representation for higher-level downstream tasks. Our method achieves an average of 95.97% IoU coverage on unseen YCB objects while just being trained on primitive shapes.
Amir-Hossein Shahidzadeh, Seong Jong Yoo, Pavan Mantripragada, Chahat Deep Singh, Cornelia Fermüller, Yiannis Aloimonos
ICRA5
2024 Vector Symbolic Sub-objects Classifiers as Manifold Analogues
abstract
Vector Symbolic Architectures (VSAs) generally consist of a hyper-algebra that is defined on points in a space of vectors. However, such a space is typically defined in usual ways, such as Euclidean Spaces for real vectors, Hamming Spaces for binary vectors, and so on. In any empirical setting, such as Artificial Intelligence, Machine Learning, Data Science, etc., observations in these spaces tend to produce subspaces of these broader spaces. As a result, these empirically-derived subspaces enable VSAs to make predictions through how severely the subspaces deviate from what is expected. Thus, it is desirable to be able to understand how such subspaces behave. As an analogy, when one observes terrain, they should find good "roads and bridges" to navigate that terrain. In Category Theory, the idea of Topos describes how such roads and bridges should be placed. In this paper, we explore the relationship between VSAs and Topoi. Namely, we show how a Topos-like representation can be constructed from empirical observations of vectors and demonstrate this on a practical example using Hyperdimensional Computing (HDC) on dense binary hypervectors. Our results indicate that a Topos can be effectively constructed for a dataset and the resulting space of vectors is biased to reflect the compositional aspects of the data. We conclude that such a Topos can be used to better guide the construction of VSAs for downstream tasks, in lieu of the original space of vectors, much like manifolds in Topology.
Renato Faraone, Peter Sutor Jr., Cornelia Fermüller, Yiannis Aloimonos
IJCNN3
2024 A Comparative Study of Hough Transform and PCA for Bolt Orientation Detection
abstract
In the fields of manufacturing and robotics, accurately determining the orientation of manufacturing components, such as bolts, is a critical yet challenging problem due to the limitations of existing detection methods. This study introduces a novel methodology for addressing this issue, leveraging traditional computer vision techniques, by proposing a streamlined approach that exploits the inherent geometric properties of bolts for orientation detection. Two methods are presented to ascertain the initial axis angle of the bolt: the Progressive Probabilistic Hough Transform (PPHT) and Principal Component Analysis (PCA). These methods are used in conjunction with a novel tip direction detection approach. The results of the study demonstrate consistent accuracy in angle determination, with PPHT and PCA both achieving angular deviations below ±0.5° in the simple dataset, and PCA showing enhanced robustness in the dataset degraded by shadows, with a maximum error under ±1.5°. This research not only reaffirms the viability of fundamental computer vision techniques in modern robotic applications but also sets a precedent for simple, generalisable, and reliable orientation detection solutions. These methods effectively bridge the gap between highly specialised machine learning systems, which often require tailored, complex models and extensive training data, and more universally applicable, straightforward approaches.
Antonio Gambale, Sonya A. Coleman, Dermot Kerr, Philip J. Vance, Emmett Kerr, Cornelia Fermüller, Yiannis Aloimonos
INDIN6
2024 Fingerspelling Classification for Robot Control
abstract
Improvements to human-robot interaction methods could increase the ease of use of robots in manufacturing environments. Many of these environments are noisy and therefore preclude the use of audio communication between humans or in human-robot interactions. Therefore, this paper proposes using a gesture based communication system for robot control. To that end, the VGG16 and VGG19 convolutional neural network (CNN) structures are used for gesture classification along with 3 datasets of American Sign Language (ASL) fingerspelling images. The model performance is evaluated, and modifications made to their parameters to improve performance, before applying them to robot control tasks. The results show that with parameter tuning, test accuracies of up to, 100% are achievable.
Kevin McCready, Sonya A. Coleman, Dermot Kerr, Nazmul H. Siddique, Emmett Kerr, Yiannis Aloimonos, Cornelia Fermüller
INDIN7
2024 Active Human Pose Estimation via an Autonomous UAV Agent
abstract
One of the core activities of an active observer involves moving to secure a "better" view of the scene, where the definition of "better" is task-dependent. This paper focuses on the task of human pose estimation from videos capturing a person’s activity. Self-occlusions within the scene can complicate or even prevent accurate human pose estimation. To address this, relocating the camera to a new vantage point is necessary to clarify the view, thereby improving 2D human pose estimation. This paper formalizes the process of achieving an improved viewpoint. Our proposed solution to this challenge comprises three main components: a NeRF-based Drone-View Data Generation Framework, an On-Drone Network for Camera View Error Estimation, and a Combined Planner for devising a feasible motion plan to reposition the camera based on the predicted errors for camera views. The Data Generation Framework utilizes NeRF-based methods to generate a comprehensive dataset of human poses and activities, enhancing the drone’s adaptability in various scenarios. The Camera View Error Estimation Network is designed to evaluate the current human pose and identify the most promising next viewing angles for the drone, ensuring a reliable and precise pose estimation from those angles. Finally, the combined planner incorporates these angles while considering the drone’s physical and environmental limitations, employing efficient algorithms to navigate safe and effective flight paths. This system represents a significant advancement in active 2D human pose estimation for an autonomous UAV agent, offering substantial potential for applications in aerial cinematography by improving the performance of autonomous human pose estimation and maintaining the operational safety and efficiency of UAVs.
Jingxi Chen, Botao He, Chahat Deep Singh, Cornelia Fermüller, Yiannis Aloimonos
IROS4
2024 Interactive-FAR: Interactive, Fast and Adaptable Routing for Navigation Among Movable Obstacles in Complex Unknown Environments
abstract
This paper introduces a real-time algorithm for navigating complex unknown environments cluttered with movable obstacles. Our algorithm achieves fast, adaptable routing by actively attempting to manipulate obstacles during path planning and adjusting the global plan from sensor feedback. The main contributions include an improved dynamic Directed Visibility Graph (DV-graph) for rapid global path searching, a real-time interaction planning method that adapts online from new sensory perceptions, and a comprehensive framework designed for interactive navigation in complex unknown or partially known environments. Our algorithm is capable of replanning the global path in several milliseconds. It can also attempt to move obstacles, update their affordances, and adapt strategies accordingly. Extensive experiments validate that our algorithm reduces the travel time by 33%, achieves up to 49% higher path efficiency, and runs faster than traditional methods by orders of magnitude in complex environments. It has been demonstrated to be the most efficient solution in terms of speed and efficiency for interactive navigation in environments of such complexity. We also open-source our code in the docker demo1to facilitate future research.
Botao He, Guofei Chen, Ji Zhang 0003, Cornelia Fermüller, Yiannis Aloimonos
IROS5
2024 MARVIS: Motion & Geometry Aware Real and Virtual Image Segmentation
abstract
Tasks such as autonomous navigation, 3D reconstruction, and object recognition near the water surfaces are crucial in marine robotics applications. However, challenges arise due to dynamic disturbances, e.g., light reflections and refraction from the random air-water interface, irregular liquid flow, and similar factors, which can lead to potential failures in perception and navigation systems. Traditional computer vision algorithms struggle to differentiate between real and virtual image regions, significantly complicating tasks. A virtual image region is an apparent representation formed by the redirection of light rays, typically through reflection or refraction, creating the illusion of an object’s presence without its actual physical location. This work proposes a novel approach for segmentation on real and virtual image regions, exploiting synthetic images combined with domain-invariant information, a Motion Entropy Kernel, and Epipolar Geometric Consistency. Our segmentation network does not need to be re-trained if the domain changes. We show this by deploying the same segmentation network in two different domains: simulation and the real world. By creating realistic synthetic images that mimic the complexities of the water surface, we provide fine-grained training data for our network (MARVIS) to discern between real and virtual images effectively. By motion & geometry-aware design choices and through comprehensive experimental analysis, we achieve state-of-the-art real-virtual image segmentation performance in unseen real world domain, achieving an IoU over 78% and a F1-Score over 86% while ensuring a small computational footprint. MARVIS offers over 43 FPS (8 FPS) inference rates on a single GPU (CPU core). Our code and dataset are available here https://github.com/jiayi-wu-umd/MARVIS.
Jiayi Wu 0005, Xiaomin Lin 0002, Shahriar Negahdaripour, Cornelia Fermüller, Yiannis Aloimonos
IROS4
2024 Temporally Consistent Atmospheric Turbulence Mitigation with Neural Representations
abstract
Atmospheric turbulence, caused by random fluctuations in the atmosphere's refractive index, introduces complex spatio-temporal distortions in imagery captured at long range. Video Atmospheric Turbulence Mitigation (ATM) aims to restore videos affected by these distortions. However, existing video ATM methods, both supervised and self-supervised, struggle to maintain temporally consistent mitigation across frames, leading to visually incoherent results. This limitation arises from the stochastic nature of atmospheric turbulence, which varies across space and time. Inspired by the observation that atmospheric turbulence induces high-frequency temporal variations, we propose ConVRT, a novel framework for consistent video restoration through turbulence. ConVRT introduces a neural video representation that explicitly decouples spatial and temporal information into a spatial content field and a temporal deformation field, enabling targeted regularization of the network's temporal representation capability. By leveraging the low-pass filtering properties of the regularized temporal representations, ConVRT effectively mitigates turbulence-induced temporal frequency variations and promotes temporal consistency. Furthermore, our training framework seamlessly integrates supervised pre-training on synthetic turbulence data with self-supervised learning on real-world videos, significantly improving the temporally consistent mitigation of ATM methods on diverse real-world data. More information can be found on our project page: https://convrt-2024.github.io/
Haoming Cai, Jingxi Chen, Brandon Yushan Feng, Weiyun Jiang, Mingyang Xie, Kevin Zhang 0003, Cornelia Fermüller, Yiannis Aloimonos, Ashok Veeraraghavan, Christopher A. Metzler
NeurIPS7
2024 Context in Human Action through Motion Complementarity
abstract
Motivated by Goldman’s Theory of Human Action - a framework in which action decomposes into 1) base physical movements, and 2) the context in which they occur - we propose a novel learning formulation for motion and context, where context is derived as the complement to motion. More specifically, we model physical movement through the adoption of Therbligs, a set of elemental physical motions centered around object manipulation. Context is modeled through the use of a contrastive mutual information loss that formulates context information as the action information not contained within movement information. We empirically prove the utility brought by this separation of representation, showing sizable improvements in action recognition and action anticipation accuracies for a variety of models. We present results over two object manipulation datasets: EPIC Kitchens 100, and 50 Salads.
Eadom Dessalene, Michael Maynord, Cornelia Fermüller, Yiannis Aloimonos
WACV3
2023 Therbligs in Action: Video Understanding through Motion Primitives
abstract
In this paper we introduce a rule-based, compositional, and hierarchical modeling of action using Therbligs as our atoms. Introducing these atoms provides us with a consistent, expressive, contact-centered representation of action. Over the atoms we introduce a differentiable method of rule-based reasoning to regularize for logical consistency. Our approach is complementary to other approaches in that the Therblig-based representations produced by our architecture augment rather than replace existing architectures' representations. We release the first Therblig-centered an-notations over two popular video datasets - EPIC Kitchens 100 and 50-Salads. We also broadly demonstrate benefits to adopting Therblig representations through evaluation on the following tasks: action segmentation, action anticipation, and action recognition - observing an average 10.5%/7.53%/6.5% relative improvement, respectively, over EPIC Kitchens and an average 8.9%/6.63%/4.8% relative improvement, respectively, over 50 Salads. Code and data will be made publicly available.
Eadom Dessalene, Michael Maynord, Cornelia Fermüller, Yiannis Aloimonos
CVPR3
2023 t-ConvESN: Temporal Convolution-Readout for Random Recurrent Neural Networks
Matthew Evanusa, Vaishnavi Patil, Michelle Girvan, Joel Goodman, Cornelia Fermüller, Yiannis Aloimonos
ICANN (6)5
2023 Mid-Vision Feedback
Michael Maynord, Eadom Dessalene, Cornelia Fermüller, Yiannis Aloimonos
ICLR3
2023 TTCDist: Fast Distance Estimation From an Active Monocular Camera Using Time-to-Contact
abstract
Distance estimation from vision is fundamental for a myriad of robotic applications such as navigation, manipu-lation, and planning. Inspired by the mammal's visual system, which gazes at specific objects, we develop two novel constraints relating time-to-contact, acceleration, and distance that we call the$\tau$-constraint and$\Phi$-constraint. They allow an active (moving) camera to estimate depth efficiently and accurately while using only a small portion of the image. The constraints are applicable to range sensing, sensor fusion, and visual servoing. We successfully validate the proposed constraints with two experiments. The first applies both constraints in a trajectory estimation task with a monocular camera and an Inertial Measurement Unit (IMU). Our methods achieve 30-70% less average trajectory error while running$25\times$and$6.2\times$faster than the popular Visual-Inertial Odometry methods VINS-Mono and ROVIO respectively. The second experiment demonstrates that when the constraints are used for feedback with efference copies the resulting closed loop system's eigenvalues are invariant to scaling of the applied control signal. We believe these results indicate the$\tau$and$\Phi$constraint's potential as the basis of robust and efficient algorithms for a multitude of robotic applications.
Levi Burner, Nitin J. Sanket, Cornelia Fermüller, Yiannis Aloimonos
ICRA3
2023 WorldGen: A Large Scale Generative Simulator
abstract
In the era of deep learning, data is the critical determining factor in the performance of neural network models. Generating large datasets suffers from various challenges such as scalability, cost efficiency and photorealism. To avoid expensive and strenuous dataset collection and annotations, researchers have inclined towards computer-generated datasets. However, a lack of photorealism and a limited amount of computer-aided data has bounded the accuracy of network predictions. To this end, we present WorldGen - an open source framework to automatically generate countless structured and unstructured 3D photorealistic scenes such as city view, object collection, and object fragmentation along with its rich ground truth annotation data. WorldGen being a generative model gives the user full access and control to features such as texture, object structure, motion, camera and lens properties for better generalizability by diminishing the data bias in the network. We demonstrate the effectiveness of WorldGen by evaluating deep optical flow. We hope such a tool can open doors for future research in a myriad of domains related to robotics and computer vision by reducing manual labor and cost for acquiring rich and high-quality data.
Chahat Deep Singh, Riya Kumari, Cornelia Fermüller, Nitin J. Sanket, Yiannis Aloimonos
ICRA3
2023 Forecasting Action Through Contact Representations From First Person Video
abstract
Human actions involving hand manipulations are structured according to the making and breaking of hand-object contact, and human visual understanding of action is reliant on anticipation of contact as is demonstrated by pioneering work in cognitive science. Taking inspiration from this, we introduce representations and models centered on contact, which we then use in action prediction and anticipation. We annotate a subset of the EPIC Kitchens dataset to include time-to-contact between hands and objects, as well as segmentations of hands and objects. Using these annotations we train the Anticipation Module, a module producing Contact Anticipation Maps and Next Active Object Segmentations - novel low-level representations providing temporal and spatial characteristics of anticipated near future action. On top of the Anticipation Module we apply Egocentric Object Manipulation Graphs (Ego-OMG), a framework for action anticipation and prediction. Ego-OMG models longer term temporal semantic relations through the use of a graph modeling transitions between contact delineated action states. Use of the Anticipation Module within Ego-OMG produces state-of-the-art results, achieving 1st and 2 place on the unseen and seen test sets, respectively, of the EPIC Kitchens Action Anticipation Challenge, and achieving state-of-the-art results on the tasks of action anticipation and action prediction over EPIC Kitchens. We perform ablation studies over characteristics of the Anticipation Module to evaluate their utility.
Eadom Dessalene, Chinmaya Devaraj, Michael Maynord, Cornelia Fermüller, Yiannis Aloimonos
IEEE Trans. Pattern Anal. Mach. Intell.4
2022 Brain-Inspired Hyperdimensional Computing for Ultra-Efficient Edge AI
abstract
Hyperdimensional Computing (HDC) is rapidly emerging as an attractive alternative to traditional deep learning algorithms. Despite the profound success of Deep Neural Networks (DNNs) in many domains, the amount of computational power and storage that they demand during training makes deploying them in edge devices very challenging if not infeasible. This, in turn, inevitably necessitates streaming the data from the edge to the cloud which raises serious concerns when it comes to availability, scalability, security, and privacy. Further, the nature of data that edge devices often receive from sensors is inherently noisy. However, DNN algorithms are very sensitive to noise, which makes accomplishing the required learning tasks with high accuracy immensely difficult. In this paper, we aim at providing a comprehensive overview of the latest advances in HDC. HDC aims at realizing real-time performance and robustness through using strategies that more closely model the human brain. HDC is, in fact, motivated by the observation that the human brain operates on high-dimensional data representations. In HDC, objects are thereby encoded with high-dimensional vectors which have thousands of elements. In this paper, we will discuss the promising robustness of HDC algorithms against noise along with the ability to learn from little data. Further, we will present the outstanding synergy between HDC and beyond von Neumann architectures and how HDC opens doors for efficient learning at the edge due to the ultra-lightweight implementation that it needs, contrary to traditional DNNs.
Hussam Amrouch, Mohsen Imani, Xun Jiao 0002, Yiannis Aloimonos, Cornelia Fermüller, Dehao Yuan, Dongning Ma, Hamza Errahmouni Barkam, Paul R. Genssler, Peter Sutor Jr.
CODES+ISSS5
2022 DiffPoseNet: Direct Differentiable Camera Pose Estimation
abstract
Current deep neural network approaches for camera pose estimation rely on scene structure for 3D motion estimation, but this decreases the robustness and thereby makes cross-dataset generalization difficult. In contrast, classical approaches to structure from motion estimate 3D motion utilizing optical flow and then compute depth. Their accuracy, however, depends strongly on the quality of the optical flow. To avoid this issue, direct methods have been proposed, which separate 3D motion from depth estimation, but compute 3D motion using only image gradients in the form of normal flow. In this paper, we introduce a network NFlowNet, for normal flow estimation which is used to enforce robust and direct constraints. In particular, normal flow is used to estimate relative camera pose based on the cheirality (depth positivity) constraint. We achieve this by formulating the optimization problem as a differentiable cheirality layer, which allows for end-to-end learning of camera pose. We perform extensive qualitative and quantitative evaluation of the proposed DiffPoseNet's sensitivity to noise and its generalization across datasets. We compare our approach to existing state-of-the-art methods on KITTI, TartanAir, and TUM-RGBD datasets.
Chethan Parameshwara, Gokul Hari, Cornelia Fermüller, Nitin J. Sanket, Yiannis Aloimonos
CVPR3
2022 Gluing Neural Networks Symbolically Through Hyperdimensional Computing
abstract
Hyperdimensional Computing affords simple, yet powerful operations to create long Hyperdimensional Vectors (hypervectors) that can efficiently encode information, be used for learning, and are dynamic enough to be modified on the fly. In this paper, we explore the notion of using binary hypervectors to directly encode the final, classifying output signals of neural networks in order to fuse differing networks together at the symbolic level. This allows multiple neural networks to work together to solve a problem, with little additional overhead. Output signals just before classification are encoded as hyper-vectors and bundled together through consensus summation to train a classification hypervector. This process can be performed iteratively and even on single neural networks by instead making a consensus of multiple classification hypervectors. We find that this outperforms the state of the art, or is on a par with it, while using very little overhead, as hypervector operations are extremely fast and efficient in comparison to the neural networks. This consensus process can learn online and even grow or lose models in real time. Hypervectors act as memories that can be stored, and even further bundled together over time, affording life long learning capabilities. Additionally, this consensus structure inherits the benefits of Hyperdimensional Computing, without sacrificing the performance of modern Machine Learning. This technique can be extrapolated to virtually any neural model, and requires little modification to employ - one simply requires recording the output signals of networks when presented with a testing example.
Peter Sutor Jr., Dehao Yuan, Douglas Summers-Stay, Cornelia Fermüller, Yiannis Aloimonos
IJCNN4
2021 3D Motion Analysis with Event-based Sensors
Cornelia Fermüller
ICPRAM1
2021 0-MMS: Zero-Shot Multi-Motion Segmentation With A Monocular Event Camera
abstract
Segmentation of moving objects in dynamic scenes is a key process in scene understanding for navigation tasks. Classical cameras suffer from motion blur in such scenarios rendering them effete. On the contrary, event cameras, because of their high temporal resolution and lack of motion blur, are tailor-made for this problem. We present an approach for monocular multi-motion segmentation, which combines bottom-up feature tracking and top-down motion compensation into a unified pipeline, which is the first of its kind to our knowledge. Using the events within a time-interval, our method segments the scene into multiple motions by splitting and merging. We further speed up our method by using the concept of motion propagation and cluster keyslices.The approach was successfully evaluated on both challenging real-world and synthetic scenarios from the EV-IMO, EED, and MOD datasets and outperformed the state-of-the-art detection rate by 12%, achieving a new state-of-the-art average detection rate of 81.06%, 94.2% and 82.35% on the aforementioned datasets. To enable further research and systematic evaluation of multi-motion segmentation, we present and open-source a new dataset/benchmark called MOD++, which includes challenging sequences and extensive data stratification in-terms of camera and object motion, velocity magnitudes, direction, and rotational speeds.
Chethan Parameshwara, Nitin J. Sanket, Chahat Deep Singh, Cornelia Fermüller, Yiannis Aloimonos
ICRA4
2021 MorphEyes: Variable Baseline Stereo For Quadrotor Navigation
abstract
Morphable design and depth-based visual control are two upcoming trends leading to advancements in the field of quadrotor autonomy. Stereo-cameras have struck the perfect balance of weight and accuracy of depth estimation but suffer from the problem of depth range being limited and dictated by the baseline chosen at design time. In this paper, we present a framework for quadrotor navigation based on a stereo camera system whose baseline can be adapted on-the-fly. We present a method to calibrate the system at a small number of discrete baselines and interpolate the parameters for the entire baseline range. We present an extensive theoretical analysis of calibration and synchronization errors. We showcase three different applications of such a system for quadrotor navigation: (a) flying through a forest, (b) flying through an unknown shaped/location static/dynamic gap, and (c) accurate 3D pose detection of an independently moving object. We show that our variable baseline system is more accurate and robust in all three scenarios. To our knowledge, this is the first work that applies the concept of morphable design to achieve a variable baseline stereo vision system on a quadrotor.
Nitin J. Sanket, Chahat Deep Singh, Varun Asthana, Cornelia Fermüller, Yiannis Aloimonos
ICRA4
2021 SpikeMS: Deep Spiking Neural Network for Motion Segmentation
abstract
Spiking Neural Networks (SNN) are the so-called third generation of neural networks which attempt to more closely match the functioning of the biological brain. They inherently encode temporal data, allowing for training with less energy usage and can be extremely energy efficient when coded on neuromorphic hardware. In addition, they are well suited for tasks involving event-based sensors, which match the event-based nature of the SNN. However, SNNs have not been as effectively applied to real-world, large-scale tasks as standard Artificial Neural Networks (ANNs) due to the algorithmic and training complexity. To exacerbate the situation further, the input representation is unconventional and requires careful analysis and deep understanding. In this paper, we propose SpikeMS, the first deep encoder-decoder SNN architecture for the real-world large-scale problem of motion segmentation using the event-based DVS camera as input. To accomplish this, we introduce a novel spatio-temporal loss formulation that includes both spike counts and classification labels in conjunction with the use of new techniques for SNN backpropagation. In addition, we show that SpikeMS is capable of incremental predictions, or predictions from smaller amounts of test data than it is trained on. This is invaluable for providing outputs even with partial input data for low-latency applications and those requiring fast predictions. We evaluated SpikeMS on challenging synthetic and real-world sequences from EV-IMO, EED and MOD datasets and achieving results on a par with a comparable ANN method, but using potentially 50 times less power.
Chethan Parameshwara, Cornelia Fermüller, Nitin J. Sanket, Matthew Evanusa, Yiannis Aloimonos
IROS3
2021 NudgeSeg: Zero-Shot Object Segmentation by Repeated Physical Interaction
abstract
Recent advances in object segmentation have demonstrated that deep neural networks excel at object segmentation for specific classes in color and depth images. However, their performance is dictated by the number of classes and objects used for training, thereby hindering generalization to never seen objects or zero-shot samples. To exacerbate the problem further, object segmentation using image frames rely on recognition and pattern matching cues. Instead, we utilize the ‘active’ nature of a robot and their ability to ‘interact’ with the environment to induce additional geometric constraints for segmenting zero-shot samples.In this paper, we present the first framework to segment unknown objects in a cluttered scene by repeatedly ‘nudging’ at the objects and moving them to obtain additional motion cues at every step using only a monochrome monocular camera. We call our framework NudgeSeg. These motion cues are used to refine the segmentation masks. We successfully test our approach to segment novel objects in various cluttered scenes and provide an extensive study with image and motion segmentation methods. We show an impressive average detection rate of over 86% on zero-shot objects.
Chahat Deep Singh, Nitin J. Sanket, Chethan Parameshwara, Cornelia Fermüller, Yiannis Aloimonos
IROS4
2021 Topology-Aware Non-Rigid Point Cloud Registration
abstract
In this paper, we introduce a non-rigid registration pipeline for pairs of unorganized point clouds that may be topologically different. Standard warp field estimation algorithms, even under robust, discontinuity-preserving regularization, tend to produce erratic motion estimates on boundaries associated with 'close-to-open' topology changes. We overcome this limitation by exploiting backward motion: in the opposite motion direction, a 'close-to-open' event becomes 'open-to-close', which is by default handled correctly. At the core of our approach lies a general, topology-agnostic warp field estimation algorithm, similar to those employed in recently introduced dynamic reconstruction systems from RGB-D input. We improve motion estimation on boundaries associated with topology changes in an efficient post-processing phase. Based on both forward and (inverted) backward warp hypotheses, we explicitly detect regions of the deformed geometry that undergo topological changes by means of local deformation criteria and broadly classify them as 'contacts' or 'separations'. Subsequently, the two motion hypotheses are seamlessly blended on a local basis, according to the type and proximity of detected events. Our method achieves state-of-the-art motion estimation accuracy on the MPI Sintel dataset. Experiments on a custom dataset with topological event annotations demonstrate the effectiveness of our pipeline in estimating motion on event boundaries, as well as promising performance in explicit topological event detection.
Konstantinos Zampogiannis, Cornelia Fermüller, Yiannis Aloimonos
IEEE Trans. Pattern Anal. Mach. Intell.2
2021 Joint direct estimation of 3D geometry and 3D motion using spatio temporal gradients
Francisco Barranco, Cornelia Fermüller, Yiannis Aloimonos, Eduardo Ros Vidal
Pattern Recognit.2
2020 Learning Visual Motion Segmentation Using Event Surfaces
abstract
Event-based cameras have been designed for scene motion perception - their high temporal resolution and spatial data sparsity converts the scene into a volume of boundary trajectories and allows to track and analyze the evolution of the scene in time. Analyzing this data is computationally expensive, and there is substantial lack of theory on dense-in-time object motion to guide the development of new algorithms; hence, many works resort to a simple solution of discretizing the event stream and converting it to classical pixel maps, which allows for application of conventional image processing methods. In this work we present a Graph Convolutional neural network for the task of scene motion segmentation by a moving camera. We convert the event stream into a 3D graph in (x,y,t) space and keep per-event temporal information. The difficulty of the task stems from the fact that unlike in metric space, the shape of an object in (x,y,t) space depends on its motion and is not the same across the dataset. We discuss properties of of the event data with respect to this 3D recognition problem, and show that our Graph Convolutional architecture is superior to PointNet++. We evaluate our method on the state of the art event-based motion segmentation dataset - EV-IMO and perform comparisons to a frame-based method proposed by its authors. Our ablation studies show that increasing the event slice width improves the accuracy, and how subsampling and edge configurations affect the network performance.
Anton Mitrokhin, Zhiyuan Hua, Cornelia Fermüller, Yiannis Aloimonos
CVPR3
2020 Network Deconvolution
Chengxi Ye, Matthew Evanusa, Anton Mitrokhin, Tom Goldstein, James A. Yorke, Cornelia Fermüller, Yiannis Aloimonos
ICLR7
2020 EVDodgeNet: Deep Dynamic Obstacle Dodging with Event Cameras
abstract
Dynamic obstacle avoidance on quadrotors requires low latency. A class of sensors that are particularly suitable for such scenarios are event cameras. In this paper, we present a deep learning based solution for dodging multiple dynamic obstacles on a quadrotor with a single event camera and on-board computation. Our approach uses a series of shallow neural networks for estimating both the ego-motion and the motion of independently moving objects. The networks are trained in simulation and directly transfer to the real world without any fine-tuning or retraining. We successfully evaluate and demonstrate the proposed approach in many real-world experiments with obstacles of different shapes and sizes, achieving an overall success rate of 70% including objects of unknown shape and a low light testing scenario. To our knowledge, this is the first deep learning - based solution to the problem of dynamic obstacle avoidance using event cameras on a quadrotor. Finally, we also extend our work to the pursuit task by merely reversing the control policy, proving that our navigation stack can cater to different scenarios.
Nitin J. Sanket, Chethan Parameshwara, Chahat Deep Singh, Ashwin V. Kuruttukulam, Cornelia Fermüller, Davide Scaramuzza 0001, Yiannis Aloimonos
ICRA5
2020 Unsupervised Learning of Dense Optical Flow, Depth and Egomotion with Event-Based Sensors
abstract
We present an unsupervised learning pipeline for dense depth, optical flow and egomotion estimation for autonomous driving applications, using the event-based output of the Dynamic Vision Sensor (DVS) as input. The backbone of our pipeline is a bioinspired encoder-decoder neural network architecture - ECN. To train the pipeline, we introduce a covariance normalization technique which resembles the lateral inhibition mechanism found in animal neural systems.Our work is the first monocular pipeline that generates dense depth and optical flow from sparse event data only, and is able to transfer from day to night scenes without any additional training. The network works in self-supervised mode and has just 150k parameters. We evaluate our pipeline on the MVSEC self driving dataset and present results for depth, optical flow and and egomotion estimation. Thanks to the efficient design, we are able to achieve inference rates of 300 FPS on a single Nvidia 1080Ti GPU. Our experiments demonstrate significant improvements upon works that used deep learning on event data, as well as the ability to perform well during both day and night.
Chengxi Ye, Anton Mitrokhin, Cornelia Fermüller, James A. Yorke, Yiannis Aloimonos
IROS3
2019 EV-IMO: Motion Segmentation Dataset and Learning Pipeline for Event Cameras
abstract
We present the first event-based learning approach for motion segmentation in indoor scenes and the first event-based dataset - EV-IMO- which includes accurate pixel-wise motion masks, egomotion and ground truth depth. Our approach is based on an efficient implementation of the SfM learning pipeline using a low parameter neural network architecture on event data. In addition to camera egomotion and a dense depth map, the network estimates independently moving object segmentation at the pixel-level and computes per-object 3D translational velocities of moving objects. We also train a shallow network with just 40k parameters, which is able to compute depth and egomotion. Our EV-IMO dataset features 32 minutes of indoor recording with up to 3 fast moving objects in the camera field of view. The objects and the camera are tracked using a VICON®motion capture system. By 3D scanning the room and the objects, ground truth of the depth map and pixel-wise object masks are obtained. We then train and evaluate our learning pipeline on EV-IMO and demonstrate that it is well suited for scene constrained robotics applications. SUPPLEMENTARY MATERIAL The supplementary video, code, trained models, appendix and a dataset will be made available at http://prg.cs.umd.edu/EV-IMO.html.
Anton Mitrokhin, Chengxi Ye, Cornelia Fermüller, Yiannis Aloimonos, Tobi Delbruck
IROS3
2019 SalientDSO: Bringing Attention to Direct Sparse Odometry
abstract
Although cluttered indoor scenes have a lot of useful high-level semantic information which can be used for mapping and localization, most visual odometry (VO) algorithms rely on the usage of geometric features such as points, lines, and planes. Lately, driven by this idea, the joint optimization of semantic labels and estimating odometry has gained popularity in the robotics community. This joint optimization method is accurate but is generally very slow. At the same time, in the vision community, direct and sparse approaches for VO have stricken the right balance between speed and accuracy. We merge the successes of these two communities and present a preprocessing method to incorporate semantic information in the form of visual saliency to direct sparse odometry (DSO)-a highly successful direct sparse VO algorithm. We also present a framework to filter the visual saliency based on scene parsing. Our framework SalientDSO relies on the widely successful deep learning-based approaches for visual saliency and scene parsing, which drives the feature selection for obtaining highly accurate and robust VO even in the presence of as few as 40 point features per frame. We provide an extensive quantitative evaluation of SalientDSO on the ICL-NUIM and the TUM monoVO data sets and show that we outperform DSO and ORB-simultaneous localization and mapping-two very popular state-of-the-art approaches in the literature. We also collect and publicly release a CVL-UMD data set which contains two indoor cluttered sequences on which we show qualitative evaluations. To the best of our knowledge, this is the first paper to use visual saliency and scene parsing to drive the feature selection in direct VO.
Huai-Jen Liang, Nitin J. Sanket, Cornelia Fermüller, Yiannis Aloimonos
IEEE Trans Autom. Sci. Eng.3
2018 Evenly Cascaded Convolutional Networks
abstract
We introduce Evenly Cascaded convolutional Network (ECN), a neural network taking inspiration from the cascade algorithm of wavelet analysis. ECN employs two feature streams - a low-level and high-level steam. At each layer these streams interact, such that low-level features are modulated using advanced perspectives from the high-level stream. ECN is evenly structured through resizing feature map dimensions by a consistent ratio, which removes the burden of ad-hoc specification of feature map dimensions. ECN produces easily interpretable features maps, a result whose intuition can be understood in the context of scale-space theory. We demonstrate that ECN's design facilitates the training process through providing easily trainable shortcuts. We report new state-of-the-art results for small networks, without the need for additional treatment such as pruning or compression - a consequence of ECN's simple structure and direct training. A 6-layered ECN design with under 500k parameters achieves 95.24% and 78.99% accuracy on CIFAR-10 and CIFAR-100 datasets, respectively, outperforming the current state-of-the-art on small parameter networks, and a 3 million parameter ECN produces results competitive to the state-of-the-art.
Chengxi Ye, Chinmaya Devaraj, Michael Maynord, Cornelia Fermüller, Yiannis Aloimonos
IEEE BigData4
2018 Real-Time Clustering and Multi-Target Tracking Using Event-Based Sensors
abstract
Clustering is crucial for many computer vision applications such as robust tracking, object detection and segmentation. This work presents a real-time clustering technique that takes advantage of the unique properties of event-based vision sensors. Since event-based sensors trigger events only when the intensity changes, the data is sparse, with low redundancy. Thus, our approach redefines the well-known mean-shift clustering method using asynchronous events instead of conventional frames. The potential of our approach is demonstrated in a multi-target tracking application using Kalman filters to smooth the trajectories. We evaluated our method on an existing dataset with patterns of different shapes and speeds, and a new dataset that we collected. The sensor was attached to the Baxter robot in an eye-in-hand setup monitoring real-world objects in an action manipulation task. Clustering accuracy achieved an F-measure of 0.95, reducing the computational cost by 88% compared to the frame-based method. The average error for tracking was 2.5 pixels and the clustering achieved a consistent number of clusters along time.
Francisco Barranco, Cornelia Fermüller, Eduardo Ros Vidal
IROS2
2018 Seeing Behind the Scene: Using Symmetry to Reason About Objects in Cluttered Environments
abstract
Symmetry is a common property shared by the majority of man-made objects. This paper presents a novel bottom-up approach for segmenting symmetric objects and recovering their symmetries from 3D pointclouds of natural scenes. Candidate rotational and reflectional symmetries are detected by fitting symmetry axes/planes to the geometry of the smooth surfaces extracted from the scene. Individual symmetries are used as constraints for the foreground segmentation problem that uses symmetry as a global grouping principle. Evaluation on a challenging dataset shows that our approach can reliably segment objects and extract their symmetries from incomplete 3D reconstructions of highly cluttered scenes, outperforming state-of-the-art methods by a wide margin.
Aleksandrs Ecins, Cornelia Fermüller, Yiannis Aloimonos
IROS2
2018 Event-Based Moving Object Detection and Tracking
abstract
Event-based vision sensors, such as the Dynamic Vision Sensor (DVS), are ideally suited for real-time motion analysis. The unique properties encompassed in the readings of such sensors provide high temporal resolution, superior sensitivity to light and low latency. These properties provide the grounds to estimate motion efficiently and reliably in the most sophisticated scenarios, but these advantages come at a price - modern event-based vision sensors have extremely low resolution, produce a lot of noise and require the development of novel algorithms to handle the asynchronous event stream. This paper presents a new, efficient approach to object tracking with asynchronous cameras. We present a novel event stream representation which enables us to utilize information about the dynamic (temporal)component of the event stream. The 3D geometry of the event stream is approximated with a parametric model to motion-compensate for the camera (without feature tracking or explicit optical flow computation), and then moving objects that don't conform to the model are detected in an iterative process. We demonstrate our framework on the task of independent motion detection and tracking, where we use the temporal model inconsistencies to locate differently moving objects in challenging situations of very fast motion.
Anton Mitrokhin, Cornelia Fermüller, Chethan Parameshwara, Yiannis Aloimonos
IROS2
2018 cilantro: A Lean, Versatile, and Efficient Library for Point Cloud Data Processing
abstract
We introduce Cilantro, an open-source C++ library for geometric and general-purpose point cloud data processing. The library provides functionality that covers low-level point cloud operations, spatial reasoning, various methods for point cloud segmentation and generic data clustering, flexible algorithms for robust or local geometric alignment, model fitting, as well as powerful visualization tools. To accommodate all kinds of workflows, Cilantro is almost fully templated, and most of its generic algorithms operate in arbitrary data dimension. At the same time, the library is easy to use and highly expressive, promoting a clean and concise coding style. Cilantro is highly optimized, has a minimal set of external dependencies, and supports rapid development of performant point cloud processing software in a wide variety of contexts.
Konstantinos Zampogiannis, Cornelia Fermüller, Yiannis Aloimonos
ACM Multimedia2
2018 Image Understanding using vision and reasoning through Scene Description Graph
Somak Aditya, Yezhou Yang, Chitta Baral, Yiannis Aloimonos, Cornelia Fermüller
Comput. Vis. Image Underst.5
2018 Prediction of Manipulation Actions
Cornelia Fermüller, Yezhou Yang, Konstantinos Zampogiannis, Francisco Barranco, Michael Pfeiffer 0001
Int. J. Comput. Vis.1
2017 Fast task-specific target detection via graph based constraints representation and checking
abstract
We present a framework for fast target detection in real-world robotics applications. Considering that an intelligent agent attends to a task-specific object target during execution, our goal is to detect the object efficiently. We propose the concept of early recognition, which influences the candidate proposal process to achieve fast and reliable detection performance. To check the target constraints efficiently, we put forward a novel policy which generates a sub-optimal checking order, and we prove that it has bounded time cost compared to the optimal checking sequence, which is not achievable in polynomial time. Experiments on two different scenarios: 1) rigid object and 2) non-rigid body part detection validate our pipeline. To show that our method is widely applicable, we further present a human-robot interaction system based on our non-rigid body part detection.
Wentao Luan, Yezhou Yang, Cornelia Fermüller, John S. Baras
ICRA3
2017 What can i do around here? Deep functional scene understanding for cognitive robots
abstract
For robots that have the capability to interact with the physical environment through their end effectors, understanding the surrounding scenes is not merely a task of image classification or object recognition. To perform actual tasks, it is critical for the robot to have a functional understanding of the visual scene. Here, we address the problem of localization and recognition of functional areas in an arbitrary indoor scene, formulated as a two-stage deep learning based detection pipeline. A new scene functionality testing-bed, which is compiled from two publicly available indoor scene datasets, is used for evaluation. Our method is evaluated quantitatively on the new dataset, demonstrating the ability to perform efficient recognition of functional areas from arbitrary indoor scenes. We also demonstrate that our detection model can be generalized to novel indoor scenes by cross validating it with images from two different datasets.
Chengxi Ye, Yezhou Yang, Ren Mao, Cornelia Fermüller, Yiannis Aloimonos
ICRA4
2016 Reliable Attribute-Based Object Recognition Using High Predictive Value Classifiers
Wentao Luan, Yezhou Yang, Cornelia Fermüller, John S. Baras
ECCV (3)3
2016 Cluttered scene segmentation using the symmetry constraint
abstract
Although modern object segmentation algorithms can deal with isolated objects in simple scenes, segmenting non-convex objects in cluttered environments remains a challenging task. We introduce a novel approach for segmenting unknown objects in partial 3D pointclouds that utilizes the powerful concept of symmetry. First, 3D bilateral symmetries in the scene are detected efficiently by extracting and matching surface normal edge curves in the pointcloud. Symmetry hypotheses are then used to initialize a segmentation process that finds points of the scene that are consistent with each of the detected symmetries. We evaluate our approach on a dataset of 3D pointcloud scans of tabletop scenes. We demonstrate that the use of the symmetry constraint enables our approach to correctly segment objects in challenging configurations and to outperform current state-of-the-art approaches.
Aleksandrs Ecins, Cornelia Fermüller, Yiannis Aloimonos
ICRA2
2016 LightNet: A Versatile, Standalone Matlab-based Environment for Deep Learning
abstract
LightNet is a lightweight, versatile, purely Matlab-based deep learning framework. The idea underlying its design is to provide an easy-to-understand, easy-to-use and efficient computational platform for deep learning research. The implemented framework supports major deep learning architectures such as Multilayer Perceptron Networks (MLP), Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN). The framework also supports both CPU and GPU computation, and the switch between them is straightforward. Different applications in computer vision, natural language processing and robotics are demonstrated as experiments.
Chengxi Ye, Chen Zhao 0009, Yezhou Yang, Cornelia Fermüller, Yiannis Aloimonos
ACM Multimedia4
2016 Guest Editorial: Special Section on CVPR 2014
abstract
The papers in this special section were presented at the IEEE Computer Vision and Pattern Recognition (CVPR), June, 2014, jointly sponsored by the IEEE and the Computer Vision Foundation.
Ronen Basri, Cornelia Fermüller, Aleix Martinez, René Vidal
IEEE Trans. Pattern Anal. Mach. Intell.2
2015 Robot Learning Manipulation Action Plans by "Watching" Unconstrained Videos from the World Wide Web
abstract
In order to advance action generation and creation in robots beyond simple learned schemas we need computational tools that allow us to automatically interpret and represent human actions. This paper presents a system that learns manipulation action plans by processing unconstrained videos from the World Wide Web. Its goal is to robustly generate the sequence of atomic actions of seen longer actions in video in order to acquire knowledge for robots. The lower level of the system consists of two convolutional neural network (CNN) based recognition modules, one for classifying the hand grasp type and the other for object recognition. The higher level is a probabilistic manipulation action grammar based parsing module that aims at generating visual sentences for robot manipulation. Experiments conducted on a publicly available unconstrained video dataset show that the system is able to learn manipulation actions by ``watching'' unconstrained videos with high accuracy.
Yezhou Yang, Cornelia Fermüller, Yiannis Aloimonos
AAAI3
2015 Learning the Semantics of Manipulation Action
abstract
In this paper we present a formal computational framework for modeling manipulation actions. The introduced formalism leads to semantics of manipulation action and has applications to both observing and understanding human manipulation actions as well as executing them with a robotic mechanism (e.g. a humanoid robot). It is based on a Combinatory Categorial Grammar. The goal of the introduced framework is to: (1) represent manipulation actions with both syntax and semantic parts, where the semantic part employs $\lambda$-calculus; (2) enable a probabilistic semantic parsing schema to learn the $\lambda$-calculus representation of manipulation action from an annotated action corpus of videos; (3) use (1) and (2) to develop a system that visually observes manipulation actions and understands their meaning while it can reason beyond observations using propositional logic and axiom schemata. The experiments conducted on a public available large manipulation action dataset validate the theoretical framework and our implementation.
Yezhou Yang, Yiannis Aloimonos, Cornelia Fermüller, Eren Erdal Aksoy
ACL (1)3
2015 Fast 2D border ownership assignment
abstract
A method for efficient border ownership assignment in 2D images is proposed. Leveraging on recent advances using Structured Random Forests (SRF) for boundary detection [8], we impose a novel border ownership structure that detects both boundaries and border ownership at the same time. Key to this work are features that predict ownership cues from 2D images. To this end, we use several different local cues: shape, spectral properties of boundary patches, and semi-global grouping cues that are indicative of perceived depth. For shape, we use HoG-like descriptors that encode local curvature (convexity and concavity). For spectral properties, such as extremal edges [28], we first learn an orthonormal basis spanned by the top K eigenvectors via PCA over common types of contour tokens [23]. For grouping, we introduce a novel mid-level descriptor that captures patterns near edges and indicates ownership information of the boundary. Experimental results over a subset of the Berkeley Segmentation Dataset (BSDS) [24] and the NYU Depth V2 [34] dataset show that our method's performance exceeds current state-of-the-art multi-stage approaches that use more complex features.
Ching Lik Teo, Cornelia Fermüller, Yiannis Aloimonos
CVPR2
2015 Grasp type revisited: A modern perspective on a classical feature for vision
abstract
The grasp type provides crucial information about human action. However, recognizing the grasp type from unconstrained scenes is challenging because of the large variations in appearance, occlusions and geometric distortions. In this paper, first we present a convolutional neural network to classify functional hand grasp types. Experiments on a public static scene hand data set validate good performance of the presented method. Then we present two applications utilizing grasp type classification: (a) inference of human action intention and (b) fine level manipulation action segmentation. Experiments on both tasks demonstrate the usefulness of grasp type as a cognitive feature for computer vision. This study shows that the grasp type is a powerful symbolic representation for action understanding, and thus opens new avenues for future research.
Yezhou Yang, Cornelia Fermüller, Yiannis Aloimonos
CVPR2
2015 Contour Detection and Characterization for Asynchronous Event Sensors
abstract
The bio-inspired, asynchronous event-based dynamic vision sensor records temporal changes in the luminance of the scene at high temporal resolution. Since events are only triggered at significant luminance changes, most events occur at the boundary of objects and their parts. The detection of these contours is an essential step for further interpretation of the scene. This paper presents an approach to learn the location of contours and their border ownership using Structured Random Forests on event-based features that encode motion, timing, texture, and spatial orientations. The classifier integrates elegantly information over time by utilizing the classification results previously computed. Finally, the contour detection and boundary assignment are demonstrated in a layer-segmentation of the scene. Experimental results demonstrate good performance in boundary detection and segmentation.
Francisco Barranco, Ching Lik Teo, Cornelia Fermüller, Yiannis Aloimonos
ICCV3
2015 Detection and Segmentation of 2D Curved Reflection Symmetric Structures
abstract
Symmetry, as one of the key components of Gestalt theory, provides an important mid-level cue that serves as input to higher visual processes such as segmentation. In this work, we propose a complete approach that links the detection of curved reflection symmetries to produce symmetry-constrained segments of structures/regions in real images with clutter. For curved reflection symmetry detection, we leverage on patch-based symmetric features to train a Structured Random Forest classifier that detects multiscaled curved symmetries in 2D images. Next, using these curved symmetries, we modulate a novel symmetry-constrained foreground-background segmentation by their symmetry scores so that we enforce global symmetrical consistency in the final segmentation. This is achieved by imposing a pairwise symmetry prior that encourages symmetric pixels to have the same labels over a MRF-based representation of the input image edges, and the final segmentation is obtained via graph-cuts. Experimental results over four publicly available datasets containing annotated symmetric structures: 1) SYMMAX-300 [38], 2) BSD-Parts, 3) Weizmann Horse (both from [18]) and 4) NY-roads [35] demonstrate the approach's applicability to different environments with state-of-the-art performance.
Ching Lik Teo, Cornelia Fermüller, Yiannis Aloimonos
ICCV2
2015 Affordance detection of tool parts from geometric features
abstract
As robots begin to collaborate with humans in everyday workspaces, they will need to understand the functions of tools and their parts. To cut an apple or hammer a nail, robots need to not just know the tool's name, but they must localize the tool's parts and identify their functions. Intuitively, the geometry of a part is closely related to its possible functions, or its affordances. Therefore, we propose two approaches for learning affordances from local shape and geometry primitives: 1) superpixel based hierarchical matching pursuit (S-HMP); and 2) structured random forests (SRF). Moreover, since a part can be used in many ways, we introduce a large RGB-Depth dataset where tool parts are labeled with multiple affordances and their relative rankings. With ranked affordances, we evaluate the proposed methods on 3 cluttered scenes and over 105 kitchen, workshop and garden tools, using ranked correlation and a weighted F-measure score [26]. Experimental results over sequences containing clutter, occlusions, and viewpoint changes show that the approaches return precise predictions that could be used by a robot. S-HMP achieves high accuracy but at a significant computational cost, while SRF provides slightly less accurate predictions but in real-time. Finally, we validate the effectiveness of our approaches on the Cornell Grasping Dataset [25] for detecting graspable regions, and achieve state-of-the-art performance.
Austin Myers, Ching Lik Teo, Cornelia Fermüller, Yiannis Aloimonos
ICRA3
2015 Learning the spatial semantics of manipulation actions through preposition grounding
abstract
In this paper, we introduce an abstract representation for manipulation actions that is based on the evolution of the spatial relations between involved objects. Object tracking in RGBD streams enables straightforward and intuitive ways to model spatial relations in 3D space. Reasoning in 3D overcomes many of the limitations of similar previous approaches, while providing significant flexibility in the desired level of abstraction. At each frame of a manipulation video, we evaluate a number of spatial predicates for all object pairs and treat the resulting set of sequences (Predicate Vector Sequences, PVS) as an action descriptor. As part of our representation, we introduce a symmetric, time-normalized pairwise distance measure that relies on finding an optimal object correspondence between two actions. We experimentally evaluate the method on the classification of various manipulation actions in video, performed at different speeds and timings and involving different objects. The results demonstrate that the proposed representation is remarkably descriptive of the high-level manipulation semantics.
Konstantinos Zampogiannis, Yezhou Yang, Cornelia Fermüller, Yiannis Aloimonos
ICRA3
2015 The Cognitive Dialogue: A new model for vision implementing common sense reasoning
Yiannis Aloimonos, Cornelia Fermüller
Image Vis. Comput.2
2014 Shadow free segmentation in still images using local density measure
abstract
Over the last decades several approaches were introduced to deal with cast shadows in background subtraction applications. However, very few algorithms exist that address the same problem for still images. In this paper we propose a figure ground segmentation algorithm to segment objects in still images affected by shadows. Instead of modeling the shadow directly in the segmentation process our approach works actively by first segmenting an object and then testing the resulting boundary for the presence of shadows and resegmenting again with modified segmentation parameters. In order to get better shadow boundary detection results we introduce a novel image preprocessing technique based on the notion of the image density map. This map improves the illumination invariance of classical filter-bank based texture description methods. We demonstrate that this texture feature improves shadow detection results. The resulting segmentation algorithm achieves good results on a new figure ground segmentation dataset with challenging illumination conditions.
Aleksandrs Ecins, Cornelia Fermüller, Yiannis Aloimonos
ICCP2
2014 Contour Motion Estimation for Asynchronous Event-Driven Cameras
abstract
This paper compares image motion estimation with asynchronous event-based cameras to Computer Vision approaches using as input frame-based video sequences. Since dynamic events are triggered at significant intensity changes, which often are at the border of objects, we refer to the event-based image motion as “contour motion.” Algorithms are presented for the estimation of accurate contour motion from local spatio-temporal information for two camera models: the dynamic vision sensor (DVS), which asynchronously records temporal changes of the luminance, and a family of new sensors which combine DVS data with intensity signals. These algorithms take advantage of the high temporal resolution of the DVS and achieve robustness using a multiresolution scheme in time. It is shown that, because of the coupling of velocity and luminance information in the event distribution, the image motion estimation problem becomes much easier with the new sensors which provide both events and image intensity than with the DVS alone. Experiments on synthesized data from computer vision benchmarks show that our algorithm on combined data outperforms computer vision methods in accuracy and can achieve real-time performance, and experiments on real data confirm the feasibility of the approach. Given that current image motion (or so-called optic flow) methods cannot estimate well at object boundaries, the approach presented here could be used complementary to optic flow techniques, and can provide new avenues for computer vision motion research.
Francisco Barranco, Cornelia Fermüller, Yiannis Aloimonos
Proc. IEEE2
2013 Action Attribute Detection from Sports Videos with Contextual Constraints
abstract
In this paper, we are interested in detecting action attributes from sports videos for event understanding and video analysis. Action attribute is a middle layer between low level motion features and high level action classes, which includes various motion patterns of human limbs and bodies and the interaction between human and objects. Successfully detecting action attributes provides a richer video description that facilitates many other important tasks, such action classification, video understanding, automatic video transcript, etc. A naive approach to deal with this challenging problem is to train a classifier for each attribute and then use them to detect attributes in novel videos independently. However, this independence assumption is often too strong, and as we show in our experiments, produces a large number of false positives in practice. We propose a novel approach that incorporates the contextual constraints for activity attribute detection. The temporal contexts within an attribute and the co-occurrence contexts between different attributes are modelled by a factorial conditional random field, which encourages agreement between different time points and attributes. The effectiveness of our methods are clearly illustrated by the experimental evaluations.
Xiaodong Yu 0002, Ching Lik Teo, Yezhou Yang, Cornelia Fermüller, Yiannis Aloimonos
BMVC4
2013 Detection of Manipulation Action Consequences (MAC)
abstract
The problem of action recognition and human activity has been an active research area in Computer Vision and Robotics. While full-body motions can be characterized by movement and change of posture, no characterization, that holds invariance, has yet been proposed for the description of manipulation actions. We propose that a fundamental concept in understanding such actions, are the consequences of actions. There is a small set of fundamental primitive action consequences that provides a systematic high-level classification of manipulation actions. In this paper a technique is developed to recognize these action consequences. At the heart of the technique lies a novel active tracking and segmentation method that monitors the changes in appearance and topological structure of the manipulated object. These are then used in a visual semantic graph (VSG) based procedure applied to the time sequence of the monitored object to recognize the action consequence. We provide a new dataset, called Manipulation Action Consequences (MAC 1.0), which can serve as test bed for other studies on this topic. Several experiments on this dataset demonstrates that our method can robustly track objects and detect their deformations and division during the manipulation. Quantitative tests prove the effectiveness and efficiency of the method.
Yezhou Yang, Cornelia Fermüller, Yiannis Aloimonos
CVPR2
2013 Embedding high-level information into low level vision: Efficient object search in clutter
abstract
The ability to search visually for objects of interest in cluttered environments is crucial for robots performing tasks in a multitude of environments. In this work, we propose a novel visual search algorithm that integrates high-level information of the target object - specifically its size and shape, with a recently introduced visual operator that rapidly clusters potential edges based on their coherence in belonging to a possible object. The output is a set of fixation points that indicate the potential location of the target object in the image. The proposed approach outperforms purely bottom-up approaches - saliency maps of Itti et al. [15], and kernel descriptors of Bo et al. [2], over two large datasets of objects in clutter collected using an RGB-Depth camera.
Ching Lik Teo, Austin Myers, Cornelia Fermüller, Yiannis Aloimonos
ICRA3
2013 Robots with language: Multi-label visual recognition using NLP
abstract
There has been a recent interest in utilizing contextual knowledge to improve multi-label visual recognition for intelligent agents like robots. Natural Language Processing (NLP) can give us labels, the correlation of labels, and the ontological knowledge about them, so we can automate the acquisition of contextual knowledge. In this paper we show how to use tools from NLP in conjunction with Vision to improve visual recognition. There are two major approaches: First, different language databases organize words according to various semantic concepts. Using these, we can build special purpose databases that can predict the labels involved given a certain context. Here we build a knowledge base for the purpose of describing common daily activities. Second, statistical language tools can provide the correlations of different labels. We show a way to learn a language model from large corpus data that exploits these correlations and propose a general optimization scheme to integrate the language model into the system. Experiments conducted on three multi-label everyday recognition tasks support the effectiveness and efficiency of our approach, with significant gains in recognition accuracies when correlation information is used.
Yezhou Yang, Ching Lik Teo, Cornelia Fermüller, Yiannis Aloimonos
ICRA3
2013 Minimalist plans for interpreting manipulation actions
abstract
Humans attribute meaning to actions, and can recognize, imitate, predict, compose from parts, and analyse complex actions performed by other humans. We have built a model of action representation and understanding which takes as input perceptual data of humans performing manipulatory actions and finds a semantic interpretation of it. It achieves this by representing actions as minimal plans based on a few primitives. The motivation for our approach is to have a description, that abstracts away the variations in the way humans perform actions. The model can be used to represent complex activities on the basis of simple actions. The primitives of these minimal plans are embodied in the physicality of the system doing the analysis. The model understands an action under observation by recognising which plan is occurring. Using primitives thus rooted in its own physical structure, the model has a semanticist and causal understanding of what it observes. Using plans, the model considers actions as well as complex activities in terms of causality, compositions, and goal achievement, enabling it to perform complex tasks like prediction of primitives, separation of interleaved actions and filtering of perceptual input. We use our model over an action dataset involving humans using hand tools on objects in a constrained universe to understand an activity it has not seen before in terms of actions whose plans it knows of. The model thus illustrates a novel approach of understanding human actions by a robot.
Anupam Guha, Yezhou Yang, Cornelia Fermüller, Yiannis Aloimonos
IROS3
2012 The image torque operator: A new tool for mid-level vision
abstract
Contours are a powerful cue for semantic image understanding. Objects and parts of objects in the image are delineated from their surrounding by closed contours which make up their boundary. In this paper we introduce a new bottom-up visual operator to capture the concept of closed contours, which we call the `Torque' operator. Its computation is inspired by the mechanical definition of torque or moment of force, and applied to image edges. The torque operator takes as input edges and computes over regions of different size a measure of how well the edges are aligned to form a closed, convex contour. We explore fundamental properties of this measure and demonstrate that it can be made a useful tool for visual attention, segmentation, and boundary edge detection by verifying its benefits on these applications.
Morimichi Nishigaki, Cornelia Fermüller, Daniel DeMenthon
CVPR2
2012 Contour-based recognition
abstract
Contour is an important cue for object recognition. In this paper, built upon the concept of torque in image space, we propose a new contour-related feature to detect and describe local contour information in images. There are two components for our proposed feature: One is a contour patch detector for detecting image patches with interesting information of object contour, which we call the Maximal/Minimal Torque Patch (MTP) detector. The other is a contour patch descriptor for characterizing a contour patch by sampling the torque values, which we call the Multi-scale Torque (MST) descriptor. Experiments for object recognition on the Caltech-101 dataset showed that the proposed contour feature outperforms other contour-related features and is on a par with many other types of features. When combing our descriptor with the complementary SIFT descriptor, impressive recognition results are observed.
Yong Xu 0007, Yuhui Quan, Zhuming Zhang, Hui Ji 0002, Cornelia Fermüller, Morimichi Nishigaki, Daniel DeMenthon
CVPR5
2012 Towards a Watson that sees: Language-guided action recognition for robots
abstract
For robots of the future to interact seamlessly with humans, they must be able to reason about their surroundings and take actions that are appropriate to the situation. Such reasoning is only possible when the robot has knowledge of how the World functions, which must either be learned or hard-coded. In this paper, we propose an approach that exploits language as an important resource of high-level knowledge that a robot can use, akin to IBM's Watson in Jeopardy!. In particular, we show how language can be leveraged to reduce the ambiguity that arises from recognizing actions involving hand-tools from video data. Starting from the premise that tools and actions are intrinsically linked, with one explaining the existence of the other, we trained a language model over a large corpus of English newswire text so that we can extract this relationship directly. This model is then used as a prior to select the best tool and action that explains the video. We formalize the approach in the context of 1) an unsupervised recognition and 2) a supervised classification scenario by an EM formulation for the former and integrating language features for the latter. Results are validated over a new hand-tool action dataset, and comparisons with state of the art STIP features showed significantly improved results when language is used. In addition, we discuss the implications of these results and how it provides a framework for integrating language into vision on other robotic applications.
Ching Lik Teo, Yezhou Yang, Hal Daumé III, Cornelia Fermüller, Yiannis Aloimonos
ICRA4
2012 Using a minimal action grammar for activity understanding in the real world
abstract
There is good reason to believe that humans use some kind of recursive grammatical structure when we recognize and perform complex manipulation activities. We have built a system to automatically build a tree structure from observations of an actor performing such activities. The activity trees that result form a framework for search and understanding, tying action to language. We explore and evaluate the system by performing experiments over a novel complex activity dataset taken using synchronized Kinect and SR4000 Time of Flight cameras. Processing of the combined 3D and 2D image data provides the necessary terminals and events to build the tree from the bottom-up. Experimental results highlight the contribution of the action grammar in: 1) providing a robust structure for complex activity recognition over real data and 2) disambiguating interleaved activities from within the same sequence.
Douglas Summers-Stay, Ching Lik Teo, Yezhou Yang, Cornelia Fermüller, Yiannis Aloimonos
IROS4
2012 Scale-space texture description on SIFT-like textons
Yong Xu 0007, Si-Bin Huang, Hui Ji 0002, Cornelia Fermüller
Comput. Vis. Image Underst.4
2011 Active scene recognition with vision and language
abstract
This paper presents a novel approach to utilizing high level knowledge for the problem of scene recognition in an active vision framework, which we call active scene recognition. In traditional approaches, high level knowledge is used in the post-processing to combine the outputs of the object detectors to achieve better classification performance. In contrast, the proposed approach employs high level knowledge actively by implementing an interaction between a reasoning module and a sensory module (Figure 1). Following this paradigm, we implemented an active scene recognizer and evaluated it with a dataset of 20 scenes and 100+ objects. We also extended it to the analysis of dynamic scenes for activity recognition with attributes. Experiments demonstrate the effectiveness of the active paradigm in introducing attention and additional constraints into the sensing process.
Xiaodong Yu 0002, Cornelia Fermüller, Ching Lik Teo, Yezhou Yang, Yiannis Aloimonos
ICCV2
2010 Learning shift-invariant sparse representation of actions
abstract
A central problem in the analysis of motion capture (MoCap) data is how to decompose motion sequences into primitives. Ideally, a description in terms of primitives should facilitate the recognition, synthesis, and characterization of actions. We propose an unsupervised learning algorithm for automatically decomposing joint movements in human motion capture (MoCap) sequences into shift-invariant basis functions. Our formulation models the time series data of joint movements in actions as a sparse linear combination of short basis functions (snippets), which are executed (or “activated”) at different positions in time. Given a set of MoCap sequences of different actions, our algorithm finds the decomposition of MoCap sequences in terms of basis functions and their activations in time. Using the tools of L1minimization, the procedure alternately solves two large convex minimizations: Given the basis functions, a variant of Orthogonal Matching Pursuit solves for the activations, and given the activations, the Split Bregman Algorithm solves for the basis functions. Experiments demonstrate the power of the decomposition in a number of applications, including action recognition, retrieval, MoCap data compression, and as a tool for classification in the diagnosis of Parkinson (a motion disorder disease).
Cornelia Fermüller, Yiannis Aloimonos, Hui Ji 0002
CVPR2
2010 An Experimental Study of Color-Based Segmentation Algorithms Based on the Mean-Shift Concept
Konstantinos Bitsakos, Cornelia Fermüller, Yiannis Aloimonos
ECCV (2)2
2009 Combining powerful local and global statistics for texture description
abstract
A texture descriptor is proposed, which combines local highly discriminative features with the global statistics of fractal geometry to achieve high descriptive power, but also invariance to geometric and illumination transformations. As local measurements SIFT features are estimated densely at multiple window sizes and discretized. On each of the discretized measurements the fractal dimension is computed to obtain the so-called multifractal spectrum, which is invariant to geometric transformations and illumination changes. Finally to achieve robustness to scale changes, a multi-scale representation of the multifractal spectrum is developed using a framelet system, that is, a redundant tight wavelet frame system. Experiments on classification demonstrate that the descriptor outperforms existing methods on the UIUC as well as the UMD high-resolution dataset.
Yong Xu 0007, Si-Bin Huang, Hui Ji 0002, Cornelia Fermüller
CVPR4
2009 Real-time shape retrieval for robotics using skip Tri-Grams
abstract
The real time requirement is an additional constraint on many intelligent applications in robotics, such as shape recognition and retrieval using a mobile robot platform. In this paper, we present a scalable approach for efficiently retrieving closed contour shapes. The contour of an object is represented by piecewise linear segments. A skip Tri-Gram is obtained by selecting three segments in the clockwise order while allowing a constant number of segments to be ¿skipped¿ in between. The main idea is to use skip Tri-Grams of the segments to implicitly encode the distant dependency of the shape. All skip Tri-Grams are used for efficiently retrieving closed contour shapes without pairwise matching feature points from two shapes. The retrieval is at least an order of magnitude faster than other state-of-the-art algorithms. We score 80% in the Bullseye retrieval test on the whole MPEG 7 shape dataset. We further test the algorithm using a mobile robot platform in an indoor environment. 8 objects are used for testing from different viewing directions, and we achieve 82% accuracy.
Konstantinos Bitsakos, Cornelia Fermüller, Yiannis Aloimonos
IROS3
2009 Active segmentation for robotics
abstract
The semantic robots of the immediate future are robots that will be able to find and recognize objects in any environment. They need the capability of segmenting objects in their visual field. In this paper, we propose a novel approach to segmentation based on the operation of fixation by an active observer. Our approach is different from current approaches: while existing works attempt to segment the whole scene at once into many areas, we segment only one image region, specifically the one containing the fixation point. Furthermore, our solution integrates monocular cues (color, texture) with binocular cues (stereo disparities and optical flow). Experiments with real imagery collected by our active robot and from the known databases demonstrate the promise of the approach.
Ajay K. Mishra, Yiannis Aloimonos, Cornelia Fermüller
IROS3
2009 Viewpoint Invariant Texture Description Using Fractal Analysis
Yong Xu 0007, Hui Ji 0002, Cornelia Fermüller
Int. J. Comput. Vis.3
2009 Robust Wavelet-Based Super-Resolution Reconstruction: Theory and Algorithm
abstract
We present an analysis and algorithm for the problem of super-resolution imaging, that is the reconstruction of HR (high-resolution) images from a sequence of LR (low-resolution) images. Super-resolution reconstruction entails solutions to two problems. One is the alignment of image frames. The other is the reconstruction of a HR image from multiple aligned LR images. Both are important for the performance of super-resolution imaging. Image alignment is addressed with a new batch algorithm, which simultaneously estimates the homographies between multiple image frames by enforcing the surface normal vectors to be the same. This approach can handle longer video sequences quite well. Reconstruction is addressed with a wavelet-based iterative reconstruction algorithm with an efficient denoising scheme. The technique is based on a new analysis of video formation. At a high level our method could be described as a better-conditioned iterative back projection scheme with an efficient regularization criteria in each iteration step. Experiments with both simulated and real data demonstrate that our approach has better performance than existing super-resolution methods. It can remove even large amounts of mixed noise without creating artifacts.
Hui Ji 0002, Cornelia Fermüller
IEEE Trans. Pattern Anal. Mach. Intell.2
2008 Measuring 1st order stretchwith a single filter
abstract
We analytically develop a filter that is able to measure the linear stretch of the transformation around a point, and present results of applying it to real signals. We show that this method is a real-time alternative solution for measuring local signal transformations. Experimentally, this method can accurately measure stretch, however, it is sensitive to shift.
Konstantinos Bitsakos, Justin Domke, Cornelia Fermüller, Yiannis Aloimonos
ICASSP3
2008 Bilateral symmetry of object silhouettes under perspective projection
abstract
Symmetry is an important property of objects and is exhibited in different forms e.g., bilateral, rotational, etc. This paper presents an algorithm for computing the bilateral symmetry of silhouettes of shallow objects under perspective distortion, exploiting the invariance of the cross ratio to projective transformations. The basic idea is to use the cross ratio to compute a number of midpoints of cross sections and then fit a straight line through them. The goodness-of-fit determines the likelihood of the line to be the axis of symmetry. We analytically estimate the midpointpsilas location as a function of the vanishing point for a given object silhouette. Hence finding the symmetry axis amounts to a 2D search in the space of vanishing points. We present experiments on two datasets as well as Internet images of symmetric objects that validate our approach.
Konstantinos Bitsakos, Hyoungjune Yi, Cornelia Fermüller
ICPR4
2007 Object Detection Using Shape Codebook
Xiaodong Yu 0002, Cornelia Fermüller, David S. Doermann
BMVC3
2007 Combining motion from texture and lines for visual navigation
abstract
Two novel methods for computing 3D structure information from video for a piecewise planar scene are presented. The first method is based on a new line constraint, which clearly separates the estimation of distance from the estimation of slant. The second method exploits the concepts of phase correlation to compute from the change of image frequencies of a textured plane, distance and slant information. The two different estimates together with structure estimates from classical image motion are combined and integrated over time using an extended Kalman filter. The estimation of the scene structure is demonstrated experimentally in a motion control algorithm that allows the robot to move along a corridor. We demonstrate the efficacy of each individual method and their combination and show that the method allows for visual navigation in textured as well as un-textured environments.
Konstantinos Bitsakos, Cornelia Fermüller
IROS3
2006 A Projective Invariant for Textures
abstract
Image texture analysis has received a lot of attention in the past years. Researchers have developed many texture signatures based on texture measurements, for the purpose of uniquely characterizing the texture. Existing texture signatures, in general, are not invariant to 3D transforms such as view-point changes and non-rigid deformations of the texture surface, which is a serious limitation for many applications. In this paper, we introduce a new texture signature, called the multifractal spectrum (MFS). It provides an efficient framework combining global spatial invariance and local robust measurements. The MFS is invariant under the bi-Lipschitz map, which includes view-point changes and non-rigid deformations of the texture surface, as well as local affine illumination changes. Experiments demonstrate that the MFS captures the essential structure of textures with quite low dimension.
Yong Xu 0007, Hui Ji 0002, Cornelia Fermüller
CVPR (2)3
2006 Wavelet-Based Super-Resolution Reconstruction: Theory and Algorithm
Hui Ji 0002, Cornelia Fermüller
ECCV (4)2
2006 A 3D Shape Constraint on Video
abstract
We propose to combine the information from multiple motion fields by enforcing a constraint on the surface normals (3D shape) of the scene in view. The fact that the shape vectors in the different views are related only by rotation can be formulated as a rank = 3 constraint. This constraint is implemented in an algorithm which solves 3D motion and structure estimation as a practical constrained minimization. Experiments demonstrate its usefulness as a tool in structure from motion providing very accurate estimates of 3D motion.
Hui Ji 0002, Cornelia Fermüller
IEEE Trans. Pattern Anal. Mach. Intell.2
2005 Integration of Motion Fields through Shape
abstract
Structure from motion from single flow fields has been studied intensively, but the integration of information from multiple flow fields has not received much attention. Here we address this problem by enforcing constraints on the shape (surface normals) of the scene in view, as opposed to constraints on the structure (depth). The advantage of integrating shape is two-fold. First, we do not need to estimate feature correspondences over multiple frames, but we only need to match patches. Second, the shape vectors in the different views are related only by rotation. This constraint on shape can be combined easily with motion estimation, thus formulating motion and structure estimation from multiple views as a practical constrained minimization problem using a rank-3 constraint. Based on this constraint, we develop a 3D motion technique, which locates through color and motion segmentation, planar patches in the scene, matches patches over multiple frames, and estimates the motion between multiple frames and the shape of the selected scene patches using the image gradients. Experiments evaluate the accuracy of the 3D motion estimation and demonstrate the motion and shape estimation of the technique by super-resolving an image sequence.
Hui Ji 0002, Cornelia Fermüller
CVPR (2)2
2005 Motion Segmentation Using Occlusions
abstract
We examine the key role of occlusions in finding independently moving objects instantaneously in a video obtained by a moving camera with a restricted field of view. In this problem, the image motion is caused by the combined effect of camera motion (egomotion), structure (depth), and the independent motion of scene entities. For a camera with a restricted field of view undergoing a small motion between frames, there exists, in general, a set of 3D camera motions compatible with the observed flow field even if only a small amount of noise is present, leading to ambiguous 3D motion estimates. If separable sets of solutions exist, motion-based clustering can detect one category of moving objects. Even if a single inseparable set of solutions is found, we show that occlusion information can be used to find ordinal depth, which is critical in identifying a new class of moving objects. In order to find ordinal depth, occlusions must not only be known, but they must also be filled (grouped) with optical flow from neighboring regions. We present a novel algorithm for filling occlusions and deducing ordinal depth under general circumstances. Finally, we describe another category of moving objects which is detected using cardinal comparisons between structure from motion and structure estimates from another source (e.g., stereo).
Abhijit S. Ogale, Cornelia Fermüller, Yiannis Aloimonos
IEEE Trans. Pattern Anal. Mach. Intell.2
2004 Bias in Shape Estimation
Hui Ji 0002, Cornelia Fermüller
ECCV (3)2
2004 Compound eye sensor for 3D ego motion estimation
abstract
We describe a compound eye vision sensor for 3D ego motion computation. Inspired by eyes of insects, we show that the compound eye sampling geometry is optimal for 3D camera motion estimation. This optimality allows us to estimate the 3D camera motion in a scene-independent and robust manner by utilizing linear equations. The mathematical model of the new sensor can be implemented in analog networks resulting in a compact computational sensor for instantaneous 3D ego motion measurements in full six degrees of freedom.
Jan Neumann, Cornelia Fermüller, Yiannis Aloimonos, Vladimir Brajovic
IROS2
2004 A hierarchy of cameras for 3D photography
Jan Neumann, Cornelia Fermüller, Yiannis Aloimonos
Comput. Vis. Image Underst.2
2003 Polydioptric Camera Design and 3D Motion Estimation
abstract
Most cameras used in computer vision applications are still based on the pinhole principle inspired by our own eyes. It has been found though that this is not necessarily the optimal image formation principle for processing visual information using a machine. We describe how to find the optimal camera for 3D motion estimation by analyzing the structure of the space formed by the light rays passing through a volume of space. Every camera corresponds to a sampling pattern in light ray space, thus the question of camera design can be rephrased as finding the optimal sampling pattern with regard to a given task. This framework suggests that large field-of-view multi-perspective (polydioptric) cameras are the optimal image sensors for 3D motion estimation. We conclude by proposing design principles for polydioptric cameras and describe an algorithm for such a camera that estimates its 3D motion in a scene independent and robust manner.
Jan Neumann, Cornelia Fermüller, Yiannis Aloimonos
CVPR (2)2
2003 Eye Design in the Plenoptic Space of Light Rays
abstract
Natural eye designs are optimized with regard to the tasks the eye-carrying organism has to perform for survival. This optimization has been performed by the process of natural evolution over many millions of years. Every eye captures a subset of the space of light rays. The information contained in this subset and the accuracy to which the eye can extract the necessary information determines an upper limit on how well an organism can perform a given task. In this work we propose a new methodology for camera design. By interpreting eyes as sample patterns in light ray space we can phrase the problem of eye design in a signal processing framework. This allows us to develop mathematical criteria for optimal eye design, which in turn enables us to build the best eye for a given task without the trial and error phase of natural evolution. The principle is evaluated on the task of 3D ego-motion estimation.
Jan Neumann, Cornelia Fermüller, Yiannis Aloimonos
ICCV2
2003 New eyes for robotics
abstract
This paper describes an imaging system that has been designed to facilitate robotic tasks of motion. The system consists of a number of cameras in a network arranged so that they sample different parts of the visual sphere. This geometric configuration has provable advantages compared to small field of view cameras for the estimation of the system's own motion and consequently the estimation of shape models from the individual cameras. The reason is that inherent ambiguities of confusion between translation and rotation disappear. Pairs of cameras may also be arranged in multiple stereo configurations which provide additional advantages for segmentation. Algorithms for the calibration of the system and the 3D motion estimation are provided.
Patrick Baker, Abhijit S. Ogale, Cornelia Fermüller, Yiannis Aloimonos
IROS3
2003 Plenoptic video geometry
Jan Neumann, Cornelia Fermüller
Vis. Comput.2
2002 Self-Calibration from Image Derivatives
Tomás Brodský, Cornelia Fermüller
Int. J. Comput. Vis.2
2002 Visual space-time geometry - A tool for perception and the imagination
abstract
Although the fundamental ideas underlying research efforts in the field of computer vision have not radically changed in the past two decades, there has been a transformation in the way work in this field is conducted. This is primarily due to the emergence of a number of tools, of both a practical and a theoretical nature. One such tool, celebrated throughout the nineties, is the geometry of visual space-time. It is known under a variety of headings, such as multiple view geometry, structure from motion, and model building. It is a mathematical theory relating multiple views (images) of a scene taken at different viewpoints to three-dimensional models of the (possibly dynamic) scene. This mathematical theory gave rise to algorithms that take as input images (or video) and provide as output a model of the scene. Such algorithms are one of the biggest successes of the field and they have many applications in other disciplines, such as graphics (image-based rendering, motion capture) and robotics (navigation). One of the difficulties, however is that the current tools cannot yet be fully automated, and they do not provide very accurate results. More research is required for automation and high precision. During the past few years we have investigated a number of basic questions underlying the structure from motion problem. Our investigations resulted in a small number of principles that characterize the problem. These principles, which give rise to automatic procedures and point to new avenues for studying the next level of the structure from motion problem, are the subject of this paper.
Cornelia Fermüller, Patrick Baker, Yiannis Aloimonos
Proc. IEEE1
2001 A Spherical Eye from Multiple Cameras (Makes Better Models of the World)
abstract
The paper describes an imaging system that has been designed specifically for the purpose of recovering egomotion and structure from video. The system consists of six cameras in a network arranged so that they sample different parts of the visual sphere. This geometric configuration has provable advantages compared to small field of view cameras for the estimation of the system's own motion and consequently the estimation of shape models from the individual cameras. The reason is that inherent ambiguities of confusion between translation and rotation disappear. We provide algorithms for the calibration of the system and 3D motion estimation. The calibration is based on a new geometric constraint that relates the images of lines parallel in space to the rotation between the cameras. The 3D motion estimation uses a constraint relating structure directly to image gradients.
Patrick Baker, Cornelia Fermüller, Yiannis Aloimonos, Robert Pless
CVPR (1)2
2001 The Statistics of Optical Flow
Cornelia Fermüller, David Shulman, Yiannis Aloimonos
Comput. Vis. Image Underst.1
2000 The Statistics of Optical Flow: Implications for the Process of Correspondence in Vision
abstract
This paper studies the three major categories of flow estimation methods: gradient-based, energy-based, and correlation methods; it analyzes different ways of compounding 1D motion estimates (image gradients, spatio-temporal frequency triplets, local correlation estimates) into 2D velocity estimates, including linear and nonlinear methods. Correcting for the bias would require knowledge of the noise parameters. In many situations, however, these are difficult to estimate accurately, as they change with the dynamic imagery in unpredictable and complex ways. Thus, the bias really is a problem inherent to optical flow estimation. We argue that the bias is also integral to the human visual system. It is the cause of the illusory perception of motion in the Ouchi pattern and also explains various psychophysical studies of the perception of moving plaids. Finally, the implication of the analysis is that flow or correspondence can be estimated very accurately only when feedback is utilized.
Cornelia Fermüller, Yiannis Aloimonos
ICPR1
2000 New eyes for building models from video
Cornelia Fermüller, Yiannis Aloimonos, Tomás Brodský
Comput. Geom.1
2000 Structure from Motion: Beyond the Epipolar Constraint
Tomás Brodský, Cornelia Fermüller, Yiannis Aloimonos
Int. J. Comput. Vis.2
2000 Observability of 3D Motion
Cornelia Fermüller, Yiannis Aloimonos
Int. J. Comput. Vis.1
1999 Shape from Video
abstract
This paper presents a novel technique for recovering the shape of a static scene from a video sequence due to a rigidly moving camera. The solution procedure consists of two stages. In the first stage, the rigid motion of the camera at each instant in time is recovered. This provides the transformation between successive viewing positions. The solution is achieved through new constraints which relate 3D motion and shape directly to the image derivatives. These constraints allow to combine the processes of 3D motion estimation and segmentation by exploiting the geometry and statistics inherent in the data. In the second stage the scene surfaces are reconstructed through an optimization procedure which utilizes data from all the frames of the video sequence. A number of experimental results demonstrate the potential of the approach.
Tomás Brodský, Cornelia Fermüller, Yiannis Aloimonos
CVPR2
1999 Motion Segmentation: A Synergistic Approach
abstract
Since estimation of camera motion requires knowledge of independent motion, and moving object detection and localization requires knowledge about the camera motion, the two problems of motion estimation and segmentation need to be solved together in a synergistic manner. This paper provides an approach to treating both these problems simultaneously. The technique introduced here is based on a novel concept, "scene ruggedness" which parameterizes the variation in estimated scene depth with the error in the underlying three-dimensional (3D) motion. The idea is that incorrect 3D motion estimates cause distortions in the estimated depth map, and as a result smooth scene patches are computed as rugged surfaces. The correct 3D motion can be distinguished, as it does not cause any distortion and thus gives rise to the background patches with the least depth variation between depth discontinuities, with the locations corresponding to independent motion being rugged. The algorithm presented employs a binocular observer whose nature is exploited in the extraction of depth discontinuities, a step that facilitates the overall procedure, but the technique can be extended to a monocular observer in a variety of ways.
Cornelia Fermüller, Tomás Brodský, Yiannis Aloimonos
CVPR1
1999 Statistical Biases in Optic Flow
abstract
The computation of optical flow from image derivatives is biased in regions of non uniform gradient distributions. A least-squares or total least squares approach to computing optic flow from image derivatives even in regions of consistent flow can lead to a systematic bias dependent upon the direction of the optic flow, the distribution of the gradient directions, and the distribution of the image noise. The bias a consistent underestimation of length and a directional error. Similar results hold for various methods of computing optical flow in the spatiotemporal frequency domain. The predicted bias in the optical flow is consistent with psychophysical evidence of human judgment of the velocity of moving plaids, and provides an explanation of the Ouchi illusion. Correction of the bias requires accurate estimates of the noise distribution; the failure of the human visual system to make these corrections illustrates both the difficulty of the task and the feasibility of using this distorted optic flow or undistorted normal flow in tasks requiring higher lever processing.
Cornelia Fermüller, Robert Pless, Yiannis Aloimonos
CVPR1
1998 Toward Motion Picture Grammars
Ruud M. Bolle, Yiannis Aloimonos, Cornelia Fermüller
ACCV (2)3
1998 Simultaneous Estimation of Viewing Geometry and Structure
Tomás Brodský, Cornelia Fermüller, Yiannis Aloimonos
ECCV (1)2
1998 What Is Computed by Structure from Motion Algorithms?
Cornelia Fermüller, Yiannis Aloimonos
ECCV (1)1
1998 Self-Calibration from Image Derivatives
abstract
This study investigates the problem of estimating the calibration parameters from image motion fields induced by a rigidly moving camera with unknown calibration parameters, where the image formation is modeled with a linear pinhole-camera model. The equations obtained show the flow to be clearly separated into a component due to the translation and the calibration parameters and a component due to the rotation and the calibration parameters. A set of parameters encoding the latter component are linearly related to the flow, and from these parameters the calibration can be determined. However, as for discrete motion, in the general case it is not possible, to decouple image measurements from two frames only into their translational and rotational component. Geometrically, the ambiguity takes the form of a part of the rotational component being parallel to the translational component, and thus the scene can be reconstructed only up to a projective transformation. In general, for a full calibration at least four successive image frames are necessary with the 3D-rotation changing between the measurements. The geometric analysis gives rise to a direct self-calibration method that avoids computation of optical flow or point correspondences and uses only normal flow measurements. In this technique the direction of translation is estimated employing in a novel way smoothness constraints. Then the calibration parameters are estimated from the rotational components of several flow fields using Levenberg-Marquardt parameter estimation, iterative in the calibration parameters only. The technique proposed does not require calibration objects in the scene or special camera motions and it also avoids the computation of exact correspondence. This makes it suitable for the calibration of active vision systems which have to acquire knowledge about their intrinsic parameters while they perform other tasks, or as a tool for analyzing image sequences in large video databases.
Tomás Brodský, Cornelia Fermüller, Yiannis Aloimonos
ICCV2
1998 Which Shape from Motion?
abstract
In a practical situation, the rigid transformation relating different views is recovered with errors. In such a case, the recovered depth of the scene contains errors, and consequently a distorted version of visual space is computed. What then are meaningful shape representations that can be computed from the images? The result presented in this paper states that if the rigid transformation between different views is estimated in a way that gives rise to a minimum number of negative depth values, then at the center of the image affine shape can be correctly computed. This result is obtained by exploiting properties of the distortion function. The distortion model turns out to be a very powerful tool in the analysis and design of 3D motion and shape estimation algorithms, and as a byproduct of our analysis we present a computational explanation of psychophysical results demonstrating human visual space distortion from motion information.
Cornelia Fermüller, Yiannis Aloimonos
ICCV1
1998 Effects of Errors in the Viewing Geometry on Shape Estimation
Loong Fah Cheong, Cornelia Fermüller, Yiannis Aloimonos
Comput. Vis. Image Underst.2
1998 Directions of Motion Fields are Hardly Ever Ambiguous
Tomás Brodský, Cornelia Fermüller, Yiannis Aloimonos
Int. J. Comput. Vis.2
1998 Ambiguity in Structure from Motion: Sphere versus Plane
Cornelia Fermüller, Yiannis Aloimonos
Int. J. Comput. Vis.1
1997 The confounding of translation and rotation in reconstruction from multiple views
abstract
If 3D rigid motion is estimated with some error a distorted version of the scene structure will in turn be computed. Of computational interest are these regions in space where the distortions are such that the depths become negative, because in order to be visible the scene has to lie in front of the image. The stability analysis for the structure-from-motion problem presented in this paper investigates the optimal relationship between the errors in the estimated translational and rotational parameters of a rigid motion, that results in the estimation of a minimum number of negative depth values. The input used is the value of the flow along some direction, which is more general than optic flow or correspondence. For a planar retina it is shown that the optimal configuration is achieved when the projections of the translational and rotational errors on the image plane are perpendicular. Furthermore, the projection of the actual and the estimated translation lie on a line passing through the image center. For a spherical retina given a rotational error, the optimal translation is the correct one, while given a translational error. The optimal rotational error is normal to the translational one at an equal distance from the real and estimated translations. The proofs, besides illuminating the confounding of translation and rotation in structure from motion, have an important application to ecological optics, explaining differences of planar and spherical eye or camera designs in motion and shape estimation.
Cornelia Fermüller, Yiannis Aloimonos
CVPR1
1997 On the Geometry of Visual Correspondence
Cornelia Fermüller, Yiannis Aloimonos
Int. J. Comput. Vis.1
1996 Directions of Motion Fields are Hardly Ever Ambiguous
Tomás Brodský, Cornelia Fermüller, Yiannis Aloimonos
ECCV (2)2
1996 Spatiotemporal Representations for Visual Navigation
Loong Fah Cheong, Cornelia Fermüller, Yiannis Aloimonos
ECCV (1)2
1995 Video Representations
Ruud M. Bolle, Yiannis Aloimonos, Cornelia Fermüller
ACCV3
1995 Global Rigidity Constraints in Image Displacement Fields
abstract
Image displacement fields-optical flow fields, stereo disparity fields, normal flow fields-due to rigid motion possess a global geometric structure which is independent of the scene in view. Motion vectors of certain lengths and directions are constrained to lie on the imaging surface at particular loci whose location and form depends solely on the 3D motion parameters. If optical flow fields or stereo disparity fields are considered, then equal vectors are shown to lie on conic sections. Similarly, for normal motion fields, equal vectors lie within regions whose boundaries also constitute conics. By studying various properties of these curves and regions and their relationships, a characterization of the structure of rigid motion fields is given. The goal of this paper is to introduce a concept underlying the global structure of image displacement fields. This concept gives rise to various constraints that could form the basis of algorithms for the recovery of visual information from multiple views.>
Cornelia Fermüller, Yiannis Aloimonos
ICCV1
1995 Representations for Active Vision
Cornelia Fermüller, Yiannis Aloimonos
IJCAI1
1995 Passive navigation as a pattern recognition problem
Cornelia Fermüller
Int. J. Comput. Vis.1
1995 Qualitative egomotion
Cornelia Fermüller, Yiannis Aloimonos
Int. J. Comput. Vis.1
1995 Vision and action
Cornelia Fermüller, Yiannis Aloimonos
Image Vis. Comput.1
1994 A Syntactic Approach to Scale-Space-Based Corner Description
abstract
Planar curves are described by information about corners integrated over various levels of resolution. The detection of corners takes place on a digital representation. To compensate for ambiguities arising from sampling problems due to the discreteness, results about the local behavior of curvature extrema in continuous scale-space are employed.>
Cornelia Fermüller, Walter G. Kropatsch
IEEE Trans. Pattern Anal. Mach. Intell.1
1993 Global 3D motion estimation
abstract
It is shown how a monocular observer can estimate its 3D motion relative to the scene by using normal flow measurements in a global and qualitative way. It is proved that local normal flow measurements form global patterns in the image plane. The position of these patterns is related to the 3D motion parameters. By locating some of these patterns, which depend only on subsets of the motion parameters, through a simple search technique, the 3D motion parameters can be found. The proposed algorithmic procedure is very robust, since it is not affected by small perturbations in the normal flow measurements. The direction of translation and the axis of rotation can be estimated with up to 100% error in the image measurements.>
Cornelia Fermüller
CVPR1
1993 Recognizing 3-D Motion
Cornelia Fermüller, Yiannis Aloimonos
IJCAI1
1993 The role of fixation in visual motion analysis
Cornelia Fermüller, Yiannis Aloimonos
Int. J. Comput. Vis.1
1992 Multi-resolution shape description by corners
abstract
A robust method for describing planar curves in multiple resolution using curvature information is presented. The method is developed by taking into account the discrete nature of digital images as well as the discrete aspect of a multiresolution structure (pyramid). The main contribution lies in the robustness of the technique, which is due to the additional information that is extracted from observing the behavior of corners in the whole pyramid. Furthermore, the resulting algorithm is conceptually simple and easily parallelizable. Theoretical results are developed analyzing the curvature of continuous curves in scale-space and showing the behavior of curvature extrema under varying scale. The results are used to eliminate any ambiguities that might arise from sampling problems due to the discreteness of the representation. Experimental results demonstrate the potential of the method.>
Cornelia Fermüller, Walter G. Kropatsch
CVPR1
1992 Perceptual computational advantages of tracking
abstract
The paradigm of active vision advocates studying visual problems in the form of modules that are directly related to a visual task for observers that are active. It is argued that in many cases when an object is moving in an unrestricted manner (translation and rotation) in the 3D world only the motion's translational components are of interest. For a monocular observer, using only the normal flow-the spatiotemporal derivatives of the image intensity function-the authors solve the problem of computing the direction of translation. Their strategy uses fixation and tracking. Fixation simplifies much of the computation by placing the object at the center of the visual field, and the main advantage of tracking is the accumulation of information over time. The authors show how tracking is accomplished using normal flow measurements and use it for two different tasks in the solution process. First, it serves as a tool to compensate for the lack of existence of an optical flow field and thus to estimate the translation parallel to the image plane; and second, it gathers information about the motion component perpendicular to the image plane.>
Cornelia Fermüller, Yiannis Aloimonos
ICPR (1)1
1992 Hierarchical curve representation
abstract
Presents a robust method for describing planar curves in multiple resolution using curvature information. The method is developed by taking into account the discrete nature of digital images as well as the discrete aspect of a multiresolution structure (pyramid). The authors deal with the robustness of the technique, which is due to the additional information that is extracted from observing the behavior of corners in the pyramid. Furthermore the resulting algorithm is conceptually simple and easily parallelizable. They develop theoretical results, analyzing the curvature of continuous curves in scale-space, which show the behavior of curvature extrema under varying scale. These results are used to eliminate any ambiguities that might arise from sampling problems due to the discreteness of the representation. Finally, experimental results demonstrate the potential of the method. >
Cornelia Fermüller, Walter G. Kropatsch
ICPR (3)1