VLDB 2026 Research / reviewers in the wild / expert
Yiannis Aloimonos
dblp:a/YiannisAloimonos · also John Aloimonos
· DBLP profile ↗
180ranked-venue papers
18as first author
43since 2021 · last 2026
0000-0002-8152-4281ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 158 · 16 first-author · 38 since 2021Graphics, computer vision, multimedia, augmented reality and games · 69 · 7 first-author · 7 since 2021Systems, architecture and hardware · 47 · 24 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Databases, data management, data science and information retrieval · 2Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HiCL: Hippocampal-Inspired Continual LearningabstractWe propose HiCL, a novel hippocampal-inspired dual-memory continual learning architecture designed to mitigate catastrophic forgetting by using elements inspired by the hippocampal circuitry. Our system encodes inputs through a grid-cell-like layer, followed by sparse pattern separation using a dentate gyrus-inspired module with top-k sparsity. Episodic memory traces are maintained in a CA3-like autoassociative memory. Task-specific processing is dynamically managed via a DG-gated mixture-of-experts mechanism, wherein inputs are routed to experts based on cosine similarity between their normalized sparse DG representations and learned task-specific DG prototypes computed through online exponential moving averages. This biologically grounded yet mathematically principled gating strategy enables differentiable, scalable task-routing without relying on a separate gating network, and enhances the model's adaptability and efficiency in learning multiple sequential tasks. Cortical outputs are consolidated using Elastic Weight Consolidation weighted by inter-task similarity. Crucially, we incorporate prioritized replay of stored patterns to reinforce essential past experiences. Evaluations on standard continual learning benchmarks demonstrate the effectiveness of our architecture in reducing task interference, achieving near state-of-the-art results in continual learning tasks at lower computational costs. Kushal Kapoor, Wyatt Mackey, Yiannis Aloimonos, Xiaomin Lin 0002 |
AAAI | 3 |
| 2026 | Guest Editorial: Artificial Intelligence Generated Content (AIGC) for Industrial Manufacturing
Huaping Liu 0001, Weiwei Wan, Jason Gu, Valeria Villani, Giulia Pedrielli, Yiannis Aloimonos |
IEEE Trans Autom. Sci. Eng. | 6 |
| 2025 | Repurposing Pre-trained Video Diffusion Models for Event-based Video InterpolationabstractVideo Frame Interpolation aims to recover realistic missing frames between observed frames, generating a high-frame-rate video from a low-frame-rate video. However, without additional guidance, the large motion between frames makes this problem ill-posed. Event-based Video Frame Interpolation (EVFI) addresses this challenge by using sparse, high-temporal-resolution event measurements as motion guidance. This guidance allows EVFI methods to significantly outperform frame-only methods. However, to date, EVFI methods have relied on a limited set of paired event-frame training data, severely limiting their performance and generalization capabilities. In this work, we overcome the limited data challenge by adapting pre-trained video diffusion models trained on internet-scale datasets to EVFI. We experimentally validate our approach on real-world EVFI datasets, including a new one that we introduce. Our method outperforms existing methods and generalizes across cameras far better than existing approaches. Jingxi Chen, Brandon Yushan Feng, Haoming Cai, Tianfu Wang 0007, Levi Burner, Dehao Yuan, Cornelia Fermüller, Christopher A. Metzler, Yiannis Aloimonos |
CVPR | 9 |
| 2025 | NatSGLD: A Dataset with Speech, Gesture, Logic, and Demonstration for Robot Learning in Natural Human-Robot InteractionabstractRecent advances in multimodal Human-Robot Interaction (HRI) datasets emphasize the integration of speech and gestures, allowing robots to absorb explicit knowledge and tacit understanding. However, existing datasets primarily focus on elementary tasks like object pointing and pushing, limiting their applicability to complex domains. They prioritize simpler human command data but place less emphasis on training robots to correctly interpret tasks and respond appropriately. To address these gaps, we present the NatSGLD dataset, which was collected using a Wizard of Oz (WoZ) method, where participants interacted with a robot they believed to be autonomous. NatSGLD records humans' multimodal commands (speech and gestures), each paired with a demonstration trajectory and a Linear Temporal Logic (LTL) formula that provides a ground-truth interpretation of the commanded tasks. This dataset serves as a foundational resource for research at the intersection of HRI and machine learning. By providing multimodal inputs and detailed annotations, NatSGLD enables exploration in areas such as multimodal instruction following, plan recognition, and human-advisable reinforcement learning from demonstrations. We release the dataset and code under the MIT License at https://www.snehesh.com/natsgld/to support future HRI research. Snehesh Shrestha, Yantian Zha, Saketh Banagiri, Ge Gao 0001, Yiannis Aloimonos, Cornelia Fermüller |
HRI | 5 |
| 2025 | Learning Normal Flow Directly from Events
Dehao Yuan, Levi Burner, Jiayi Wu 0005, Jingxi Chen, Yiannis Aloimonos, Cornelia Fermüller |
ICCV | 6 |
| 2025 | Air-FAR: Fast and Adaptable Routing for Aerial Navigation in Large-Scale Complex Unknown EnvironmentsabstractThis paper presents a novel approach for realtime 3D navigation in large-scale complex environments by introducing a hierarchical 3D visibility graph (V-graph) and an efficient path search method. The proposed algorithm addresses the computational challenges of V-graph construction and shortest path search on the graph simultaneously. By introducing hierarchical 3D V-graph construction with heuristic visibility update, the 3D V-graph is constructed in$O\left(K \cdot n^{2} \log n\right)$time, which guarantees real-time performance. The proposed iterative divide-and-conquer path search method can achieve near-optimal path solutions within the constraints of realtime operations. The algorithm ensures efficient 3D V-graph construction and path search. Extensive simulated and realworld environments validated that our algorithm reduces the travel time by 42%, achieves up to 24.8% higher trajectory efficiency, and runs faster than most benchmarks by orders of magnitude in complex environments. The code and developed simulator have been open-sourced to facilitate future research. Botao He, Guofei Chen, Cornelia Fermüller, Yiannis Aloimonos, Ji Zhang 0003 |
ICRA | 4 |
| 2025 | ODYSSEE: Oyster Detection Yielded by Sensor Systems on Edge ElectronicsabstractOysters are a vital keystone species in coastal ecosystems, providing significant economic, environmental, and cultural benefits. As the importance of oysters grows, so does the relevance of autonomous systems for their detection and monitoring. However, current monitoring strategies often rely on destructive methods. While manual identification of oysters from video footage is non-destructive, it is time-consuming, requires expert input, and is further complicated by the challenges of the underwater environment. To address these challenges, we propose a novel pipeline using stable diffusion to augment a collected real dataset with photorealistic synthetic data. This method enhances the dataset used to train a YOLOv10-based vision model. The model is then deployed and tested on an edge platform; Aqua2, an Autonomous Underwater Vehicle (AUV), achieving a state-of-the-art 0.657 mAP@50 for oyster detection. Xiaomin Lin 0002, Vivek Mange, Arjun Suresh, Bernhard Neuberger, Aadi Palnitkar, Brendan Campbell, Kleio Baxevani, Jeremy Mallette, Alhim Vera, Markus Vincze, Ioannis M. Rekleitis, Herbert G. Tanner, Yiannis Aloimonos |
ICRA | 14 |
| 2025 | Discovering Object Attributes by Prompting Large Language Models With Perception-Action ApisabstractThere has been a lot of interest in grounding natural language to physical entities through visual context. While Vision Language Models (VLMs) can ground linguistic instructions to visual sensory information, they struggle with grounding non-visual attributes, like the weight of an object. Our key insight is that non-visual attribute detection can be effectively achieved by active perception guided by visual reasoning. To this end, we present a perception-action API that consists of VLMs and Large Language Models (LLMs) as backbones, together with a set of robot control functions. When prompted with this API and a natural language query, an LLM generates a program to actively identify attributes given an input image. Offline testing on the Odd-One-Out ($\mathbf{O}^{\mathbf{3}}$) dataset demonstrates that our framework outperforms vanilla VLMs in detecting attributes like relative object location, size, and weight. Online testing in realistic household scenes on AI2THOR and a real robot demonstration on a DJI RoboMaster EP robot highlight the efficacy of our approach. Angelos Mavrogiannis, Dehao Yuan, Yiannis Aloimonos |
ICRA | 3 |
| 2025 | FeelAnyForce: Estimating Contact Force Feedback from Tactile Sensation for Vision-Based Tactile SensorsabstractIn this paper, we tackle the problem of estimating 3D contact forces using vision-based tactile sensors. In particular, our goal is to estimate contact forces over a large range (up to 15 N) on any objects while generalizing across different vision-based tactile sensors. Thus, we collected a dataset of over 200K indentations using a robotic arm that pressed various indenters onto a GelSight Mini sensor mounted on a force sensor and then used the data to train a multi-head transformer for force regression. Strong generalization is achieved via accurate data collection and multi-objective optimization that leverages depth contact images. Despite being trained only on primitive shapes and textures, the regressor achieves a mean absolute error of 4% on a dataset of unseen real-world objects. We further evaluate our approach's generalization capability to other GelSight mini and DIGIT sensors, and propose a reproducible calibration procedure for other sensors. Finally, the method was evaluated on real-world tasks, including weighing objects and controlling the deformation of delicate objects. Supplementary material and demo are available at http://prg.cs.umd.edu/FeelAnyForce. Amir-Hossein Shahidzadeh, Gabriele M. Caddeo, Koushik Alapati, Lorenzo Natale, Cornelia Fermüller, Yiannis Aloimonos |
ICRA | 6 |
| 2025 | ViewActive: Active viewpoint optimization from a single imageabstractWhen observing objects, humans benefit from their spatial visualization and mental rotation ability to envision potential optimal viewpoints based on the current observation. This capability is crucial for enabling robots to achieve efficient and robust scene perception during operation, as optimal viewpoints provide essential and informative features for accurately representing scenes in 2D images, thereby enhancing downstream tasks.To endow robots with this human-like active viewpoint optimization capability, we propose ViewActive, a modernized machine learning approach drawing inspiration from aspect graph, which provides viewpoint optimization guidance based solely on the current 2D image input. Specifically, we introduce the 3D Viewpoint Quality Field (VQF), a compact and consistent representation for viewpoint quality distribution similar to an aspect graph, composed of three general-purpose viewpoint quality metrics: self-occlusion ratio, occupancy-aware surface normal entropy, and visual entropy. We utilize pre-trained image encoders to extract robust visual and semantic features, which are then decoded into the 3D VQF, allowing our model to generalize effectively across diverse objects, including unseen categories. The lightweight ViewActive network (72 FPS on a single GPU) significantly enhances the performance of state-of-the-art object recognition pipelines and can be integrated into real-time motion planning for robotic applications. Our code and dataset are available here https://github.com/jiayi-wu-umd/ViewActive. Jiayi Wu 0005, Xiaomin Lin 0002, Botao He, Cornelia Fermüller, Yiannis Aloimonos |
IROS | 5 |
| 2024 | CodedEvents: Optimal Point-Spread-Function Engineering for 3D-Tracking with Event CamerasabstractPoint-spread-function (PSF) engineering is a well-established computational imaging technique that uses phase masks and other optical elements to embed extra information (e.g., depth) into the images captured by conventional CMOS image sensors. To date, however, PSF-engineering has not been applied to neuromorphic event cameras; a powerful new image sensing technology that responds to changes in the log-intensity of light. This paper establishes theoretical limits (Cramér Rao bounds) on 3D point localization and tracking with PSF-engineered event cameras. Using these bounds, we first demonstrate that existing Fisher phase masks are already near-optimal for localizing static flashing point sources (e.g., blinking fluorescent molecules). We then demonstrate that existing designs are sub-optimal for tracking moving point sources and proceed to use our theory to design optimal phase masks and binary amplitude masks for this task. To overcome the non-convexity of the design problem, we leverage novel implicit neural representation based parameterizations of the phase and amplitude masks. We demonstrate the efficacy of our designs through extensive simulations. We also validate our method with a simple prototype. Sachin Shah, Matthew A. Chan 0002, Haoming Cai, Jingxi Chen, Sakshum Kulshrestha, Chahat Deep Singh, Yiannis Aloimonos, Christopher A. Metzler |
CVPR | 7 |
| 2024 | Decodable and Sample Invariant Continuous Object EncoderabstractWe propose Hyper-Dimensional Function Encoding (HDFE). Given samples of a continuous object (e.g. a function), HDFE produces an explicit vector representation of the given object, invariant to the sample distribution and density. Sample distribution and density invariance enables HDFE to consistently encode continuous objects regardless of their sampling, and therefore allows neural networks to receive continuous objects as inputs for machine learning tasks, such as classification and regression. Besides, HDFE does not require any training and is proved to map the object into an organized embedding space, which facilitates the training of the downstream tasks. In addition, the encoding is decodable, which enables neural networks to regress continuous objects
by regressing their encodings. Therefore, HDFE serves as an interface for processing continuous objects.
We apply HDFE to function-to-function mapping, where vanilla HDFE achieves competitive performance with the state-of-the-art algorithm. We apply HDFE to point cloud surface normal estimation, where a simple replacement from PointNet to HDFE leads to 12\% and 15\% error reductions in two benchmarks.
In addition, by integrating HDFE into the PointNet-based SOTA network, we improve the SOTA baseline by 2.5\% and 1.7\% on the same benchmarks. Dehao Yuan, Furong Huang, Cornelia Fermüller, Yiannis Aloimonos |
ICLR | 4 |
| 2024 | A Linear Time and Space Local Point Cloud Geometry Encoder via Vectorized Kernel Mixture (VecKM)abstractWe propose VecKM, a local point cloud geometry encoder that is descriptive and efficient to compute. VecKM leverages a unique approach by vectorizing a kernel mixture to represent the local point cloud. Such representation’s descriptiveness is supported by two theorems that validate its ability to reconstruct and preserve the similarity of the local shape. Unlike existing encoders downsampling the local point cloud, VecKM constructs the local geometry encoding using all neighboring points, producing a more descriptive encoding. Moreover, VecKM is efficient to compute and scalable to large point cloud inputs: VecKM reduces the memory cost from $(n^2+nKd)$ to $(nd+np)$; and reduces the major runtime cost from computing $nK$ MLPs to $n$ MLPs, where $n$ is the size of the point cloud, $K$ is the neighborhood size, $d$ is the encoding dimension, and $p$ is a marginal factor. The efficiency is due to VecKM’s unique factorizable property that eliminates the need of explicitly grouping points into neighbors. In the normal estimation task, VecKM demonstrates not only 100x faster inference speed but also highest accuracy and strongest robustness. In classification and segmentation tasks, integrating VecKM as a preprocessing module achieves consistently better performance than the PointNet, PointNet++, and point transformer baselines, and runs consistently faster by up to 10 times. Dehao Yuan, Cornelia Fermüller, Tahseen Rabbani, Furong Huang, Yiannis Aloimonos |
ICML | 5 |
| 2024 | UIVNAV: Underwater Information-driven Vision-based Navigation via Imitation LearningabstractAutonomous navigation in the underwater environment is challenging due to limited visibility, dynamic changes, and the lack of a cost-efficient, accurate localization system. We introduce UIVNAV, a novel end-to-end underwater navigation solution designed to navigate robots over Objects of Interest (OOI) while avoiding obstacles, all without relying on localization. UIVNAVutilizes imitation learning and draws inspiration from the navigation strategies employed by human divers, who do not rely on localization. UIVNAVconsists of the following phases: (1) generating an intermediate representation (IR) and (2) training the navigation policy based on human-labeled IR. By training the navigation policy on IR instead of raw data, the second phase is domain-invariant — the navigation policy does not need to be retrained if the domain or the OOI changes. We demonstrate this within simulation by deploying the same navigation policy to survey two distinct Objects of Interest (OOIs): oyster and rock reefs. We compared our method with complete coverage and random walk methods, showing that our approach is more efficient in gathering information for OOIs while avoiding obstacles. The results show that UIVNAVchooses to visit the areas with larger area sizes of oysters or rocks with no prior information about the environment or localization. Moreover, a robot using UIVNAVcompared to complete coverage method surveys on average 36% more oysters when traveling the same distances. We also demonstrate the feasibility of real-time deployment of UIVNAVin pool experiments with BlueROV underwater robot for surveying a bed of oyster shells. Xiaomin Lin 0002, Nare Karapetyan, Kaustubh Joshi 0002, Tianchen Liu, Nikhil Chopra, Miao Yu 0007, Pratap Tokekar, Yiannis Aloimonos |
ICRA | 8 |
| 2024 | Cook2LTL: Translating Cooking Recipes to LTL Formulae using Large Language ModelsabstractCooking recipes are challenging to translate to robot plans as they feature rich linguistic complexity, temporally-extended interconnected tasks, and an almost infinite space of possible actions. Our key insight is that combining a source of cooking domain knowledge with a formalism that captures the temporal richness of cooking recipes could enable the extraction of unambiguous, robot-executable plans. In this work, we use Linear Temporal Logic (LTL) as a formal language expressive enough to model the temporal nature of cooking recipes. Leveraging a pretrained Large Language Model (LLM), we present Cook2LTL, a system that translates instruction steps from an arbitrary cooking recipe found on the internet to a set of LTL formulae, grounding high-level cooking actions to a set of primitive actions that are executable by a manipulator in a kitchen environment. Cook2LTL makes use of a caching scheme that dynamically builds a queryable action library at runtime. We instantiate Cook2LTL in a realistic simulation environment (AI2-THOR), and evaluate its performance across a series of cooking recipes. We demonstrate that our system significantly decreases LLM API calls (−51%), latency (−59%), and cost (−42%) compared to a baseline that queries the LLM for every newly encountered action at runtime. Angelos Mavrogiannis, Christoforos I. Mavrogiannis, Yiannis Aloimonos |
ICRA | 3 |
| 2024 | AcTExplore: Active Tactile Exploration on Unknown ObjectsabstractTactile exploration plays a crucial role in understanding object structures for fundamental robotics tasks such as grasping and manipulation. However, efficiently exploring such objects using tactile sensors is challenging, primarily due to the large-scale unknown environments and limited sensing coverage of these sensors. To this end, we present AcTExplore, an active tactile exploration method driven by reinforcement learning for object reconstruction at scales that automatically explores the object surfaces in a limited number of steps. Through sufficient exploration, our algorithm incrementally collects tactile data and reconstructs 3D shapes of the objects as well, which can serve as a representation for higher-level downstream tasks. Our method achieves an average of 95.97% IoU coverage on unseen YCB objects while just being trained on primitive shapes. Amir-Hossein Shahidzadeh, Seong Jong Yoo, Pavan Mantripragada, Chahat Deep Singh, Cornelia Fermüller, Yiannis Aloimonos |
ICRA | 6 |
| 2024 | Vector Symbolic Sub-objects Classifiers as Manifold AnaloguesabstractVector Symbolic Architectures (VSAs) generally consist of a hyper-algebra that is defined on points in a space of vectors. However, such a space is typically defined in usual ways, such as Euclidean Spaces for real vectors, Hamming Spaces for binary vectors, and so on. In any empirical setting, such as Artificial Intelligence, Machine Learning, Data Science, etc., observations in these spaces tend to produce subspaces of these broader spaces. As a result, these empirically-derived subspaces enable VSAs to make predictions through how severely the subspaces deviate from what is expected. Thus, it is desirable to be able to understand how such subspaces behave. As an analogy, when one observes terrain, they should find good "roads and bridges" to navigate that terrain. In Category Theory, the idea of Topos describes how such roads and bridges should be placed. In this paper, we explore the relationship between VSAs and Topoi. Namely, we show how a Topos-like representation can be constructed from empirical observations of vectors and demonstrate this on a practical example using Hyperdimensional Computing (HDC) on dense binary hypervectors. Our results indicate that a Topos can be effectively constructed for a dataset and the resulting space of vectors is biased to reflect the compositional aspects of the data. We conclude that such a Topos can be used to better guide the construction of VSAs for downstream tasks, in lieu of the original space of vectors, much like manifolds in Topology. Renato Faraone, Peter Sutor Jr., Cornelia Fermüller, Yiannis Aloimonos |
IJCNN | 4 |
| 2024 | A Comparative Study of Hough Transform and PCA for Bolt Orientation DetectionabstractIn the fields of manufacturing and robotics, accurately determining the orientation of manufacturing components, such as bolts, is a critical yet challenging problem due to the limitations of existing detection methods. This study introduces a novel methodology for addressing this issue, leveraging traditional computer vision techniques, by proposing a streamlined approach that exploits the inherent geometric properties of bolts for orientation detection. Two methods are presented to ascertain the initial axis angle of the bolt: the Progressive Probabilistic Hough Transform (PPHT) and Principal Component Analysis (PCA). These methods are used in conjunction with a novel tip direction detection approach. The results of the study demonstrate consistent accuracy in angle determination, with PPHT and PCA both achieving angular deviations below ±0.5° in the simple dataset, and PCA showing enhanced robustness in the dataset degraded by shadows, with a maximum error under ±1.5°. This research not only reaffirms the viability of fundamental computer vision techniques in modern robotic applications but also sets a precedent for simple, generalisable, and reliable orientation detection solutions. These methods effectively bridge the gap between highly specialised machine learning systems, which often require tailored, complex models and extensive training data, and more universally applicable, straightforward approaches. Antonio Gambale, Sonya A. Coleman, Dermot Kerr, Philip J. Vance, Emmett Kerr, Cornelia Fermüller, Yiannis Aloimonos |
INDIN | 7 |
| 2024 | Fingerspelling Classification for Robot ControlabstractImprovements to human-robot interaction methods could increase the ease of use of robots in manufacturing environments. Many of these environments are noisy and therefore preclude the use of audio communication between humans or in human-robot interactions. Therefore, this paper proposes using a gesture based communication system for robot control. To that end, the VGG16 and VGG19 convolutional neural network (CNN) structures are used for gesture classification along with 3 datasets of American Sign Language (ASL) fingerspelling images. The model performance is evaluated, and modifications made to their parameters to improve performance, before applying them to robot control tasks. The results show that with parameter tuning, test accuracies of up to, 100% are achievable. Kevin McCready, Sonya A. Coleman, Dermot Kerr, Nazmul H. Siddique, Emmett Kerr, Yiannis Aloimonos, Cornelia Fermüller |
INDIN | 6 |
| 2024 | Active Human Pose Estimation via an Autonomous UAV AgentabstractOne of the core activities of an active observer involves moving to secure a "better" view of the scene, where the definition of "better" is task-dependent. This paper focuses on the task of human pose estimation from videos capturing a person’s activity. Self-occlusions within the scene can complicate or even prevent accurate human pose estimation. To address this, relocating the camera to a new vantage point is necessary to clarify the view, thereby improving 2D human pose estimation. This paper formalizes the process of achieving an improved viewpoint. Our proposed solution to this challenge comprises three main components: a NeRF-based Drone-View Data Generation Framework, an On-Drone Network for Camera View Error Estimation, and a Combined Planner for devising a feasible motion plan to reposition the camera based on the predicted errors for camera views. The Data Generation Framework utilizes NeRF-based methods to generate a comprehensive dataset of human poses and activities, enhancing the drone’s adaptability in various scenarios. The Camera View Error Estimation Network is designed to evaluate the current human pose and identify the most promising next viewing angles for the drone, ensuring a reliable and precise pose estimation from those angles. Finally, the combined planner incorporates these angles while considering the drone’s physical and environmental limitations, employing efficient algorithms to navigate safe and effective flight paths. This system represents a significant advancement in active 2D human pose estimation for an autonomous UAV agent, offering substantial potential for applications in aerial cinematography by improving the performance of autonomous human pose estimation and maintaining the operational safety and efficiency of UAVs. Jingxi Chen, Botao He, Chahat Deep Singh, Cornelia Fermüller, Yiannis Aloimonos |
IROS | 5 |
| 2024 | Interactive-FAR: Interactive, Fast and Adaptable Routing for Navigation Among Movable Obstacles in Complex Unknown EnvironmentsabstractThis paper introduces a real-time algorithm for navigating complex unknown environments cluttered with movable obstacles. Our algorithm achieves fast, adaptable routing by actively attempting to manipulate obstacles during path planning and adjusting the global plan from sensor feedback. The main contributions include an improved dynamic Directed Visibility Graph (DV-graph) for rapid global path searching, a real-time interaction planning method that adapts online from new sensory perceptions, and a comprehensive framework designed for interactive navigation in complex unknown or partially known environments. Our algorithm is capable of replanning the global path in several milliseconds. It can also attempt to move obstacles, update their affordances, and adapt strategies accordingly. Extensive experiments validate that our algorithm reduces the travel time by 33%, achieves up to 49% higher path efficiency, and runs faster than traditional methods by orders of magnitude in complex environments. It has been demonstrated to be the most efficient solution in terms of speed and efficiency for interactive navigation in environments of such complexity. We also open-source our code in the docker demo1to facilitate future research. Botao He, Guofei Chen, Ji Zhang 0003, Cornelia Fermüller, Yiannis Aloimonos |
IROS | 6 |
| 2024 | MARVIS: Motion & Geometry Aware Real and Virtual Image SegmentationabstractTasks such as autonomous navigation, 3D reconstruction, and object recognition near the water surfaces are crucial in marine robotics applications. However, challenges arise due to dynamic disturbances, e.g., light reflections and refraction from the random air-water interface, irregular liquid flow, and similar factors, which can lead to potential failures in perception and navigation systems. Traditional computer vision algorithms struggle to differentiate between real and virtual image regions, significantly complicating tasks. A virtual image region is an apparent representation formed by the redirection of light rays, typically through reflection or refraction, creating the illusion of an object’s presence without its actual physical location. This work proposes a novel approach for segmentation on real and virtual image regions, exploiting synthetic images combined with domain-invariant information, a Motion Entropy Kernel, and Epipolar Geometric Consistency. Our segmentation network does not need to be re-trained if the domain changes. We show this by deploying the same segmentation network in two different domains: simulation and the real world. By creating realistic synthetic images that mimic the complexities of the water surface, we provide fine-grained training data for our network (MARVIS) to discern between real and virtual images effectively. By motion & geometry-aware design choices and through comprehensive experimental analysis, we achieve state-of-the-art real-virtual image segmentation performance in unseen real world domain, achieving an IoU over 78% and a F1-Score over 86% while ensuring a small computational footprint. MARVIS offers over 43 FPS (8 FPS) inference rates on a single GPU (CPU core). Our code and dataset are available here https://github.com/jiayi-wu-umd/MARVIS. Jiayi Wu 0005, Xiaomin Lin 0002, Shahriar Negahdaripour, Cornelia Fermüller, Yiannis Aloimonos |
IROS | 5 |
| 2024 | Embodiment: Self-Supervised Depth Estimation Based on Camera ModelsabstractDepth estimationn is a critical topic for robotics and vision-related tasks. In monocular depth estimation, in comparison with supervised learning that requires expensive ground truth labeling, self-supervised methods possess great potential due to no labeling cost. However, self-supervised learning still has a large gap with supervised learning in 3D reconstruction and depth estimation performance. Meanwhile, scaling is also a major issue for monocular unsupervised depth estimation, which commonly still needs ground truth scale from GPS, LiDAR, or existing maps to correct. In the era of deep learning, existing methods primarily rely on exploring image relationships to train unsupervised neural networks, while the physical properties of the camera itself—such as intrinsics and extrinsics—are often overlooked. These physical properties are not just mathematical parameters; they are embodiments of the camera’s interaction with the physical world. By embedding these physical properties into the depth learning model, we can calculate depth priors for ground regions and regions connected to the ground based on physical principles, providing free supervision signals without the need for additional sensors. This approach is not only easy to implement but also enhances the effects of all unsupervised methods by embedding the camera’s physical properties into the model, thereby achieving an embodied understanding of the real world. Jinchang Zhang, Praveen Kumar Reddy, Xue-Iuan Wong, Yiannis Aloimonos, Guoyu Lu 0001 |
IROS | 4 |
| 2024 | Temporally Consistent Atmospheric Turbulence Mitigation with Neural RepresentationsabstractAtmospheric turbulence, caused by random fluctuations in the atmosphere's refractive index, introduces complex spatio-temporal distortions in imagery captured at long range. Video Atmospheric Turbulence Mitigation (ATM) aims to restore videos affected by these distortions. However, existing video ATM methods, both supervised and self-supervised, struggle to maintain temporally consistent mitigation across frames, leading to visually incoherent results. This limitation arises from the stochastic nature of atmospheric turbulence, which varies across space and time. Inspired by the observation that atmospheric turbulence induces high-frequency temporal variations, we propose ConVRT, a novel framework for consistent video restoration through turbulence. ConVRT introduces a neural video representation that explicitly decouples spatial and temporal information into a spatial content field and a temporal deformation field, enabling targeted regularization of the network's temporal representation capability. By leveraging the low-pass filtering properties of the regularized temporal representations, ConVRT effectively mitigates turbulence-induced temporal frequency variations and promotes temporal consistency. Furthermore, our training framework seamlessly integrates supervised pre-training on synthetic turbulence data with self-supervised learning on real-world videos, significantly improving the temporally consistent mitigation of ATM methods on diverse real-world data. More information can be found on our project page: https://convrt-2024.github.io/ Haoming Cai, Jingxi Chen, Brandon Yushan Feng, Weiyun Jiang, Mingyang Xie, Kevin Zhang 0003, Cornelia Fermüller, Yiannis Aloimonos, Ashok Veeraraghavan, Christopher A. Metzler |
NeurIPS | 8 |
| 2024 | Context in Human Action through Motion ComplementarityabstractMotivated by Goldman’s Theory of Human Action - a framework in which action decomposes into 1) base physical movements, and 2) the context in which they occur - we propose a novel learning formulation for motion and context, where context is derived as the complement to motion. More specifically, we model physical movement through the adoption of Therbligs, a set of elemental physical motions centered around object manipulation. Context is modeled through the use of a contrastive mutual information loss that formulates context information as the action information not contained within movement information. We empirically prove the utility brought by this separation of representation, showing sizable improvements in action recognition and action anticipation accuracies for a variety of models. We present results over two object manipulation datasets: EPIC Kitchens 100, and 50 Salads. Eadom Dessalene, Michael Maynord, Cornelia Fermüller, Yiannis Aloimonos |
WACV | 4 |
| 2023 | Therbligs in Action: Video Understanding through Motion PrimitivesabstractIn this paper we introduce a rule-based, compositional, and hierarchical modeling of action using Therbligs as our atoms. Introducing these atoms provides us with a consistent, expressive, contact-centered representation of action. Over the atoms we introduce a differentiable method of rule-based reasoning to regularize for logical consistency. Our approach is complementary to other approaches in that the Therblig-based representations produced by our architecture augment rather than replace existing architectures' representations. We release the first Therblig-centered an-notations over two popular video datasets - EPIC Kitchens 100 and 50-Salads. We also broadly demonstrate benefits to adopting Therblig representations through evaluation on the following tasks: action segmentation, action anticipation, and action recognition - observing an average 10.5%/7.53%/6.5% relative improvement, respectively, over EPIC Kitchens and an average 8.9%/6.63%/4.8% relative improvement, respectively, over 50 Salads. Code and data will be made publicly available. Eadom Dessalene, Michael Maynord, Cornelia Fermüller, Yiannis Aloimonos |
CVPR | 4 |
| 2023 | t-ConvESN: Temporal Convolution-Readout for Random Recurrent Neural Networks
Matthew Evanusa, Vaishnavi Patil, Michelle Girvan, Joel Goodman, Cornelia Fermüller, Yiannis Aloimonos |
ICANN (6) | 6 |
| 2023 | Mid-Vision Feedback
Michael Maynord, Eadom Dessalene, Cornelia Fermüller, Yiannis Aloimonos |
ICLR | 4 |
| 2023 | TTCDist: Fast Distance Estimation From an Active Monocular Camera Using Time-to-ContactabstractDistance estimation from vision is fundamental for a myriad of robotic applications such as navigation, manipu-lation, and planning. Inspired by the mammal's visual system, which gazes at specific objects, we develop two novel constraints relating time-to-contact, acceleration, and distance that we call the$\tau$-constraint and$\Phi$-constraint. They allow an active (moving) camera to estimate depth efficiently and accurately while using only a small portion of the image. The constraints are applicable to range sensing, sensor fusion, and visual servoing. We successfully validate the proposed constraints with two experiments. The first applies both constraints in a trajectory estimation task with a monocular camera and an Inertial Measurement Unit (IMU). Our methods achieve 30-70% less average trajectory error while running$25\times$and$6.2\times$faster than the popular Visual-Inertial Odometry methods VINS-Mono and ROVIO respectively. The second experiment demonstrates that when the constraints are used for feedback with efference copies the resulting closed loop system's eigenvalues are invariant to scaling of the applied control signal. We believe these results indicate the$\tau$and$\Phi$constraint's potential as the basis of robust and efficient algorithms for a multitude of robotic applications. Levi Burner, Nitin J. Sanket, Cornelia Fermüller, Yiannis Aloimonos |
ICRA | 4 |
| 2023 | OysterNet: Enhanced Oyster Detection Using SimulationabstractOysters play a pivotal role in the bay living ecosystem and are considered the living filters for the ocean. In recent years, oyster reefs have undergone major devastation caused by commercial over-harvesting, requiring preservation to maintain ecological balance. The foundation of this preservation is to estimate the oyster density which requires accurate oyster detection. However, systems for accurate oyster detection require large datasets obtaining which is an expensive and labor-intensive task in underwater environments. To this end, we present a novel method to mathematically model oysters and render images of oysters in simulation to boost the detection performance with minimal real data. Utilizing our synthetic data along with real data for oyster detection, we obtain up to 35.1 % boost in performance as compared to using only real data with our OysterNet network. We also improve the state-of-the-art by 12.7%. This shows that using underlying geometrical properties of objects can help to enhance recognition task accuracy on limited datasets successfully and we hope more researchers adopt such a strategy for hard-to-obtain datasets. Xiaomin Lin 0002, Nitin J. Sanket, Nare Karapetyan, Yiannis Aloimonos |
ICRA | 4 |
| 2023 | WorldGen: A Large Scale Generative SimulatorabstractIn the era of deep learning, data is the critical determining factor in the performance of neural network models. Generating large datasets suffers from various challenges such as scalability, cost efficiency and photorealism. To avoid expensive and strenuous dataset collection and annotations, researchers have inclined towards computer-generated datasets. However, a lack of photorealism and a limited amount of computer-aided data has bounded the accuracy of network predictions. To this end, we present WorldGen - an open source framework to automatically generate countless structured and unstructured 3D photorealistic scenes such as city view, object collection, and object fragmentation along with its rich ground truth annotation data. WorldGen being a generative model gives the user full access and control to features such as texture, object structure, motion, camera and lens properties for better generalizability by diminishing the data bias in the network. We demonstrate the effectiveness of WorldGen by evaluating deep optical flow. We hope such a tool can open doors for future research in a myriad of domains related to robotics and computer vision by reducing manual labor and cost for acquiring rich and high-quality data. Chahat Deep Singh, Riya Kumari, Cornelia Fermüller, Nitin J. Sanket, Yiannis Aloimonos |
ICRA | 5 |
| 2023 | Detecting Olives with Synthetic or Real Data? Olive the AboveabstractModern robotics has enabled the advancement in yield estimation for precision agriculture. However, when applied to the olive industry, the high variation of olive colors and their similarity to the background leaf canopy presents a challenge. Labeling several thousands of very dense olive grove images for segmentation is a labor-intensive task. This paper presents a novel approach to detecting olives without the need to manually label data. In this work, we present the world's first olive detection dataset comprised of synthetic and real olive tree images. This is accomplished by generating an auto-labeled photorealistic 3D model of an olive tree. Its geometry is then simplified for lightweight rendering purposes. In addition, experiments are conducted with a mix of synthetically generated and real images, yielding an improvement of up to 66% compared to when only using a small sample of real data. When access to real, human-labeled data is limited, a combination of mostly synthetic data and a small amount of real data can enhance olive detection. Yianni Karabatis, Xiaomin Lin 0002, Nitin J. Sanket, Michail G. Lagoudakis, Yiannis Aloimonos |
IROS | 5 |
| 2023 | Forecasting Action Through Contact Representations From First Person VideoabstractHuman actions involving hand manipulations are structured according to the making and breaking of hand-object contact, and human visual understanding of action is reliant on anticipation of contact as is demonstrated by pioneering work in cognitive science. Taking inspiration from this, we introduce representations and models centered on contact, which we then use in action prediction and anticipation. We annotate a subset of the EPIC Kitchens dataset to include time-to-contact between hands and objects, as well as segmentations of hands and objects. Using these annotations we train the Anticipation Module, a module producing Contact Anticipation Maps and Next Active Object Segmentations - novel low-level representations providing temporal and spatial characteristics of anticipated near future action. On top of the Anticipation Module we apply Egocentric Object Manipulation Graphs (Ego-OMG), a framework for action anticipation and prediction. Ego-OMG models longer term temporal semantic relations through the use of a graph modeling transitions between contact delineated action states. Use of the Anticipation Module within Ego-OMG produces state-of-the-art results, achieving 1st and 2 place on the unseen and seen test sets, respectively, of the EPIC Kitchens Action Anticipation Challenge, and achieving state-of-the-art results on the tasks of action anticipation and action prediction over EPIC Kitchens. We perform ablation studies over characteristics of the Anticipation Module to evaluate their utility. Eadom Dessalene, Chinmaya Devaraj, Michael Maynord, Cornelia Fermüller, Yiannis Aloimonos |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Brain-Inspired Hyperdimensional Computing for Ultra-Efficient Edge AIabstractHyperdimensional Computing (HDC) is rapidly emerging as an attractive alternative to traditional deep learning algorithms. Despite the profound success of Deep Neural Networks (DNNs) in many domains, the amount of computational power and storage that they demand during training makes deploying them in edge devices very challenging if not infeasible. This, in turn, inevitably necessitates streaming the data from the edge to the cloud which raises serious concerns when it comes to availability, scalability, security, and privacy. Further, the nature of data that edge devices often receive from sensors is inherently noisy. However, DNN algorithms are very sensitive to noise, which makes accomplishing the required learning tasks with high accuracy immensely difficult. In this paper, we aim at providing a comprehensive overview of the latest advances in HDC. HDC aims at realizing real-time performance and robustness through using strategies that more closely model the human brain. HDC is, in fact, motivated by the observation that the human brain operates on high-dimensional data representations. In HDC, objects are thereby encoded with high-dimensional vectors which have thousands of elements. In this paper, we will discuss the promising robustness of HDC algorithms against noise along with the ability to learn from little data. Further, we will present the outstanding synergy between HDC and beyond von Neumann architectures and how HDC opens doors for efficient learning at the edge due to the ultra-lightweight implementation that it needs, contrary to traditional DNNs. Hussam Amrouch, Mohsen Imani, Xun Jiao 0002, Yiannis Aloimonos, Cornelia Fermüller, Dehao Yuan, Dongning Ma, Hamza Errahmouni Barkam, Paul R. Genssler, Peter Sutor Jr. |
CODES+ISSS | 4 |
| 2022 | DiffPoseNet: Direct Differentiable Camera Pose EstimationabstractCurrent deep neural network approaches for camera pose estimation rely on scene structure for 3D motion estimation, but this decreases the robustness and thereby makes cross-dataset generalization difficult. In contrast, classical approaches to structure from motion estimate 3D motion utilizing optical flow and then compute depth. Their accuracy, however, depends strongly on the quality of the optical flow. To avoid this issue, direct methods have been proposed, which separate 3D motion from depth estimation, but compute 3D motion using only image gradients in the form of normal flow. In this paper, we introduce a network NFlowNet, for normal flow estimation which is used to enforce robust and direct constraints. In particular, normal flow is used to estimate relative camera pose based on the cheirality (depth positivity) constraint. We achieve this by formulating the optimization problem as a differentiable cheirality layer, which allows for end-to-end learning of camera pose. We perform extensive qualitative and quantitative evaluation of the proposed DiffPoseNet's sensitivity to noise and its generalization across datasets. We compare our approach to existing state-of-the-art methods on KITTI, TartanAir, and TUM-RGBD datasets. Chethan Parameshwara, Gokul Hari, Cornelia Fermüller, Nitin J. Sanket, Yiannis Aloimonos |
CVPR | 5 |
| 2022 | Gluing Neural Networks Symbolically Through Hyperdimensional ComputingabstractHyperdimensional Computing affords simple, yet powerful operations to create long Hyperdimensional Vectors (hypervectors) that can efficiently encode information, be used for learning, and are dynamic enough to be modified on the fly. In this paper, we explore the notion of using binary hypervectors to directly encode the final, classifying output signals of neural networks in order to fuse differing networks together at the symbolic level. This allows multiple neural networks to work together to solve a problem, with little additional overhead. Output signals just before classification are encoded as hyper-vectors and bundled together through consensus summation to train a classification hypervector. This process can be performed iteratively and even on single neural networks by instead making a consensus of multiple classification hypervectors. We find that this outperforms the state of the art, or is on a par with it, while using very little overhead, as hypervector operations are extremely fast and efficient in comparison to the neural networks. This consensus process can learn online and even grow or lose models in real time. Hypervectors act as memories that can be stored, and even further bundled together over time, affording life long learning capabilities. Additionally, this consensus structure inherits the benefits of Hyperdimensional Computing, without sacrificing the performance of modern Machine Learning. This technique can be extrapolated to virtually any neural model, and requires little modification to employ - one simply requires recording the output signals of networks when presented with a testing example. Peter Sutor Jr., Dehao Yuan, Douglas Summers-Stay, Cornelia Fermüller, Yiannis Aloimonos |
IJCNN | 5 |
| 2021 | 0-MMS: Zero-Shot Multi-Motion Segmentation With A Monocular Event CameraabstractSegmentation of moving objects in dynamic scenes is a key process in scene understanding for navigation tasks. Classical cameras suffer from motion blur in such scenarios rendering them effete. On the contrary, event cameras, because of their high temporal resolution and lack of motion blur, are tailor-made for this problem. We present an approach for monocular multi-motion segmentation, which combines bottom-up feature tracking and top-down motion compensation into a unified pipeline, which is the first of its kind to our knowledge. Using the events within a time-interval, our method segments the scene into multiple motions by splitting and merging. We further speed up our method by using the concept of motion propagation and cluster keyslices.The approach was successfully evaluated on both challenging real-world and synthetic scenarios from the EV-IMO, EED, and MOD datasets and outperformed the state-of-the-art detection rate by 12%, achieving a new state-of-the-art average detection rate of 81.06%, 94.2% and 82.35% on the aforementioned datasets. To enable further research and systematic evaluation of multi-motion segmentation, we present and open-source a new dataset/benchmark called MOD++, which includes challenging sequences and extensive data stratification in-terms of camera and object motion, velocity magnitudes, direction, and rotational speeds. Chethan Parameshwara, Nitin J. Sanket, Chahat Deep Singh, Cornelia Fermüller, Yiannis Aloimonos |
ICRA | 5 |
| 2021 | Detecting and Counting OystersabstractOysters are an essential species in the Chesapeake Bay living ecosystem. Oysters are filter feeders and considered the vacuum cleaners of the Chesapeake Bay that can considerably improve the Bay's water quality. Many oyster restoration programs have been initiated in the past decades and continued to date. Advancements in robotics and artificial intelligence have opened new opportunities for aquaculture. Drone-like ROVs with high maneuverability are getting more affordable and, if equipped with proper sensory devices, can monitor the oysters. This work presents our efforts for videography of the Chesapeake bay bottom using an ROV, constructing a database of oysters, implementing Mask R-CNN for detecting oysters, and counting their number in a video by tracking them. Behzad Sadrfaridpour, Yiannis Aloimonos, Miao Yu 0007, Donald Webster |
ICRA | 2 |
| 2021 | MorphEyes: Variable Baseline Stereo For Quadrotor NavigationabstractMorphable design and depth-based visual control are two upcoming trends leading to advancements in the field of quadrotor autonomy. Stereo-cameras have struck the perfect balance of weight and accuracy of depth estimation but suffer from the problem of depth range being limited and dictated by the baseline chosen at design time. In this paper, we present a framework for quadrotor navigation based on a stereo camera system whose baseline can be adapted on-the-fly. We present a method to calibrate the system at a small number of discrete baselines and interpolate the parameters for the entire baseline range. We present an extensive theoretical analysis of calibration and synchronization errors. We showcase three different applications of such a system for quadrotor navigation: (a) flying through a forest, (b) flying through an unknown shaped/location static/dynamic gap, and (c) accurate 3D pose detection of an independently moving object. We show that our variable baseline system is more accurate and robust in all three scenarios. To our knowledge, this is the first work that applies the concept of morphable design to achieve a variable baseline stereo vision system on a quadrotor. Nitin J. Sanket, Chahat Deep Singh, Varun Asthana, Cornelia Fermüller, Yiannis Aloimonos |
ICRA | 5 |
| 2021 | SpikeMS: Deep Spiking Neural Network for Motion SegmentationabstractSpiking Neural Networks (SNN) are the so-called third generation of neural networks which attempt to more closely match the functioning of the biological brain. They inherently encode temporal data, allowing for training with less energy usage and can be extremely energy efficient when coded on neuromorphic hardware. In addition, they are well suited for tasks involving event-based sensors, which match the event-based nature of the SNN. However, SNNs have not been as effectively applied to real-world, large-scale tasks as standard Artificial Neural Networks (ANNs) due to the algorithmic and training complexity. To exacerbate the situation further, the input representation is unconventional and requires careful analysis and deep understanding. In this paper, we propose SpikeMS, the first deep encoder-decoder SNN architecture for the real-world large-scale problem of motion segmentation using the event-based DVS camera as input. To accomplish this, we introduce a novel spatio-temporal loss formulation that includes both spike counts and classification labels in conjunction with the use of new techniques for SNN backpropagation. In addition, we show that SpikeMS is capable of incremental predictions, or predictions from smaller amounts of test data than it is trained on. This is invaluable for providing outputs even with partial input data for low-latency applications and those requiring fast predictions. We evaluated SpikeMS on challenging synthetic and real-world sequences from EV-IMO, EED and MOD datasets and achieving results on a par with a comparable ANN method, but using potentially 50 times less power. Chethan Parameshwara, Cornelia Fermüller, Nitin J. Sanket, Matthew Evanusa, Yiannis Aloimonos |
IROS | 6 |
| 2021 | NudgeSeg: Zero-Shot Object Segmentation by Repeated Physical InteractionabstractRecent advances in object segmentation have demonstrated that deep neural networks excel at object segmentation for specific classes in color and depth images. However, their performance is dictated by the number of classes and objects used for training, thereby hindering generalization to never seen objects or zero-shot samples. To exacerbate the problem further, object segmentation using image frames rely on recognition and pattern matching cues. Instead, we utilize the ‘active’ nature of a robot and their ability to ‘interact’ with the environment to induce additional geometric constraints for segmenting zero-shot samples.In this paper, we present the first framework to segment unknown objects in a cluttered scene by repeatedly ‘nudging’ at the objects and moving them to obtain additional motion cues at every step using only a monochrome monocular camera. We call our framework NudgeSeg. These motion cues are used to refine the segmentation masks. We successfully test our approach to segment novel objects in various cluttered scenes and provide an extensive study with image and motion segmentation methods. We show an impressive average detection rate of over 86% on zero-shot objects. Chahat Deep Singh, Nitin J. Sanket, Chethan Parameshwara, Cornelia Fermüller, Yiannis Aloimonos |
IROS | 5 |
| 2021 | Topology-Aware Non-Rigid Point Cloud RegistrationabstractIn this paper, we introduce a non-rigid registration pipeline for pairs of unorganized point clouds that may be topologically different. Standard warp field estimation algorithms, even under robust, discontinuity-preserving regularization, tend to produce erratic motion estimates on boundaries associated with 'close-to-open' topology changes. We overcome this limitation by exploiting backward motion: in the opposite motion direction, a 'close-to-open' event becomes 'open-to-close', which is by default handled correctly. At the core of our approach lies a general, topology-agnostic warp field estimation algorithm, similar to those employed in recently introduced dynamic reconstruction systems from RGB-D input. We improve motion estimation on boundaries associated with topology changes in an efficient post-processing phase. Based on both forward and (inverted) backward warp hypotheses, we explicitly detect regions of the deformed geometry that undergo topological changes by means of local deformation criteria and broadly classify them as 'contacts' or 'separations'. Subsequently, the two motion hypotheses are seamlessly blended on a local basis, according to the type and proximity of detected events. Our method achieves state-of-the-art motion estimation accuracy on the MPI Sintel dataset. Experiments on a custom dataset with topological event annotations demonstrate the effectiveness of our pipeline in estimating motion on event boundaries, as well as promising performance in explicit topological event detection. Konstantinos Zampogiannis, Cornelia Fermüller, Yiannis Aloimonos |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | Joint direct estimation of 3D geometry and 3D motion using spatio temporal gradients
Francisco Barranco, Cornelia Fermüller, Yiannis Aloimonos, Eduardo Ros Vidal |
Pattern Recognit. | 3 |
| 2020 | Learning Visual Motion Segmentation Using Event SurfacesabstractEvent-based cameras have been designed for scene motion perception - their high temporal resolution and spatial data sparsity converts the scene into a volume of boundary trajectories and allows to track and analyze the evolution of the scene in time. Analyzing this data is computationally expensive, and there is substantial lack of theory on dense-in-time object motion to guide the development of new algorithms; hence, many works resort to a simple solution of discretizing the event stream and converting it to classical pixel maps, which allows for application of conventional image processing methods. In this work we present a Graph Convolutional neural network for the task of scene motion segmentation by a moving camera. We convert the event stream into a 3D graph in (x,y,t) space and keep per-event temporal information. The difficulty of the task stems from the fact that unlike in metric space, the shape of an object in (x,y,t) space depends on its motion and is not the same across the dataset. We discuss properties of of the event data with respect to this 3D recognition problem, and show that our Graph Convolutional architecture is superior to PointNet++. We evaluate our method on the state of the art event-based motion segmentation dataset - EV-IMO and perform comparisons to a frame-based method proposed by its authors. Our ablation studies show that increasing the event slice width improves the accuracy, and how subsampling and edge configurations affect the network performance. Anton Mitrokhin, Zhiyuan Hua, Cornelia Fermüller, Yiannis Aloimonos |
CVPR | 4 |
| 2020 | Network Deconvolution
Chengxi Ye, Matthew Evanusa, Anton Mitrokhin, Tom Goldstein, James A. Yorke, Cornelia Fermüller, Yiannis Aloimonos |
ICLR | 8 |
| 2020 | EVDodgeNet: Deep Dynamic Obstacle Dodging with Event CamerasabstractDynamic obstacle avoidance on quadrotors requires low latency. A class of sensors that are particularly suitable for such scenarios are event cameras. In this paper, we present a deep learning based solution for dodging multiple dynamic obstacles on a quadrotor with a single event camera and on-board computation. Our approach uses a series of shallow neural networks for estimating both the ego-motion and the motion of independently moving objects. The networks are trained in simulation and directly transfer to the real world without any fine-tuning or retraining. We successfully evaluate and demonstrate the proposed approach in many real-world experiments with obstacles of different shapes and sizes, achieving an overall success rate of 70% including objects of unknown shape and a low light testing scenario. To our knowledge, this is the first deep learning - based solution to the problem of dynamic obstacle avoidance using event cameras on a quadrotor. Finally, we also extend our work to the pursuit task by merely reversing the control policy, proving that our navigation stack can cater to different scenarios. Nitin J. Sanket, Chethan Parameshwara, Chahat Deep Singh, Ashwin V. Kuruttukulam, Cornelia Fermüller, Davide Scaramuzza 0001, Yiannis Aloimonos |
ICRA | 7 |
| 2020 | Unsupervised Learning of Dense Optical Flow, Depth and Egomotion with Event-Based SensorsabstractWe present an unsupervised learning pipeline for dense depth, optical flow and egomotion estimation for autonomous driving applications, using the event-based output of the Dynamic Vision Sensor (DVS) as input. The backbone of our pipeline is a bioinspired encoder-decoder neural network architecture - ECN. To train the pipeline, we introduce a covariance normalization technique which resembles the lateral inhibition mechanism found in animal neural systems.Our work is the first monocular pipeline that generates dense depth and optical flow from sparse event data only, and is able to transfer from day to night scenes without any additional training. The network works in self-supervised mode and has just 150k parameters. We evaluate our pipeline on the MVSEC self driving dataset and present results for depth, optical flow and and egomotion estimation. Thanks to the efficient design, we are able to achieve inference rates of 300 FPS on a single Nvidia 1080Ti GPU. Our experiments demonstrate significant improvements upon works that used deep learning on event data, as well as the ability to perform well during both day and night. Chengxi Ye, Anton Mitrokhin, Cornelia Fermüller, James A. Yorke, Yiannis Aloimonos |
IROS | 5 |
| 2019 | EV-IMO: Motion Segmentation Dataset and Learning Pipeline for Event CamerasabstractWe present the first event-based learning approach for motion segmentation in indoor scenes and the first event-based dataset - EV-IMO- which includes accurate pixel-wise motion masks, egomotion and ground truth depth. Our approach is based on an efficient implementation of the SfM learning pipeline using a low parameter neural network architecture on event data. In addition to camera egomotion and a dense depth map, the network estimates independently moving object segmentation at the pixel-level and computes per-object 3D translational velocities of moving objects. We also train a shallow network with just 40k parameters, which is able to compute depth and egomotion. Our EV-IMO dataset features 32 minutes of indoor recording with up to 3 fast moving objects in the camera field of view. The objects and the camera are tracked using a VICON®motion capture system. By 3D scanning the room and the objects, ground truth of the depth map and pixel-wise object masks are obtained. We then train and evaluate our learning pipeline on EV-IMO and demonstrate that it is well suited for scene constrained robotics applications. SUPPLEMENTARY MATERIAL The supplementary video, code, trained models, appendix and a dataset will be made available at http://prg.cs.umd.edu/EV-IMO.html. Anton Mitrokhin, Chengxi Ye, Cornelia Fermüller, Yiannis Aloimonos, Tobi Delbruck |
IROS | 4 |
| 2019 | SalientDSO: Bringing Attention to Direct Sparse OdometryabstractAlthough cluttered indoor scenes have a lot of useful high-level semantic information which can be used for mapping and localization, most visual odometry (VO) algorithms rely on the usage of geometric features such as points, lines, and planes. Lately, driven by this idea, the joint optimization of semantic labels and estimating odometry has gained popularity in the robotics community. This joint optimization method is accurate but is generally very slow. At the same time, in the vision community, direct and sparse approaches for VO have stricken the right balance between speed and accuracy. We merge the successes of these two communities and present a preprocessing method to incorporate semantic information in the form of visual saliency to direct sparse odometry (DSO)-a highly successful direct sparse VO algorithm. We also present a framework to filter the visual saliency based on scene parsing. Our framework SalientDSO relies on the widely successful deep learning-based approaches for visual saliency and scene parsing, which drives the feature selection for obtaining highly accurate and robust VO even in the presence of as few as 40 point features per frame. We provide an extensive quantitative evaluation of SalientDSO on the ICL-NUIM and the TUM monoVO data sets and show that we outperform DSO and ORB-simultaneous localization and mapping-two very popular state-of-the-art approaches in the literature. We also collect and publicly release a CVL-UMD data set which contains two indoor cluttered sequences on which we show qualitative evaluations. To the best of our knowledge, this is the first paper to use visual saliency and scene parsing to drive the feature selection in direct VO. Huai-Jen Liang, Nitin J. Sanket, Cornelia Fermüller, Yiannis Aloimonos |
IEEE Trans Autom. Sci. Eng. | 4 |
| 2018 | Evenly Cascaded Convolutional NetworksabstractWe introduce Evenly Cascaded convolutional Network (ECN), a neural network taking inspiration from the cascade algorithm of wavelet analysis. ECN employs two feature streams - a low-level and high-level steam. At each layer these streams interact, such that low-level features are modulated using advanced perspectives from the high-level stream. ECN is evenly structured through resizing feature map dimensions by a consistent ratio, which removes the burden of ad-hoc specification of feature map dimensions. ECN produces easily interpretable features maps, a result whose intuition can be understood in the context of scale-space theory. We demonstrate that ECN's design facilitates the training process through providing easily trainable shortcuts. We report new state-of-the-art results for small networks, without the need for additional treatment such as pruning or compression - a consequence of ECN's simple structure and direct training. A 6-layered ECN design with under 500k parameters achieves 95.24% and 78.99% accuracy on CIFAR-10 and CIFAR-100 datasets, respectively, outperforming the current state-of-the-art on small parameter networks, and a 3 million parameter ECN produces results competitive to the state-of-the-art. Chengxi Ye, Chinmaya Devaraj, Michael Maynord, Cornelia Fermüller, Yiannis Aloimonos |
IEEE BigData | 5 |
| 2018 | An Embodied Intelligent Tutor for Literal Concepts' Recognition
Marietta Sionti, Thomas Schack, Yiannis Aloimonos |
CogSci | 3 |
| 2018 | Seeing Behind the Scene: Using Symmetry to Reason About Objects in Cluttered EnvironmentsabstractSymmetry is a common property shared by the majority of man-made objects. This paper presents a novel bottom-up approach for segmenting symmetric objects and recovering their symmetries from 3D pointclouds of natural scenes. Candidate rotational and reflectional symmetries are detected by fitting symmetry axes/planes to the geometry of the smooth surfaces extracted from the scene. Individual symmetries are used as constraints for the foreground segmentation problem that uses symmetry as a global grouping principle. Evaluation on a challenging dataset shows that our approach can reliably segment objects and extract their symmetries from incomplete 3D reconstructions of highly cluttered scenes, outperforming state-of-the-art methods by a wide margin. Aleksandrs Ecins, Cornelia Fermüller, Yiannis Aloimonos |
IROS | 3 |
| 2018 | Event-Based Moving Object Detection and TrackingabstractEvent-based vision sensors, such as the Dynamic Vision Sensor (DVS), are ideally suited for real-time motion analysis. The unique properties encompassed in the readings of such sensors provide high temporal resolution, superior sensitivity to light and low latency. These properties provide the grounds to estimate motion efficiently and reliably in the most sophisticated scenarios, but these advantages come at a price - modern event-based vision sensors have extremely low resolution, produce a lot of noise and require the development of novel algorithms to handle the asynchronous event stream. This paper presents a new, efficient approach to object tracking with asynchronous cameras. We present a novel event stream representation which enables us to utilize information about the dynamic (temporal)component of the event stream. The 3D geometry of the event stream is approximated with a parametric model to motion-compensate for the camera (without feature tracking or explicit optical flow computation), and then moving objects that don't conform to the model are detected in an iterative process. We demonstrate our framework on the task of independent motion detection and tracking, where we use the temporal model inconsistencies to locate differently moving objects in challenging situations of very fast motion. Anton Mitrokhin, Cornelia Fermüller, Chethan Parameshwara, Yiannis Aloimonos |
IROS | 4 |
| 2018 | cilantro: A Lean, Versatile, and Efficient Library for Point Cloud Data ProcessingabstractWe introduce Cilantro, an open-source C++ library for geometric and general-purpose point cloud data processing. The library provides functionality that covers low-level point cloud operations, spatial reasoning, various methods for point cloud segmentation and generic data clustering, flexible algorithms for robust or local geometric alignment, model fitting, as well as powerful visualization tools. To accommodate all kinds of workflows, Cilantro is almost fully templated, and most of its generic algorithms operate in arbitrary data dimension. At the same time, the library is easy to use and highly expressive, promoting a clean and concise coding style. Cilantro is highly optimized, has a minimal set of external dependencies, and supports rapid development of performant point cloud processing software in a wide variety of contexts. Konstantinos Zampogiannis, Cornelia Fermüller, Yiannis Aloimonos |
ACM Multimedia | 3 |
| 2018 | Combining Knowledge and Reasoning through Probabilistic Soft Logic for Image Puzzle Solving
Somak Aditya, Yezhou Yang, Chitta Baral, Yiannis Aloimonos |
UAI | 4 |
| 2018 | Image Understanding using vision and reasoning through Scene Description Graph
Somak Aditya, Yezhou Yang, Chitta Baral, Yiannis Aloimonos, Cornelia Fermüller |
Comput. Vis. Image Underst. | 4 |
| 2017 | What can i do around here? Deep functional scene understanding for cognitive robotsabstractFor robots that have the capability to interact with the physical environment through their end effectors, understanding the surrounding scenes is not merely a task of image classification or object recognition. To perform actual tasks, it is critical for the robot to have a functional understanding of the visual scene. Here, we address the problem of localization and recognition of functional areas in an arbitrary indoor scene, formulated as a two-stage deep learning based detection pipeline. A new scene functionality testing-bed, which is compiled from two publicly available indoor scene datasets, is used for evaluation. Our method is evaluated quantitatively on the new dataset, demonstrating the ability to perform efficient recognition of functional areas from arbitrary indoor scenes. We also demonstrate that our detection model can be generalized to novel indoor scenes by cross validating it with images from two different datasets. Chengxi Ye, Yezhou Yang, Ren Mao, Cornelia Fermüller, Yiannis Aloimonos |
ICRA | 5 |
| 2016 | Cluttered scene segmentation using the symmetry constraintabstractAlthough modern object segmentation algorithms can deal with isolated objects in simple scenes, segmenting non-convex objects in cluttered environments remains a challenging task. We introduce a novel approach for segmenting unknown objects in partial 3D pointclouds that utilizes the powerful concept of symmetry. First, 3D bilateral symmetries in the scene are detected efficiently by extracting and matching surface normal edge curves in the pointcloud. Symmetry hypotheses are then used to initialize a segmentation process that finds points of the scene that are consistent with each of the detected symmetries. We evaluate our approach on a dataset of 3D pointcloud scans of tabletop scenes. We demonstrate that the use of the symmetry constraint enables our approach to correctly segment objects in challenging configurations and to outperform current state-of-the-art approaches. Aleksandrs Ecins, Cornelia Fermüller, Yiannis Aloimonos |
ICRA | 3 |
| 2016 | LightNet: A Versatile, Standalone Matlab-based Environment for Deep LearningabstractLightNet is a lightweight, versatile, purely Matlab-based deep learning framework. The idea underlying its design is to provide an easy-to-understand, easy-to-use and efficient computational platform for deep learning research. The implemented framework supports major deep learning architectures such as Multilayer Perceptron Networks (MLP), Convolutional Neural Networks (CNN) and Recurrent Neural Networks (RNN). The framework also supports both CPU and GPU computation, and the switch between them is straightforward. Different applications in computer vision, natural language processing and robotics are demonstrated as experiments. Chengxi Ye, Chen Zhao 0009, Yezhou Yang, Cornelia Fermüller, Yiannis Aloimonos |
ACM Multimedia | 5 |
| 2015 | Robot Learning Manipulation Action Plans by "Watching" Unconstrained Videos from the World Wide WebabstractIn order to advance action generation and creation in robots beyond simple learned schemas we need computational tools that allow us to automatically interpret and represent human actions. This paper presents a system that learns manipulation action plans by processing unconstrained videos from the World Wide Web. Its goal is to robustly generate the sequence of atomic actions of seen longer actions in video in order to acquire knowledge for robots. The lower level of the system consists of two convolutional neural network (CNN) based recognition modules, one for classifying the hand grasp type and the other for object recognition. The higher level is a probabilistic manipulation action grammar based parsing module that aims at generating visual sentences for robot manipulation. Experiments conducted on a publicly available unconstrained video dataset show that the system is able to learn manipulation actions by ``watching'' unconstrained videos with high accuracy. Yezhou Yang, Cornelia Fermüller, Yiannis Aloimonos |
AAAI | 4 |
| 2015 | Learning the Semantics of Manipulation ActionabstractIn this paper we present a formal computational framework for modeling manipulation actions. The introduced formalism leads to semantics of manipulation action and has applications to both observing and understanding human manipulation actions as well as executing them with a robotic mechanism (e.g. a humanoid robot). It is based on a Combinatory Categorial Grammar. The goal of the introduced framework is to: (1) represent manipulation actions with both syntax and semantic parts, where the semantic part employs $\lambda$-calculus; (2) enable a probabilistic semantic parsing schema to learn the $\lambda$-calculus representation of manipulation action from an annotated action corpus of videos; (3) use (1) and (2) to develop a system that visually observes manipulation actions and understands their meaning while it can reason beyond observations using propositional logic and axiom schemata. The experiments conducted on a public available large manipulation action dataset validate the theoretical framework and our implementation. Yezhou Yang, Yiannis Aloimonos, Cornelia Fermüller, Eren Erdal Aksoy |
ACL (1) | 2 |
| 2015 | Fast 2D border ownership assignmentabstractA method for efficient border ownership assignment in 2D images is proposed. Leveraging on recent advances using Structured Random Forests (SRF) for boundary detection [8], we impose a novel border ownership structure that detects both boundaries and border ownership at the same time. Key to this work are features that predict ownership cues from 2D images. To this end, we use several different local cues: shape, spectral properties of boundary patches, and semi-global grouping cues that are indicative of perceived depth. For shape, we use HoG-like descriptors that encode local curvature (convexity and concavity). For spectral properties, such as extremal edges [28], we first learn an orthonormal basis spanned by the top K eigenvectors via PCA over common types of contour tokens [23]. For grouping, we introduce a novel mid-level descriptor that captures patterns near edges and indicates ownership information of the boundary. Experimental results over a subset of the Berkeley Segmentation Dataset (BSDS) [24] and the NYU Depth V2 [34] dataset show that our method's performance exceeds current state-of-the-art multi-stage approaches that use more complex features. Ching Lik Teo, Cornelia Fermüller, Yiannis Aloimonos |
CVPR | 3 |
| 2015 | Grasp type revisited: A modern perspective on a classical feature for visionabstractThe grasp type provides crucial information about human action. However, recognizing the grasp type from unconstrained scenes is challenging because of the large variations in appearance, occlusions and geometric distortions. In this paper, first we present a convolutional neural network to classify functional hand grasp types. Experiments on a public static scene hand data set validate good performance of the presented method. Then we present two applications utilizing grasp type classification: (a) inference of human action intention and (b) fine level manipulation action segmentation. Experiments on both tasks demonstrate the usefulness of grasp type as a cognitive feature for computer vision. This study shows that the grasp type is a powerful symbolic representation for action understanding, and thus opens new avenues for future research. Yezhou Yang, Cornelia Fermüller, Yiannis Aloimonos |
CVPR | 4 |
| 2015 | Contour Detection and Characterization for Asynchronous Event SensorsabstractThe bio-inspired, asynchronous event-based dynamic vision sensor records temporal changes in the luminance of the scene at high temporal resolution. Since events are only triggered at significant luminance changes, most events occur at the boundary of objects and their parts. The detection of these contours is an essential step for further interpretation of the scene. This paper presents an approach to learn the location of contours and their border ownership using Structured Random Forests on event-based features that encode motion, timing, texture, and spatial orientations. The classifier integrates elegantly information over time by utilizing the classification results previously computed. Finally, the contour detection and boundary assignment are demonstrated in a layer-segmentation of the scene. Experimental results demonstrate good performance in boundary detection and segmentation. Francisco Barranco, Ching Lik Teo, Cornelia Fermüller, Yiannis Aloimonos |
ICCV | 4 |
| 2015 | Detection and Segmentation of 2D Curved Reflection Symmetric StructuresabstractSymmetry, as one of the key components of Gestalt theory, provides an important mid-level cue that serves as input to higher visual processes such as segmentation. In this work, we propose a complete approach that links the detection of curved reflection symmetries to produce symmetry-constrained segments of structures/regions in real images with clutter. For curved reflection symmetry detection, we leverage on patch-based symmetric features to train a Structured Random Forest classifier that detects multiscaled curved symmetries in 2D images. Next, using these curved symmetries, we modulate a novel symmetry-constrained foreground-background segmentation by their symmetry scores so that we enforce global symmetrical consistency in the final segmentation. This is achieved by imposing a pairwise symmetry prior that encourages symmetric pixels to have the same labels over a MRF-based representation of the input image edges, and the final segmentation is obtained via graph-cuts. Experimental results over four publicly available datasets containing annotated symmetric structures: 1) SYMMAX-300 [38], 2) BSD-Parts, 3) Weizmann Horse (both from [18]) and 4) NY-roads [35] demonstrate the approach's applicability to different environments with state-of-the-art performance. Ching Lik Teo, Cornelia Fermüller, Yiannis Aloimonos |
ICCV | 3 |
| 2015 | Affordance detection of tool parts from geometric featuresabstractAs robots begin to collaborate with humans in everyday workspaces, they will need to understand the functions of tools and their parts. To cut an apple or hammer a nail, robots need to not just know the tool's name, but they must localize the tool's parts and identify their functions. Intuitively, the geometry of a part is closely related to its possible functions, or its affordances. Therefore, we propose two approaches for learning affordances from local shape and geometry primitives: 1) superpixel based hierarchical matching pursuit (S-HMP); and 2) structured random forests (SRF). Moreover, since a part can be used in many ways, we introduce a large RGB-Depth dataset where tool parts are labeled with multiple affordances and their relative rankings. With ranked affordances, we evaluate the proposed methods on 3 cluttered scenes and over 105 kitchen, workshop and garden tools, using ranked correlation and a weighted F-measure score [26]. Experimental results over sequences containing clutter, occlusions, and viewpoint changes show that the approaches return precise predictions that could be used by a robot. S-HMP achieves high accuracy but at a significant computational cost, while SRF provides slightly less accurate predictions but in real-time. Finally, we validate the effectiveness of our approaches on the Cornell Grasping Dataset [25] for detecting graspable regions, and achieve state-of-the-art performance. Austin Myers, Ching Lik Teo, Cornelia Fermüller, Yiannis Aloimonos |
ICRA | 4 |
| 2015 | Learning the spatial semantics of manipulation actions through preposition groundingabstractIn this paper, we introduce an abstract representation for manipulation actions that is based on the evolution of the spatial relations between involved objects. Object tracking in RGBD streams enables straightforward and intuitive ways to model spatial relations in 3D space. Reasoning in 3D overcomes many of the limitations of similar previous approaches, while providing significant flexibility in the desired level of abstraction. At each frame of a manipulation video, we evaluate a number of spatial predicates for all object pairs and treat the resulting set of sequences (Predicate Vector Sequences, PVS) as an action descriptor. As part of our representation, we introduce a symmetric, time-normalized pairwise distance measure that relies on finding an optimal object correspondence between two actions. We experimentally evaluate the method on the classification of various manipulation actions in video, performed at different speeds and timings and involving different objects. The results demonstrate that the proposed representation is remarkably descriptive of the high-level manipulation semantics. Konstantinos Zampogiannis, Yezhou Yang, Cornelia Fermüller, Yiannis Aloimonos |
ICRA | 4 |
| 2015 | The Cognitive Dialogue: A new model for vision implementing common sense reasoning
Yiannis Aloimonos, Cornelia Fermüller |
Image Vis. Comput. | 1 |
| 2014 | Shadow free segmentation in still images using local density measureabstractOver the last decades several approaches were introduced to deal with cast shadows in background subtraction applications. However, very few algorithms exist that address the same problem for still images. In this paper we propose a figure ground segmentation algorithm to segment objects in still images affected by shadows. Instead of modeling the shadow directly in the segmentation process our approach works actively by first segmenting an object and then testing the resulting boundary for the presence of shadows and resegmenting again with modified segmentation parameters. In order to get better shadow boundary detection results we introduce a novel image preprocessing technique based on the notion of the image density map. This map improves the illumination invariance of classical filter-bank based texture description methods. We demonstrate that this texture feature improves shadow detection results. The resulting segmentation algorithm achieves good results on a new figure ground segmentation dataset with challenging illumination conditions. Aleksandrs Ecins, Cornelia Fermüller, Yiannis Aloimonos |
ICCP | 3 |
| 2014 | Contour Motion Estimation for Asynchronous Event-Driven CamerasabstractThis paper compares image motion estimation with asynchronous event-based cameras to Computer Vision approaches using as input frame-based video sequences. Since dynamic events are triggered at significant intensity changes, which often are at the border of objects, we refer to the event-based image motion as “contour motion.” Algorithms are presented for the estimation of accurate contour motion from local spatio-temporal information for two camera models: the dynamic vision sensor (DVS), which asynchronously records temporal changes of the luminance, and a family of new sensors which combine DVS data with intensity signals. These algorithms take advantage of the high temporal resolution of the DVS and achieve robustness using a multiresolution scheme in time. It is shown that, because of the coupling of velocity and luminance information in the event distribution, the image motion estimation problem becomes much easier with the new sensors which provide both events and image intensity than with the DVS alone. Experiments on synthesized data from computer vision benchmarks show that our algorithm on combined data outperforms computer vision methods in accuracy and can achieve real-time performance, and experiments on real data confirm the feasibility of the approach. Given that current image motion (or so-called optic flow) methods cannot estimate well at object boundaries, the approach presented here could be used complementary to optic flow techniques, and can provide new avenues for computer vision motion research. Francisco Barranco, Cornelia Fermüller, Yiannis Aloimonos |
Proc. IEEE | 3 |
| 2013 | Action Attribute Detection from Sports Videos with Contextual ConstraintsabstractIn this paper, we are interested in detecting action attributes from sports videos for event understanding and video analysis. Action attribute is a middle layer between low level motion features and high level action classes, which includes various motion patterns of human limbs and bodies and the interaction between human and objects. Successfully detecting action attributes provides a richer video description that facilitates many other important tasks, such action classification, video understanding, automatic video transcript, etc. A naive approach to deal with this challenging problem is to train a classifier for each attribute and then use them to detect attributes in novel videos independently. However, this independence assumption is often too strong, and as we show in our experiments, produces a large number of false positives in practice. We propose a novel approach that incorporates the contextual constraints for activity attribute detection. The temporal contexts within an attribute and the co-occurrence contexts between different attributes are modelled by a factorial conditional random field, which encourages agreement between different time points and attributes. The effectiveness of our methods are clearly illustrated by the experimental evaluations. Xiaodong Yu 0002, Ching Lik Teo, Yezhou Yang, Cornelia Fermüller, Yiannis Aloimonos |
BMVC | 5 |
| 2013 | Detection of Manipulation Action Consequences (MAC)abstractThe problem of action recognition and human activity has been an active research area in Computer Vision and Robotics. While full-body motions can be characterized by movement and change of posture, no characterization, that holds invariance, has yet been proposed for the description of manipulation actions. We propose that a fundamental concept in understanding such actions, are the consequences of actions. There is a small set of fundamental primitive action consequences that provides a systematic high-level classification of manipulation actions. In this paper a technique is developed to recognize these action consequences. At the heart of the technique lies a novel active tracking and segmentation method that monitors the changes in appearance and topological structure of the manipulated object. These are then used in a visual semantic graph (VSG) based procedure applied to the time sequence of the monitored object to recognize the action consequence. We provide a new dataset, called Manipulation Action Consequences (MAC 1.0), which can serve as test bed for other studies on this topic. Several experiments on this dataset demonstrates that our method can robustly track objects and detect their deformations and division during the manipulation. Quantitative tests prove the effectiveness and efficiency of the method. Yezhou Yang, Cornelia Fermüller, Yiannis Aloimonos |
CVPR | 3 |
| 2013 | Embedding high-level information into low level vision: Efficient object search in clutterabstractThe ability to search visually for objects of interest in cluttered environments is crucial for robots performing tasks in a multitude of environments. In this work, we propose a novel visual search algorithm that integrates high-level information of the target object - specifically its size and shape, with a recently introduced visual operator that rapidly clusters potential edges based on their coherence in belonging to a possible object. The output is a set of fixation points that indicate the potential location of the target object in the image. The proposed approach outperforms purely bottom-up approaches - saliency maps of Itti et al. [15], and kernel descriptors of Bo et al. [2], over two large datasets of objects in clutter collected using an RGB-Depth camera. Ching Lik Teo, Austin Myers, Cornelia Fermüller, Yiannis Aloimonos |
ICRA | 4 |
| 2013 | Robots with language: Multi-label visual recognition using NLPabstractThere has been a recent interest in utilizing contextual knowledge to improve multi-label visual recognition for intelligent agents like robots. Natural Language Processing (NLP) can give us labels, the correlation of labels, and the ontological knowledge about them, so we can automate the acquisition of contextual knowledge. In this paper we show how to use tools from NLP in conjunction with Vision to improve visual recognition. There are two major approaches: First, different language databases organize words according to various semantic concepts. Using these, we can build special purpose databases that can predict the labels involved given a certain context. Here we build a knowledge base for the purpose of describing common daily activities. Second, statistical language tools can provide the correlations of different labels. We show a way to learn a language model from large corpus data that exploits these correlations and propose a general optimization scheme to integrate the language model into the system. Experiments conducted on three multi-label everyday recognition tasks support the effectiveness and efficiency of our approach, with significant gains in recognition accuracies when correlation information is used. Yezhou Yang, Ching Lik Teo, Cornelia Fermüller, Yiannis Aloimonos |
ICRA | 4 |
| 2013 | Minimalist plans for interpreting manipulation actionsabstractHumans attribute meaning to actions, and can recognize, imitate, predict, compose from parts, and analyse complex actions performed by other humans. We have built a model of action representation and understanding which takes as input perceptual data of humans performing manipulatory actions and finds a semantic interpretation of it. It achieves this by representing actions as minimal plans based on a few primitives. The motivation for our approach is to have a description, that abstracts away the variations in the way humans perform actions. The model can be used to represent complex activities on the basis of simple actions. The primitives of these minimal plans are embodied in the physicality of the system doing the analysis. The model understands an action under observation by recognising which plan is occurring. Using primitives thus rooted in its own physical structure, the model has a semanticist and causal understanding of what it observes. Using plans, the model considers actions as well as complex activities in terms of causality, compositions, and goal achievement, enabling it to perform complex tasks like prediction of primitives, separation of interleaved actions and filtering of perceptual input. We use our model over an action dataset involving humans using hand tools on objects in a constrained universe to understand an activity it has not seen before in terms of actions whose plans it knows of. The model thus illustrates a novel approach of understanding human actions by a robot. Anupam Guha, Yezhou Yang, Cornelia Fermüller, Yiannis Aloimonos |
IROS | 4 |
| 2012 | Segmenting "simple" objects using RGB-DabstractSegmenting “simple” objects using low-level visual cues is an important capability for a vision system to learn in an unsupervised manner. We define a “simple” object as a compact region enclosed by depth and/or contact boundary in the scene. We propose a segmentation process to extract all the “simple” objects that builds on the fixation-based segmentation framework [1] that segments a region given a point anywhere inside it. In this work, we augment that framework with a fixation strategy to automatically select points inside the “simple” objects and a post-segmentation process to select only the regions corresponding to the “simple” objects in the scene. A novel characteristic of our approach is the incorporation of border ownership, the knowledge about the object side of a boundary pixel. We evaluate the process on a publicly available RGB-D dataset [2] and find that the proposed method successfully extracts 91.4% of all objects in the dataset. Ajay K. Mishra, Ashish Shrivastava 0001, Yiannis Aloimonos |
ICRA | 3 |
| 2012 | Towards a Watson that sees: Language-guided action recognition for robotsabstractFor robots of the future to interact seamlessly with humans, they must be able to reason about their surroundings and take actions that are appropriate to the situation. Such reasoning is only possible when the robot has knowledge of how the World functions, which must either be learned or hard-coded. In this paper, we propose an approach that exploits language as an important resource of high-level knowledge that a robot can use, akin to IBM's Watson in Jeopardy!. In particular, we show how language can be leveraged to reduce the ambiguity that arises from recognizing actions involving hand-tools from video data. Starting from the premise that tools and actions are intrinsically linked, with one explaining the existence of the other, we trained a language model over a large corpus of English newswire text so that we can extract this relationship directly. This model is then used as a prior to select the best tool and action that explains the video. We formalize the approach in the context of 1) an unsupervised recognition and 2) a supervised classification scenario by an EM formulation for the former and integrating language features for the latter. Results are validated over a new hand-tool action dataset, and comparisons with state of the art STIP features showed significantly improved results when language is used. In addition, we discuss the implications of these results and how it provides a framework for integrating language into vision on other robotic applications. Ching Lik Teo, Yezhou Yang, Hal Daumé III, Cornelia Fermüller, Yiannis Aloimonos |
ICRA | 5 |
| 2012 | Using a minimal action grammar for activity understanding in the real worldabstractThere is good reason to believe that humans use some kind of recursive grammatical structure when we recognize and perform complex manipulation activities. We have built a system to automatically build a tree structure from observations of an actor performing such activities. The activity trees that result form a framework for search and understanding, tying action to language. We explore and evaluate the system by performing experiments over a novel complex activity dataset taken using synchronized Kinect and SR4000 Time of Flight cameras. Processing of the combined 3D and 2D image data provides the necessary terminals and events to build the tree from the bottom-up. Experimental results highlight the contribution of the action grammar in: 1) providing a robust structure for complex activity recognition over real data and 2) disambiguating interleaved activities from within the same sequence. Douglas Summers-Stay, Ching Lik Teo, Yezhou Yang, Cornelia Fermüller, Yiannis Aloimonos |
IROS | 5 |
| 2012 | Active Visual SegmentationabstractAttention is an integral part of the human visual system and has been widely studied in the visual attention literature. The human eyes fixate at important locations in the scene, and every fixation point lies inside a particular region of arbitrary shape and size, which can either be an entire object or a part of it. Using that fixation point as an identification marker on the object, we propose a method to segment the object of interest by finding the "optimal" closed contour around the fixation point in the polar space, avoiding the perennial problem of scale in the Cartesian space. The proposed segmentation process is carried out in two separate steps: First, all visual cues are combined to generate the probabilistic boundary edge map of the scene; second, in this edge map, the "optimal" closed contour around a given fixation point is found. Having two separate steps also makes it possible to establish a simple feedback between the mid-level cue (regions) and the low-level visual cues (edges). In fact, we propose a segmentation refinement process based on such a feedback process. Finally, our experiments show the promise of the proposed method as an automatic segmentation framework for a general purpose visual system. Ajay K. Mishra, Yiannis Aloimonos, Loong Fah Cheong, Ashraf A. Kassim |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2011 | Corpus-Guided Sentence Generation of Natural Images
Yezhou Yang, Ching Lik Teo, Hal Daumé III, Yiannis Aloimonos |
EMNLP | 4 |
| 2011 | Active scene recognition with vision and languageabstractThis paper presents a novel approach to utilizing high level knowledge for the problem of scene recognition in an active vision framework, which we call active scene recognition. In traditional approaches, high level knowledge is used in the post-processing to combine the outputs of the object detectors to achieve better classification performance. In contrast, the proposed approach employs high level knowledge actively by implementing an interaction between a reasoning module and a sensory module (Figure 1). Following this paradigm, we implemented an active scene recognizer and evaluated it with a dataset of 20 scenes and 100+ objects. We also extended it to the analysis of dynamic scenes for activity recognition with attributes. Experiments demonstrate the effectiveness of the active paradigm in introducing attention and additional constraints into the sensing process. Xiaodong Yu 0002, Cornelia Fermüller, Ching Lik Teo, Yezhou Yang, Yiannis Aloimonos |
ICCV | 5 |
| 2010 | Learning shift-invariant sparse representation of actionsabstractA central problem in the analysis of motion capture (MoCap) data is how to decompose motion sequences into primitives. Ideally, a description in terms of primitives should facilitate the recognition, synthesis, and characterization of actions. We propose an unsupervised learning algorithm for automatically decomposing joint movements in human motion capture (MoCap) sequences into shift-invariant basis functions. Our formulation models the time series data of joint movements in actions as a sparse linear combination of short basis functions (snippets), which are executed (or “activated”) at different positions in time. Given a set of MoCap sequences of different actions, our algorithm finds the decomposition of MoCap sequences in terms of basis functions and their activations in time. Using the tools of L1minimization, the procedure alternately solves two large convex minimizations: Given the basis functions, a variant of Orthogonal Matching Pursuit solves for the activations, and given the activations, the Split Bregman Algorithm solves for the basis functions. Experiments demonstrate the power of the decomposition in a number of applications, including action recognition, retrieval, MoCap data compression, and as a tool for classification in the diagnosis of Parkinson (a motion disorder disease). Cornelia Fermüller, Yiannis Aloimonos, Hui Ji 0002 |
CVPR | 3 |
| 2010 | An Experimental Study of Color-Based Segmentation Algorithms Based on the Mean-Shift Concept
Konstantinos Bitsakos, Cornelia Fermüller, Yiannis Aloimonos |
ECCV (2) | 3 |
| 2010 | Attribute-Based Transfer Learning for Object Categorization with Zero/One Training Example
Xiaodong Yu 0002, Yiannis Aloimonos |
ECCV (5) | 2 |
| 2010 | Moving obstacle detection using cameras for driver assistance systemabstractMoving obstacles have potentially higher risks of collision than stationary obstacles in traffic. Therefore, it is meaningful to detect moving obstacles by sensors equipped on a car for drive assistance applications. We propose two algorithms to detect moving obstacles using camera(s) depending on the relative motion between cameras and obstacles. Since camera(s) moves along with the subjective car, it is challenging to find actually moving obstacles in image sequences. The first algorithm identifies moving obstacle regions in images by checking conflicts between image motion and epipolar constraint when obstacles move in different direction from cameras' motion. The second algorithm identifies moving obstacle regions in images by finding disparity differences between stereo and motion especially when obstacles move in same direction as cameras' motion. Experiments show not only qualitative performance of our detection algorithm, but also quantitative accuracy of egomotion and optical flow estimation in our algorithm. Morimichi Nishigaki, Yiannis Aloimonos |
Intelligent Vehicles Symposium | 2 |
| 2009 | Active segmentation with fixationabstractThe human visual system observes and understands a scene/image by making a series of fixations. Every “fixation point” lies inside a particular region of arbitrary shape and size in the scene which can either be an object or just a part of it. We define as a basic segmentation problem the task of segmenting that region containing the “fixation point”. Segmenting this region is equivalent to finding the enclosing contour - a connected set of boundary edge fragments in the edge map of the scene - around the fixation. We present here a novel algorithm that finds this bounding contour and achieves the segmentation of one object, given the fixation. The proposed segmentation framework combines monocular cues (color/intensity/texture) with stereo and/or motion, in a cue independent manner. We evaluate the performance of the proposed algorithm on challenging videos and stereo pairs. Although the proposed algorithm is more suitable for an active observer capable of fixating at different locations in the scene, it applies to a single image as well. In fact, we show that even with monocular cues alone, the introduced algorithm performs as well or better than a number of image segmentation algorithms, when applied to challenging inputs. Ajay K. Mishra, Yiannis Aloimonos, Loong Fah Cheong |
ICCV | 2 |
| 2009 | Real-time shape retrieval for robotics using skip Tri-GramsabstractThe real time requirement is an additional constraint on many intelligent applications in robotics, such as shape recognition and retrieval using a mobile robot platform. In this paper, we present a scalable approach for efficiently retrieving closed contour shapes. The contour of an object is represented by piecewise linear segments. A skip Tri-Gram is obtained by selecting three segments in the clockwise order while allowing a constant number of segments to be ¿skipped¿ in between. The main idea is to use skip Tri-Grams of the segments to implicitly encode the distant dependency of the shape. All skip Tri-Grams are used for efficiently retrieving closed contour shapes without pairwise matching feature points from two shapes. The retrieval is at least an order of magnitude faster than other state-of-the-art algorithms. We score 80% in the Bullseye retrieval test on the whole MPEG 7 shape dataset. We further test the algorithm using a mobile robot platform in an indoor environment. 8 objects are used for testing from different viewing directions, and we achieve 82% accuracy. Konstantinos Bitsakos, Cornelia Fermüller, Yiannis Aloimonos |
IROS | 4 |
| 2009 | Active segmentation for roboticsabstractThe semantic robots of the immediate future are robots that will be able to find and recognize objects in any environment. They need the capability of segmenting objects in their visual field. In this paper, we propose a novel approach to segmentation based on the operation of fixation by an active observer. Our approach is different from current approaches: while existing works attempt to segment the whole scene at once into many areas, we segment only one image region, specifically the one containing the fixation point. Furthermore, our solution integrates monocular cues (color, texture) with binocular cues (stereo disparities and optical flow). Experiments with real imagery collected by our active robot and from the known databases demonstrate the promise of the approach. Ajay K. Mishra, Yiannis Aloimonos, Cornelia Fermüller |
IROS | 2 |
| 2009 | Image Transformations and BlurringabstractSince cameras blur the incoming light during measurement, different images of the same surface do not contain the same information about that surface. Thus, in general, corresponding points in multiple views of a scene have different image intensities. While multiple-view geometry constrains the locations of corresponding points, it does not give relationships between the signals at corresponding locations. This paper offers an elementary treatment of these relationships. We first develop the notion of "ideal" and "real" images, corresponding to, respectively, the raw incoming light and the measured signal. This framework separates the filtering and geometric aspects of imaging. We then consider how to synthesize one view of a surface from another; if the transformation between the two views is affine, it emerges that this is possible if and only if the singular values of the affine matrix are positive. Next, we consider how to combine the information in several views of a surface into a single output image. By developing a new tool called "frequency segmentation," we show how this can be done despite not knowing the blurring kernel. Justin Domke, Yiannis Aloimonos |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2008 | Who killed the directed model?abstractPrior distributions are useful for robust low-level vision, and undirected models (e.g. Markov Random Fields) have become a central tool for this purpose. Though sometimes these priors can be specified by hand, this becomes difficult in large models, which has motivated learning these models from data. However, maximum likelihood learning of undirected models is extremely difficult- essentially all known methods require approximations and/or high computational cost. Conversely, directed models are essentially trivial to learn from data, but have not received much attention for low-level vision. We compare the two formalisms of directed and undirected models, and conclude that there is no a priori reason to believe one better represents low-level vision quantities. We formulate two simple directed priors, for natural images and stereo disparity, to empirically test if the undirected formalism is superior. We find in both cases that a simple directed model can achieve results similar to the best learnt undirected models with significant speedups in training time, suggesting that directed models are an attractive choice for tractable learning. Justin Domke, Alap Karapurkar, Yiannis Aloimonos |
CVPR | 3 |
| 2008 | Measuring 1st order stretchwith a single filterabstractWe analytically develop a filter that is able to measure the linear stretch of the transformation around a point, and present results of applying it to real signals. We show that this method is a real-time alternative solution for measuring local signal transformations. Experimentally, this method can accurately measure stretch, however, it is sensitive to shift. Konstantinos Bitsakos, Justin Domke, Cornelia Fermüller, Yiannis Aloimonos |
ICASSP | 4 |
| 2007 | Multiple View Image Reconstruction: A Harmonic ApproachabstractThis paper presents a new constraint connecting the signals in multiple views of a surface. The constraint arises from a harmonic analysis of the geometry of the imaging process and it gives rise to a new technique for multiple view image reconstruction. Given several views of a surface from different positions, fundamentally different information is present in each image, owing to the fact that cameras measure the incoming light only after the application of a low-pass filter. Our analysis shows how the geometry of the imaging is connected to this filtering. This leads to a technique for constructing a single output image containing all the information present in the input images. Justin Domke, Yiannis Aloimonos |
CVPR | 2 |
| 2007 | Signals on Pencils of LinesabstractThis paper proposes the "epipolar pencil transformation" (EPT). This is a tool for comparing the signals in different images, with no use of feature detection, yet taking advantage of the constraints given by epipolar geometry. The idea is to develop a descriptor for each point, summarizing the signals on the pencil of lines intersecting at that point. To compute the EPT, first find compact descriptors for each line, then combine these appropriately for each pencil. Given the EPT for two images, computing the epipolar geometry reduces to a closest pairs problem- select one pencil from each set such that the L1distance (in descriptor space) is minimized. By this reduction to a high dimensional closest pairs problem, recent advances in computational geometry can be used to efficiently identify the best global solution. This technique is robust, as each potential solution is evaluated by comparing the signals for all the lines passing through the two hypothesized epipoles. At the same time, as the closest pairs algorithm performs a global search, the solution is not distracted by local minima. The EPT is used here both for the problem of two-view rigid motion, and many-view place recognition. Justin Domke, Yiannis Aloimonos |
ICCV | 2 |
| 2007 | A Roadmap to the Integration of Early Visual Modules
Abhijit S. Ogale, Yiannis Aloimonos |
Int. J. Comput. Vis. | 2 |
| 2006 | Deformation and Viewpoint Invariant Color HistogramsabstractWe develop a theoretical basis for creating color histograms that are in-variant under deformation or changes in viewpoint. The gradients in differ-ent color channels weight the influence of a pixel on the histogram so as to cancel out the changes induced by deformations. Experiments show these histograms to be invariant under a variety of distortions and changes in view-point. 1 Justin Domke, Yiannis Aloimonos |
BMVC | 2 |
| 2006 | Integration of Visual and Inertial Information for Egomotion: a Stochastic ApproachabstractWe present a probabilistic framework for visual correspondence, inertial measurements and egomotion. First, we describe a simple method based on Gabor filters to produce correspondence probability distributions. Next, we generate a noise model for inertial measurements. Probability distributions over the motions are then computed directly from the correspondence distributions and the inertial measurements. We investigate combining the inertial and visual information for a single distribution over the motions. We find that with smaller amounts of correspondence information, fusion of the visual data with the inertial sensor results in much better egomotion estimation. This is essentially because inertial measurements decrease the "translation-rotation" ambiguity. However, when more correspondence information is used, this ambiguity is reduced to such a degree that the inertial measurements provide negligible improvement in accuracy. This suggests that inertial and visual information are more closely integrated in a compositional sense Justin Domke, Yiannis Aloimonos |
ICRA | 2 |
| 2006 | A sensory grammar for inferring behaviors in sensor networksabstractThe ability of a sensor network to parse out observable activities into a set of distinguishable actions is a powerful feature that can potentially enable many applications of sensor networks to everyday life situations. In this paper we introduce a framework that uses a hierarchy of Probabilistic Context Free Grammars (PCFGs) to perform such parsing. The power of the framework comes from the hierarchical organization of grammars that allows the use of simple local sensor measurements for reasoning about more macroscopic behaviors. Our presentation describes how to use a set of phonemes to construct grammars and how to achieve distributed operation using a messaging model. The proposed framework is flexible. It can be mapped to a network hierarchy or can be applied sequentially and across the network to infer behaviors as they unfold in space and time. We demonstrate this functionality by inferring simple motion patterns using a sequence of simple direction vectors obtained from our camera sensor network testbed. Dimitrios Lymberopoulos, Abhijit S. Ogale, Andreas Savvides, Yiannis Aloimonos |
IPSN | 4 |
| 2006 | Understanding visuo-motor primitives for motion synthesis and analysisabstractAbstract The problem addressed in this paper concerns the representation of human movement in terms of atomic visuo‐motor primitives considering both generation and perception of movement. We introduce the concept of kinetology, the phonology of human movement, and five principles on which such a system should be based: compactness, view‐invariance, reproducibility, selectivity, and reconstructivity. We propose visuo‐motor primitives and demonstrate their kinetological properties. Further evaluation is accomplished with experiments on compression and decompression. Our long‐term goal is to demonstrate that action has a space characterized by a visuo‐motor language. Copyright © 2006 John Wiley & Sons, Ltd. Gutemberg Guerra-Filho, Yiannis Aloimonos |
Comput. Animat. Virtual Worlds | 2 |
| 2005 | Robust Contrast Invariant Stereo CorrespondenceabstractA stereo pair of cameras attached to a robot will inevitably yield images with different contrast. Even if we assume that the camera hardware is identical, due to slightly different points of view, the amount of light entering the two cameras is also different, causing dynamically adjusted internal parameters such as aperture, exposure and gain to be different. Due to the difficulty of obtaining and maintaining precise intensity or color calibration between the two cameras, contrast invariance becomes an extremely desirable property of stereo correspondence algorithms. The problem of achieving point correspondence between a stereo pair of images is often addressed by using the intensity or color differences as a local matching metric, which is sensitive to contrast changes. We present an algorithm for contrast invariant stereo matching which relies on multiple spatial frequency channels for local matching. A fast global framework uses the local matching to compute the correspondences and find the occlusions. We demonstrate that the use of multiple frequency channels allows the algorithm to yield good results even in the presence of significant amounts of noise. Abhijit S. Ogale, Yiannis Aloimonos |
ICRA | 2 |
| 2005 | Shape and the Stereo Correspondence Problem
Abhijit S. Ogale, Yiannis Aloimonos |
Int. J. Comput. Vis. | 2 |
| 2005 | Motion Segmentation Using OcclusionsabstractWe examine the key role of occlusions in finding independently moving objects instantaneously in a video obtained by a moving camera with a restricted field of view. In this problem, the image motion is caused by the combined effect of camera motion (egomotion), structure (depth), and the independent motion of scene entities. For a camera with a restricted field of view undergoing a small motion between frames, there exists, in general, a set of 3D camera motions compatible with the observed flow field even if only a small amount of noise is present, leading to ambiguous 3D motion estimates. If separable sets of solutions exist, motion-based clustering can detect one category of moving objects. Even if a single inseparable set of solutions is found, we show that occlusion information can be used to find ordinal depth, which is critical in identifying a new class of moving objects. In order to find ordinal depth, occlusions must not only be known, but they must also be filled (grouped) with optical flow from neighboring regions. We present a novel algorithm for filling occlusions and deducing ordinal depth under general circumstances. Finally, we describe another category of moving objects which is detected using cardinal comparisons between structure from motion and structure estimates from another source (e.g., stereo). Abhijit S. Ogale, Cornelia Fermüller, Yiannis Aloimonos |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2004 | Stereo Correspondence with Slanted Surfaces: Critical Implications of Horizontal Slant
Abhijit S. Ogale, Yiannis Aloimonos |
CVPR (1) | 2 |
| 2004 | Structure from Motion of Parallel Lines
Patrick Baker, Yiannis Aloimonos |
ECCV (4) | 2 |
| 2004 | Compound eye sensor for 3D ego motion estimationabstractWe describe a compound eye vision sensor for 3D ego motion computation. Inspired by eyes of insects, we show that the compound eye sampling geometry is optimal for 3D camera motion estimation. This optimality allows us to estimate the 3D camera motion in a scene-independent and robust manner by utilizing linear equations. The mathematical model of the new sensor can be implemented in analog networks resulting in a compact computational sensor for instantaneous 3D ego motion measurements in full six degrees of freedom. Jan Neumann, Cornelia Fermüller, Yiannis Aloimonos, Vladimir Brajovic |
IROS | 3 |
| 2004 | A hierarchy of cameras for 3D photography
Jan Neumann, Cornelia Fermüller, Yiannis Aloimonos |
Comput. Vis. Image Underst. | 3 |
| 2003 | Polydioptric Camera Design and 3D Motion EstimationabstractMost cameras used in computer vision applications are still based on the pinhole principle inspired by our own eyes. It has been found though that this is not necessarily the optimal image formation principle for processing visual information using a machine. We describe how to find the optimal camera for 3D motion estimation by analyzing the structure of the space formed by the light rays passing through a volume of space. Every camera corresponds to a sampling pattern in light ray space, thus the question of camera design can be rephrased as finding the optimal sampling pattern with regard to a given task. This framework suggests that large field-of-view multi-perspective (polydioptric) cameras are the optimal image sensors for 3D motion estimation. We conclude by proposing design principles for polydioptric cameras and describe an algorithm for such a camera that estimates its 3D motion in a scene independent and robust manner. Jan Neumann, Cornelia Fermüller, Yiannis Aloimonos |
CVPR (2) | 3 |
| 2003 | Eye Design in the Plenoptic Space of Light RaysabstractNatural eye designs are optimized with regard to the tasks the eye-carrying organism has to perform for survival. This optimization has been performed by the process of natural evolution over many millions of years. Every eye captures a subset of the space of light rays. The information contained in this subset and the accuracy to which the eye can extract the necessary information determines an upper limit on how well an organism can perform a given task. In this work we propose a new methodology for camera design. By interpreting eyes as sample patterns in light ray space we can phrase the problem of eye design in a signal processing framework. This allows us to develop mathematical criteria for optimal eye design, which in turn enables us to build the best eye for a given task without the trial and error phase of natural evolution. The principle is evaluated on the task of 3D ego-motion estimation. Jan Neumann, Cornelia Fermüller, Yiannis Aloimonos |
ICCV | 3 |
| 2003 | New eyes for roboticsabstractThis paper describes an imaging system that has been designed to facilitate robotic tasks of motion. The system consists of a number of cameras in a network arranged so that they sample different parts of the visual sphere. This geometric configuration has provable advantages compared to small field of view cameras for the estimation of the system's own motion and consequently the estimation of shape models from the individual cameras. The reason is that inherent ambiguities of confusion between translation and rotation disappear. Pairs of cameras may also be arranged in multiple stereo configurations which provide additional advantages for segmentation. Algorithms for the calibration of the system and the 3D motion estimation are provided. Patrick Baker, Abhijit S. Ogale, Cornelia Fermüller, Yiannis Aloimonos |
IROS | 4 |
| 2003 | Computational video
Yiannis Aloimonos |
Vis. Comput. | 1 |
| 2002 | Harmonic Computational Geometry: A new tool for visual correspondence
Yiannis Aloimonos |
BMVC | 1 |
| 2002 | Spatio-Temporal Stereo Using Multi-Resolution Subdivision Surfaces
Jan Neumann, Yiannis Aloimonos |
Int. J. Comput. Vis. | 2 |
| 2002 | Visual space-time geometry - A tool for perception and the imaginationabstractAlthough the fundamental ideas underlying research efforts in the field of computer vision have not radically changed in the past two decades, there has been a transformation in the way work in this field is conducted. This is primarily due to the emergence of a number of tools, of both a practical and a theoretical nature. One such tool, celebrated throughout the nineties, is the geometry of visual space-time. It is known under a variety of headings, such as multiple view geometry, structure from motion, and model building. It is a mathematical theory relating multiple views (images) of a scene taken at different viewpoints to three-dimensional models of the (possibly dynamic) scene. This mathematical theory gave rise to algorithms that take as input images (or video) and provide as output a model of the scene. Such algorithms are one of the biggest successes of the field and they have many applications in other disciplines, such as graphics (image-based rendering, motion capture) and robotics (navigation). One of the difficulties, however is that the current tools cannot yet be fully automated, and they do not provide very accurate results. More research is required for automation and high precision. During the past few years we have investigated a number of basic questions underlying the structure from motion problem. Our investigations resulted in a small number of principles that characterize the problem. These principles, which give rise to automatic procedures and point to new avenues for studying the next level of the structure from motion problem, are the subject of this paper. Cornelia Fermüller, Patrick Baker, Yiannis Aloimonos |
Proc. IEEE | 3 |
| 2001 | A Spherical Eye from Multiple Cameras (Makes Better Models of the World)abstractThe paper describes an imaging system that has been designed specifically for the purpose of recovering egomotion and structure from video. The system consists of six cameras in a network arranged so that they sample different parts of the visual sphere. This geometric configuration has provable advantages compared to small field of view cameras for the estimation of the system's own motion and consequently the estimation of shape models from the individual cameras. The reason is that inherent ambiguities of confusion between translation and rotation disappear. We provide algorithms for the calibration of the system and 3D motion estimation. The calibration is based on a new geometric constraint that relates the images of lines parallel in space to the rotation between the cameras. The 3D motion estimation uses a constraint relating structure directly to image gradients. Patrick Baker, Cornelia Fermüller, Yiannis Aloimonos, Robert Pless |
CVPR (1) | 3 |
| 2001 | The Statistics of Optical Flow
Cornelia Fermüller, David Shulman, Yiannis Aloimonos |
Comput. Vis. Image Underst. | 3 |
| 2000 | The Statistics of Optical Flow: Implications for the Process of Correspondence in VisionabstractThis paper studies the three major categories of flow estimation methods: gradient-based, energy-based, and correlation methods; it analyzes different ways of compounding 1D motion estimates (image gradients, spatio-temporal frequency triplets, local correlation estimates) into 2D velocity estimates, including linear and nonlinear methods. Correcting for the bias would require knowledge of the noise parameters. In many situations, however, these are difficult to estimate accurately, as they change with the dynamic imagery in unpredictable and complex ways. Thus, the bias really is a problem inherent to optical flow estimation. We argue that the bias is also integral to the human visual system. It is the cause of the illusory perception of motion in the Ouchi pattern and also explains various psychophysical studies of the perception of moving plaids. Finally, the implication of the analysis is that flow or correspondence can be estimated very accurately only when feedback is utilized. Cornelia Fermüller, Yiannis Aloimonos |
ICPR | 2 |
| 2000 | New eyes for building models from video
Cornelia Fermüller, Yiannis Aloimonos, Tomás Brodský |
Comput. Geom. | 2 |
| 2000 | Structure from Motion: Beyond the Epipolar Constraint
Tomás Brodský, Cornelia Fermüller, Yiannis Aloimonos |
Int. J. Comput. Vis. | 3 |
| 2000 | Observability of 3D Motion
Cornelia Fermüller, Yiannis Aloimonos |
Int. J. Comput. Vis. | 2 |
| 2000 | Detecting Independent Motion: The Statistics of Temporal ContinuityabstractWe consider a problem central in aerial visual surveillance applications; detection and tracking of small, independently moving objects in long and noisy video sequences. We directly use spatiotemporal image intensity gradient measurements to compute an exact model of background motion. This allows the creation of accurate mosaics over many frames, and the definition of a constraint violation function which acts as an indicator of independent motion. A novel temporal integration method maintains confidence measures over long subsequences without computing the optic flow, requiring object models, or using a Kalman filter. The mosaic acts as a stable feature frame, allowing precise localization of the independently moving objects. We present a statistical analysis of the effects of image noise on the constraint violation measure and find a good match between the predicted probability distribution function and the measured sample frequencies in a test sequence. Robert Pless, Tomás Brodský, Yiannis Aloimonos |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 1999 | Shape from VideoabstractThis paper presents a novel technique for recovering the shape of a static scene from a video sequence due to a rigidly moving camera. The solution procedure consists of two stages. In the first stage, the rigid motion of the camera at each instant in time is recovered. This provides the transformation between successive viewing positions. The solution is achieved through new constraints which relate 3D motion and shape directly to the image derivatives. These constraints allow to combine the processes of 3D motion estimation and segmentation by exploiting the geometry and statistics inherent in the data. In the second stage the scene surfaces are reconstructed through an optimization procedure which utilizes data from all the frames of the video sequence. A number of experimental results demonstrate the potential of the approach. Tomás Brodský, Cornelia Fermüller, Yiannis Aloimonos |
CVPR | 3 |
| 1999 | Motion Segmentation: A Synergistic ApproachabstractSince estimation of camera motion requires knowledge of independent motion, and moving object detection and localization requires knowledge about the camera motion, the two problems of motion estimation and segmentation need to be solved together in a synergistic manner. This paper provides an approach to treating both these problems simultaneously. The technique introduced here is based on a novel concept, "scene ruggedness" which parameterizes the variation in estimated scene depth with the error in the underlying three-dimensional (3D) motion. The idea is that incorrect 3D motion estimates cause distortions in the estimated depth map, and as a result smooth scene patches are computed as rugged surfaces. The correct 3D motion can be distinguished, as it does not cause any distortion and thus gives rise to the background patches with the least depth variation between depth discontinuities, with the locations corresponding to independent motion being rugged. The algorithm presented employs a binocular observer whose nature is exploited in the extraction of depth discontinuities, a step that facilitates the overall procedure, but the technique can be extended to a monocular observer in a variety of ways. Cornelia Fermüller, Tomás Brodský, Yiannis Aloimonos |
CVPR | 3 |
| 1999 | Statistical Biases in Optic FlowabstractThe computation of optical flow from image derivatives is biased in regions of non uniform gradient distributions. A least-squares or total least squares approach to computing optic flow from image derivatives even in regions of consistent flow can lead to a systematic bias dependent upon the direction of the optic flow, the distribution of the gradient directions, and the distribution of the image noise. The bias a consistent underestimation of length and a directional error. Similar results hold for various methods of computing optical flow in the spatiotemporal frequency domain. The predicted bias in the optical flow is consistent with psychophysical evidence of human judgment of the velocity of moving plaids, and provides an explanation of the Ouchi illusion. Correction of the bias requires accurate estimates of the noise distribution; the failure of the human visual system to make these corrections illustrates both the difficulty of the task and the feasibility of using this distorted optic flow or undistorted normal flow in tasks requiring higher lever processing. Cornelia Fermüller, Robert Pless, Yiannis Aloimonos |
CVPR | 3 |
| 1999 | Independent Motion: The Importance of HistoryabstractWe consider a problem central in aerial visual surveillance applications-detection and tracking of small, independently moving objects in long and noisy video sequences. We directly use spatiotemporal image intensity gradient measurements to compute an exact model of background motion. This allows the creation of accurate mosaics over many frames and the definition of a constraint violation function which acts as an indication of independent motion. A novel temporal integration method maintains confidence measures over long subsequences without computing the optic flow, requiring object models, or using a Kalman filler. The mosaic acts as a stable feature frame, allowing precise localization of the independently moving objects. We present a statistical analysis of the effects of image noise on the constraint violation measure and find a good match between the predicted probability distribution function and the measured sample frequencies in a test sequence. Robert Pless, Tomás Brodský, Yiannis Aloimonos |
CVPR | 3 |
| 1998 | Toward Motion Picture Grammars
Ruud M. Bolle, Yiannis Aloimonos, Cornelia Fermüller |
ACCV (2) | 2 |
| 1998 | Changes in Surface Convexity and Topology Caused by Distortions of Stereoscopic Visual Space
Gregory Baratoff, Yiannis Aloimonos |
ECCV (2) | 2 |
| 1998 | Simultaneous Estimation of Viewing Geometry and Structure
Tomás Brodský, Cornelia Fermüller, Yiannis Aloimonos |
ECCV (1) | 3 |
| 1998 | What Is Computed by Structure from Motion Algorithms?
Cornelia Fermüller, Yiannis Aloimonos |
ECCV (1) | 2 |
| 1998 | Self-Calibration from Image DerivativesabstractThis study investigates the problem of estimating the calibration parameters from image motion fields induced by a rigidly moving camera with unknown calibration parameters, where the image formation is modeled with a linear pinhole-camera model. The equations obtained show the flow to be clearly separated into a component due to the translation and the calibration parameters and a component due to the rotation and the calibration parameters. A set of parameters encoding the latter component are linearly related to the flow, and from these parameters the calibration can be determined. However, as for discrete motion, in the general case it is not possible, to decouple image measurements from two frames only into their translational and rotational component. Geometrically, the ambiguity takes the form of a part of the rotational component being parallel to the translational component, and thus the scene can be reconstructed only up to a projective transformation. In general, for a full calibration at least four successive image frames are necessary with the 3D-rotation changing between the measurements. The geometric analysis gives rise to a direct self-calibration method that avoids computation of optical flow or point correspondences and uses only normal flow measurements. In this technique the direction of translation is estimated employing in a novel way smoothness constraints. Then the calibration parameters are estimated from the rotational components of several flow fields using Levenberg-Marquardt parameter estimation, iterative in the calibration parameters only. The technique proposed does not require calibration objects in the scene or special camera motions and it also avoids the computation of exact correspondence. This makes it suitable for the calibration of active vision systems which have to acquire knowledge about their intrinsic parameters while they perform other tasks, or as a tool for analyzing image sequences in large video databases. Tomás Brodský, Cornelia Fermüller, Yiannis Aloimonos |
ICCV | 3 |
| 1998 | Which Shape from Motion?abstractIn a practical situation, the rigid transformation relating different views is recovered with errors. In such a case, the recovered depth of the scene contains errors, and consequently a distorted version of visual space is computed. What then are meaningful shape representations that can be computed from the images? The result presented in this paper states that if the rigid transformation between different views is estimated in a way that gives rise to a minimum number of negative depth values, then at the center of the image affine shape can be correctly computed. This result is obtained by exploiting properties of the distortion function. The distortion model turns out to be a very powerful tool in the analysis and design of 3D motion and shape estimation algorithms, and as a byproduct of our analysis we present a computational explanation of psychophysical results demonstrating human visual space distortion from motion information. Cornelia Fermüller, Yiannis Aloimonos |
ICCV | 2 |
| 1998 | Effects of Errors in the Viewing Geometry on Shape Estimation
Loong Fah Cheong, Cornelia Fermüller, Yiannis Aloimonos |
Comput. Vis. Image Underst. | 3 |
| 1998 | Directions of Motion Fields are Hardly Ever Ambiguous
Tomás Brodský, Cornelia Fermüller, Yiannis Aloimonos |
Int. J. Comput. Vis. | 3 |
| 1998 | Ambiguity in Structure from Motion: Sphere versus Plane
Cornelia Fermüller, Yiannis Aloimonos |
Int. J. Comput. Vis. | 2 |
| 1997 | The confounding of translation and rotation in reconstruction from multiple viewsabstractIf 3D rigid motion is estimated with some error a distorted version of the scene structure will in turn be computed. Of computational interest are these regions in space where the distortions are such that the depths become negative, because in order to be visible the scene has to lie in front of the image. The stability analysis for the structure-from-motion problem presented in this paper investigates the optimal relationship between the errors in the estimated translational and rotational parameters of a rigid motion, that results in the estimation of a minimum number of negative depth values. The input used is the value of the flow along some direction, which is more general than optic flow or correspondence. For a planar retina it is shown that the optimal configuration is achieved when the projections of the translational and rotational errors on the image plane are perpendicular. Furthermore, the projection of the actual and the estimated translation lie on a line passing through the image center. For a spherical retina given a rotational error, the optimal translation is the correct one, while given a translational error. The optimal rotational error is normal to the translational one at an equal distance from the real and estimated translations. The proofs, besides illuminating the confounding of translation and rotation in structure from motion, have an important application to ecological optics, explaining differences of planar and spherical eye or camera designs in motion and shape estimation. Cornelia Fermüller, Yiannis Aloimonos |
CVPR | 2 |
| 1997 | On the Geometry of Visual Correspondence
Cornelia Fermüller, Yiannis Aloimonos |
Int. J. Comput. Vis. | 2 |
| 1996 | Directions of Motion Fields are Hardly Ever Ambiguous
Tomás Brodský, Cornelia Fermüller, Yiannis Aloimonos |
ECCV (2) | 3 |
| 1996 | Spatiotemporal Representations for Visual Navigation
Loong Fah Cheong, Cornelia Fermüller, Yiannis Aloimonos |
ECCV (1) | 3 |
| 1996 | Early detection of independent motion from active control of normal image flow patternsabstractAn important initial step in interpreting a dynamic scene is to detect moving objects in the environment. This paper presents a novel solution to the problem of early motion detection by a moving observer. The solution requires the observer to be active in the acquisition of images thereby controlling the optical flow pattern due to egomotion. A theoretical analysis is done based on geometric considerations to establish conditions that are necessary and sufficient to guarantee motion detection at a point. The detection problem is posed in terms of locally computable image quantities (the normal image flow) which this makes it implementable in real time. The performance of the technique can be improved by imposing any applicable constraint; this is demonstrated for the detection of the motions of "compact" objects satisfying a size bound. The goal is to design a flexible and efficient early motion detection strategy that can be tailored to the needs of a particular navigation system. Rajeev Sharma, Yiannis Aloimonos |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 1995 | Video Representations
Ruud M. Bolle, Yiannis Aloimonos, Cornelia Fermüller |
ACCV | 2 |
| 1995 | Global Rigidity Constraints in Image Displacement FieldsabstractImage displacement fields-optical flow fields, stereo disparity fields, normal flow fields-due to rigid motion possess a global geometric structure which is independent of the scene in view. Motion vectors of certain lengths and directions are constrained to lie on the imaging surface at particular loci whose location and form depends solely on the 3D motion parameters. If optical flow fields or stereo disparity fields are considered, then equal vectors are shown to lie on conic sections. Similarly, for normal motion fields, equal vectors lie within regions whose boundaries also constitute conics. By studying various properties of these curves and regions and their relationships, a characterization of the structure of rigid motion fields is given. The goal of this paper is to introduce a concept underlying the global structure of image displacement fields. This concept gives rise to various constraints that could form the basis of algorithms for the recovery of visual information from multiple views.> Cornelia Fermüller, Yiannis Aloimonos |
ICCV | 2 |
| 1995 | Representations for Active Vision
Cornelia Fermüller, Yiannis Aloimonos |
IJCAI | 2 |
| 1995 | Guest editorial: Qualitative vision
Yiannis Aloimonos |
Int. J. Comput. Vis. | 1 |
| 1995 | Qualitative egomotion
Cornelia Fermüller, Yiannis Aloimonos |
Int. J. Comput. Vis. | 2 |
| 1995 | Vision and action
Cornelia Fermüller, Yiannis Aloimonos |
Image Vis. Comput. | 2 |
| 1995 | Direct motion stereo for passive navigationabstractWe address the problem of motion recovery for a head-eye system from stereo image sequences. Two types of motions, the translation of the vehicle and the panning motion of the head, are considered. We show how these motions and the depth map of the scene can be estimated directly from the measurements of image gradients and time derivatives in a sequence of stereo images. There is no need to estimate image motion, track a scene feature over time, or establish point correspondences in a stereo image pair. We present the results of various experiments with real scenes. Shahriar Negahdaripour, Brian Y. Hayashi, Yiannis Aloimonos |
IEEE Trans. Robotics Autom. | 3 |
| 1994 | Estimating the heading direction using normal flow
Yiannis Aloimonos, Zoran Duric |
Int. J. Comput. Vis. | 1 |
| 1994 | How normal flow constrains relative depth for an active observer
Liuqing Huang, Yiannis Aloimonos |
Image Vis. Comput. | 2 |
| 1993 | Action Representation and Purpose: Re-evaluating the Foundations of Computational Vision
Michael J. Black, Yiannis Aloimonos, Ian Horswill, Jitendra Malik, Giulio Sandini, Michael J. Tarr |
IJCAI | 2 |
| 1993 | Recognizing 3-D Motion
Cornelia Fermüller, Yiannis Aloimonos |
IJCAI | 2 |
| 1993 | The role of fixation in visual motion analysis
Cornelia Fermüller, Yiannis Aloimonos |
Int. J. Comput. Vis. | 2 |
| 1993 | Report on Workshop on High Performance Computing and Communications for Grand Challenge Applications: Computer Vision, Speech and Natural Language Processing, and Artificial IntelligenceabstractThe findings of a workshop, the goals of which were to identify applications, research problems, and designs of high performance computing and communications (HPCC) systems for supporting applications are discussed. In computer vision, the main scientific issues are machine learning, surface reconstruction, inverse optics and integration, model acquisition, and perception and action. In speech and natural language processing (SNLP), issues were identified statistical analysis in corpus-based speech and language understanding, search strategies for language analysis, auditory and vocal-tract modeling, integration of multiple levels of speech and language analyses, and connectionist systems. In AI, important issues that need immediate attention include the development of efficient machine learning and heuristic search methods that can adapt to different architectural configurations, and the design and construction of scalable and verifiable knowledge bases, active memories, and artificial neural networks.> Benjamin W. Wah, Thomas S. Huang, Aravind K. Joshi, Dan I. Moldovan, Yiannis Aloimonos, Ruzena Bajcsy, Dana H. Ballard, Doug DeGroot, Kenneth A. De Jong, Charles R. Dyer, Scott E. Fahlman, Ralph Grishman, Lynette Hirschman, Richard E. Korf, Stephen E. Levinson, Daniel P. Miranker, N. H. Morgan, Sergei Nirenburg, Tomaso A. Poggio, Edward M. Riseman, Craig Stanfil, Salvatore J. Stolfo, Steven L. Tanimoto, Charles C. Weems |
IEEE Trans. Knowl. Data Eng. | 5 |
| 1993 | Probabilistic analysis of some navigation strategies in a dynamic environmentabstractThe problem of efficient path planning for a point robot in a partially known dynamic environment is considered. The static known part of the environment consists of point shelters distributed in planar terrain, and the dynamic, unknown part is abstracted in the form of alarms that cause the robot to leave its current (preplanned) path and divert to the nearest shelter. We give a probabilistic analysis of the expected times for the dynamic paths generated when the alarms follow a Poisson distribution with parameter lambda . A case study with three shelters serves to illustrate the dependence of the expected travel times on lambda for two alternate static paths. Two different strategies are presented for the general case of n shelters and shown to be superior for different ranges of values of the alarm rate lambda (very low and very high values, respectively). We also discuss some ways of generalizing the approach and possible applications.> Rajeev Sharma, David M. Mount, Yiannis Aloimonos |
IEEE Trans. Syst. Man Cybern. | 3 |
| 1992 | Exploratory active vision: theoryabstractAn active approach to the integration of shape from x modules-here shape from shading and shape from texture-is proposed. The question of what constitutes a good motion for the active observer is addressed. Generally, the role of the visual system is to provide depth information to an autonomous robot; a trajectory module will then interpret it to determine a motion for the robot, which in turn will affect the visual information received. It is suggested that the motion can also be chosen so as to improve the performance of the visual system.> Jean-Yves Hervé, Yiannis Aloimonos |
CVPR | 2 |
| 1992 | The geometry of visual interceptionabstractUnder the traditional paradigm of considering vision as a recovery problem, visual interception is just another application of the structure-from-motion module. However, the inherent difficulties of three-dimensional reconstruction have delayed any real-time applications. The authors offer a robust solution under the active qualitative vision paradigm. From the image intensity function, they obtain the locomotive intrinsics of the agent and the target. Based on this relative information, they present a control strategy that decides in real time whether the velocity of the agent should be increased or decreased at any time instant, thus guiding the agent to intercept the target. The problem of visual interception can thus be solved by simple computation without correspondence.> Liuqing Huang, Yiannis Aloimonos |
CVPR | 2 |
| 1992 | Visual motion analysis under interceptive behaviorabstractThe development of the visual processes that would be needed by a mobile robot system for visually intercepting a moving target is considered. Many relevant active visual processes are proposed that provide robust input for qualitative motion control strategies. The processes for detecting independent motion and for monitoring progress toward the moving target are summarized.> Rajeev Sharma, Yiannis Aloimonos |
CVPR | 2 |
| 1992 | Active Egomotion Estimation: A Qualitative Approach
Yiannis Aloimonos, Zoran Duric |
ECCV | 1 |
| 1992 | Perceptual computational advantages of trackingabstractThe paradigm of active vision advocates studying visual problems in the form of modules that are directly related to a visual task for observers that are active. It is argued that in many cases when an object is moving in an unrestricted manner (translation and rotation) in the 3D world only the motion's translational components are of interest. For a monocular observer, using only the normal flow-the spatiotemporal derivatives of the image intensity function-the authors solve the problem of computing the direction of translation. Their strategy uses fixation and tracking. Fixation simplifies much of the computation by placing the object at the center of the visual field, and the main advantage of tracking is the accumulation of information over time. The authors show how tracking is accomplished using normal flow measurements and use it for two different tasks in the solution process. First, it serves as a tool to compensate for the lack of existence of an optical flow field and thus to estimate the translation parallel to the image plane; and second, it gathers information about the motion component perpendicular to the image plane.> Cornelia Fermüller, Yiannis Aloimonos |
ICPR (1) | 2 |
| 1992 | Object recognition by a robotic agent: the purposive approachabstractStudies the problem of object recognition by considering it in the context of an agent operating in an environment, where the agent's intentions translate into a set of behaviors. In this context, an object can fulfil a function; if the agent recognizes this, it has in effect recognized the object. To perform this type of recognition one needs on one hand a definition of the desired function, and on the other the means of determining whether the object can fulfil that function. To illustrate this approach the authors describe the visual recognition abilities that might be needed by an autonomous cleaning robot.> Ehud Rivlin, Yiannis Aloimonos, Azriel Rosenfeld |
ICPR (1) | 2 |
| 1992 | Purposive, qualitative, active vision
Yiannis Aloimonos |
CVGIP Image Underst. | 1 |
| 1992 | Optimal Visual Motion Estimation: A NoteabstractThe problem of estimating 3D motion in an optimal manner using correspondences of features in two views is analyzed. The importance of having an optimal estimator is twofold: first, for the estimation itself and, second, for the bound it offers on how much sensitivity one can expect from a two-frame, point-based motion algorithm. The optimal estimator turns out to be nonlinear, and for that reason, techniques that provide very good initial guesses for the iterative computation of the optimal estimator are developed.> Minas E. Spetsakis, Yiannis Aloimonos |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1992 | Coordinated motion planning: the warehouseman's problem with constraints on free spaceabstractThe warehouseman's problem, namely, the coordinated motion planning for multiple independent objects confined in a room, is addressed. Several constraints are presented under which the problem becomes tractable, and polynomial time algorithms that guarantee a coordinated rearrangement under the given conditions are obtained. The concept of temporary storage space (TSS) is introduced as a general way of constraining free space. Each algorithm presented has a different set of constraints on the possible sizes and/or relative placements of the square blocks. For each case, an adequate TSS is proposed that guarantees rearrangement of n blocks through algorithms having O(n/sup 2/) running time. The practical utility of the presented techniques is also discussed in the light of the complexity of motion coordination.> Rajeev Sharma, Yiannis Aloimonos |
IEEE Trans. Syst. Man Cybern. | 2 |
| 1991 | A response to "ignorance, myopia, and naiveté in computer vision systems" by R. C. Jain and T. O. Binford
Yiannis Aloimonos, Azriel Rosenfeld |
CVGIP Image Underst. | 1 |
| 1991 | A multi-frame approach to visual motion perception
Minas E. Spetsakis, Yiannis Aloimonos |
Int. J. Comput. Vis. | 2 |
| 1991 | On the visual mathematics of tracking
Yiannis Aloimonos, Dimitris P. Tsakiris |
Image Vis. Comput. | 1 |
| 1990 | Tracking In A Complex Visual Environment
Yiannis Aloimonos, Dimitris P. Tsakiris |
ECCV | 1 |
| 1990 | Purposive and qualitative active visionabstractThe traditional view of the problem of computer vision as a recovery problem is questioned, and the paradigm of purposive-qualitative vision is offered as an alternative. This paradigm considers vision as a general recognition problem (recognition of objects, patterns or situations). To demonstrate the usefulness of the framework, the design of the Medusa of CVL is described. It is noted that this machine can perform complex visual tasks without reconstructing the world. If it is provided with intentions, knowledge of the environment, and planning capabilities, it can perform highly sophisticated navigational tasks. It is explained why the traditional structure from motion problem cannot be solved in some cases and why there is reason to be pessimistic about the optimal performance of a structure from motion module. New directions for future research on this problem in the recovery paradigm, e.g., research on stability or robustness, are suggested.> Yiannis Aloimonos |
ICPR (1) | 1 |
| 1990 | Approximate constrained motion planningabstractThe problem of finding a collision-free path connecting two points (start and goal) in the presence of obstacles, with constraints on the curvature of the path, is examined. This problem of curvature-constrained motion planning arises when, for example, a vehicle with constraints on its steering mechanism needs to be maneuvered through obstacles. Though no lower bound on the difficulty of the problem in 2-D is known, exact algorithms given to date for the reachability questions are exponential. It is shown that a variation of the problem is NP-hard. Notably, however, the same variation to polynomially solvable motion planning problems does not make them intractable. In addition, it is proven that epsilon -approximations to this problem cannot exist unless the underlying decision problem is polynomially solvable. An algorithm which is expected to find a desired path, when one exists, with a required probability is presented. Results indicate that a variable-size discretization is necessary for the task, linking the required probability to the size of the discretization locally.> Anup Basu, Yiannis Aloimonos |
ICRA | 2 |
| 1990 | Structure from motion using line correspondences
Minas E. Spetsakis, Yiannis Aloimonos |
Int. J. Comput. Vis. | 2 |
| 1990 | Perspective approximations
Yiannis Aloimonos |
Image Vis. Comput. | 1 |
| 1990 | Correspondenceless Stereo and Motion: Planar SurfacesabstractIt is shown that a binocular observer can recover the depth and three-dimensional motion of a rigid planar patch without using any correspondences between the left and right image frames (static) or between the successive dynamic frames (dynamic). Uniqueness and robustness issues are studied with respect to this problem and experimental results are given from the application of the theory to real images.> Yiannis Aloimonos, Jean-Yves Hervé |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1989 | Obstacle Avoidance Using Flow Field DivergenceabstractThe use of certain measures of flow field divergence is investigated as a qualitative cue for obstacle avoidance during visual navigation. It is shown that a quantity termed the directional divergence of the 2-D motion field can be used as a reliable indicator of the presence of obstacles in the visual field of an observer undergoing generalized rotational and translational motion. The necessary measurements can be robustly obtained from real image sequences. Experimental results are presented showing that the system responds as expected to divergence in real-world image sequences, and the use of the system to navigate between obstacles is demonstrated.> Randal C. Nelson, Yiannis Aloimonos |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 1988 | Optimal Computing Of Structure From Motion Using Point Correspondences In Two FramesabstractOne of the problems associated with any approach to the structure from motion problem using point correspondence, i.e. recovering the structure of a moving object from its successive images, is the use of least squares on dependent variables. We formulate the problem as a quadratic minimization problem with a non-linear constraint. Then we derive the condition for i,he solution to be optimal under the assumption of Gaussian noise in the input, in the Maximum Likelihood Principle sense. This constraint minimization reduces to the solution of a nonlinear system which in the presence of modest noise is easy to approximate. We present two efficient ways to approximate it and we discuss some inherent limitations of the structure from motion problem when two frames are used that should be taken into account in robotics applications that involve dynamic imagery. In addition, our formulation introduces a framework in which previous works on the subject become special cases. Minas E. Spetsakis, Yiannis Aloimonos |
ICCV | 2 |
| 1988 | Shape from patterns: Regularization
Yiannis Aloimonos, Michael Swain |
Int. J. Comput. Vis. | 1 |
| 1988 | Active vision
Yiannis Aloimonos, Isaac Weiss, Amit Bandyopadhyay |
Int. J. Comput. Vis. | 1 |
| 1988 | Visual shape computationabstractPerceptual processes responsible for computing shape from several cues, including shading, texture, contour, and stereo, are examined. It is noted that these computational problems, as well as that of computing shaping from motion, are ill-posed in the sense of Hadamard. It is suggested that regularization theory can be used along with a priori knowledge to restrict the space of possible solutions, and thus restore the problem's well-prosedness. Some alternative methods are outlined, and the idea of active vision is explored briefly in connection with the problem.> Yiannis Aloimonos |
Proc. IEEE | 1 |
| 1987 | Closed Form Solution to the Structure from Motion Problem from Line Correspondences
Minas E. Spetsakis, Yiannis Aloimonos |
AAAI | 2 |
| 1987 | Determining three dimensional transformation parameters from images: TheoryabstractWe present a theory for the determination of the three dimensional transformation parameters of an object, from its images. The input to this process is the image intensity function and its temporal derivative. In particular, our results are: 1) If the structure of the transforming object in view is known, then the transformation parameters are determined from the solution of a linear system. Rigid motion is a special ease of our theory. 2)If the structure of the object in view is not known, then both the structure and transformation parameters may be computed through a hill climbing or simulated annealing algorithm. E. Ito, Yiannis Aloimonos |
ICRA | 2 |
| 1987 | Combining Sources of Information in Vision I. Computing Shape from Shading and Motion
Yiannis Aloimonos |
IJCAI | 1 |
| 1987 | A Robust Algorithm for Determining the Translation of a Rigidly Moving Surface without Correspondence, for Robotics Applications
Anup Basu, Yiannis Aloimonos |
IJCAI | 2 |
| 1987 | Texture, contour, shape, and motion
Yiannis Aloimonos, Mike Swain, Paul B. Chou, Anup Basu |
Pattern Recognit. Lett. | 2 |
| 1986 | Determining the 3-D Motion of a Rigid Surface Patch Without Correspondence under Perspective Projection
Yiannis Aloimonos, Isidore Rigoutsos |
AAAI | 1 |