Supun Samarasekera

dblp:26/4413 · DBLP profile ↗
← Back
51ranked-venue papers
0as first author
11since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 37 · 8 since 2021Artificial intelligence and machine learning · 28 · 6 since 2021Systems, architecture and hardware · 10 · 3 since 2021Human-computer interaction and ubiquitous computing · 9 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2026 DUDA: Distilled Unsupervised Domain Adaptation for Lightweight Semantic Segmentation
abstract
Unsupervised Domain Adaptation (UDA) is essential for enabling semantic segmentation in new domains without requiring costly pixel-wise annotations. State-of-the-art (SOTA) UDA methods primarily use self-training with architecturally identical teacher and student networks, relying on Exponential Moving Average (EMA) updates. However, these approaches face substantial performance degradation with lightweight models due to inherent architectural inflexibility leading to low-quality pseudo-labels. To address this, we propose Distilled Unsupervised Domain Adaptation (DUDA), a novel framework that combines EMA-based self-training with knowledge distillation (KD). Our method employs an auxiliary student network to bridge the architectural gap between heavyweight and lightweight models for EMA-based updates, resulting in improved pseudo-label quality. DUDA employs a strategic fusion of UDA and KD, incorporating innovative elements such as gradual distillation from large to small networks, inconsistency loss prioritizing poorly adapted classes, and learning with multiple teachers. Extensive experiments across four UDA benchmarks demonstrate DUDA's superiority in achieving SOTA performance with lightweight models, often surpassing the performance of heavyweight models from other approaches.
Beomseok Kang, Niluthpol Chowdhury Mithun, Abhinav Rajvanshi, Han-Pang Chiu, Supun Samarasekera
WACV5
2024 Unsupervised Domain Adaptation for Semantic Segmentation with Pseudo Label Self-Refinement
abstract
Deep learning-based solutions for semantic segmentation suffer from significant performance degradation when tested on data with different characteristics than what was used during the training. Adapting the models using annotated data from the new domain is not always practical. Unsupervised Domain Adaptation (UDA) approaches are crucial in deploying these models in the actual operating conditions. Recent state-of-the-art (SOTA) UDA methods employ a teacher-student self-training approach, where a teacher model is used to generate pseudo-labels for the new data which in turn guide the training process of the student model. Though this approach has seen a lot of success, it suffers from the issue of noisy pseudo-labels being propagated in the training process. To address this issue, we propose an auxiliary pseudo-label refinement network (PRN) for online refining of the pseudo labels and also localizing the pixels whose predicted labels are likely to be noisy. Being able to improve the quality of pseudo labels and select highly reliable ones, PRN helps self-training of segmentation models to be robust against pseudo label noise propagation during different stages of adaptation. We evaluate our approach on benchmark datasets with three different domain shifts, and our approach consistently performs significantly better than the previous state-of-the-art methods.
Xingchen Zhao, Niluthpol Chowdhury Mithun, Abhinav Rajvanshi, Han-Pang Chiu, Supun Samarasekera
WACV5
2023 C-SFDA: A Curriculum Learning Aided Self-Training Framework for Efficient Source Free Domain Adaptation
abstract
Unsupervised domain adaptation (UDA) approaches focus on adapting models trained on a labeled source domain to an unlabeled target domain. In contrast to UDA, source-free domain adaptation (SFDA) is a more practical setup as access to source data is no longer required during adaptation. Recent state-of-the-art (SOTA) methods on SFDA mostly focus on pseudo-label refinement based self-training which generally suffers from two issues: i) inevitable occurrence of noisy pseudo-labels that could lead to early training time memorization, ii) refinement process requires maintaining a memory bank which creates a significant burden in resource constraint scenarios. To address these concerns, we propose C-SFDA, a curriculum learning aided self-training framework for SFDA that adapts efficiently and reliably to changes across domains based on selective pseudo-labeling. Specifically, we employ a curriculum learning scheme to promote learning from a restricted amount of pseudo labels selected based on their reliabilities. This simple yet effective step successfully prevents label noise propagation during different stages of adaptation and eliminates the need for costly memory-bank based label refinement. Our extensive experimental evaluations on both image recognition and semantic segmentation tasks confirm the effectiveness of our method. C-SFDA is also applicable to online test-time domain adaptation and outperforms previous SOTA methods in this task.
Nazmul Karim, Niluthpol Chowdhury Mithun, Abhinav Rajvanshi, Han-Pang Chiu, Supun Samarasekera, Nazanin Rahnavard
CVPR5
2023 Cross-View Visual Geo-Localization for Outdoor Augmented Reality
abstract
Precise estimation of global orientation and location is critical to ensure a compelling outdoor Augmented Reality (AR) experience. We address the problem of geo-pose estimation by cross-view matching of query ground images to a geo-referenced aerial satellite image database. Recently, neural network-based methods have shown state-of-the-art performance in cross-view matching. However, most of the prior works focus only on location estimation, ignoring orientation, which cannot meet the requirements in outdoor AR applications. We propose a new transformer neural network-based model and a modified triplet ranking loss for joint location and orientation estimation. Experiments on several benchmark cross-view geo-localization datasets show that our model achieves state-of-the-art performance. Furthermore, we present an approach to extend the single image query-based geo-localization approach by utilizing temporal information from a navigation pipeline for robust continuous geo-localization. Experimentation on several large-scale real-world video sequences demonstrates that our approach enables high-precision and stable AR insertion.
Niluthpol Chowdhury Mithun, Kshitij Minhas, Han-Pang Chiu, Taragay Oskiper, Mikhail Sizintsev, Supun Samarasekera, Rakesh Kumar 0001
VR6
2022 Semantically-aware Spatio-temporal Reasoning Agent for Vision-and-Language Navigation in Continuous Environments
abstract
This paper presents a novel approach for the Vision-and-Language Navigation (VLN) task in continuous 3D environments, which requires an autonomous agent to follow natural language instructions in unseen environments. Existing end-to-end learning-based VLN methods struggle at this task as they focus mostly on utilizing raw visual observations and lack the semantic spatio-temporal reasoning capabilities which is crucial in generalizing to new environments. In this regard, we present a hybrid transformer-recurrence model which focuses on combining classical semantic mapping techniques with a learning-based method. Our method creates a temporal semantic memory by building a top-down local ego-centric semantic map and performs cross-modal grounding to align map and language modalities to enable effective learning of VLN policy. Empirical results in a photo-realistic long-horizon simulation environment show that the proposed approach outperforms a variety of state-of-the-art methods and baselines with over 22% relative improvement in SPL in prior unseen environments.
Muhammad Zubair Irshad, Niluthpol Chowdhury Mithun, Zachary Seymour, Han-Pang Chiu, Supun Samarasekera, Rakesh Kumar 0001
ICPR5
2022 GraphMapper: Efficient Visual Navigation by Scene Graph Generation
abstract
Understanding the geometric relationships between objects in a scene is a core capability in enabling both humans and autonomous agents to navigate in new environments. A sparse, unified representation of the scene topology will allow agents to act efficiently to move through their environment, communicate the environment state with others, and utilize the representation for diverse downstream tasks. To this end, we propose a method to train an autonomous agent to learn to accumulate a 3D scene graph representation of its environment by simultaneously learning to navigate through said environment. We demonstrate that our approach, GraphMapper, enables the learning of effective navigation policies through fewer interactions with the environment than vision-based systems alone. Further, we show that GraphMapper can act as a modular scene encoder to operate alongside existing Learning-based solutions to not only increase navigational efficiency but also generate intermediate scene representations that are useful for other future tasks.
Zachary Seymour, Niluthpol Chowdhury Mithun, Han-Pang Chiu, Supun Samarasekera, Rakesh Kumar 0001
ICPR4
2022 Towards Safe, Realistic Testbed for Robotic Systems with Human Interaction
abstract
Simulation has been a necessary, safe testbed for robotics systems (RS). However, testing in simulation alone is not enough for robotic systems operating in close proximity, or interacting directly with, humans, because simulated humans are very limited. Furthermore, testing with real humans can be unsafe and costly. As recent advances in machine learning are being brought to physical robotic systems, how to collect data as well as evaluate them with human interactions safely yet realistically is a critical question. This paper presents a Mixed-Reality (MR) system toward human-centered development of robotic systems emphasizing benefits as a data collection and testbed tool. MR testbeds allow humans to interact with various levels of virtuality to maintain both realism and safety. We detail the advantages and limitations of these different levels of realism or virtualization, and report our MR-based RS testbed implemented using off-the-shelf MR devices with the Unity game engine and ROS. We demonstrate our testbed in a multi-robot, multi-person tracking and monitoring application. We share our vision and insights earned during the development and data collection.
Bhoram Lee, Jonathan Brookshire, Rhys Yahata, Supun Samarasekera
ICRA4
2022 Ranging-Aided Ground Robot Navigation Using UWB Nodes at Unknown Locations
abstract
Ranging information from ultra-wideband (UWB) ranging radios can be used to improve estimated navigation accuracy of a ground robot with other on-board sensors. However, all ranging-aided navigation methods demand the locations of ranging nodes to be known, which is not suitable for time-pressed situations, dynamic cluttered environments, or collaborative navigation applications. This paper describes a new ranging-aided navigation approach that does not require the locations of ranging radios. Our approach formulates relative pose constraints using ranging readings. The formulation is based on geometric relationships between each stationary ranging node and two ranging antennas on the moving robot across time. Our experiments show that estimated navigation accuracy of the ground robot is substantially enhanced with ranging information using our approach under a variety of scenarios, when ranging nodes are placed at unknown locations. We analyze and compare our performance with a traditional ranging-aided method, which requires mapping the positions of ranging nodes. We also demonstrate the applicability of our approach for collaborative navigation in large-scale unknown environments, by using ranging information from one mobile robot to improve navigation estimation of the other robot. This application does not require the installation of ranging nodes at fixed locations.
Abhinav Rajvanshi, Han-Pang Chiu, Alex Krasner, Mikhail Sizintsev, Glenn Murray, Supun Samarasekera
IROS6
2022 SIGNAV: Semantically-Informed GPS-Denied Navigation and Mapping in Visually-Degraded Environments
abstract
Understanding the perceived scene during navigation enables intelligent robot behaviors. Current vision-based semantic SLAM (Simultaneous Localization and Mapping) systems provide these capabilities. However, their performance decreases in visually-degraded environments, that are common places for critical robotic applications, such as search and rescue missions. In this paper, we present SIGNAV, a real-time semantic SLAM system to operate in perceptually-challenging situations. To improve the robustness for navigation in dark environments, SIGNAV leverages a multi-sensor navigation architecture to fuse vision with additional sensing modalities, including an inertial measurement unit (IMU), LiDAR, and wheel odometry. A new 2.5D semantic segmentation method is also developed to combine both images and LiDAR depth maps to generate semantic labels of 3D mapped points in real time. We demonstrate that the navigation accuracy from SIGNAV in a variety of indoor environments under both normal lighting and dark conditions. SIGNAV also provides semantic scene understanding capabilities in visually-degraded environments. We also show the benefits of semantic information to SIGNAV’s performance.
Alex Krasner, Mikhail Sizintsev, Abhinav Rajvanshi, Han-Pang Chiu, Niluthpol Chowdhury Mithun, Kevin Kaighn, Phillip Miller, Ryan Villamil, Supun Samarasekera
WACV9
2021 MaAST: Map Attention with Semantic Transformers for Efficient Visual Navigation
abstract
Visual navigation for autonomous agents is a core task in the fields of computer vision and robotics. Learning-based methods, such as deep reinforcement learning, have the potential to outperform the classical solutions developed for this task; however, they come at a significantly increased computational load. Through this work, we design a novel approach that focuses on performing better or comparable to the existing learning-based solutions but under a clear time/computational budget. To this end, we propose a method to encode vital scene semantics such as traversable paths, unexplored areas, and observed scene objects–alongside raw visual streams such as RGB, depth, and semantic segmentation masks—into a semantically informed, top-down egocentric map representation. Further, to enable the effective use of this information, we introduce a novel 2-D map attention mechanism, based on the successful multi-layer Transformer networks. We conduct experiments on 3-D reconstructed indoor PointGoal visual navigation and demonstrate the effectiveness of our approach. We show that by using our novel attention schema and auxiliary rewards to better utilize scene semantics, we outperform multiple baselines trained with only raw inputs or implicit semantic information while operating with an 80% decrease in the agent’s experience.
Zachary Seymour, Kowshik Thopalli, Niluthpol Chowdhury Mithun, Han-Pang Chiu, Supun Samarasekera, Rakesh Kumar 0001
ICRA5
2021 Long-Range Augmented Reality with Dynamic Occlusion Rendering
abstract
Proper occlusion based rendering is very important to achieve realism in all indoor and outdoor Augmented Reality (AR) applications. This paper addresses the problem of fast and accurate dynamic occlusion reasoning by real objects in the scene for large scale outdoor AR applications. Conceptually, proper occlusion reasoning requires an estimate of depth for every point in augmented scene which is technically hard to achieve for outdoor scenarios, especially in the presence of moving objects. We propose a method to detect and automatically infer the depth for real objects in the scene without explicit detailed scene modeling and depth sensing (e.g. without using sensors such as 3D-LiDAR). Specifically, we employ instance segmentation of color image data to detect real dynamic objects in the scene and use either a top-down terrain elevation model or deep learning based monocular depth estimation model to infer their metric distance from the camera for proper occlusion reasoning in real time. The realized solution is implemented in a low latency real-time framework for video-see-though AR and is directly extendable to optical-see-through AR. We minimize latency in depth reasoning and occlusion rendering by doing semantic object tracking and prediction in video frames.
Mikhail Sizintsev, Niluthpol Chowdhury Mithun, Han-Pang Chiu, Supun Samarasekera, Rakesh Kumar 0001
IEEE Trans. Vis. Comput. Graph.4
2020 RGB2LIDAR: Towards Solving Large-Scale Cross-Modal Visual Localization
abstract
We study an important, yet largely unexplored problem of large-scale cross-modal visual localization by matching ground RGB images to a geo-referenced aerial LIDAR 3D point cloud (rendered as depth images). Prior works were demonstrated on small datasets and did not lend themselves to scaling up for large-scale applications. To enable large-scale evaluation, we introduce a new dataset containing over 550K pairs (covering 143 km2 area) of RGB and aerial LIDAR depth images. We propose a novel joint embedding based method that effectively combines the appearance and semantic cues from both modalities to handle drastic cross-modal variations. Experiments on the proposed dataset show that our model achieves a strong result of a median rank of 5 in matching across a large test set of 50K location pairs collected from a 14km^2 area. This represents a significant advancement over prior works in performance and scale. We conclude with qualitative results to highlight the challenging nature of this task and the benefits of the proposed model. Our work provides a foundation for further research in cross-modal visual localization.
Niluthpol Chowdhury Mithun, Karan Sikka, Han-Pang Chiu, Supun Samarasekera, Rakesh Kumar 0001
ACM Multimedia4
2019 Semantically-Aware Attentive Neural Embeddings for 2D Long-Term Visual Localization
Zachary Seymour, Karan Sikka, Han-Pang Chiu, Supun Samarasekera, Rakesh Kumar 0001
BMVC4
2018 Augmented Reality Driving Using Semantic Geo-Registration
abstract
We propose a new approach that utilizes semantic information to register 2D monocular video frames to the world using 3D georeferenced data, for augmented reality driving applications. The geo-registration process uses our predicted vehicle pose to generate a rendered depth map for each frame, allowing 3D graphics to be convincingly blended with the real world view. We also estimate absolute depth values for dynamic objects, up to 120 meters, based on the rendered depth map and update the rendered depth map to reflect scene changes over time. This process also creates opportunistic global heading measurements, which are fused with other sensors, to improve estimates of the 6 degrees-of- freedom global pose of the vehicle over state-of-the-art outdoor augmented reality systems [5]-, [19]. We evaluate the navigation accuracy and depth map quality of our system on a driving vehicle within various large-scale environments for producing realistic augmentations.
Han-Pang Chiu, Varun Murali, Ryan Villamil, G. Drew Kessler, Supun Samarasekera, Rakesh Kumar 0001
VR5
2015 Augmented Reality Scout: Joint Unaided-Eye and Telescopic-Zoom System for Immersive Team Training
abstract
In this paper we present a dual, wide area, collaborative augmented reality (AR) system that consists of standard live view augmentation, e.g., from helmet, and zoomed-in view augmentation, e.g., from binoculars. The proposed advanced scouting capability allows long range high precision augmentation of live unaided and zoomed-in imagery with aerial and terrain based synthetic objects, vehicles, people and effects. The inserted objects must appear stable in the display and not jitter or drift as the user moves around and examines the scene. The AR insertions for the binocs must work instantly when they are picked up anywhere as the user moves around. The design of both AR modules is based on using two different cameras with wide and narrow field of view (FoV) lenses. The wide FoV gives context and enables the recovery of location and orientation of the prop in 6 degrees of freedom (DoF) much more robustly, whereas the narrow FoV is used for the actual augmentation and increased precision in tracking. Furthermore, narrow camera in unaided eye and wide camera on the binoculars are jointly used for global yaw (heading) correction. We present our navigation algorithms using monocular cameras in combination with IMU and GPS in an Extended Kalman Filter (EKF) framework to obtain robust and real-time pose estimation for precise augmentation and cooperative tracking.
Taragay Oskiper, Mikhail Sizintsev, Vlad Branzoi, Supun Samarasekera, Rakesh Kumar 0001
ISMAR4
2015 AR-Weapon: Live Augmented Reality Based First-Person Shooting System
abstract
This paper introduces a user-worn Augmented Reality (AR) based first-person weapon shooting system (AR-Weapon), suitable for both training and gaming. Different from existing AR-based first-person shooting systems, AR-Weapon does not use fiducial markers placed in the scene for tracking. Instead it uses natural scene features observed by the tracking camera from the live view of the world. The AR-Weapon system estimates 6-degrees of freedom orientation and location of the weapon and of the user operating it, thus allowing the weapon to fire simulated projectiles for both direct fire and non-line of sight during live runs. In addition, stereo cameras are used to compute depth and provide dynamic occlusion reasoning. Using the 6-DOF head and weapon tracking, dynamic occlusion reasoning and a terrain model of the environment, the fully virtual projectiles and synthetic avatars are displayed on the user's head mounted Optical-See-Through (OST) display overlaid over the live view of the real world. Since the projectiles, weapon characteristics and virtual enemy combatants are all simulated they can easily be changed to vary scenarios, new projectile types and future weapons. In this paper, we present the technical algorithms, system design and experiment results for a prototype AR-Weapon system.
Vlad Branzoi, Mikhail Sizintsev, Nicholas Vitovitch, Taragay Oskiper, Ryan Villamil, Ali Chaudhry, Supun Samarasekera, Rakesh Kumar 0001
WACV8
2015 Augmented Reality Binoculars
abstract
In this paper we present an augmented reality binocular system to allow long range high precision augmentation of live telescopic imagery with aerial and terrain based synthetic objects, vehicles, people and effects. The inserted objects must appear stable in the display and must not jitter and drift as the user pans around and examines the scene with the binoculars. The design of the system is based on using two different cameras with wide field of view and narrow field of view lenses enclosed in a binocular shaped shell. Using the wide field of view gives us context and enables us to recover the 3D location and orientation of the binoculars much more robustly, whereas the narrow field of view is used for the actual augmentation as well as to increase precision in tracking. We present our navigation algorithm that uses the two cameras in combination with an inertial measurement unit and global positioning system in an extended Kalman filter and provides jitter free, robust and real-time pose estimation for precise augmentation. We have demonstrated successful use of our system as part of information sharing example as well as a live simulated training system for observer training, in which fixed and rotary wing aircrafts, ground vehicles, and weapon effects are combined with real world scenes.
Taragay Oskiper, Mikhail Sizintsev, Vlad Branzoi, Supun Samarasekera, Rakesh Kumar 0001
IEEE Trans. Vis. Comput. Graph.4
2014 Constrained optimal selection for multi-sensor robot navigation using plug-and-play factor graphs
abstract
This paper proposes a real-time navigation approach that is able to integrate many sensor types while fulfilling performance needs and system constraints. Our approach uses a plug-and-play factor graph framework, which extends factor graph formulation to encode sensor measurements with different frequencies, latencies, and noise distributions. It provides a flexible foundation for plug-and-play sensing, and can incorporate new evolving sensors. A novel constrained optimal selection mechanism is presented to identify the optimal subset of active sensors to use, during initialization and when any sensor condition changes. This mechanism constructs candidate subsets of sensors based on heuristic rules and a ternary tree expansion algorithm. It quickly decides the optimal subset among candidates by maximizing observability coverage on state variables, while satisfying resource constraints and accuracy demands. Experimental results demonstrate that our approach selects subsets of sensors to provide satisfactory navigation solutions under various conditions, on large-scale real data sets using many sensors.
Han-Pang Chiu, Xun S. Zhou, Luca Carlone, Frank Dellaert, Supun Samarasekera, Rakesh Kumar 0001
ICRA5
2014 Precise vision-aided aerial navigation
abstract
This paper proposes a novel vision-aided navigation approach that continuously estimates precise 3D absolute pose for aerial vehicles, using only inertial measurements and monocular camera observations. Our approach is able to provide accurate navigation solutions under long-term GPS outage, by tightly incorporating absolute geo-registered information into two kinds of visual measurements: 2D-3D tie-points, and geo-registered feature tracks. 2D-3D tie-points are established by finding feature correspondences to align an aerial video frame to a 2D geo-referenced image rendered from the 3D terrain database. These measurements provide global information to correct accumulated error in navigation estimation. Geo-registered feature tracks are generated by associating features across consecutive frames. They enable the propagation of 3D geo-referenced values to further improve the pose estimation. All sensor measurements are fully optimized in a smoother-based inference framework, which achieves efficient relinearization and real-time estimation of navigation states and their covariances over a constant-length of sliding window. Experimental results demonstrate that our approach provides accurate and consistent aerial navigation solutions on several large-scale GPS-denied scenarios.
Han-Pang Chiu, Aveek Das, Phillip Miller, Supun Samarasekera, Rakesh Kumar 0001
IROS4
2014 Augmented reality binoculars on the move
abstract
In this paper, we expand our previous work on augmented reality (AR) binoculars to support wider range of user motion - up to thousand square meters compared to only a few square meters as before. We present our latest improvements and additions to our pose estimation pipeline and demonstrate stable registration of objects on the real world scenery while the binoculars are undergoing significant amount of parallax-inducing translation.
Taragay Oskiper, Mikhail Sizintsev, Vlad Branzoi, Supun Samarasekera, Rakesh Kumar 0001
ISMAR4
2014 AR-mentor: Augmented reality based mentoring system
abstract
AR-Mentor is a wearable real time Augmented Reality (AR) mentoring system that is configured to assist in maintenance and repair tasks of complex machinery, such as vehicles, appliances, and industrial machinery. The system combines a wearable Optical-See-Through (OST) display device with high precision 6-Degree-Of-Freedom (DOF) pose tracking and a virtual personal assistant (VPA) with natural language, verbal conversational interaction, providing guidance to the user in the form of visual, audio and locational cues. The system is designed to be heads-up and hands-free allowing the user to freely move about the maintenance or training environment and receive globally aligned and context aware visual and audio instructions (animations, symbolic icons, text, multimedia content, speech). The user can interact with the system, ask questions and get clarifications and specific guidance for the task at hand. A pilot application with AR-Mentor was successfully built to instruct a novice to perform an advanced 33-step maintenance task on a training vehicle. The initial live training tests demonstrate that AR-Mentor is able to help and serve as an assistant to an instructor, freeing him/her to cover more students and to focus on higher-order teaching.
Vlad Branzoi, Michael Wolverton, Glenn Murray, Nicholas Vitovitch, Louise Yarnall, Girish Acharya, Supun Samarasekera, Rakesh Kumar 0001
ISMAR8
2013 Robust vision-aided navigation using Sliding-Window Factor graphs
abstract
This paper proposes a navigation algorithm that provides a low-latency solution while estimating the full nonlinear navigation state. Our approach uses Sliding-Window Factor Graphs, which extend existing incremental smoothing methods to operate on the subset of measurements and states that exist inside a sliding time window. We split the estimation into a fast short-term smoother, a slower but fully global smoother, and a shared map of 3D landmarks. A novel three-stage visual feature model is presented that takes advantage of both smoothers to optimize the 3D landmark map, while minimizing the computation required for processing tracked features in the short-term smoother. This three-stage model is formulated based on the maturity of the estimation of the 3D location of the underlying landmark in the map. Long-range associations are used as global measurements from matured landmarks in the short-term smoother and loop closure constraints in the long-term smoother. Experimental results demonstrate our approach provides highly-accurate solutions on large-scale real data sets using multiple sensors in GPS-denied settings.
Han-Pang Chiu, Frank Dellaert, Supun Samarasekera, Rakesh Kumar 0001
ICRA4
2013 Augmented Reality binoculars
abstract
In this paper we present an augmented reality binocular system to allow long range high precision augmentation of live telescopic imagery with aerial and terrain based synthetic objects, vehicles, people and effects. The inserted objects must appear stable in the display and must not jitter and drift as the user pans around and examines the scene with the binoculars. The design of the system is based on using two different cameras with wide field of view, and narrow field of view lenses enclosed in a binocular shaped shell. Using the wide field of view gives us context and enables us to recover the 3D location and orientation of the binoculars much more robustly, whereas the narrow field of view is used for the actual augmentation as well as to increase precision in tracking. We present our navigation algorithm that uses the two cameras in combination with an IMU and GPS in an Extended Kalman Filter (EKF) and provides jitter free, robust and real-time pose estimation for precise augmentation. We have demonstrated successful use of our system as part of a live simulated training system for observer training, in which fixed and rotary wing aircrafts, ground vehicles, and weapon effects are combined with real world scenes.
Taragay Oskiper, Mikhail Sizintsev, Vlad Branzoi, Supun Samarasekera, Rakesh Kumar 0001
ISMAR4
2012 Long-Range Pedestrian Detection using stereo and a cascade of convolutional network classifiers
abstract
In this paper, we present a system for detecting pedestrians at long ranges using a combination of stereo-based detection, classification using deep learning, and a cascade of specialized classifiers that can reduce false positives and computational load. Specifically, we use stereo to perform detection of vertical structures which are further filtered based on edge responses. A convolutional neural network was then designed to support the classification of pedestrians using both appearance and stereo disparity-based features. A second convolutional network classifier was trained specifically for the case of long-range detections using appearance only. We further speed up the classifier using a cascade approach and multi-threading. The system was deployed on two robots, one using a high resolution stereo pair with 180 degree fisheye lenses and the other using 80 degree FOV lenses. Results are demonstrated on a large dataset captured in a variety of environments.
Zsolt Kira, Raia Hadsell, Garbis Salgian, Supun Samarasekera
IROS4
2012 Multi-sensor navigation algorithm using monocular camera, IMU and GPS for large scale augmented reality
abstract
Camera tracking system for augmented reality applications that can operate both indoors and outdoors is described. The system uses a monocular camera, a MEMS-type inertial measurement unit (IMU) with 3-axis gyroscopes and accelerometers, and GPS unit to accurately and robustly track the camera motion in 6 degrees of freedom (with correct scale) in arbitrary indoor or outdoor scenes. IMU and camera fusion is performed in a tightly coupled manner by an error-state extended Kalman filter (EKF) such that each visually tracked feature contributes as an individual measurement as opposed to the more traditional approaches where camera pose estimates are first extracted by means of feature tracking and then used as measurement updates in a filter framework. Robustness in feature tracking and hence in visual measurement generation is achieved by IMU aided feature matching and a two-point relative pose estimation method, to remove outliers from the raw feature point matches. Landmark matching to contain long-term drift in orientation via on the fly user generated geo-tiepoint mechanism is described.
Taragay Oskiper, Supun Samarasekera, Rakesh Kumar 0001
ISMAR2
2011 Vehicle tracking across nonoverlapping cameras using joint kinematic and appearance features
abstract
We describe a vehicle tracking algorithm using input from a network of nonoverlapping cameras. Our algorithm is based on a novel statistical formulation that uses joint kinematic and image appearance information to link local tracks of the same vehicles into global tracks with longer persistence. The algorithm can handle significant spatial separation between the cameras and is robust to challenging tracking conditions such as high traffic density, or complex road infrastructure. In these cases, traditional tracking formulations based on MHT, or JPDA algorithms, may fail to produce track associations across cameras due to the weak predictive models employed. We make several new contributions in this paper. Firstly, we model kinematic constraints between any two local tracks using road networks and transit time distributions. The transit time distributions are calculated dynamically as convolutions of normalized transit time distributions that are learned and adapted separately for individual roads. Secondly, we present a complete statistical tracker formulation, which combines kinematic and appearance likelihoods within a multi-hypothesis framework. We have extensively evaluated the algorithm proposed using a network of ground-based cameras with narrow field of view. The tracking results obtained on a large ground-truthed dataset demonstrate the effectiveness of the algorithm proposed.
Bogdan Matei, Harpreet Sawhney, Supun Samarasekera
CVPR3
2011 High-precision localization using visual landmarks fused with range data
abstract
Visual landmark matching with a pre-built landmark database is a popular technique for localization. Traditionally, landmark database was built with visual odometry system, and the 3D information of each visual landmark is reconstructed from video. Due to the drift of the visual odometry system, a global consistent landmark database is difficult to build, and the inaccuracy of each 3D landmark limits the performance of landmark matching. In this paper, we demonstrated that with the use of precise 3D Li-dar range data, we are able to build a global consistent database of high precision 3D visual landmarks, which improves the landmark matching accuracy dramatically. In order to further improve the accuracy and robustness, landmark matching is fused with a multi-stereo based visual odometry system to estimate the camera pose in two aspects. First, a local visual odometry trajectory based consistency check is performed to reject some bad landmark matchings or those with large errors, and then a kalman filtering is used to further smooth out some landmark matching errors. Finally, a disk-cache-mechanism is proposed to obtain the real-time performance when the size of the landmark grows for a large-scale area. A week-long real time live marine training experiments have demonstrated the high-precision and robustness of our proposed system.
Han-Pang Chiu, Taragay Oskiper, Saad Ali, Raia Hadsell, Supun Samarasekera, Rakesh Kumar 0001
CVPR6
2011 A graph traversal based algorithm for obstacle detection using lidar or stereo
abstract
We present a novel computationally efficient approach to obstacle detection that is applicable to both structured (e.g. indoor, road) and unstructured (e.g. off-road, grassy terrain) environments. In contrast to previous works that attempt to explicitly identify obstacles, we explicitly detect scene regions that are traversable - safe for the robot to go to - from its current position. Traversability is defined on a 2D grid of cells. Given 3D points, we map them to individual cells and compute histograms of elevations of the points in each cell. This elevation information is then used in a graph based algorithm to label all traversable cells. In this manner, positive and negative obstacles, as well as unknown regions are implicitly detected and avoided. Our notion of traversability does not make any flat-world assumptions and does not need sensor pitch-roll compensation. It also accounts for overhanging structures like tree branches. We demonstrate that our approach can be used with both lidar and stereo sensors even though the two sensors differ in their resolution and accuracy. We present several results from our real-time implementation on realistic environments using both lidar and stereo.
Sujit Kuthirummal, Aveek Das, Supun Samarasekera
IROS3
2011 Tightly-coupled robust vision aided inertial navigation algorithm for augmented reality using monocular camera and IMU
abstract
Odometry component of a camera tracking system for augmented reality applications is described. The system uses a MEMS-type inertial measurement unit (IMU) with 3-axis gyroscopes and accelerometers and a monocular camera to accurately and robustly track the camera motion in 6 degrees of freedom (with correct scale) in arbitrary indoor or outdoor scenes. Tight coupling of IMU and camera is achieved by an error-state extended Kalman filter (EKF) which performs sensor fusion for inertial navigation at a deep level such that each visually tracked feature contributes as an individual measurement as opposed to the more traditional approaches where camera pose estimates are first extracted by means of feature tracking and then used as measurement updates in a filter framework. Robustness, on the other hand, is achieved by using a geometric hypothesize-and-test architecture based on the five-point relative pose estimation method, rather than a Mahalanobis distance type gating mechanism derived from the Kalman filter state prediction, to select the inlier tracks and remove outliers from the raw feature point matches which would otherwise corrupt the filter since tracks are directly used as measurements.
Taragay Oskiper, Supun Samarasekera, Rakesh Kumar 0001
ISMAR2
2011 Stable vision-aided navigation for large-area augmented reality
abstract
In this paper, we present a unified approach for a drift-free and jitter-reduced vision-aided navigation system. This approach is based on an error-state Kalman filter algorithm using both relative (local) measurements obtained from image based motion estimation through visual odometry, and global measurements as a result of landmark matching through a pre-built visual landmark database. To improve the accuracy in pose estimation for augmented reality applications, we capture the 3D local reconstruction uncertainty of each landmark point as a covariance matrix and implicity rely more on closer points in the filter. We conduct a number of experiments aimed at evaluating different aspects of our Kalman filter framework, and show our approach can provide highly-accurate and stable pose both indoors and outdoors over large areas. The results demonstrate both the long term stability and the overall accuracy of our algorithm as intended to provide a solution to the camera tracking problem in augmented reality applications.
Taragay Oskiper, Han-Pang Chiu, Supun Samarasekera, Rakesh Kumar 0001
VR4
2010 Robust visual path following for heterogeneous mobile platforms
abstract
We present an innovative path following system based upon multi-camera visual odometry and visual landmark matching. This technology enables reliable mobile robot navigation in real world scenarios including GPS-denied environments both indoors and outdoors. We recover paths in full 3D, making it applicable to both on and off-road ground vehicles. Our controller relies on pose updates from visual odometry, allowing us to achieve path following even when only a joystick drive interface to the base robot platform is available. We experimentally investigate two specific applications of our technology to autonomous navigation on ground vehicles - non line-of-sight leader-following (between heterogeneous platforms) and retro-traverse to home base. For safety and reliability we add dynamic short range obstacle detection and reactive avoidance capabilities to our controller. We show the results for end-to-end real time implementation of this technology using current off-the-shelf computing and network resources in challenging environments.
Aveek Das, Oleg Naroditsky, Supun Samarasekera, Rakesh Kumar 0001
ICRA4
2010 Multi-modal sensor fusion algorithm for ubiquitous infrastructure-free localization in vision-impaired environments
abstract
In this paper, we present a unified approach for a camera tracking system based on an error-state Kalman filter algorithm. The filter uses relative (local) measurements obtained from image based motion estimation through visual odometry, as well as global measurements produced by landmark matching through a pre-built visual landmark database and range measurements obtained from radio frequency (RF) ranging radios. We show our results by using the camera poses output by our system to render views from a 3D graphical model built upon the same coordinate frame as the landmark database which also forms the global coordinate system and compare them to the actual video images. These results help demonstrate both the long term stability and the overall accuracy of our algorithm as intended to provide a solution to the GPS denied ubiquitous camera tracking problem under both vision-aided and vision-impaired conditions.
Taragay Oskiper, Han-Pang Chiu, Supun Samarasekera, Rakesh Kumar 0001
IROS4
2009 VideoTrek: A vision system for a tag-along robot
abstract
We present a system that combines multiple visual navigation techniques to achieve GPS-denied, non-line-of-sight SLAM capability for heterogeneous platforms. Our approach builds on several layers of vision algorithms, including sparse frame-to-frame structure from motion (visual odometry), a Kalman filter for fusion with inertial measurement unit (IMU) data and a distributed visual landmark matching capability with geometric consistency verification. We apply these techniques to implement a tag-along robot, where a human operator leads the way and a robot autonomously follows. We show results for a real-time implementation of such a system with real field constraints on CPU power and network resources.
Oleg Naroditsky, Aveek Das, Supun Samarasekera, Taragay Oskiper, Rakesh Kumar 0001
CVPR4
2008 Matching vehicles under large pose transformations using approximate 3D models and piecewise MRF model
abstract
We propose a robust object recognition method based on approximate 3D models that can effectively match objects under large viewpoint changes and partial occlusion. The specific problem we solve is: given two views of an object, determine if the views are for the same or different object. Our domain of interest is vehicles, but the approach can be generalized to other man-made rigid objects. A key contribution of our approach is the use of approximate models with locally and globally constrained rendering to determine matching objects. We utilize a compact set of 3D models to provide geometry constraints and transfer appearance features for object matching across disparate viewpoints. The closest model from the set, together with its poses with respect to the data, is used to render an object both at pixel (local) level and region/part (global) level. Especially, symmetry and semantic part ownership are used to extrapolate appearance information. A piecewise Markov Random Field (MRF) model is employed to combine observations obtained from local pixel and global region level. Belief Propagation (BP) with reduced memory requirement is employed to solve the MRF model effectively. No training is required, and a realistic object image in a disparate viewpoint can be obtained from as few as just one image. Experimental results on vehicle data from multiple sensor platforms demonstrate the efficacy of our method.
Yanlin Guo, Cen Rao, Supun Samarasekera, Janet Kim, Rakesh Kumar 0001, Harpreet Sawhney
CVPR3
2008 Building segmentation for densely built urban regions using aerial LIDAR data
abstract
We present a novel building segmentation system for densely built areas, containing thousands of buildings per square kilometer. We employ solely sparse LIDAR (Light/Laser Detection Ranging) 3D data, captured from an aerial platform, with resolution less than one point per square meter. The goal of our work is to create segmented and delineated buildings as well as structures on top of buildings without requiring scanning for the sides of buildings. Building segmentation is a critical component in many applications such as 3D visualization, robot navigation and cartography. LIDAR has emerged in recent years as a more robust alternative to 2D imagery because it acquires 3D structure directly, without the shortcomings of stereo in un- textured regions and at depth discontinuities. Our main technical contributions in this paper are: (i) a ground segmentation algorithm which can handle both rural regions, and heavily urbanized areas, where the ground is 20% or less of the data, (ii) a building segmentation technique, which is robust to buildings in close proximity to each other, sparse measurements and nearby structured vegetation clutter, and (Hi) an algorithm for estimating the orientation of a boundary contour of a building, based on minimizing the number of vertices in a rectilinear approximation to the building outline, which can cope with significant quantization noise in the outline measurements. We have applied the proposed building segmentation system to several urban regions with areas of hundreds of square kilometers each, obtaining average segmentation speeds of less than three minutes per km2on a standard Pentium processor. Extensive qualitative results obtained by overlaying the 3D segmented regions onto 2D imagery indicate accurate performance of our system.
Bogdan Matei, Harpreet Sawhney, Supun Samarasekera, Janet Kim, Rakesh Kumar 0001
CVPR3
2008 Real-time global localization with a pre-built visual landmark database
abstract
In this paper, we study how to build a vision-based system for global localization with accuracies within 10cm. for robots and humans operating both indoors and outdoors over wide areas covering many square kilometers. In particular, we study the parameters of building a landmark database rapidly and utilizing that database online for real-time accurate global localization. Although the accuracy of traditional short-term motion based visual odometry systems has improved significantly in recent years, these systems alone cannot solve the drift problem over large areas. Landmark based localization combined with visual odometry is a viable solution to the large scale localization problem. However, a systematic study of the specification and use of such a landmark database has not been undertaken. We propose techniques to build and optimize a landmark database systematically and efficiently using visual odometry. First, topology inference is utilized to find overlapping images in the database. Second, bundle adjustment is used to refine the accuracy of each 3D landmark. Finally, the database is optimized to balance the size of the database with achievable accuracy. Once the landmark database is obtained, a new real-time global localization methodology that works both indoors and outdoors is proposed. We present results of our study on both synthetic and real datasets that help us determine critical design parameters for the landmark database and the achievable accuracies of our proposed system.
Taragay Oskiper, Supun Samarasekera, Rakesh Kumar 0001, Harpreet Sawhney
CVPR3
2008 Toward a sentient environment: real-time wide area multiple human tracking with identities
Manoj Aggarwal, Thomas Germano, Ian Roth, Alexandar Knowles, Rakesh Kumar 0001, Harpreet Sawhney, Supun Samarasekera
Mach. Vis. Appl.8
2007 Visual Odometry System Using Multiple Stereo Cameras and Inertial Measurement Unit
abstract
Over the past decade, tremendous amount of research activity has focused around the problem of localization in GPS denied environments. Challenges with localization are highlighted in human wearable systems where the operator can freely move through both indoors and outdoors. In this paper, we present a robust method that addresses these challenges using a human wearable system with two pairs of backward and forward looking stereo cameras together with an inertial measurement unit (IMU). This algorithm can run in real-time with 15 Hz update rate on a dual-core 2 GHz laptop PC and it is designed to be a highly accurate local (relative) pose estimation mechanism acting as the front-end to a simultaneous localization and mapping (SLAM) type method capable of global corrections through landmark matching. Extensive tests of our prototype system so far, reveal that without any global landmark matching, we achieve between 0.5% and 1% accuracy in localizing a person over a 500 meter travel indoors and outdoors. To our knowledge, such performance results with a real time system have not been reported before.
Taragay Oskiper, Supun Samarasekera, Rakesh Kumar 0001
CVPR3
2007 Ten-fold Improvement in Visual Odometry Using Landmark Matching
abstract
Our goal is to create a visual odometry system for robots and wearable systems such that localization accuracies of centimeters can be obtained for hundreds of meters of distance traveled. Existing systems have achieved approximately a 1% to 5% localization error rate whereas our proposed system achieves close to 0.1% error rate, a ten-fold reduction. Traditional visual odometry systems drift over time as the frame-to-frame errors accumulate. In this paper, we propose to improve visual odometry using visual landmarks in the scene. First, a dynamic local landmark tracking technique is proposed to track a set of local landmarks across image frames and select an optimal set of tracked local landmarks for pose computation. As a result, the error associated with each pose computation is minimized to reduce the drift significantly. Second, a global landmark based drift correction technique is proposed to recognize previously visited locations and use them to correct drift accumulated during motion. At each visited location along the route, a set of distinctive visual landmarks is automatically extracted and inserted into a landmark database dynamically. We integrate the landmark based approach into a navigation system with 2 stereo pairs and a low-cost inertial measurement unit (IMU) for increased robustness. We demonstrate that a real-time visual odometry system using local and global landmarks can precisely locate a user within 1 meter over 1000 meters in unknown indoor/outdoor environments with challenging situations such as climbing stairs, opening doors, moving foreground objects etc..
Taragay Oskiper, Supun Samarasekera, Rakesh Kumar 0001, Harpreet Sawhney
ICCV3
2001 Aerial video surveillance and exploitation
abstract
There is growing interest in performing aerial surveillance using video cameras. Compared to traditional framing cameras, video cameras provide the capability to observe ongoing activity within a scene and to automatically control the camera to track the activity. However, the high data rates and relatively small field of view of video cameras present new technical challenges that must be overcome before such cameras can be widely used. In this paper, we present a framework and details of the key components for real-time, automatic exploitation of aerial video for surveillance applications. The framework involves separating an aerial video into the natural components corresponding to the scene. Three major components of the scene are the static background geometry, moving objects, and appearance of the static and dynamic components of the scene. In order to delineate videos into these scene components, we have developed real time, image-processing techniques for 2-D/3-D frame-to-frame alignment, change detection, camera control, and tracking of independently moving objects in cluttered scenes. The geo-location of video and tracked objects is estimated by registration of the video to controlled reference imagery, elevation maps, and site models. Finally static, dynamic and reprojected mosaics may be constructed for compression, enhanced visualization, and mapping applications.
Rakesh Kumar 0001, Harpreet Sawhney, Supun Samarasekera, Steven C. Hsu, Yanlin Guo, Keith J. Hanna, Art Pope, Richard P. Wildes, David J. Hirvonen, Michael W. Hansen, Peter Burt
Proc. IEEE3
2000 Multi-View 3D Analysis with Applications for Augmented Reality and Enhanced Video Visualization
abstract
The article presents methods for 3D scene geometry recovery/refinement and pose estimation from motion imagery in two representative scenarios. First, we present a method for pose estimation and scene geometry recovery from extended sequences without prior knowledge of the scheme. Second, we discuss how to recover camera poses when a rough scene model is provided. We show how to extend and refine the scene model using the recovered poses. Finally, we present applications of the above techniques for 3D imagery manipulation such as enhanced visualization for video and 3D insertion of synthetic objects in the imagery.
Yanlin Guo, Steven C. Hsu, Supun Samarasekera, Harpreet Sawhney, Rakesh Kumar 0001
CVPR3
2000 Pose Estimation, Model Refinement, and Enhanced Visualization Using Video
abstract
In this paper we present methods for exploitation and enhanced visualization of video given a prior coarse untextured polyhedral model of a scene. Since it is necessary to estimate the 3D poses of the moving camera, we develop an algorithm where tracked features are used to predict the pose between frames and the predicted poses are refined by a coarse to fine process of aligning projected 3D model line segments to oriented image gradient energy pyramids. The estimated poses can be used to update the model with information derived from video, and to re-project and visualize the video from different points of view with a larger scene context. Via image registration, we update the placement of objects in the model and the 3D shape of new or erroneously modeled objects, then map video texture to the model. Experimental results are presented for long aerial and ground level videos of a large-scale urban scene.
Stephen C. Hsu, Supun Samarasekera, Rakesh Kumar 0001, Harpreet Sawhney
CVPR2
2000 3D Manipulation of Motion Imagery
abstract
We present a set of automatic methods for the recovery and refinement of 3D scene geometry and camera poses from motion imagery. First, we present a two-frame "direct" method, which simultaneously estimates both relative pose between the cameras and 3D scene geometry using information from the images alone. Second, we discuss how to automatically estimate pose and scene geometry using extended sequences. Third, we present methods for recovery of pose when a scene model is known a priori. We show how the recovered pose can be use to extend and refine the scene model. Finally, we present applications of the above techniques for 3D imagery manipulation such as enhanced visualization of video, synthetic view generation and insertion of synthetic objects in the imagery.
Rakesh Kumar 0001, Harpreet Sawhney, Yanlin Guo, Steven C. Hsu, Supun Samarasekera
ICIP5
2000 Registration of Highly-Oblique and Zoomed in Aerial Video to Reference Imagery
abstract
We present methods for estimating precise geo-coordinates of objects observed in highly oblique video from an airborne camera in real time. High precision is achieved by registering observed video frames in real time to rendered views of the scene using stored reference imagery. The reference imagery includes previously collected satellite images and terrain maps that have been precisely aligned to geo-coordinates. We present methods that work for both nadir and highly oblique video imagery and under a variety of conditions where any one frame may not have enough information for accurate geo-registration.
Rakesh Kumar 0001, Supun Samarasekera, Steven C. Hsu, Keith J. Hanna
ICPR2
1998 Auto Cameraman Via Collaborative Sensing Agents
Yuntao Cui, Supun Samarasekera, Michael Greiffenhagen
ACCV (1)3
1998 Content based Active Video Data Acquisition via Automated Cameramen
abstract
We propose to actively apply content based operation to video data acquisition. Our goal is to develop an automated cameraman to replace the human operator who acquires data in a content sensitive (smart) manner. This auto cameramen is capable of (1) constantly monitoring a global surrounding; (2) automatically keeping track of important visual events; (3) dynamically, based on the detected visual events, determining the video acquisition strategy; and (4) actively generating continuous video clips that are visual events coherent.
Yuntao Cui, Supun Samarasekera
ICIP (2)3
1998 Multimedia applications of computer vision
abstract
The technical foundation for many applications of computer vision to multimedia applications is efficient and robust image motion estimation. These algorithms enable the creation of algorithms for mosaic construction, registration of video to a database multi-sensor registration, 3D estimation and representation, and video content indexing and retrieval. Demonstrations on the above topics will be shown at WACV'98 by the Media Vision Group of Sarnoff Corporation.
Jane C. Asmuth, Douglas Dixon, Keith J. Hanna, Steven C. Hsu, Rakesh Kumar 0001, Vince Paragano, Art Pope, Supun Samarasekera, Harpreet Sawhney
WACV8
1998 User-Steered Image Segmentation Paradigms: Live Wire and Live Lane
abstract
In multidimensional image analysis, there are, and will continue to be, situations wherein automatic image segmentation methods fail, calling for considerable user assistance in the process. The main goals of segmentation research for such situations ought to be (i) to provideeffective controlto the user on the segmentation processwhileit is being executed, and (ii) to minimize the total user's time required in the process. With these goals in mind, we present in this paper two paradigms, referred to aslive wireandlive lane, for practical image segmentation in large applications. For both approaches, we think of the pixel vertices and oriented edges as forming a graph, assign a set of features to each oriented edge to characterize its ``boundariness,'' and transform feature values to costs. We provide training facilities and automatic optimal feature and transform selection methods so that these assignments can be made with consistent effectiveness in any application. In live wire, the user first selects an initial point on the boundary. For any subsequent point indicated by the cursor, an optimal path from the initial point to the current point is found and displayed in real time. The user thus has a live wire on hand which is moved by moving the cursor. If the cursor goes close to the boundary, the live wire snaps onto the boundary. At this point, if the live wire describes the boundary appropriately, the user deposits the cursor which now becomes the new starting point and the process continues. A few points (live-wire segments) are usually adequate to segment the whole 2D boundary. In live lane, the user selects only the initial point. Subsequent points are selected automatically as the cursor is moved within a lane surrounding the boundary whose width changes as a function of the speed and acceleration of cursor motion. Live-wire segments are generated and displayed in real time between successive points. The users get the feeling that the curve snaps onto the boundary as and while they roughly mark in the vicinity of the boundary. We describe formal evaluation studies to compare the utility of the new methods with that of manual tracing based on speed and repeatability of tracing and on data taken from a large ongoing application. The studies indicate that the new methods are statistically significantly more repeatable and 1.5–2.5 times faster than manual tracing.
Alexandre X. Falcão, Jayaram K. Udupa, Supun Samarasekera, Shoba Sharma, Bruce Elliot Hirsch, Roberto A. Lotufo
Graph. Model. Image Process.3
1997 Multiple Sclerosis Lesion Quantification Using Fuzzy-Connectedness Principles
abstract
Multiple sclerosis (MS) is a disease of the white matter. Magnetic resonance imaging (MRI) is proven to be a sensitive method of monitoring the progression of this disease and of its changes due to treatment protocols. Quantification of the severity of the disease through estimation of MS lesion volume via MR imaging is vital for understanding and monitoring the disease and its treatment. This paper presents a novel methodology and a system that can be routinely used for segmenting and estimating the volume of MS lesions via dual-echo fast spin-echo MR imagery. A recently developed concept of fuzzy objects forms the basis of this methodology. An operator indicates a few points in the images by pointing to the white matter, the grey matter, and the cerebro-spinal fluid (CSF). Each of these objects is then detected as a fuzzy connected set. The holes in the union of these objects correspond to potential lesion sites which are utilized to detect each potential lesion as a three-dimensional (3-D) fuzzy connected object. These objects are presented to the operator who indicates acceptance/rejection through the click of a mouse button. The number and volume of accepted lesions is then computed and output. Based on several evaluation studies, we conclude that the methodology is highly reliable and consistent, with a coefficient of variation (due to subjective operator actions) of 0.9% (based on 20 patient studies, three operators, and two trials) for volume and a mean false-negative volume fraction of 1.3%, with a 95% confidence interval of 0%-2.8% (based on ten patient studies).
Jayaram K. Udupa, Luogang Wei, Supun Samarasekera, Yukio Miki, Mark A. van Buchem, Robert I. Grossman
IEEE Trans. Medical Imaging3
1996 Fuzzy Connectedness and Object Definition: Theory, Algorithms, and Applications in Image Segmentation
abstract
Images are by nature fuzzy. Approaches to object information extraction from images should attempt to use this fact and retain fuzziness as realistically as possible. In past image segmentation research, the notion of “hanging togetherness” of image elements specified by their fuzzy connectedness has been lacking. We present a theory of fuzzy objects forn-dimensional digital spaces based on a notion of fuzzy connectedness of image elements. Although our definitions lead to problems of enormous combinatorial complexity, the theoretical results allow us to reduce this dramatically, leading us to practical algorithms for fuzzy object extraction. We present algorithms for extracting a specified fuzzy object and for identifying all fuzzy objects present in the image data. We demonstrate the utility of the theory and algorithms in image segmentation based on several practical examples all drawn from medical imaging.
Jayaram K. Udupa, Supun Samarasekera
CVGIP Graph. Model. Image Process.2
1992 Design and performance of a prototype analog neural computer
abstract
A prototype programmable analog neural computer and selected applications are described. The machine is assembled from over 100 custom VLSI modules containing neurons, synapses, routing switches and programmable synaptic time constants. Connection symmetry and modular construction allow expansion to arbitrary size. The network runs in real time analog mode, however connection architecture as well as neuron and synapse parameters are controlled by a digital host that monitors also the network performance through an A/D interface. Programming and monitoring software has been developed and several application examples including the dynamic decomposition of acoustical patterns are described. The machine is intended for real time, real world computations including ATR. In current configuration its maximal speed is equivalent to that of a digital machine capable of more than 1011 flops. A much larger machine is currently under development.
Paul Mueller, Jan Van der Spiegel, Vincent Agami, David Blackman, Peter Chance, Christopher Donham, Ralph Etienne-Cummings, Jason Flinn, Mike Massa, Supun Samarasekera
Neurocomputing11