EDBT 2026 Demo / reviewers in the wild / expert
B. S. Manjunath
dblp:m/BSManjunath · also Bangalore S. Manjunath
· DBLP profile ↗
196ranked-venue papers
10as first author
15since 2021 · last 2026
0000-0003-2804-3611ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 160 · 6 first-author · 11 since 2021Artificial intelligence and machine learning · 51 · 7 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 3 since 2021Security and privacy · 7 · 1 since 2021Computer networks · 3Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Markovian Reeb Graphs for Simulating Spatiotemporal Patterns of Life
Anantajit Subrahmanya, Chandrakanth Gudavalli, Connor Levenson, B. S. Manjunath |
ICPR (2) | 4 |
| 2025 | METAREG: Robust Camera Parameter Estimation by Leveraging Noisy Camera ExtrinsicsabstractNovel view synthesis methods, such as Neural Radiance Fields (NeRFs) and 3D Gaussian Splatting (3DGS), rely on Structure-from-Motion (SfM) pipelines like COLMAP for camera parameter estimation. However, these pipelines are prone to errors due to factors like doppelgangers and perspective distortion. While edge devices (e.g., mobile phones) capture images with embedded GPS and IMU data, this meta-data is often noisy. We introduce MetaReg, a robust pipeline that improves camera parameter estimation by leveraging noisy GPS metadata. MetaReg enhances COLMAP with pre-and post-processing steps: the preprocessing stage estimates image overlap using metadata to reduce unnecessary image matching, while the post-processing stage aligns estimated camera coordinates with the world coordinate system using noisy GPS priors. Experiments on challenging datasets demonstrate that MetaReg significantly improves camera parameter estimation, enhancing robustness and accuracy. Chandrakanth Gudavalli, Tajuddin Manhar Mohammed, Ananth Vishnu Bhaskar, Elliot Staudt, Cheng Peng 0008, Abhay Yadav, Rama Chellappa, Shivkumar Chandrasekaran, B. S. Manjunath |
ICIP | 9 |
| 2025 | DDS: Decoupled Dynamic Scene-Graph Generation NetworkabstractScene-graph generation involves creating a structural representation of the relationships between objects in a scene by predicting subject-object-relation triplets from input data. Existing methods show poor performance in detecting triplets outside of a predefined set, primarily due to their reliance on dependent feature learning. To address this issue we propose DDS– a decoupled dynamic scene-graph generation network- that consists of two independent branches that can disentangle extracted features. The key innovation of the current paper is the decoupling of the features representing the relationships from those of the objects, which enables the detection of novel object-relationship combinations. The DDS model is evaluated on three datasets and outperforms previous methods by a significant margin, especially in detecting previously unseen triplets. A S. M. Iftekhar, Raphael Ruschel, Suya You, B. S. Manjunath |
WACV | 5 |
| 2025 | Generalizable Deepfake Detection With Phase-Based Motion AnalysisabstractWe propose PhaseForensics, a DeepFake (DF) video detection method that uses a phase-based motion representation of facial temporal dynamics. Existing methods that rely on temporal information across video frames for DF detection have many advantages over the methods that only utilize the per-frame features. However, these temporal DF detection methods still show limited cross-dataset generalization and robustness to common distortions due to factors such as error-prone motion estimation, inaccurate landmark tracking, or the susceptibility of the pixel intensity-based features to adversarial distortions and the cross-dataset domain shifts. Our key insight to overcome these issues is to leverage the temporal phase variations in the band-pass frequency components of a face region across video frames. This not only enables a robust estimate of the temporal dynamics in the facial regions, but is also less prone to cross-dataset variations. Furthermore, we show that the band-pass filters used to compute the local per-frame phase form an effective defense against the perturbations commonly seen in gradient-based adversarial attacks. Overall, with PhaseForensics, we show improved distortion and adversarial robustness, and state-of-the-art cross-dataset generalization, with 92.4% video-level AUC on the challenging CelebDFv2 benchmark (a recent state of-the-art method, FTCN, compares at 86.9%). Ekta Prashnani, Michael Goebel, B. S. Manjunath |
IEEE Trans. Image Process. | 3 |
| 2024 | WildlifeMapper: Aerial Image Analysis for Multi-Species Detection and IdentificationabstractWe introduce WildlifeMapper (WM), a flexible model designed to detect, locate, and identify multiple species in aerial imagery. It addresses the limitations of traditional, labor-intensive wildlife population assessments that are central to advancing environmental conservation efforts worldwide. While a number of methods exist to automate this process, they are often limited in their ability to general-ize to different species or landscapes due to the dominance of homogeneous backgrounds and/or poorly captured local image structures. WM introduces two novel modules that help to capture the local structure and context of objects of interest to accurately localize and identify them, achieving a state-of-the-art (SOTA) detection rate of 0.56 mAP. Fur-ther, we introduce a large aerial imagery dataset with more than 11 k Images and 28k annotations verified by domain experts. WM also achieves SOTA performance on 3 other publicly available aerial survey datasets collected across 4 different countries, improving mAP by 42%. Source code and trained models are available at Github11https://github.com/UCSB-VRL/WildlifeMapper. Chandrakanth Gudavalli, Connor Levenson, Lacey Hughey, Jared A. Stabach, Irene Amoke, Gordon Ojwang, Joseph Mukeka, Stephen Mwiu, Joseph Ogutu, Howard Frederick, B. S. Manjunath |
CVPR | 13 |
| 2024 | CIMGEN: Controlled Satellite Image Manipulation by Finetuning Pretrained Generative Models on Limited Data
Chandrakanth Gudavalli, Erik Rosten, Lakshmanan Nataraj, Shivkumar Chandrasekaran, B. S. Manjunath |
ICPR (21) | 5 |
| 2024 | ReeSPOT: Reeb Graph Models Semantic Patterns of Normalcy in Human Trajectories
S. Shailja, Chandrakanth Gudavalli, Connor Levenson, Amil Khan, B. S. Manjunath |
ICPR (7) | 6 |
| 2023 | MethaneMapper: Spectral Absorption Aware Hyperspectral Transformer for Methane DetectionabstractMethane (CH4) is the chief contributor to global climate change. Recent Airborne Visible-Infrared Imaging Spectrometer-Next Generation (AVIRIS-NG) has been very useful in quantitative mapping of methane emissions. Existing methods for analyzing this data are sensitive to local terrain conditions, often require manual inspection from domain experts, prone to significant error and hence are not scalable. To address these challenges, we propose a novel end-to-end spectral absorption wavelength aware transformer network, MethaneMapper, to detect and quantify the emissions. MethaneMapper introduces two novel modules that help to locate the most relevant methane plume regions in the spectral domain and uses them to localize these accurately. Thorough evaluation shows that MethaneMapper achieves 0.63 mAP in detection and reduces the model size (by 5×) compared to the current state of the art. In addition, we also introduce a large-scale dataset of methane plume segmentation mask for over 1200 AVIRIS-NG flight lines from 2015–2022. It contains over 4000 methane plume sites. Our dataset will provide researchers the opportunity to develop and advance new methods for tackling this challenging green-house gas detection problem with significant broader social impact. Dataset and source code link11https://github.com/UCSB-Vrl/methaneMapper-Spectral-Absorption-aware-Hyperspectral-Transformer-for-Methane-Detection. Ivan Arevalo, A S. M. Iftekhar, B. S. Manjunath |
CVPR | 4 |
| 2023 | A robust approach to 3D neuron shape representation for quantification and classificationabstractWe consider the problem of finding an accurate representation of neuron shapes, extracting sub-cellular features, and classifying neurons based on neuron shapes. In neuroscience research, the skeleton representation is often used as a compact and abstract representation of neuron shapes. However, existing methods are limited to getting and analyzing "curve" skeletons which can only be applied for tubular shapes. This paper presents a 3D neuron morphology analysis method for more general and complex neuron shapes. First, we introduce the concept of skeleton mesh to represent general neuron shapes and propose a novel method for computing mesh representations from 3D surface point clouds. A skeleton graph is then obtained from skeleton mesh and is used to extract sub-cellular features. Finally, an unsupervised learning method is used to embed the skeleton graph for neuron classification. Extensive experiment results are provided and demonstrate the robustness of our method to analyze neuron morphology. Jiaxiang Jiang, Michael Goebel, Cezar Borba, William Smith 0001, B. S. Manjunath |
BMC Bioinform. | 5 |
| 2023 | Context-Driven Detection of Invertebrate Species in Deep-Sea VideoabstractAbstract Each year, underwater remotely operated vehicles (ROVs) collect thousands of hours of video of unexplored ocean habitats revealing a plethora of information regarding biodiversity on Earth. However, fully utilizing this information remains a challenge as proper annotations and analysis require trained scientists’ time, which is both limited and costly. To this end, we present a Dataset for Underwater Substrate and Invertebrate Analysis (DUSIA), a benchmark suite and growing large-scale dataset to train, validate, and test methods for temporally localizing four underwater substrates as well as temporally and spatially localizing 59 underwater invertebrate species. DUSIA currently includes over ten hours of footage across 25 videos captured in 1080p at 30 fps by an ROV following pre-planned transects across the ocean floor near the Channel Islands of California. Each video includes annotations indicating the start and end times of substrates across the video in addition to counts of species of interest. Some frames are annotated with precise bounding box locations for invertebrate species of interest, as seen in Fig. 1. To our knowledge, DUSIA is the first dataset of its kind for deep sea exploration, with video from a moving camera, that includes substrate annotations and invertebrate species that are present at significant depths where sunlight does not penetrate. Additionally, we present the novel context-driven object detector (CDD) where we use explicit substrate classification to influence an object detection network to simultaneously predict a substrate and species class influenced by that substrate. We also present a method for improving training on partially annotated bounding box frames. Finally, we offer a baseline method for automating the counting of invertebrate species of interest. R. Austin McEver, Connor Levenson, A S. M. Iftekhar, B. S. Manjunath |
Int. J. Comput. Vis. | 5 |
| 2023 | Resampling Estimation Based RPC Metadata Verification in Satellite ImageryabstractRecent advances in machine learning and computer vision have made it simple to manipulate a variety of media, including satellite images. Most of the commercially available satellite images go through the process of orthorectification to remove potential distortions due to terrain variations. This orthorectification process typically involves the use of rational polynomial coefficients (RPC) that geometrically remap the pixels in the original image to the rectified image. This paper proposes the first method to verify the authenticity of RPC metadata in an orthorectified satellite image. The steps include calculating the Residual Discrete Fourier Transform (DFT) pattern from the image using a linear predictor based residual spectral analysis and comparing with Expected Residual DFT pattern that is obtained using the RPC metadata associated with the image. If the metadata associated with orthorectified image is correct, then the Residual-DFT pattern (which represents image data) and the Expected-Residual-DFT pattern (which represents metadata) should be similar. We use SSIM (Structural Similarity Index Metric) to quantify the similarity and thereby verify if the data has been tampered or not. Detailed experimental results demonstrate that our method achieves over 97% accuracy in the majority of binary tampering detection tests. Chandrakanth Gudavalli, Michael Goebel, Tejaswi Nanjundaswamy, Lakshmanan Nataraj, Shivkumar Chandrasekaran, B. S. Manjunath |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2023 | ReeBundle: A Method for Topological Modeling of White Matter Pathways Using Diffusion MRIabstractTractography can generate millions of complex curvilinear fibers (streamlines) in 3D that exhibit the geometry of white matter pathways in the brain. Common approaches to analyzing white matter connectivity are based on adjacency matrices that quantify connection strength but do not account for any topological information. A critical element in neurological and developmental disorders is the topological deterioration and irregularities in streamlines. In this paper, we propose a novel Reeb graph-based method "ReeBundle" that efficiently encodes the topology and geometry of white matter fibers. Given the trajectories of neuronal fiber pathways (neuroanatomical bundle), we re-bundle the streamlines by modeling their spatial evolution to capture geometrically significant events (akin to a fingerprint). ReeBundle parameters control the granularity of the model and handle the presence of improbable streamlines commonly produced by tractography. Further, we propose a new Reeb graph-based distance metric that quantifies topological differences for automated quality control and bundle comparison. We show the practical usage of our method using two datasets: (1) For International Society for Magnetic Resonance in Medicine (ISMRM) dataset, ReeBundle handles the morphology of the white matter tract configurations due to branching and local ambiguities in complicated bundle tracts like anterior and posterior commissures; (2) For the longitudinal repeated measures in the Cognitive Resilience and Sleep History (CRASH) dataset, repeated scans of a given subject acquired weeks apart lead to provably similar Reeb graphs that differ significantly from other subjects, thus highlighting ReeBundle's potential for clinical fingerprinting of brain regions. S. Shailja, Vikram Bhagavatula, Matthew Cieslak, Jean M. Vettel, Scott T. Grafton, B. S. Manjunath |
IEEE Trans. Medical Imaging | 6 |
| 2022 | LOCL: Learning Object-Attribute Composition using Localization
A S. M. Iftekhar, Ekta Prashnani, B. S. Manjunath |
BMVC | 4 |
| 2021 | A Computational Geometry Approach for Modeling Neuronal Fiber Pathways
S. Shailja, Angela Zhang, B. S. Manjunath |
MICCAI (8) | 3 |
| 2021 | StressNet: Detecting Stress in Thermal VideosabstractPrecise measurement of physiological signals is critical for the effective monitoring of human vital signs. Recent developments in computer vision have demonstrated that signals such as pulse rate and respiration rate can be extracted from digital video of humans, increasing the possibility of contact-less monitoring. This paper presents a novel approach to obtaining physiological signals and classifying stress states from thermal video. The proposed network–"StressNet"–features a hybrid emission representation model that models the direct emission and absorption of heat by the skin and underlying blood vessels. This results in an information-rich feature representation of the face, which is used by spatio-temporal network for reconstructing the ISTI ( Initial Systolic Time Interval : a measure of change in cardiac sympathetic activity that is considered to be a quantitative index of stress in humans). The reconstructed ISTI signal is fed into a stress-detection model to detect and classify the individual’s stress state (i.e. stress or no stress). A detailed evaluation demonstrates that Stress-Net achieves estimated the ISTI signal with 95% accuracy and detect stress with average precision of 0.842. A S. M. Iftekhar, Michael Goebel, Tom Bullock, Mary H. MacLean, Michael B. Miller, Tyler Santander, Barry Giesbrecht, Scott T. Grafton, B. S. Manjunath |
WACV | 10 |
| 2020 | VSGNet: Spatial Attention Network for Detecting Human Object Interactions Using Graph ConvolutionsabstractComprehensive visual understanding requires detection frameworks that can effectively learn and utilize object interactions while analyzing objects individually. This is the main objective in Human-Object Interaction (HOI) detection task. In particular, relative spatial reasoning and structural connections between objects are essential cues for analyzing interactions, which is addressed by the proposed Visual-Spatial-Graph Network (VSGNet) architecture. VSGNet extracts visual features from the human-object pairs, refines the features with spatial configurations of the pair, and utilizes the structural connections between the pair via graph convolutions. The performance of VSGNet is thoroughly evaluated using the Verbs in COCO (V-COCO) dataset. Experimental results indicate that VSGNet outperforms state-of-the-art solutions by 8% or 4 mAP. Oytun Ulutan, A S. M. Iftekhar, B. S. Manjunath |
CVPR | 3 |
| 2020 | Vision-Based Gesture Recognition in Human-Robot Teams Using Synthetic DataabstractBuilding successful collaboration between humans and robots requires efficient, effective, and natural communication. Here we study a RGB-based deep learning approach for controlling robots through gestures (e.g., "follow me"). To address the challenge of collecting high-quality annotated data from human subjects, synthetic data is considered for this domain. We contribute a dataset of gestures that includes real videos with human subjects and synthetic videos from our custom simulator. A solution is presented for gesture recognition based on the state-of-the-art I3D model. Comprehensive testing was conducted to optimize the parameters for this model. Finally, to gather insight on the value of synthetic data, several experiments are described that systematically study the properties of synthetic data (e.g., gesture variations, character variety, generalization to new gestures). We discuss practical implications for the design of effective human-robot collaboration and the usefulness of synthetic data for deep learning. Celso de Melo, Brandon Rothrock, Prudhvi Gurram, Oytun Ulutan, B. S. Manjunath |
IROS | 5 |
| 2020 | Deep Remote Sensing Methods for Methane Detection in Overhead Hyperspectral ImageryabstractEffective analysis of hyperspectral imagery is essential for gathering fast and actionable information of large areas affected by atmospheric and green house gases. Existing methods, which process hyperspectral data to detect amorphous gases such as CH4require manual inspection from domain experts and annotation of massive datasets. These methods do not scale well and are prone to human errors due to the plumes’ small pixel-footprint signature. The proposed Hyperspectral Mask-RCNN (H-mrcnn) uses principled statistics, signal processing, and deep neural networks to address these limitations. H-mrcnn introduces fast algorithms to analyze large-area hyper-spectral information and methods to autonomously represent and detect CH4plumes. H-mrcnn processes information by match-filtering sliding windows of hyperspectral data across the spectral bands. This process produces information-rich features that are both effective plume representations and gas concentration analogs. The optimized matched-filtering stage processes spectral data, which is spatially sampled to train an ensemble of gas detectors. The ensemble outputs are fused to estimate a natural and accurate plume mask. Thorough evaluation demonstrates that H-mrcnn matches the manual and experience-dependent annotation process of experts by 85% (IOU). H-mrcnn scales to larger datasets, reduces the manual data processing and labeling time (×12), and produces rapid actionable information about gas plumes. Carlos Torres 0001, Oytun Ulutan, Alana K. Ayasse, Dar A. Roberts, B. S. Manjunath |
WACV | 6 |
| 2020 | Actor Conditioned Attention Maps for Video Action DetectionabstractWhile observing complex events with multiple actors, humans do not assess each actor separately, but infer from the context. The surrounding context provides essential information for understanding actions. To this end, we propose to replace region of interest(RoI) pooling with an attention module, which ranks each spatio-temporal region's relevance to a detected actor instead of cropping. We refer to these as Actor-Conditioned Attention Maps (ACAM), which amplify/dampen the features extracted from the entire scene. The resulting actor-conditioned features focus the model on regions that are relevant to the conditioned actor. For actor localization, we leverage pre-trained object detectors, which transfer better. The proposed model is efficient and our action detection pipeline achieves near real-time performance. Experimental results on AVA 2.1 and JHMDB demonstrate the effectiveness of attention maps, with improvements of 7 mAP on AVA and 4 mAP on JHMDB. Oytun Ulutan, Swati Rallapalli, Mudhakar Srivatsa, Carlos Torres 0001, B. S. Manjunath |
WACV | 5 |
| 2020 | Superpixel Embedding NetworkabstractSuperpixel segmentation is a fundamental computer vision technique that finds application in a multitude of high level computer vision tasks. Most state-of-the-art superpixel segmentation methods are unsupervised in nature and thus cannot fully utilize frequently occurring texture patterns or incorporate multiscale context. In this paper, we show that superpixel segmentation can be improved by leveraging the superior modeling power of deep convolutional autoencoders in a fully unsupervised manner. We pose the superpixel segmentation problem as one of manifold learning where pixels that belong to similar texture patterns are assigned near identical embedding vectors. The proposed deep network is able to learn image-wide and dataset-wide feature patterns and the relationships between them. This knowledge is used to segment and group pixels in a way that is consistent with a more global definition of pattern coherence. Experiments demonstrate that the superpixels obtained from the embeddings learned by the proposed method outperform the state-of-theart superpixel segmentation methods for boundary precision and recall values. Additionally, we find that semantic edges obtained from the superpixel embeddings to be significantly better than the contemporary unsupervised approaches. Utkarsh Gaur, B. S. Manjunath |
IEEE Trans. Image Process. | 2 |
| 2020 | How Do Drivers Allocate Their Potential Attention? Driving Fixation Prediction via Convolutional Neural NetworksabstractThe traffic driving environment is a complex and dynamic changing scene in which drivers have to pay close attention to salient and important targets or regions for safe driving. Modeling drivers' eye movements and attention allocation in traffic driving can also help guiding unmanned intelligent vehicles. However, until now, few studies have modeled drivers' true fixations and allocations while driving. To this end, we collect an eye tracking dataset from a total of 28 experienced drivers viewing 16 traffic driving videos. Based on the multiple drivers' attention allocation dataset, we propose a convolutional-deconvolutional neural network (CDNN) to predict the drivers' eye fixations. The experimental results indicate that the proposed CDNN outperforms the state-of-the-art saliency models and predicts drivers' attentional locations more accurately. The proposed CDNN can predict the major fixation location and shows excellent detection of secondary important information or regions that cannot be ignored during driving if they exist. Compared with the present object detection models in autonomous and assisted driving systems, our human-like driving model does not detect all of the objects appearing in the driving scenes, but it provides the most relevant regions or targets, which can largely reduce the interference of irrelevant scene information. Tao Deng 0002, Long Qin 0002, Thuyen Ngo, B. S. Manjunath |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2019 | Accurate 3D Cell Segmentation Using Deep Features and CRF RefinementabstractWe consider the problem of accurately identifying cell boundaries and labeling individual cells in confocal microscopy images, specifically, 3D image stacks of cells with tagged cell membranes. Precise identification of cell boundaries, their shapes, and quantifying inter-cellular space leads to a better understanding of cell morphogenesis. Towards this, we outline a cell segmentation method that uses a deep neural network architecture to extract a confidence map of cell boundaries, followed by a 3D watershed algorithm and a final refinement using a conditional random field. In addition to improving the accuracy of segmentation compared to other state-of-the-art methods, the proposed approach also generalizes well to different datasets without the need to retrain the network for each dataset. Detailed experimental results are provided, and the source code is available on GitHub1. Jiaxiang Jiang, Po-Yu Kao, Samuel A. Belteton, Daniel Szymanski, B. S. Manjunath |
ICIP | 5 |
| 2019 | Caesar: cross-camera complex activity recognitionabstractDetecting activities from video taken with a single camera is an active research area for ML-based machine vision. In this paper, we examine the next research frontier: near real-time detection of complex activities spanning multiple (possibly wireless) cameras, a capability applicable to surveillance tasks. We argue that a system for such complex activity detection must employ a hybrid design: one in which rule-based activity detection must complement neural network based detection. Moreover, to be practical, such a system must scale well to multiple cameras and have low end-to-end latency. Caesar, our edge computing based system for complex activity detection, provides an extensible vocabulary of activities to allow users to specify complex actions in terms of spatial and temporal relationships between actors, objects, and activities. Caesar converts these specifications to graphs, efficiently monitors camera feeds, partitions processing between cameras and the edge cluster, retrieves minimal information from cameras, carefully schedules neural network invocation, and efficiently matches specification graphs to the underlying data in order to detect complex activities. Our evaluations show that Caesar can reduce wireless bandwidth, on-board camera memory, and detection latency by an order of magnitude while achieving good precision and recall for all complex activities on a public multi-camera dataset. Pradipta Ghosh, Oytun Ulutan, B. S. Manjunath, Kevin S. Chan, Ramesh Govindan |
SenSys | 4 |
| 2019 | Hybrid LSTM and Encoder-Decoder Architecture for Detection of Image ForgeriesabstractWith advanced image journaling tools, one can easily alter the semantic meaning of an image by exploiting certain manipulation techniques such as copy clone, object splicing, and removal, which mislead the viewers. In contrast, the identification of these manipulations becomes a very challenging task as manipulated regions are not visually apparent. This paper proposes a high-confidence manipulation localization architecture that utilizes resampling features, long short-term memory (LSTM) cells, and an encoder-decoder network to segment out manipulated regions from non-manipulated ones. Resampling features are used to capture artifacts, such as JPEG quality loss, upsampling, downsampling, rotation, and shearing. The proposed network exploits larger receptive fields (spatial maps) and frequency-domain correlation to analyze the discriminative characteristics between the manipulated and non-manipulated regions by incorporating the encoder and LSTM network. Finally, the decoder network learns the mapping from low-resolution feature maps to pixel-wise predictions for image tamper localization. With the predicted mask provided by the final layer (softmax) of the proposed architecture, end-to-end training is performed to learn the network parameters through back-propagation using the ground-truth masks. Furthermore, a large image splicing dataset is introduced to guide the training process. The proposed method is capable of localizing image manipulations at the pixel level with high precision, which is demonstrated through rigorous experimentation on three diverse datasets. Jawadul H. Bappy, Cody Simons, Lakshmanan Nataraj, B. S. Manjunath, Amit K. Roy-Chowdhury |
IEEE Trans. Image Process. | 4 |
| 2018 | An Order Preserving Bilinear Model for Person Detection in Multi-Modal DataabstractWe propose a new order preserving bilinear framework that exploits low-resolution video for person detection in a multi-modal setting using deep neural networks. In this setting cameras are strategically placed such that less robust sensors, e.g. geophones that monitor seismic activity, are located within the field of views (FOVs) of cameras. The primary challenge is being able to leverage sufficient information from videos where there are less than 40 pixels on targets, while also taking advantage of less discriminative information from other modalities, e.g. seismic. Unlike state-of-the-art methods, our bilinear framework retains spatio-temporal order when computing the vector outer products between pairs of features. Despite the high dimensionality of these outer products, we demonstrate that our order preserving bilinear framework yields better performance than recent orderless bilinear models and alternative fusion methods. Code is available at https://github.com/oulutan/OP-Bilinear-Model. Oytun Ulutan, Benjamin S. Riggan, Nasser M. Nasrabadi, B. S. Manjunath |
WACV | 4 |
| 2018 | A Multiview Multimodal System for Monitoring Patient SleepabstractClinical observations indicate that during critical care at the hospitals, a patient's sleep positioning and motion have a significant effect on recovery rate. Unfortunately, there is no formal medical protocol to record, quantify, and analyze motion of patients. There are very few clinical studies that use manual analysis of sleep poses and motion recordings to support medical benefits of patient positioning and motion monitoring. Manual processes do not scale, are prone to human errors, and put strain on an already taxed healthcare workforce. This study introduces multimodal, multiview motion analysis and summarization for healthcare (MASH). MASH is an autonomous system, which addresses these issues by monitoring healthcare environments and enabling the recording and analysis of patient sleep-pose patterns. MASH uses three RGB-D cameras to monitor patients in a medical intensive care unit (ICU) room. The proposed algorithms estimate pose direction at different temporal resolutions and use keyframes to efficiently represent pose transition dynamics. MASH combines deep features computed from the data with a modified version of hidden Markov model (HMM) to flexibly model pose duration and summarize patient motion. The performance is evaluated in ideal (BC: bright and clear/occlusion-free) and natural (DO: dark and occluded) scenarios at two motion resolutions and in two environments: a mock-up and a medical ICU. The usage of deep features is evaluated and their performance compared with engineered features. Experimental results using deep features in DO scenes increase performance from $\text{86.7}\%$ to $\text{93.6}\%$, while matching the classification performance of engineered features in BC scenes. The performance of MASH is compared with HMM and C3D. The overall overtime tracing and summarization error rate across all methods increased when transitioning from the mock-up to the the medical ICU data. The proposed keyframe estimation helps achieve a $\text{78}\%$ transition classification accuracy. Carlos Torres 0001, Jeffrey C. Fried, Kenneth Rose, B. S. Manjunath |
IEEE Trans. Multim. | 4 |
| 2017 | Exploiting Spatial Structure for Localizing Manipulated Image RegionsabstractThe advent of high-tech journaling tools facilitates an image to be manipulated in a way that can easily evade state-of-the-art image tampering detection approaches. The recent success of the deep learning approaches in different recognition tasks inspires us to develop a high confidence detection framework which can localize manipulated regions in an image. Unlike semantic object segmentation where all meaningful regions (objects) are segmented, the localization of image manipulation focuses only the possible tampered region which makes the problem even more challenging. In order to formulate the framework, we employ a hybrid CNN-LSTM model to capture discriminative features between manipulated and non-manipulated regions. One of the key properties of manipulated regions is that they exhibit discriminative features in boundaries shared with neighboring non-manipulated pixels. Our motivation is to learn the boundary discrepancy, i.e., the spatial structure, between manipulated and non-manipulated regions with the combination of LSTM and convolution layers. We perform end-to-end training of the network to learn the parameters through back-propagation given ground-truth mask information. The overall framework is capable of detecting different types of image manipulations, including copy-move, removal and splicing. Our model shows promising results in localizing manipulated regions, which is demonstrated through rigorous experimentation on three diverse datasets. Jawadul H. Bappy, Amit K. Roy-Chowdhury, Jason Bunk, Lakshmanan Nataraj, B. S. Manjunath |
ICCV | 5 |
| 2017 | Weakly Supervised Manifold Learning for Dense Semantic Object CorrespondenceabstractThe goal of the semantic object correspondence problem is to compute dense association maps for a pair of images such that the same object parts get matched for very different appearing object instances. Our method builds on the recent findings that deep convolutional neural networks (DCNNs) implicitly learn a latent model of object parts even when trained for classification. We also leverage a key correspondence problem insight that the geometric structure between object parts is consistent across multiple object instances. These two concepts are then combined in the form of a novel optimization scheme. This optimization learns a feature embedding by rewarding for projecting features closer on the manifold if they have low feature-space distance. Simultaneously, the optimization penalizes feature clusters whose geometric structure is inconsistent with the observed geometric structure of object parts. In this manner, by accounting for feature space similarities and feature neighborhood context together, a manifold is learned where features belonging to semantically similar object parts cluster together. We also describe transferring these embedded features to the sister tasks of semantic keypoint classification and localization task via a Siamese DCNN. We provide qualitative results on the Pascal VOC 2012 images and quantitative results on the Pascal Berkeley dataset where we improve on the state of the art by over 5% on classification and over 9% on localization tasks. Utkarsh Gaur, B. S. Manjunath |
ICCV | 2 |
| 2017 | Saccade gaze prediction using a recurrent neural networkabstractWe present a model that generates close-to-human gaze sequences for a given image in the free viewing task. The proposed approach leverages recent advances in image recognition using convolutional neural networks and sequence modeling with recurrent neural networks. Feature maps from convolutional neural networks are used as inputs to a recurrent neural network. The recurrent neural network acts like a visual working memory that integrates the scene information and outputs a sequence of saccades. The model is trained end-to-end with real-world human eye-tracking data using back propagation and adaptive stochastic gradient descent. Overall, the proposed model is simple compared to the state-of-the-art methods while offering favorable performance on a standard eye-tracking data set. Thuyen Ngo, B. S. Manjunath |
ICIP | 2 |
| 2017 | Beyond Spatial Auto-Regressive Models: Predicting Housing Prices with Satellite ImageryabstractWhen modeling geo-spatial data, it is critical to capture spatial correlations for achieving high accuracy. Spatial Auto-Regression (SAR) is a common tool used to model such data, where the spatial contiguity matrix (W) encodes thespatial correlations. However, the efficacy of SAR is limited by two factors. First, it depends on the choice of contiguity matrix, which is typically not learnt from data, but instead, is assumed to be known apriori. Second, it assumes that the observations can be explained by linear models. In this paper, we propose a Convolutional Neural Network (CNN) framework to model geo-spatial data (specifically housing prices), to learn the spatial correlations automatically. We show that neighborhood information embedded in satellite imagery can be leveraged to achieve the desired spatial smoothing. An additional upside of our framework is the relaxation of linear assumption on the data. Specific challenges we tackle while implementing our framework include, (i) how much of the neighborhood is relevant while estimating housing prices? (ii) what is the right approach to capture multiple resolutions of satellite imagery? and (iii) what other data-sources can help improve the estimation of spatial correlations? We demonstrate a marked improvement of 57% on top of the SAR baseline through the use of features from deep neural networks for the cities of London, Birmingham and Liverpool. Archith J. Bency, Swati Rallapalli, Raghu K. Ganti, Mudhakar Srivatsa, B. S. Manjunath |
WACV | 5 |
| 2017 | Search Tracker: Human-Derived Object Tracking in the Wild Through Large-Scale Search and RetrievalabstractHumans use context and scene knowledge to easily localize moving objects in conditions of complex illumination changes, scene clutter, and occlusions. In this paper, we present a method to leverage human knowledge in the form of annotated video libraries in a novel search and retrieval-based setting to track objects in unseen video sequences. For every video sequence, a document that represents motion information is generated. Documents of the unseen video are queried against the library at multiple scales to find videos with similar motion characteristics. This provides us with coarse localization of objects in the unseen video. We further adapt these retrieved object locations to the new video using an efficient warping scheme. The proposed method is validated on in-the-wild video surveillance data sets where we outperform state-of-the-art appearance-based trackers. We also introduce a new challenging data set with complex object appearance changes. Archith J. Bency, S. Karthikeyan 0001, Carter De Leo, Santhoshkumar Sunderrajan, B. S. Manjunath |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2016 | UAVSensor Fusion with Latent-Dynamic Conditional Random Fields in Coronal Plane EstimationabstractWe present a real-time body orientation estimation in a micro-Unmanned Air Vehicle video stream. This work is part ofafully autonomous UAVsystem which can maneuver to face a single individual in challenging outdoor environments. Our body orientation estimation consists of the following steps: (a) obtaining a set ofvisual appearance models for each body orientation, where each model is tagged with a set of scene information (obtained from sensors), (b) exploiting the mutual information of on-board sensors using latent-dynamic conditional random fields (WCRF), (c) Characterizing each visual appearance model with the most discriminative sensor information, (d) fast estimation ofbody orientation during the test flights given theWCRF parameters and the corresponding sensor readings. The key aspects of our approach is to add sparsity to the sensor readings with latent variables followed by long range dependency analysis. Experimental results obtained over real-time video streams demonstrate a significant improvement in both speed (l5-fps) and accuracy (72%) compared to the state of the art techniques that only rely on visual data. Video demonstration ofour autonomous flights (both from ground view and aerial view) are included in the supplementary material. Amir M. Rahimi, Raphael Ruschel, B. S. Manjunath |
CVPR | 3 |
| 2016 | Weakly Supervised Localization Using Deep Feature Maps
Archith J. Bency, Heesung Kwon, Hyungtae Lee, S. Karthikeyan 0001, B. S. Manjunath |
ECCV (1) | 5 |
| 2016 | Membrane segmentation via active learning with deep networksabstractSegmentation is a key component of several bio-medical image processing systems. Recently, segmentation methods based on supervised learning such as deep convolutional networks have enjoyed immense success for natural image datasets and biological datasets alike. These methods require large volumes of data to avoid overfitting which limits their applicability. In this work, we present a transfer learning mechanism based on active learning which allows us to utilize pre-trained deep networks for segmenting new domains with limited labelled data. We introduce a novel optimization criterion to allow feedback on the most uncertain, yet abundant image patterns thus provisioning for an expert in the loop albeit with minimum amount of guidance. Our experiments demonstrate the effectiveness of the proposed method in improving segmentation performance with very limited labelled data. Utkarsh Gaur, Matthew Kourakis, Erin Newman-Smith, William Smith 0001, B. S. Manjunath |
ICIP | 5 |
| 2016 | Eye-CU: Sleep pose classification for healthcare using multimodal multiview dataabstractManual analysis of body poses of bed-ridden patients requires staff to continuously track and record patient poses. Two limitations in the dissemination of pose-related therapies are scarce human resources and unreliable automated systems. This work addresses these issues by introducing a new method and a new system for robust automated classification of sleep poses in an Intensive Care Unit (ICU) environment. The new method, coupled-constrained Least-Squares (cc-LS), uses multimodal and multiview (MM) data and finds the set of modality trust values that minimizes the difference between expected and estimated labels. The new system, Eye-CU, is an affordable multi-sensor modular system for unobtrusive data collection and analysis in healthcare. Experimental results indicate that the performance of cc-LS matches the performance of existing methods in ideal scenarios. This method outperforms the latest techniques in challenging scenarios by 13% for those with poor illumination and by 70% for those with both poor illumination and occlusions. Results also show that a reduced Eye-CU configuration can classify poses without pressure information with only a slight drop in its performance. Carlos Torres 0001, Victor Fragoso, Scott D. Hammond, Jeffrey C. Fried, B. S. Manjunath |
WACV | 5 |
| 2016 | CellECT: cell evolution capturing toolabstractBACKGROUND: Robust methods for the segmentation and analysis of cells in 3D time sequences (3D+t) are critical for quantitative cell biology. While many automated methods for segmentation perform very well, few generalize reliably to diverse datasets. Such automated methods could significantly benefit from at least minimal user guidance. Identification and correction of segmentation errors in time-series data is of prime importance for proper validation of the subsequent analysis. The primary contribution of this work is a novel method for interactive segmentation and analysis of microscopy data, which learns from and guides user interactions to improve overall segmentation. RESULTS: We introduce an interactive cell analysis application, called CellECT, for 3D+t microscopy datasets. The core segmentation tool is watershed-based and allows the user to add, remove or modify existing segments by means of manipulating guidance markers. A confidence metric learns from the user interaction and highlights regions of uncertainty in the segmentation for the user's attention. User corrected segmentations are then propagated to neighboring time points. The analysis tool computes local and global statistics for various cell measurements over the time sequence. Detailed results on two large datasets containing membrane and nuclei data are presented: a 3D+t confocal microscopy dataset of the ascidian Phallusia mammillata consisting of 18 time points, and a 3D+t single plane illumination microscopy (SPIM) dataset consisting of 192 time points. Additionally, CellECT was used to segment a large population of jigsaw-puzzle shaped epidermal cells from Arabidopsis thaliana leaves. The cell coordinates obtained using CellECT are compared to those of manually segmented cells. CONCLUSIONS: CellECT provides tools for convenient segmentation and analysis of 3D+t membrane datasets by incorporating human interaction into automated algorithms. Users can modify segmentation results through the help of guidance markers, and an adaptive confidence metric highlights problematic regions. Segmentations can be propagated to multiple time points, and once a segmentation is available for a time sequence cells can be analyzed to observe trends. The segmentation and analysis tools presented here generalize well to membrane or cell wall volumetric time series datasets. Diana L. Delibaltov, Utkarsh Gaur, Jennifer Kim, Matthew Kourakis, Erin Newman-Smith, William Smith 0001, Samuel A. Belteton, Daniel Szymanski, B. S. Manjunath |
BMC Bioinform. | 9 |
| 2016 | Context-Aware Hypergraph Modeling for Re-identification and SummarizationabstractTracking and re-identification in wide-area camera networks is a challenging problem due to non-overlapping visual fields, varying imaging conditions, and appearance changes. We consider the problem of person re-identification and tracking, and propose a novel clothing context-aware color extraction method that is robust to such changes. Annotated samples are used to learn color drift patterns in a non-parametric manner using the random forest distance (RFD) function. The color drift patterns are automatically transferred to associate objects across different views using a unified graph matching framework . A hypergraph representation is used to link related objects for search and re-identification. A diverse hypergraph ranking technique is proposed for person-focused network summarization . The proposed algorithm is validated on a wide-area camera network consisting of ten cameras on bike paths. Also, the proposed algorithm is compared with the state of the art person re-identification algorithms on the VIPeR dataset . Santhoshkumar Sunderrajan, B. S. Manjunath |
IEEE Trans. Multim. | 2 |
| 2015 | An end-to-end system for content-based video retrieval using behavior, actions, and appearance with interactive query refinementabstractWe describe a system for content-based retrieval from large surveillance video archives, using behavior, action and appearance of objects. Objects are detected, tracked, and classified into broad categories. Their behavior and appearance are characterized by action detectors and descriptors, which are indexed in an archive. Queries can be posed as video exemplars, and the results can be refined through relevance feedback. The contributions of our system include the fusion of behavior and action detectors with appearance for matching; the improvement of query results through interactive query refinement (IQR), which learns a discriminative classifier online based on user feedback; and reasonable performance on low resolution, poor quality video. The system operates on video from ground cameras and aerial platforms, both RGB and IR. Performance is evaluated on publicly-available surveillance datasets, showing that subtle actions can be detected under difficult conditions, with reasonable improvement from IQR. Anthony Hoogs, A. G. Amitha Perera, Roderic Collins, Arslan Basharat, Keith Fieldhouse, Chuck Atkins, Linus Sherrill, Benjamin Boeckel, Russell Blue, Matthew Woehlke, C. Greco, Zhaohui Sun, Eran Swears, Naresh P. Cuntoor, J. Luck, B. Drew, D. Hanson, D. Rowley, J. Kopaz, T. Rude, D. Keefe, Amit Srivastava, Saurabh Khanwalkar, Chia-Chih Chen, Jake K. Aggarwal, Larry Davis 0001, Yaser Yacoob, Dong Liu 0001, Shih-Fu Chang, Bi Song, Amit K. Roy-Chowdhury, Kenneth Sullivan, Jelena Tesic, Shivkumar Chandrasekaran, B. S. Manjunath, K. Reddy, Mubarak Shah, K. Chang, Tsuhan Chen, Mita Desai |
AVSS | 37 |
| 2015 | Eye tracking assisted extraction of attentionally important objects from videosabstractVisual attention is a crucial indicator of the relative importance of objects in visual scenes to human viewers. In this paper, we propose an algorithm to extract objects which attract visual attention from videos. As human attention is naturally biased towards high level semantic objects in visual scenes, this information can be valuable to extract salient objects. The proposed algorithm extracts dominant visual tracks using eye tracking data from multiple subjects on a video sequence by a combination of mean-shift clustering and Hungarian algorithm. These visual tracks guide a generic object search algorithm to get candidate object locations and extents in every frame. Further, we propose a novel multiple object extraction algorithm by constructing a spatio-temporal mixed graph over object candidates. Bounding box based object extraction inference is performed using binary linear integer programming on a cost function defined over the graph. Finally, the object boundaries are refined using grabcut segmentation. The proposed technique outperforms state-of-the-art video segmentation using eye tracking prior and obtains favorable object extraction over algorithms which do not utilize eye tracking data. S. Karthikeyan 0001, Thuyen Ngo, Miguel P. Eckstein, B. S. Manjunath |
CVPR | 4 |
| 2015 | Weakly Supervised Graph Based Semantic Segmentation by Learning Communities of Image-PartsabstractWe present a weakly-supervised approach to semantic segmentation. The goal is to assign pixel-level labels given only partial information, for example, image-level labels. This is an important problem in many application scenarios where it is difficult to get accurate segmentation or not feasible to obtain detailed annotations. The proposed approach starts with an initial coarse segmentation, followed by a spectral clustering approach that groups related image parts into communities. A community-driven graph is then constructed that captures spatial and feature relationships between communities while a label graph captures correlations between image labels. Finally, mapping the image level labels to appropriate communities is formulated as a convex optimization problem. The proposed approach does not require location information for image level labels and can be trained using partially labeled datasets. Compared to the state-of-the-art weakly supervised approaches, we achieve a significant performance improvement of 9% on MSRC-21 dataset and 11% on LabelMe dataset, while being more than 300 times faster. Niloufar Pourian, S. Karthikeyan 0001, B. S. Manjunath |
ICCV | 3 |
| 2015 | Search and retrieval of multi-modal data associated with image-partsabstractWe present a novel framework for querying multi-modal data from a heterogeneous database containing images, textual tags, and GPS coordinates. We construct a bi-layer graph structure using localized image-parts and associated GPS locations and textual tags from the database. The first layer graphs capture similar data points from a single modality using a spectral clustering algorithm. The second layer of our multi-modal network allows one to integrate the relationships between clusters of different modalities. The proposed network model enables us to use flexible multi-modal queries on the database. Niloufar Pourian, S. Karthikeyan 0001, B. S. Manjunath |
ICIP | 3 |
| 2015 | Features we trust!abstractWe investigate the problem of image classification within a supervised learning framework that exploits implicit mutual information in different visual features and their associated classifiers. In our proposed two stage hierarchical processing, visual features are first clustered with the objective of maximizing diversity. Majority vote within each cluster is used to enforce diversity. Many partitioning variations are evaluated using K-nearest neighbor to obtain the highest inter-cluster entropy. In the second step, a richer measure of discrimination is obtained using a fully connected conditional random fields (CRF) over clusters. The unary and interaction potentials are defined over mutual information within each cluster and inter-dependencies across clusters respectively. Experimenting over five distinct datasets, we demonstrate an average performance gain of 30% compared with state of the art techniques. Amir M. Rahimi, Lakshmanan Nataraj, B. S. Manjunath |
ICIP | 3 |
| 2015 | Sleep Pose Recognition in an ICU Using Multimodal Data and Environmental Feedback
Carlos Torres 0001, Scott D. Hammond, Jeffrey C. Fried, B. S. Manjunath |
ICVS | 4 |
| 2015 | SATTVA: SpArsiTy inspired classificaTion of malware VAriantsabstractThere is an alarming increase in the amount of malware that is generated today. However, several studies have shown that most of these new malware are just variants of existing ones. Fast detection of these variants plays an effective role in thwarting new attacks. In this paper, we propose a novel approach to detect malware variants using a sparse representation framework. Exploiting the fact that most malware variants have small differences in their structure, we model a new/unknown malware sample as a sparse linear combination of other malware in the training set. Lakshmanan Nataraj, S. Karthikeyan 0001, B. S. Manjunath |
IH&MMSec | 3 |
| 2015 | Retrieval of Images with Objects of Specific Size, Location, and Spatial ConfigurationabstractAn approach to image retrieval using spatial configurations is presented. The goal is to search the database for images that contain similar objects (image-patches) with a given configuration, size and position. The proposed approach consists of creating localized representations robust to segmentation variations, and a sub-graph matching method to compare the query with the database items. Localized object representations are created using a community detection method that groups visually similar segments. Extensive experimental results on three challenging datasets are provided to demonstrate the feasibility of the approach. Niloufar Pourian, B. S. Manjunath |
WACV | 2 |
| 2015 | Characterizing spatial distributions of astrocytes in the mammalian retinaabstractMOTIVATION: In addition to being involved in retinal vascular growth, astrocytes play an important role in diseases and injuries, such as glaucomatous neuro-degeneration and retinal detachment. Studying astrocytes, their morphological cell characteristics and their spatial relationships to the surrounding vasculature in the retina may elucidate their role in these conditions. RESULTS: Our results show that in normal healthy retinas, the distribution of observed astrocyte cells does not follow a uniform distribution. The cells are significantly more densely packed around the blood vessels than a uniform distribution would predict. We also show that compared with the distribution of all cells, large cells are more dense in the vicinity of veins and toward the optic nerve head whereas smaller cells are often more dense in the vicinity of arteries. We hypothesize that since veinal astrocytes are known to transport toxic metabolic waste away from neurons they may be more critical than arterial astrocytes and therefore require larger cell bodies to process waste more efficiently. AVAILABILITY AND IMPLEMENTATION: A 1/8th size down-sampled version of the seven retinal image mosaics described in this article can be found on BISQUE (Kvilekval et al., 2010) at http://bisque.ece.ucsb.edu/client_service/view?resource=http://bisque.ece.ucsb.edu/data_service/dataset/6566968. Aruna Jammalamadaka, Panuakdet Suwannatat, Steven K. Fisher, B. S. Manjunath, Tobias Höllerer, Gabriel Luna |
Bioinform. | 4 |
| 2015 | PixNet: A Localized Feature Representation for Classification and Visual SearchabstractThis paper presents a novel localized visual image feature motivated by image segmentation. The proposed feature embeds relative spatial information by learning different image parts while having a compact representation. First, an attributed graph representation of an image is created based on segmentation and localized image features. Subsequently, communities of image regions are discovered based on their spatial and visual characteristics over all images. The community detection problem is modeled as a spectral graph partitioning problem. This results in finding meaningful image part groupings . A histogram of communities forms a robust and spatially localized representation for each image in the database. Such a region-based representation enables one to search for queries that might not have been possible with global image representations. We apply this representation to image classification and search and retrieval tasks. Extensive experiments on three challenging datasets, including the large-scale ImageNet dataset, demonstrate that the proposed representation achieves promising results compared to the current state-of-the-art methods. Niloufar Pourian, B. S. Manjunath |
IEEE Trans. Multim. | 2 |
| 2014 | Synapse classification and localization in Electron Micrographs
Vignesh Jagadeesh, James R. Anderson 0002, Bryan W. Jones, Robert Marc, Steven K. Fisher, B. S. Manjunath |
Pattern Recognit. Lett. | 6 |
| 2014 | Multi-Label Learning With Fused Multimodal Bi-Relational GraphabstractThe problem of multi-label image classification using multiple feature modalities is considered in this work. Given a collection of images with partial labels, we first model the association between different feature modalities and the images labels. These associations are then propagated with a graph diffusion kernel to classify the unlabeled images. Towards this objective, a novel Fused Multimodal Bi-relational Graph representation is proposed, with multiple graphs corresponding to different feature modalities, and one graph corresponding to the image labels. Such a representation allows for effective exploitation of both feature complementariness and label correlation. This contrasts with previous work where these two factors are considered in isolation. Furthermore, we provide a solution to learn the weight for each image graph by estimating the discriminative power of the corresponding feature modality. Experimental results with our proposed method on two standard multi-label image datasets are very promising. Jiejun Xu, Vignesh Jagadeesh, B. S. Manjunath |
IEEE Trans. Multim. | 3 |
| 2014 | Calibrating a wide-area camera network with non-overlapping views using mobile devicesabstractIn a wide-area camera network, cameras are often placed such that their views do not overlap. Collaborative tasks such as tracking and activity analysis still require discovering the network topology including the extrinsic calibration of the cameras. This work addresses the problem of calibrating a fixed camera in a wide-area camera network in a global coordinate system so that the results can be shared across calibrations. We achieve this by using commonly available mobile devices such as smartphones. At least one mobile device takes images that overlap with a fixed camera's view and records the GPS position and 3D orientation of the device when an image is captured. These sensor measurements (including the image, GPS position, and device orientation) are fused in order to calibrate the fixed camera. This article derives a novel maximum likelihood estimation formulation for finding the most probable location and orientation of a fixed camera. This formulation is solved in a distributed manner using a consensus algorithm. We evaluate the efficacy of the proposed methodology with several simulated and real-world datasets. Thomas Kuo, Zefeng Ni, Santhoshkumar Sunderrajan, B. S. Manjunath |
ACM Trans. Sens. Networks | 4 |
| 2014 | Multicamera video summarization and anomaly detection from activity motifsabstractCamera network systems generate large volumes of potentially useful data, but extracting value from multiple, related videos can be a daunting task for a human reviewer. Multicamera video summarization seeks to make this task more tractable by generating a reduced set of output summary videos that concisely capture important portions of the input set. We present a system that approaches summarization at the level of detected activity motifs and shortens the input videos by compacting the representation of individual activities. Additionally, redundancy is removed across camera views by omitting from the summary activity occurrences that can be predicted by other occurrences. The system also detects anomalous events within a unified framework and can highlight them in the summary. Our contributions are a method for selecting useful parts of an activity to present to a viewer using activity motifs and a novel framework to score the importance of activity occurrences and allow transfer of importance between temporally related activities without solving the correspondence problem. We provide summarization results for a two camera network, an eleven camera network, and data from PETS 2001. We also include results from Amazon Mechanical Turk human experiments to evaluate how our visualization decisions affect task performance. Carter De Leo, B. S. Manjunath |
ACM Trans. Sens. Networks | 2 |
| 2013 | SigMal: a static signal processing based malware triageabstractIn this work, we propose SigMal, a fast and precise malware detection framework based on signal processing techniques. SigMal is designed to operate with systems that process large amounts of binary samples. It has been observed that many samples received by such systems are variants of previously-seen malware, and they retain some similarity at the binary level. Previous systems used this notion of malware similarity to detect new variants of previously-seen malware. SigMal improves the state-of-the-art by leveraging techniques borrowed from signal processing to extract noise-resistant similarity signatures from the samples. SigMal uses an efficient nearest-neighbor search technique, which is scalable to millions of samples. We evaluate SigMal on 1.2 million recent samples, both packed and unpacked, observed over a duration of three months. In addition, we also used a constant dataset of known benign executables. Our results show that SigMal can classify 50% of the recent incoming samples with above 99% precision. We also show that SigMal could have detected, on average, 70 malware samples per day before any antivirus vendor detected them. Dhilung Kirat, Lakshmanan Nataraj, Giovanni Vigna, B. S. Manjunath |
ACSAC | 4 |
| 2013 | From Where and How to What We SeeabstractEye movement studies have confirmed that overt attention is highly biased towards faces and text regions in images. In this paper we explore a novel problem of predicting face and text regions in images using eye tracking data from multiple subjects. The problem is challenging as we aim to predict the semantics (face/text/background) only from eye tracking data without utilizing any image information. The proposed algorithm spatially clusters eye tracking data obtained in an image into different coherent groups and subsequently models the likelihood of the clusters containing faces and text using a fully connected Markov Random Field (MRF). Given the eye tracking data from a test image, it predicts potential face/head (humans, dogs and cats) and text locations reliably. Furthermore, the approach can be used to select regions of interest for further analysis by object detectors for faces and text. The hybrid eye position/object detector approach achieves better detection performance and reduced computation time compared to using only the object detection algorithm. We also present a new eye tracking dataset on 300 images selected from ICDAR, Street-view, Flickr and Oxford-IIIT Pet Dataset from 15 subjects. S. Karthikeyan 0001, Vignesh Jagadeesh, Renuka Shenoy, Miguel P. Eckstein, B. S. Manjunath |
ICCV | 5 |
| 2013 | Camera Alignment Using Trajectory Intersections in Unsynchronized VideosabstractThis paper addresses the novel and challenging problem of aligning camera views that are unsynchronized by low and/or variable frame rates using object trajectories. Unlike existing trajectory-based alignment methods, our method does not require frame-to-frame synchronization. Instead, we propose using the intersections of corresponding object trajectories to match views. To find these intersections, we introduce a novel trajectory matching algorithm based on matching Spatio-Temporal Context Graphs (STCGs). These graphs represent the distances between trajectories in time and space within a view, and are matched to an STCG from another view to find the corresponding trajectories. To the best of our knowledge, this is one of the first attempts to align views that are unsynchronized with variable frame rates. The results on simulated and real-world datasets show trajectory intersections are a viable feature for camera alignment, and that the trajectory matching method performs well in real-world scenarios. Thomas Kuo, Santhoshkumar Sunderrajan, B. S. Manjunath |
ICCV | 3 |
| 2013 | Learning top down scene context for visual attention modeling in natural imagesabstractTop down image semantics play a major role in predicting where people look in images. Present state-of-the-art approaches to model human visual attention incorporate high level object detections signifying top down image semantics in a separate channel along with other bottom up saliency channels. However, multiple objects in a scene are competing to attract our attention and this interaction is ignored in current models. To overcome this limitation, we propose a novel object context based visual attention model which incorporates the co-occurrence of multiple objects in a scene for visual attention modeling. The proposed regression based algorithm uses several high level object detectors for faces, people, cars, text and understands how their joint presence affects visual attention. Experimental results on the MIT eye tracking dataset demonstrates that the proposed method outperforms other state-of-the-art visual attention models. Shanmugavadivel Karthikeyan, Vignesh Jagadeesh, B. S. Manjunath |
ICIP | 3 |
| 2013 | Learning bottom-up text attention maps for text detection using stroke width transformabstractHumans have a remarkable ability to quickly discern regions containing text from other noisy regions in images. The primary contribution of this paper is to learn a model to mimic this behavior and aid text detection algorithms. The proposed approach utilizes multiple low level visual features which signify visually salient regions and learns a model to eventually provide a text attention map which indicates potential text regions in images. In the next stage, a text detector using stroke width transform only focusses on these selective image regions achieving dual benefits of reduced computation time and better detection performance. Experimental results on the ICDAR 2003 text detection dataset demonstrate that the proposed method outperforms the baseline implementation of stroke width transform, and the generated text attention maps compare favorably with human fixation maps on text images. S. Karthikeyan 0001, Vignesh Jagadeesh, B. S. Manjunath |
ICIP | 3 |
| 2013 | Robust multiple object tracking by detection with interacting Markov chain Monte CarloabstractThis paper presents a novel and computationally efficient multi-object tracking-by-detection algorithm with interacting particle filters. The proposed online tracking methodology could be scaled to hundreds of objects and could be completely parallelized. For every object, we have a set of two particle filters, i.e. local and global. The local particle filter models the local motion of the object. The global particle filter models the interaction with the other objects and scene. These particle filters are integrated into a unified Interacting Markov Chain Monte Carlo (IMCMC) framework. The local particle filter improves its performance by interacting with the global particle filter while they both are run in parallel. We indicate the manner in which we bring in object interaction and domain specific information into account by using global filters without further increase in complexity. Most importantly, the complexity of the proposed methodology varies linearly in the number of objects. We validated the proposed algorithms on two completely different domains 1) Pedestrian Tracking in urban scenarios 2) Biological cell tracking (Melanosomes). The proposed algorithm is found to yield favorable results compared to the existing algorithms. Santhoshkumar Sunderrajan, S. Karthikeyan 0001, B. S. Manjunath |
ICIP | 3 |
| 2013 | A Linear Program Formulation for the Segmentation of Ciona Membrane Volumes
Diana L. Delibaltov, Pratim Ghosh, Volkan Rodoplu, Michael Veeman, William Smith 0001, B. S. Manjunath |
MICCAI (1) | 6 |
| 2013 | Statistical Analysis of Dendritic Spine Distributions in Rat Hippocampal CulturesabstractBACKGROUND: Dendritic spines serve as key computational structures in brain plasticity. Much remains to be learned about their spatial and temporal distribution among neurons. Our aim in this study was to perform exploratory analyses based on the population distributions of dendritic spines with regard to their morphological characteristics and period of growth in dissociated hippocampal neurons. We fit a log-linear model to the contingency table of spine features such as spine type and distance from the soma to first determine which features were important in modeling the spines, as well as the relationships between such features. A multinomial logistic regression was then used to predict the spine types using the features suggested by the log-linear model, along with neighboring spine information. Finally, an important variant of Ripley's K-function applicable to linear networks was used to study the spatial distribution of spines along dendrites. RESULTS: Our study indicated that in the culture system, (i) dendritic spine densities were "completely spatially random", (ii) spine type and distance from the soma were independent quantities, and most importantly, (iii) spines had a tendency to cluster with other spines of the same type. CONCLUSIONS: Although these results may vary with other systems, our primary contribution is the set of statistical tools for morphological modeling of spines which can be used to assess neuronal cultures following gene manipulation such as RNAi, and to study induced pluripotent stem cells differentiated to neurons. Aruna Jammalamadaka, Sourav Banerjee, Kenneth S. Kosik, B. S. Manjunath |
BMC Bioinform. | 4 |
| 2013 | Robust Simultaneous Registration and Segmentation with Sparse Error ReconstructionabstractWe introduce a fast and efficient variational framework for Simultaneous Registration and Segmentation (SRS) applicable to a wide variety of image sequences. We demonstrate that a dense correspondence map (between consecutive frames) can be reconstructed correctly even in the presence of partial occlusion, shading, and reflections. The errors are efficiently handled by exploiting their sparse nature. In addition, the segmentation functional is reformulated using a dual Rudin-Osher-Fatemi (ROF) model for fast implementation. Moreover, nonparametric shape prior terms that are suited for this dual-ROF model are proposed. The efficacy of the proposed method is validated with extensive experiments on both indoor, outdoor natural and biological image sequences, demonstrating the higher accuracy and efficiency compared to various state-of-the-art methods. Pratim Ghosh, B. S. Manjunath |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Robust Segmentation Based Tracing Using an Adaptive Wrapper for Inducing PriorsabstractSegmentation based tracing algorithms detect the extent and borders of an object in a given frame IZ by propagating results from frames I1 ≤ z < Z. Although application specific tracers have been forthcoming, techniques that automatically adapt across applications have been less explored. We approach this problem by learning a prior model on topological dynamics that encourages segmentation transitions across frames that are most likely for a given application. Further, we augment a generic tracing technique with a locality sensitive prior derived from dense optic flow fields for deformation guidance. The proposed approach comprises two stages where the generic tracer initially yields multiple segmentation transitions when its parameters are perturbed, and the learnt topology prior subsequently propagates high scoring segmentations. Because the learnt topology model wraps around a generic tracer and adapts it by setting its free parameters, the need for careful parameter tuning is completely obviated. Through extensive experimental validation in surveillance, biological and medical image datasets, we verify the applicability of the proposed model while demonstrating good tracing performance under severe clutter. Vignesh Jagadeesh, B. S. Manjunath, Bryan W. Jones, Robert Marc, Steven K. Fisher |
IEEE Trans. Image Process. | 2 |
| 2013 | Graph-Based Topic-Focused Retrieval in Distributed Camera NetworkabstractWide-area wireless camera networks are being increasingly deployed in many urban scenarios. The large amount of data generated from these cameras pose significant information processing challenges. In this work, we focus on representation, search and retrieval of moving objects in the scene, with emphasis on local camera node video analysis. We develop a graph model that captures the relationships among objects without the need to identify global trajectories. Specifically, two types of edges are defined in the graph: object edges linking the same object across the whole network and context edges linking different objects within a spatial-temporal proximity. We propose a manifold ranking method with a greedy diversification step to order the relevant items based on similarity as well as diversity within the database. Detailed experimental results using video data from a 10-camera network covering bike paths are presented. Jiejun Xu, Vignesh Jagadeesh, Zefeng Ni, Santhoshkumar Sunderrajan, B. S. Manjunath |
IEEE Trans. Multim. | 5 |
| 2012 | Unified hypergraph for image ranking in a multimodal contextabstractImage ranking has long been studied, yet it remains a very challenging problem. Increasingly, online images come with additional metadata such as user annotations and geographic coordinates. They provide rich complementary information. We propose to combine such multimodal information through a unified hypergraph to improve image retrieval performance. Hypergraphs allow for the simultaneously capture of higher order relationships among images using different modalities, e.g. visual content, user tags, and geolocations. Each image is represented as a vertex in the hypergraph. Each hyperedge is formed by a vertex and it's k-nearest neighbors. Three types of hyperedges exist in our unified hypergraph, which are in correspondence to the three different modalities. Image ranking is then formulated as a ranking problem on a unified hypergraph. The proposed method can easily be extended to incorporate additional modalities as long as a similarity function exists to compare the features. Experimental results on large datasets are promising. Jiejun Xu, Vishwakarma Singh, Ziyu Guan, B. S. Manjunath |
ICASSP | 4 |
| 2012 | Online parameter estimation in dynamic Markov Random Fields for image sequence analysisabstractMarkov Random Fields (MRF) have proven to be extremely useful models for efficient and accurate image segmentation.Recent literature points to an increased effort towards incorporating useful priors (shape, geometry, context) in a MRF framework. However, topological priors, considered extremely crucial in biological and natural image sequences have been less explored. This work proposes a strategy wherein free parameters of the MRF are used to make it topology aware using a semantic graphical model working in conjunction with the MRF. Estimation of free parameters is constrained by prior knowledge of an object's topological dynamics encoded by the graphical model. Maximizing a regional conformance measure yields parameters for the frame under consideration. The application motivating this work is the tracing of neuronal structures across 3D serial section Transmission Electron Micrograph (ssTEM) stacks. Applicability of the proposed method is demonstrated by tracing 3D structures in ssTEM stacks. Vignesh Jagadeesh, B. S. Manjunath, Bryan W. Jones, Robert Marc, Steven K. Fisher |
ICIP | 2 |
| 2012 | Unified probabilistic framework for simultaneous detection and tracking of multiple objects with application to bio-image sequencesabstractWe present a detection based tracking algorithm for tracking melanosomes (organelles containing melanin) in time lapse image sequences imaged using bright field microscopy. Due to heavy imaging noise detecting all the melanosomes accurately in every frame is difficult. Therefore, two sets of imperfect detections are used in a unified probabilistic approach to simultaneously perform melanosome detection and tracking. We propose a novel iterative algorithm which jointly estimates the optimal set of detections and track results in every iteration from the previous tracks and detections. Our algorithm obtains significantly better tracking results than the state of the art tracking-by-detection algorithm. S. Karthikeyan 0001, Diana L. Delibaltov, Utkarsh Gaur, Mei Jiang, B. S. Manjunath |
ICIP | 6 |
| 2012 | Intra-class multi-output regression based subspace analysisabstractA common challenge when dealing with heterogenous tasks such as face expression analysis, face and object recognition is high dimensionality and extreme appearance variations within each class. To handle such scenarios, we formulate a supervised Non-negative Matrix Factorization (NMF) based subspace learning technique that simultaneously preserves the intra-class regression information (local) and enhances inter-class discrimination (global) in the low dimensional embedding. Our method leverages the multi-dimensional image labels that quantify the within class regression to learn the subspaces for recognition. In addition, our formulation includes a novel multi-output regression based NMF algorithm. Shanmugavadivel Karthikeyan, Swapna Joshi, B. S. Manjunath, Scott T. Grafton |
ICIP | 3 |
| 2012 | Scalable Tracing of Electron Micrographs by Fusing Top Down and Bottom Up Cues Using Hypergraph Diffusion
Vignesh Jagadeesh, Min-Chi Shih, B. S. Manjunath, Kenneth Rose |
MICCAI (3) | 3 |
| 2011 | Generalized subspace based high dimensional density estimationabstractOur paper presents a novel high dimensional probability density estimation technique using any dimensionality reduction method. Our method first performs subspace reduction using any matrix factorization algorithm and estimates the density in the low-dimensional space using sample-point variable bandwidth kernel density estimation. Subsequently, the high dimensional density is approximated from the low dimensional density parameters. The reconstruction error due to dimensionality reduction process is also modeled in a principled and efficient manner to obtain the high dimensional density estimate. We show the effectiveness of our technique by using two popular dimensionality reduction tools, principal component analysis and non-negative matrix factorization. This technique is applied to AT&T, Yale, Pointing'04 and CMU-PIE face recognition datasets and improved performance compared to other dimensionality reduction and density estimation algorithms is obtained. Shanmugavadivel Karthikeyan, Mehmet Emre Sargin, Swapna Joshi, B. S. Manjunath, Scott T. Grafton |
ICIP | 4 |
| 2011 | Multiple Structure Tracing in 3D Electron Micrographs
Vignesh Jagadeesh, Nhat Vu, B. S. Manjunath |
MICCAI (1) | 3 |
| 2011 | Malware images: visualization and automatic classificationabstractWe propose a simple yet effective method for visualizing and classifying malware using image processing techniques. Malware binaries are visualized as gray-scale images, with the observation that for many malware families, the images belonging to the same family appear very similar in layout and texture. Motivated by this visual similarity, a classification method using standard image features is proposed. Neither disassembly nor code execution is required for classification. Preliminary experimental results are quite promising with 98% classification accuracy on a malware database of 9,458 samples with 25 different malware families. Our technique also exhibits interesting resilience to popular obfuscation techniques such as section encryption. Lakshmanan Nataraj, S. Karthikeyan 0001, G. Jacob, B. S. Manjunath |
VizSEC | 4 |
| 2011 | Variable Length Open Contour Tracking Using a Deformable TrellisabstractThis paper focuses on contour tracking, an important problem in computer vision, and specifically on open contours that often directly represent a curvilinear object. Compelling applications are found in the field of bioimage analysis where blood vessels, dendrites, and various other biological structures are tracked over time. General open contour tracking, and biological images in particular, pose major challenges including scene clutter with similar structures (e.g., in the cell), and time varying contour length due to natural growth and shortening phenomena, which have not been adequately answered by earlier approaches based on closed and fixed end-point contours. We propose a model-based estimation algorithm to track open contours of time-varying length, which is robust to neighborhood clutter with similar structures. The method employs a deformable trellis in conjunction with a probabilistic (hidden Markov) model to estimate contour position, deformation, growth and shortening. It generates a maximum a posteriori estimate given observations in the current frame and prior contour information from previous frames. Experimental results on synthetic and real-world data demonstrate the effectiveness and performance gains of the proposed algorithm. Mehmet Emre Sargin, Alphan Altinok, B. S. Manjunath, Kenneth Rose |
IEEE Trans. Image Process. | 3 |
| 2010 | Generalized simultaneous registration and segmentationabstractSimultaneous registration and segmentation (SRS) provides a powerful framework for tracking an object of interest in an image sequence. The state-of-the-art SRS-based tracking methods assume that the illumination is maintained constant across consecutive frames. However, this assumption does not hold in many natural image sequences due to dynamic light source and shadows. We propose a generalized model for SRS-based tracking in this paper to account for non-uniform additive illumination changes. More specifically, we introduce two new terms in the SRS energy functional which address the above mentioned problem. The first term couples the shape-based cue and intensity-based cue to establish a correspondence between them. The second term compensates for the illumination change which is complementary to the first term. We demonstrate that the proposed SRS energy functional yields superior performance over the state-of-the-art SRS-based methods for various indoor and outdoor image sequences. Pratim Ghosh, Mehmet Emre Sargin, B. S. Manjunath |
CVPR | 3 |
| 2010 | Anatomical parts-based regression using non-negative matrix factorizationabstractNon-negative matrix factorization (NMF) is an excellent tool for unsupervised parts-based learning, but proves to be ineffective when parts of a whole follow a specific pattern. Analyzing such local changes is particularly important when studying anatomical transformations. We propose a supervised method that incorporates a regression constraint into the NMF framework and learns maximally changing parts in the basis images, called Regression based NMF (RNMF). The algorithm is made robust against outliers by learning the distribution of the input manifold space, where the data resides. One of our main goals is to achieve good region localization. By incorporating a gradient smoothing and independence constraint into the factorized bases, contiguous local regions are captured. We apply our technique to a synthetic dataset and structural MRI brain images of subjects with varying ages. RNMF finds the localized regions which are expected to be highly changing over age to be manifested in its significant basis and it also achieves the best performance compared to other statistical regression and dimensionality reduction techniques. Swapna Joshi, Shanmugavadivel Karthikeyan, B. S. Manjunath, Scott T. Grafton, Kent A. Kiehl |
CVPR | 3 |
| 2010 | Use of imperfectly segmented nuclei in the classification of histopathology images of breast cancerabstractMany features used in the analysis of pathology imagery are inspired by grading features defined by clinical pathologists as important for diagnosis and characterization. A large majority of these features are features of cell nuclei; as such, there is often the desire to segment the imagery into individual nuclei prior to feature extraction and further analysis. In this paper we present an analysis of the utility of imperfectly segmented cell nuclei for classification of H&E stained histopathology imagery of breast tissue. We show the object- and image-level classification performance using these imperfectly segmented nuclei in a benign versus malignant decision. Results indicate that very good classification accuracies can be achieved with imperfectly segmented nuclei and further that perfect nuclei segmentation does not necessarily guarantee better classification accuracy. Laura E. Boucheron, B. S. Manjunath, Neal R. Harvey |
ICASSP | 2 |
| 2010 | Interactive graph cut segmentation of touching neuronal structures from electron micrographsabstractA novel interactive segmentation framework comprising of a two stage s-t mincut is proposed. The framework has been designed keeping in mind the need to segment touching neuronal structures in Electron Micrograph (EM) images. The first stage undersegments the image, and groups touching structures into a single class. The second stage accepts user interaction to separate touching structures. The technique introduces user feedback through a Markov Random Field formulation. Furthermore, a method for constructing interaction potentials using an edge response function is proposed. Encouraging results, and a comparison to state of the art methods is presented. Vignesh Jagadeesh, B. S. Manjunath |
ICIP | 2 |
| 2010 | Precise localization of key-points to identify local regions for robust data hidingabstractWe propose a novel data hiding system where data is embedded in local non-overlapping regions in an image. To survive cropping, the encoder embeds the same data in multiple regions of fixed dimensions, while the decoder's challenge is to independently retrieve these regions. Salient feature points are computed on an image and the local regions are centered around them. To obtain non-overlapping regions, the points are pruned based on their corner strength and the size of the region. The decoder can retrieve the data only if it can precisely identify one or more of the same key-points. We present suitable key-point pruning methods such that even after considering a reduced number of key-points, the receiver is successful in identifying the same key-point locations as the encoder. We perform experimental comparison of various corner detectors and also study the performance of segmentation methods to obtain robust key-points. Lakshmanan Nataraj, Anindya Sarkar, B. S. Manjunath |
ICIP | 3 |
| 2010 | Distributed particle filter tracking with online multiple instance learning in a camera sensor networkabstractThis paper proposes a distributed algorithm for object tracking in a camera sensor network. At each camera node, an efficient online multiple instance learning algorithm is used to model object's appearance. This is integrated with particle filter for camera's image plane tracking. To improve the tracking accuracy, each camera node shares its particle states with others and fuses multi-camera information locally. In particular, particle weights are updated according to the fused information. Then, appearance model is updated with the re-weighted particles. The effectiveness of the proposed algorithm is demonstrated on human tracking in challenging environments. Zefeng Ni, Santhoshkumar Sunderrajan, Amir M. Rahimi, B. S. Manjunath |
ICIP | 4 |
| 2010 | Discriminative Basis Selection Using Non-negative Matrix FactorizationabstractNon-negative matrix factorization (NMF) has proven to be useful in image classification applications such as face recognition. We propose a novel discriminative basis selection method for classification of image categories based on the popular term frequency-inverse document frequency (TF-IDF) weight used in information retrieval. We extend the algorithm to incorporate color, and overcome the drawbacks of using unaligned images. Our method is able to choose visually significant bases which best discriminate between categories and thus prune the classification space to increase correct classifications. We apply our technique to ETH-80, a standard image classification benchmark dataset. Our results show that our algorithm outperforms other state-of-the-art techniques. Aruna Jammalamadaka, Swapna Joshi, Shanmugavadivel Karthikeyan, B. S. Manjunath |
ICPR | 4 |
| 2010 | Graphical Model-Based Tracking of Curvilinear Structures in Bio-image SequencesabstractTracking of curvilinear structures is a task of fundamental importance in the quantitative analysis of biological structures such as neurons, blood vessels, retinal interconnects, microtubules, etc. The state of the art HMM-based contour tracking scheme for tracking microtubules, while performing well in most scenarios, can miss the track if, during its growth, it intersects another microtubule in its neighbourhood. In this paper we present a graphical model-based tracking algorithm which propagates across frames information about the dynamics of all the microtubules. This allows the algorithm to faithfully differentiate the contour of interest from others that contribute to the clutter, and maintain tracking accuracy. We present results of experiments on real microtubule images captured using fluorescence microscopy, and show that our proposed scheme outperforms the existing HMM-based scheme. Pradeep Koulgi, Mehmet Emre Sargin, Kenneth Rose, B. S. Manjunath |
ICPR | 4 |
| 2010 | Particle Filter Tracking with Online Multiple Instance LearningabstractThis paper addresses the problem of object tracking by learning a discriminative classifier to separate the object from its background. The online-learned classifier is used to adaptively model object's appearance and its background. To solve the typical problem of erroneous training examples generated during tracking, an online multiple instance learning (MIL) algorithm is used by allowing false positive examples. In addition, particle filter is applied to make best use of the learned classifier and help to generate a better representative set of training examples for the online MIL learning. The effectiveness of the proposed algorithm is demonstrated in some challenging environdments for human tracking. Zefeng Ni, Santhoshkumar Sunderrajan, Amir M. Rahimi, B. S. Manjunath |
ICPR | 4 |
| 2010 | Object Tracking with Ratio Cycles Using Shape and Appearance CuesabstractWe present a method for object tracking over time sequence imagery. The image plane is represented with a 4-connected planar graph where vertices are associated with pixels. On each image, the outer contour of the object is localized by finding the optimal cycle in the graph such that a cost function based on temporal, appearance and shape priors is minimized. Our contribution is the particle filtering-based framework to integrate the shape cue with the temporal and appearance cues. We demonstrate that incorporating the shape prior yields promising performance improvement over temporal and appearance priors on various object tracking scenarios. Mehmet Emre Sargin, Pratim Ghosh, B. S. Manjunath, Kenneth Rose |
ICPR | 3 |
| 2010 | Bisque: a platform for bioimage analysis and managementabstractAbstract Motivation: Advances in the field of microscopy have brought about the need for better image management and analysis solutions. Novel imaging techniques have created vast stores of images and metadata that are difficult to organize, search, process and analyze. These tasks are further complicated by conflicting and proprietary image and metadata formats, that impede analyzing and sharing of images and any associated data. These obstacles have resulted in research resources being locked away in digital media and file cabinets. Current image management systems do not address the pressing needs of researchers who must quantify image data on a regular basis. Results: We present Bisque, a web-based platform specifically designed to provide researchers with organizational and quantitative analysis tools for 5D image data. Users can extend Bisque with both data model and analysis extensions in order to adapt the system to local needs. Bisque's extensibility stems from two core concepts: flexible metadata facility and an open web-based architecture. Together these empower researchers to create, develop and share novel bioimage analyses. Several case studies using Bisque with specific applications are presented as an indication of how users can expect to extend Bisque for their own purposes. Availability: Bisque is web based, cross-platform and open source. The system is also available as software-as-a-service through the Center of Bioimage Informatics at UCSB. Contact: [email protected]; [email protected] Supplementary information: The supplementary material is available at Bioinformatics online, including screen shots, metadata XML descriptions and implementation details. Kristian Kvilekval, Dmitry V. Fedorov, Boguslaw Obara, Ambuj K. Singh, B. S. Manjunath |
Bioinform. | 5 |
| 2010 | On the Length and Area Regularization for Multiphase Level Set SegmentationabstractIn this paper we introduce novel regularization techniques for level set segmentation that target specifically the problem of multiphase segmentation. When the multiphase model is used to obtain a partitioning of the image in more than two regions, a new set of issues arise with respect to the single phase case in terms of regularization strategies. For example, if smoothing or shrinking each contour individually could be a good model in the single phase case, this is not necessarily true in the multiphase scenario. In this paper, we address these issues designing enhanced length and area regularization terms, whose minimization yields evolution equations in which each level set function involved in the multiphase segmentation can “sense” the presence of the other level set functions and evolve accordingly. In other words, the coupling of the level set function, which before was limited to the data term (i.e. the proper segmentation driving force), is extended in a mathematically principled way to the regularization terms as well. The resulting regularization technique is more suitable to eliminate spurious regions and other kind of artifacts. An extensive experimental evaluation supports the model we introduce in this paper, showing improved segmentation performance with respect to traditional regularization techniques. Luca Bertelli, Shivkumar Chandrasekaran, Frédéric Gibou, B. S. Manjunath |
Int. J. Comput. Vis. | 4 |
| 2010 | Efficient and Robust Detection of Duplicate Videos in a Large DatabaseabstractWe present an efficient and accurate method for duplicate video detection in a large database using video fingerprints. We have empirically chosen the color layout descriptor, a compact and robust frame-based descriptor, to create fingerprints which are further encoded by vector quantization (VQ). We propose a new nonmetric distance measure to find the similarity between the query and a database video fingerprint and experimentally show its superior performance over other distance measures for accurate duplicate detection. Efficient search cannot be performed for high-dimensional data using a nonmetric distance measure with existing indexing techniques. Therefore, we develop novel search algorithms based on precomputed distances and new dataset pruning techniques yielding practical retrieval times. We perform experiments with a database of 38 000 videos, worth 1600 h of content. For individual queries with an average duration of 60 s (about 50% of the average database video length), the duplicate video is retrieved in 0.032 s, on Intel Xeon with CPU 2.33 GHz, with a very high accuracy of 97.5%. Anindya Sarkar, Vishwakarma Singh, Pratim Ghosh, B. S. Manjunath, Ambuj K. Singh |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2010 | Matrix embedding with pseudorandom coefficient selection and error correction for robust and secure steganographyabstractIn matrix embedding (ME)-based steganography, the host coefficients are minimally perturbed such that the transmitted bits fall in a coset of a linear code, with the syndrome conveying the hidden bits. The corresponding embedding distortion and vulnerability to steganalysis are significantly less than that of conventional quantization index modulation (QIM)-based hiding. However, ME is less robust to attacks, with a single host bit error leading to multiple decoding errors for the hidden bits. In this paper, we employ the ME-RA scheme, a combination of ME-based hiding with powerful repeat accumulate (RA) codes for error correction, to address this problem. A key contribution of this paper is to compute log likelihood ratios for RA decoding, taking into account the many-to-one mapping between the host coefficients and an encoded bit, for ME. To reduce detectability, we hide in randomized blocks, as in the recently proposed Yet Another Steganographic Scheme (YASS), replacing the QIM-based embedding in YASS by the proposed ME-RA scheme. We also show that the embedding performance can be improved by employing punctured RA codes. Through experiments based on a couple of thousand images, we show that for the same embedded data rate and a moderate attack level, the proposed ME-based method results in a lower detection rate than that obtained for QIM-based YASS. Anindya Sarkar, Upamanyu Madhow, B. S. Manjunath |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2010 | A Nonconservative Flow Field for Robust Variational Image SegmentationabstractWe introduce a robust image segmentation method based on a variational formulation using edge flow vectors. We demonstrate the nonconservative nature of this flow field, a feature that helps in a better segmentation of objects with concavities. A multiscale version of this method is developed and is shown to improve the localization of the object boundaries. We compare and contrast the proposed method with well known state-of-the-art methods. Detailed experimental results are provided on both synthetic and natural images that demonstrate that the proposed approach is quite competitive. Pratim Ghosh, Luca Bertelli, Baris Sumengen, B. S. Manjunath |
IEEE Trans. Image Process. | 4 |
| 2010 | Video Annotation Through Search and Graph Reinforcement MiningabstractUnlimited vocabulary annotation of multimedia documents remains elusive despite progress solving the problem in the case of a small, fixed lexicon. Taking advantage of the repetitive nature of modern information and online media databases with independent annotation instances, we present an approach to automatically annotate multimedia documents that uses mining techniques to discover new annotations from similar documents and to filter existing incorrect annotations. The annotation set is not limited to words that have training data or for which models have been created. It is limited only by the words in the collective annotation vocabulary of all the database documents. A graph reinforcement method driven by a particular modality (e.g., visual) is used to determine the contribution of a similar document to the annotation target. The graph supplies possible annotations of a different modality (e.g., text) that can be mined for annotations of the target. Experiments are performed using videos crawled from YouTube. A customized precision-recall metric shows that the annotations obtained using the proposed method are superior to those originally existing for the document. These extended, filtered tags are also superior to a state-of-the-art semi-supervised technique for graph reinforcement learning on the initial user-supplied annotations. Emily Moxley, Tao Mei 0001, B. S. Manjunath |
IEEE Trans. Multim. | 3 |
| 2009 | Robust dynamical model for simultaneous registration and segmentation in a variational framework: A Bayesian approachabstractWe introduce a dynamical model for simultaneous registration and segmentation in a variational framework for image sequences, where the dynamics is incorporated using a Bayesian formulation. A linear stochastic equation relating the tracked object (or a region of interest) is first derived under the assumption that the successive images in the sequence are related by a dense and possibly non-linear displacement field. This derivation allows for the use of a computationally efficient and recursive implementation of the Bayesian formulation in this framework. The contour of the tracked object returned by the dynamical model is not only close to the previously detected shape but is also consistent with the temporal statistics of the tracked object. The performance of the proposed approach is evaluated on real image sequences. It is shown that, with respect to a variety of error metrics such as F-measure, mean absolute deviation and Hausdorff distance, the proposed approach outperforms the state-of-the art approach without the dynamical model. Pratim Ghosh, Mehmet Emre Sargin, B. S. Manjunath |
ICCV | 3 |
| 2009 | Probabilistic occlusion boundary detection on spatio-temporal latticesabstractIn this paper, we present an algorithm for occlusion boundary detection. The main contribution is a probabilistic detection framework defined on spatio-temporal lattices, which enables joint analysis of image frames. For this purpose, we introduce two complementary cost functions for creating the spatio-temporal lattice and for performing global inference of the occlusion boundaries, respectively. In addition, a novel combination of low-level occlusion features is discriminatively learnt in the detection framework. Simulations on the CMU Motion Dataset provide ample evidence that proposed algorithm outperforms the leading existing methods. Mehmet Emre Sargin, Luca Bertelli, B. S. Manjunath, Kenneth Rose |
ICCV | 3 |
| 2009 | Adding Gaussian noise to "denoise" JPEG for detecting image resizingabstractA common problem affecting most image resizing detection algorithms is that they are susceptible to JPEG compression. This is because JPEG introduces periodic artifacts, as it works on 8×8 blocks. We propose a novel yet counter intuitive technique to "denoise" JPEG images by adding Gaussian noise. We add a suitable amount of Gaussian noise to a resized and JPEG compressed image so that the periodicity due to JPEG compression is suppressed while that due to resizing is retained. The controlled Gaussian noise addition works better than median filtering and weighted averaging based filtering for suppressing the JPEG induced periodicity. Lakshmanan Nataraj, Anindya Sarkar, B. S. Manjunath |
ICIP | 3 |
| 2009 | Improving the quality of depth image based rendering for 3D Video systemsabstractIn 3D video (3DV) applications, a reduced number of views plus depth maps are transmitted or stored. When there is a need to render virtual views in between the actual views, the technique of depth image based rendering (DIBR) can be used to generate the intermediate views. To address the problem of noisy depth information in 3DV systems, we propose novel methods that can be easily incorporated into DIBR to improve synthesized image quality. These include: (1) a heuristic scheme with adaptive spatting that blends multiple warped reference pixels based on their depth, warped pixel positions and camera parameters; (2) an approximation of the first scheme with up-sampling for fast processing; (3) boundary only splatting; and (4) view weighting based on hole distribution. Experiment results show that the proposed methods can improve synthesis quality significantly. Zefeng Ni, Dong Tian, Sitaram Bhagavathy, Joan Llach, B. S. Manjunath |
ICIP | 5 |
| 2009 | Double embedding in the quantization index modulation frameworkabstractQuantization index modulation (QIM) is a commonly used data hiding technique where a single bit is embedded per coefficient. Here, we propose the use of double embedding in the QIM framework where a single coefficient is modified twice, using two quantizers, to embed two bits. The motivation behind substituting single embedding with double embedding in the QIM framework for a certain steganographic scheme is to increase its hiding rate without significantly increasing the embedding distortion and the stego scheme's detectability against steganalysis. We empirically determine the best way to couple the double embedding framework with a repeat accumulate code based error correction scheme. For moderate noise levels, the use of double embedding is seen to be significantly advantageous over single embedding. Anindya Sarkar, B. S. Manjunath |
ICIP | 2 |
| 2009 | Not all tags are created equal: Learning flickr tag semantics for global annotationabstractLarge collaborative datasets offer the challenging opportunity of creating systems capable of extracting knowledge in the presence of noisy data. In this work we explore the ability to automatically learn tag semantics by mining a global georeferenced image collection crawled from Flickr with the aim of improving an automatic annotation system. We are able to categorize sets of tags as places, landmarks, and visual descriptors. By organizing our dataset of more than 1.69 million images using a quadtree we can efficiently find geographic areas with sufficient density to provide useful results for place and landmark extraction. Precision-recall curves for our techniques compared with previous existing work used to identify place tags and manual groundtruth landmark annotation show the merit of our methods applied on a world scale. Emily Moxley, Jim Kleban, Jiejun Xu, B. S. Manjunath |
ICME | 4 |
| 2009 | A biosegmentation benchmark for evaluation of bioimage analysis methodsabstractBACKGROUND: We present a biosegmentation benchmark that includes infrastructure, datasets with associated ground truth, and validation methods for biological image analysis. The primary motivation for creating this resource comes from the fact that it is very difficult, if not impossible, for an end-user to choose from a wide range of segmentation methods available in the literature for a particular bioimaging problem. No single algorithm is likely to be equally effective on diverse set of images and each method has its own strengths and limitations. We hope that our benchmark resource would be of considerable help to both the bioimaging researchers looking for novel image processing methods and image processing researchers exploring application of their methods to biology. RESULTS: Our benchmark consists of different classes of images and ground truth data, ranging in scale from subcellular, cellular to tissue level, each of which pose their own set of challenges to image analysis. The associated ground truth data can be used to evaluate the effectiveness of different methods, to improve methods and to compare results. Standard evaluation methods and some analysis tools are integrated into a database framework that is available online at http://bioimage.ucsb.edu/biosegmentation/. CONCLUSION: This online benchmark will facilitate integration and comparison of image analysis methods for bioimages. While the primary focus is on biological images, we believe that the dataset and infrastructure will be of interest to researchers and developers working with biological image analysis, image segmentation and object tracking in general. Elisa Drelie Gelasca, Boguslaw Obara, Dmitry V. Fedorov, Kristian Kvilekval, B. S. Manjunath |
BMC Bioinform. | 5 |
| 2008 | Shape prior segmentation of multiple objects with graph cutsabstractWe present a new shape prior segmentation method using graph cuts capable of segmenting multiple objects. The shape prior energy is based on a shape distance popular with level set approaches. We also present a multiphase graph cut framework to simultaneously segment multiple, possibly overlapping objects. The multiphase formulation differs from multiway cuts in that the former can account for object overlaps by allowing a pixel to have multiple labels. We then extend the shape prior energy to encompass multiple shape priors. Unlike variational methods, a major advantage of our approach is that the segmentation energy is minimized directly without having to compute its gradient, which can be a cumbersome task and often relies on approximations. Experiments demonstrate that our algorithm can cope with image noise and clutter, as well as partial occlusions and affine transformations of the shape. Nhat Vu, B. S. Manjunath |
CVPR | 2 |
| 2008 | Deformable trellis: open contour tracking in bio-image sequencesabstractThis paper presents an open contour tracking method that employs an arc-emission hidden Markov model (HMM). The algorithm encodes the shape information of the structure in a spatially deformable trellis model that is iteratively modified to account for observations in subsequent frames. As the open contour is determined on the trellis of an HMM, a dynamic programming procedure reduces the computational complexity to linear in the length of the structure (or contour). The method was developed for tracking general curvilinear structures, and tested on subcellular image sequences, where microtubules grow, shrink and undergo lateral motion from frame to frame. Microtubule length changes are modeled by the addition of appropriate transient and absorbing states to the HMM. Our results provide experimental evidence for the proposed algorithm's capability to track non-rigid curvilinear objects in challenging environments in terms of noise and clutter. Mehmet Emre Sargin, Alphan Altinok, Kenneth Rose, B. S. Manjunath |
ICASSP | 4 |
| 2008 | Reference-based probabilistic segmentation as non-rigid registration using Thin Plate SplinesabstractIn this paper we demonstrate the effectiveness of reference (or atlas)-based non-rigid registration for the segmentation of medical and biological imagery. In particular we introduce a segmentation functional exploiting feature information about the reference image and we minimize it with respect to the parameters of the non-rigid transformation, akin to a region-based maximum likelihood estimation process. The warping transformation is modeled using thin plate splines, which incorporate information about the global rigid motion and the non-rigid local displacements. Extensive experimental evaluations and comparisons with other segmentation techniques on a complex biological dataset are presented. The proposed algorithm outperforms the others in both classification rate and, in particular, localization accuracy. Luca Bertelli, Pratim Ghosh, B. S. Manjunath, Frédéric Gibou |
ICIP | 3 |
| 2008 | Evaluation and benchmark for biological image segmentationabstractThis paper describes ongoing work on creating a benchmarking and validation dataset for biological image segmentation. While the primary target is biological images, we believe that the dataset would be of help to researchers working in image segmentation and tracking in general. The motivation for creating this resource comes from the observation that while there are a large number of effective segmentation methods available in the research literature, it is difficult for the application scientists to make an informed choice as to what methods would work for her particular problem. No one single tool exists that is effective on a diverse set of application contexts and different methods have their own strengths and limitations. We describe below three different classes of data, ranging in scale from subcellular to cellular to tissue level images, each of which pose their own set of challenges to image analysis. Of particular value to the image processing researchers is that the data comes with associated ground truth information that can be used to evaluate the effectiveness of different methods. The analysis and evaluation are also integrated into a database framework that is available online at http://dough.ece.ucsb.edu. Elisa Drelie Gelasca, Jiyun Byun, Boguslaw Obara, B. S. Manjunath |
ICIP | 4 |
| 2008 | A lightweight multiview tracked person descriptor for camera sensor networksabstractWe present a simple multiple view 3D model for object tracking and identification in camera networks. Our model is composed of 8 distinct views in the interval [0, 7pi/4]. Each of the 8 parts describes the person's appearance from that particular viewpoint. The model contains both color and structure information about each view which are assembled into a single entity and is meant as a simple, lightweight object representation for use in camera sensor networks. It is versatile in that it can be gradually assembled on-line while a person is tracked. The model's ease of use and effectiveness for identification in surveillance video is demonstrated. Michael J. Quinn, Thomas Kuo, B. S. Manjunath |
ICIP | 3 |
| 2008 | Conditional iterative decoding of Two Dimensional Hidden Markov ModelsabstractTwo dimensional hidden markov models (2D-HMMs) provide substantial benefits for many computer vision and image analysis applications. Many fundamental image analysis problems, including segmentation and classification, are target applications for the 2D- HMMs. As opposed to the i.i.d. assumption of the image observations, the naturally existing spatial correlations can be readily modeled by solving the 2D-HMM decoding problem. However, computational complexity of the 2D-HMM decoding grows exponentially with the image size and is known to be NP-hard. In this paper, we present a conditional iterative decoding (CID) algorithm for the approximate decoding of 2D-HMMs. We compare the performance of the CID algorithm to the Turbo-HMM (T-HMM) decoding algorithm and show that CID gives promising results. We demonstrate the proposed algorithm on modeling spatial deformations of human faces in recognizing people across their different facial expressions. Mehmet Emre Sargin, Alphan Altinok, Kenneth Rose, B. S. Manjunath |
ICIP | 4 |
| 2008 | Estimation of optimum coding redundancy and frequency domain analysis of attacks for YASS - a randomized block based hiding schemeabstractOur recently introduced JPEG steganographic method called yet another steganographic scheme (YASS) can resist blind steganalysis by embedding data in the discrete cosine transform (DCT) domain in randomly chosen image blocks. To maximize the embedding rate for a given image and a specified attack channel, the redundancy factor used by the repeat- accumulate (RA) code based error correction framework in YASS is optimally chosen by the encoder. An efficient method is suggested for the decoder to accurately compute this redundancy factor. We also show experimentally which DCT coefficients are better suited for hiding and detection under various attacks. The effectiveness of YASS for robust steganography is demonstrated for certain attacks. Anindya Sarkar, Lakshmanan Nataraj, B. S. Manjunath, Upamanyu Madhow |
ICIP | 3 |
| 2008 | Graph cut segmentation of neuronal structures from transmission electron micrographsabstractIn many neurophysiological studies, understanding the neuronal circuitry of the brain requires detailed 3D models of the nerve cells and their synapses. Typically, researchers build the 3D models by manually tracing the 2D cross-sectional profiles of the 3D structures from serial electron micrograph (EM) stacks and then construct the models from these 2D contours. While current computer-aided techniques can reduce the tracing time, they often require extensive user interaction. We propose a segmentation framework to extract the 2D profiles that is both fast and requires a minimal amount of user interaction. The framework uses graph cuts to minimize an energy defined over the image intensity and the flux of the intensity gradient field. Furthermore, to correct segmentation errors, our framework allows for efficient and intuitive editing of the initial results. Nhat Vu, B. S. Manjunath |
ICIP | 2 |
| 2008 | An automatic method to learn and transfer the photometric appearance of partially overlapping imagesabstractThe first major contribution of this paper is a robust method to learn the photometric mapping between the overlapping portions of two registered images acquired either under different lighting conditions or different sensor modalities. Then, once such mapping is learnt, we demonstrate how it generalizes so that the photometric appearance can be transferred from one image to the other out of their overlapping area. This task is fundamental in several different contexts, such as image colorization, seamless mosaicking or change detection. After introducing the theory and discussing the algorithms, we will present several examples that confirm the efficacy of the proposed method in dealing with different types of images. Marco Zuliani, Luca Bertelli, B. S. Manjunath |
ICIP | 3 |
| 2008 | Automatic video annotation through search and miningabstractConventional approaches to video annotation predominantly focus on supervised identification of a limited set of concepts, while unsupervised annotation with infinite vocabulary remains unexplored. This work aims to exploit the overlap in content of news video to automatically annotate by mining similar videos that reinforce, filter, and improve the original annotations. The algorithm employs a two-step process of search followed by mining. Given a query video consisting of visual content and speech-recognized transcripts, similar videos are first ranked in a multimodal search. Then, the transcripts associated with these similar videos are mined to extract keywords for the query. We conducted extensive experiments over the TRECVID 2005 corpus and showed the superiority of the proposed approach to using only the mining process on the original video for annotation. This work represents the first attempt at unsupervised automatic video annotation leveraging overlapping video content. Emily Moxley, Tao Mei 0001, Xian-Sheng Hua 0001, Wei-Ying Ma, B. S. Manjunath |
ICME | 5 |
| 2008 | Drums, curve descriptors and affine invariant region matching
Marco Zuliani, Luca Bertelli, Charles S. Kenney, Shivkumar Chandrasekaran, B. S. Manjunath |
Image Vis. Comput. | 5 |
| 2008 | A Variational Framework for Multiregion Pairwise-Similarity-Based Image SegmentationabstractVariational cost functions that are based on pairwise similarity between pixels can be minimized within level set framework resulting in a binary image segmentation. In this paper we extend such cost functions and address multi-region image segmentation problem by employing a multi-phase level set framework. For multi-modal images cost functions become more complicated and relatively difficult to minimize. We extend our previous work, proposed for background/foreground separation, to the segmentation of images in more than two regions. We also demonstrate an efficient implementation of the curve evolution, which reduces the computational time significantly. Finally, we validate the proposed method on the Berkeley Segmentation Data Set by comparing its performance with other segmentation techniques. Luca Bertelli, Baris Sumengen, B. S. Manjunath, Frédéric Gibou |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2007 | Secure Steganography: Statistical Restoration of the Second Order Dependencies for Improved SecurityabstractWe present practical approaches for steganography that can provide improved security by closely matching the second-order statistics of the host rather than just the marginal distribution. The methods are based on the framework of statistical restoration, wherein a fraction of the host symbols available for hiding is actually used to restore the statistics; thus reducing the rate, but providing security against steganalysis. We establish correspondence between steganography and the earth-mover's distance (EMD), a popular distance metric used in computer vision applications. The EMD framework can be used to define the optimum flow (modifications) of the host symbols for compensation. This formulation is used for image steganography by restoring the second-order statistics of the blockwise discrete cosine transform (DCT) coefficients. Some practical limitations of this approach (such as computational complexity and difficulty in dealing with overlapping coefficient pairs) are noted, and a new method is proposed that alleviates these deficiencies by identifying the coefficients to modify based on a local compensation criterion. Experimental results on several thousand natural images demonstrate the utility of the presented methods. Anindya Sarkar, Kaushal Solanki, Upamanyu Madhow, Shivkumar Chandrasekaran, B. S. Manjunath |
ICASSP (2) | 5 |
| 2007 | Pairwise Similarities across Images for Multiple View Rigid/Non-Rigid Segmentation and RegistrationabstractA variational approach for the background/foreground segmentation of multiple views of the same scene is presented. The main novelty is the introduction of cost functions based on pairwise similarity between pixels across different images. These cost functions are minimized within a level set framework. In addition, a warping model (rigid or non-rigid) between the emerging foregrounds in the different views is imposed, thus avoiding the introduction of a specific shape term in the cost function to handle occlusions. The thin plate spline (TPS) warping is for the first time employed within the level set framework to model non- rigid deformations. The minimization of these cost functions leads to simultaneous segmentation and registration of the different views. Examples of segmentations of a variety of objects are shown and possible applications are proposed. Luca Bertelli, Marco Zuliani, B. S. Manjunath |
ICCV | 3 |
| 2007 | A Variational Approach to Exploit Prior Information in Object-Background Segregation: Application to Retinal ImagesabstractOne of the main challenges in image segmentation is to adapt prior knowledge about the objects/regions that are likely to be present in an image, in order to obtain more precise detection and recognition. Typical applications of such knowledge-based segmentation include partitioning satellite images and microscopy images, where the context is generally well defined. In particular, we present an approach that exploits the knowledge about foreground and background information given in a reference image, in segmenting images containing similar objects or regions. This problem is presented within a variational framework, where cost functions based on pair-wise pixel similarities are minimized. This is perhaps one of the first attempts in using non-shape based prior information within a segmentation framework. We validate the proposed method to segment the outer nuclear layer (ONL) in retinal images. This approach successfully segments the ONL within an image and enables further quantitative analysis. Luca Bertelli, Jiyun Byun, B. S. Manjunath |
ICIP (6) | 3 |
| 2007 | Edge Preserving Filters using Geodesic Distances on Weighted Orthogonal DomainsabstractWe introduce a framework for image enhancement, which smooths images while preserving edge information. Domain (spatial) and range (feature) information are combined in one single measure in a principled way. This measure turns out to be the geodesic distance between pixels, calculated on weighted orthogonal domains. The weight function is computed to capture the underlying structure of the image manifold, but allowing at the same time to efficiently solve, using the Fast Marching algorithm on orthogonal domains, the eikonal equation to obtain the geodesic distances. We show promising results in edge-preserving denoising of gray scale, color and texture images. Luca Bertelli, B. S. Manjunath |
ICIP (1) | 2 |
| 2007 | Tracing Curvilinear Structures in Live Cell ImagesabstractTracing of curvilinear structures is one of the fundamental tools in the quantitative analysis of biological images, for extracting information about structures such as blood vessels, neurons, microtubules, and similar entities. Due to the limitations in biological sample preparation and fluorescence imaging, typical images in live cell studies exhibit severe noise and considerable clutter. These images are manually analyzed through a laborious and approximate set of quantification tasks. In this paper, we describe a constrained optimization method for extracting curvilinear structures from live cell fluorescence images. We show that the proposed method is largely insensitive to frequent intersections, intensity variations along the curve, and generates successful traces within noisy regions. We demonstrate the results of our approach on live cell microtubule images. Mehmet Emre Sargin, Alphan Altinok, Kenneth Rose, B. S. Manjunath |
ICIP (6) | 4 |
| 2007 | Estimating Steganographic Capacity for Odd-Even Based Embedding and its Use in Individual CompensationabstractWe present a method to compute the steganographic capacity for images, with odd-even based hiding in the quantized discrete cosine transform domain. The method has been generalized for varying orders of co-occurrence statistics for statistical restoration based steganography. We further utilize this capacity estimate to hide the maximum possible data per individual frequency stream, while ensuring that the first order histograms of individual frequency coefficients remain matched. We also show that certain frequency components are more useful for steganalysis after first order statistical restoration is performed for a certain band of select frequencies. Anindya Sarkar, B. S. Manjunath |
ICIP (1) | 2 |
| 2007 | Retina Layer Segmentation and Spatial Alignment of Antibody Expression LevelsabstractThe expression levels of rod opsin and glial fibrillary acidic protein (GFAP) capture important structural changes in the retina during injury and recovery. Quantitatively measuring these expression levels in confocal micrographs requires identifying the retinal layer boundaries and spatially corresponding the layers across different images. In this paper, a method to segment the retinal layers using a parametric active contour model is presented. Then spatially aligned expression levels across different images are determined by thresholding the solution to a Dirichlet boundary value problem. Our analysis provides quantitative metrics of retinal restructuring that are needed for improving retinal therapies after injury. Nhat Vu, Pratim Ghosh, B. S. Manjunath |
ICIP (2) | 3 |
| 2006 | Activity Analysis in Microtubule Videos by Mixture of Hidden Markov ModelsabstractWe present an automated method for the tracking and dynamics modeling of microtubules -a major component of the cytoskeleton- which provides researchers with a previously unattainable level of data analysis and quantification capabilities. The proposed method improves upon the manual tracking and analysis techniques by i) increasing accuracy and quantified sample size in data collection, ii) eliminating user bias and standardizing analysis, iii) making available new features that are impractical to capture manually, iv) enabling statistical extraction of dynamics patterns from cellular processes, and v) greatly reducing required time for entire studies. An automated procedure is proposed to track each resolvable microtubule, whose aggregate activity is then modeled by mixtures of Hidden Markov Models to uncover dynamics patterns of underlying cellular and experimental conditions. Our results support manually established findings on an actual microtubule dataset and illustrate how automated analysis of spatial and temporal patterns offers previously unattainable insights to cellular processes. Alphan Altinok, Motaz A. El Saban, Austin J. Peck, Leslie Wilson, Stuart C. Feinstein, B. S. Manjunath, Kenneth Rose |
CVPR (2) | 6 |
| 2006 | Redundancy in All Pairs Fast Marching MethodabstractIn this paper, we analyze the redundancy in calculating all pairs of geodesic distances on a rectangular grid. Fast marching method is an efficient way to estimate the geodesic distances from a point. But when calculated for all the points on the grid, this introduces certain redundancy. Our analysis shows that over 90% of the distances are actually recalculated. We propose a novel solution which exploits this redundancy to reduce the number of distances evaluated using the fast marching method and enforces the symmetry of the distance matrix. Experimental results show the improved accuracy obtained with our implementation. Luca Bertelli, Baris Sumengen, B. S. Manjunath |
ICIP | 3 |
| 2006 | Multi-Focus Imaging using Local Focus Estimation and MosaickingabstractWe propose an algorithm to generate one multi-focus image from a set of images acquired at different focus settings. First images are registered to avoid large misalignments. Each image is tiled with overlapping neighborhoods. Then, for each region the tile that corresponds to the best focus is chosen to construct the multi-focus image. The overlapping tiles are then seamlessly mosaicked. Our approach is presented for images from optical microscopes and hand held consumer cameras, and demonstrates robustness to temporal changes and small misalignments. The implementation is computationally efficient and gives good results. Dmitry V. Fedorov, Baris Sumengen, B. S. Manjunath |
ICIP | 3 |
| 2006 | Provably Secure Steganography: Achieving Zero K-L Divergence using Statistical RestorationabstractIn this paper, we present a framework for the design of steganographic schemes that can provide provable security by achieving zero Kullback-Leibler divergence between the cover and the stego signal distributions, while hiding at high rates. The approach is to reserve a number of host symbols for statistical restoration: host statistics perturbed by data embedding are restored by suitably modifying the symbols from the reserved set. A dynamic embedding approach is proposed, which avoids hiding in low probability regions of the host distribution. The framework is applied to design practical schemes for image steganography, which are evaluated using supervised learning on a set of about 1000 natural images. For the presented JPEG steganography scheme, it is seen that the detector is indeed reduced to random guessing. Kaushal Solanki, Kenneth Sullivan, Upamanyu Madhow, B. S. Manjunath, Shivkumar Chandrasekaran |
ICIP | 4 |
| 2006 | Determining Achievable Rates for Secure, Zero Divergence, SteganographyabstractIn steganography (the hiding of data into innocuous covers for secret communication) it is difficult to estimate how much data can be hidden while still remaining undetectable. To measure the inherent detectability of steganography, Cachin suggested the ϵ-secure measure, where ϵ is the Kullback Leibler (K-L) divergence between the cover distribution and the distribution after hiding. At zero divergence, an optimal statistical detector can do no better than guessing; the data is undetectable. The hider's key question then is, what hiding rate can be used while maintaining zero divergence? Though work has been done on the theoretical capacity of steganography, it is often difficult to use these results in practice. We therefore examine the limits of a practical scheme known to allow embedding with zero-divergence. This scheme is independent of the embedding algorithm and therefore can be generically applied to find an achievable secure hiding rate for arbitrary cover distributions. Kenneth Sullivan, Kaushal Solanki, B. S. Manjunath, Upamanyu Madhow, Shivkumar Chandrasekaran |
ICIP | 3 |
| 2006 | Graph Partitioning Active Contours (GPAC) for Image SegmentationabstractIn this paper, we introduce new types of variational segmentation cost functions and associated active contour methods that are based on pairwise similarities or dissimilarities of the pixels. As a solution to a minimization problem, we introduce a new curve evolution framework, the graph partitioning active contours (GPAC). Using global features, our curve evolution is able to produce results close to the ideal minimization of such cost functions. New and efficient implementation techniques are also introduced in this paper. Our experiments show that GPAC solution is effective on natural images and computationally efficient. Experiments on gray-scale, color, and texture images show promising segmentation results. Baris Sumengen, B. S. Manjunath |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | Modeling and Detection of Geospatial Objects Using Texture MotifsabstractWe propose the use of texture motifs, or characteristic spatially recurrent patterns, for modeling and detecting geospatial objects. A method is proposed for learning a texture-motif model from object examples and detecting objects based on the learned model. The model is learned in a two-layered framework: the first learns the constituent "texture elements" of the motif and, the second, the spatial distribution of the elements. In the experimental session, we demonstrate the model training and selection methodology for objects given a set of training examples. The utility of such models for detecting the presence or absence of geospatial objects in large aerial image datasets comprising tens of thousands of image tiles is then emphasized Sitaram Bhagavathy, B. S. Manjunath |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2006 | 'Print and Scan' Resilient Data Hiding in ImagesabstractPrint-scan resilient data hiding finds important applications in document security and image copyright protection. This paper proposes methods to hide information into images that achieve robustness against printing and scanning with blind decoding. The selective embedding in low frequencies scheme hides information in the magnitude of selected low-frequency discrete Fourier transform coefficients. The differential quantization index modulation scheme embeds information in the phase spectrum of images by quantizing the difference in phase of adjacent frequency locations. A significant contribution of this paper is analytical and experimental modeling of the print-scan process, which forms the basis of the proposed embedding schemes. A novel approach for estimating the rotation undergone by the image during the scanning process is also proposed, which specifically exploits the knowledge of the digital halftoning scheme employed by the printer. Using the proposed methods, several hundred information bits can be embedded into images with perfect recovery against the print-scan operation. Moreover, the hidden images also survive several other attacks, such as Gaussian or median filtering, scaling or aspect ratio change, heavy JPEG compression, and rows and/or columns removal Kaushal Solanki, Upamanyu Madhow, B. S. Manjunath, Shivkumar Chandrasekaran, Ibrahim El-Khalil |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2006 | Steganalysis for Markov cover data with applications to imagesabstractThe difficult task of steganalysis, or the detection of the presence of hidden data, can be greatly aided by exploiting the correlations inherent in typical host or cover signals. In particular, several effective image steganalysis techniques are based on the strong interpixel dependencies exhibited by natural images. Thus, existing theoretical benchmarks based on independent and identically distributed (i.i.d.) models for the cover data underestimate attainable steganalysis performance and, hence, overestimate the security of the steganography technique used for hiding the data. In this paper, we investigate detection-theoretic performance benchmarks for steganalysis when the cover data are modeled as a Markov chain. The main application explored here is steganalysis of data hidden in images. While the Markov chain model does not completely capture the spatial dependencies, it provides an analytically tractable framework whose predictions are consistent with the performance of practical steganalysis algorithms that account for spatial dependencies. Numerical results are provided for image steganalysis of spread-spectrum and perturbed quantization data hiding. Kenneth Sullivan, Upamanyu Madhow, Shivkumar Chandrasekaran, B. S. Manjunath |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2005 | An Axiomatic Approach to Corner DetectionabstractThis paper presents an axiomatic approach to corner detection. In the first part of the paper we review five currently used corner detection methods (Harris-Stephens, Forstner, Shi-Tomasi, Rohr, and Kenney et al ) for graylevel images. This is followed by a discussion of extending these corner detectors to images with different pixel dimensions such as signals (pixel dimension one) and tomographic medical images (pixel dimension three) as well as different intensity dimensions such as color or LADAR images (intensity dimension three). These extensions are motivated by analyzing a particular example of optical flow in pixel and intensity space with arbitrary dimensions. Placing corner detection in a general setting enables us to state four axioms that any corner detector might reasonably be required to satisfy. Our main result is that only the Shi-Tomasi (and equivalently the Kenney et al. 2-norm detector) satisfy all four of the axioms. Charles S. Kenney, Marco Zuliani, B. S. Manjunath |
CVPR (1) | 3 |
| 2005 | Statistical restoration for robust and secure steganographyabstractWe investigate data hiding techniques that attempt to defeat steganalysis by restoring the statistics of the composite image to resemble that of the cover. The approach is to reserve a number of host symbols for statistical restoration: host statistics perturbed by data embedding are restored by suitably modifying the symbols from the reserved set. While statistical restoration has broad applicability to a variety of hiding methods, we illustrate our ideas here for quantization index modulation (QIM) based hiding. We propose a method for significantly reducing the detectability of QIM, while preserving its robustness to attacks. We next use the framework of statistical restoration to develop a method to combat steganalysis techniques which detect block-DCT embedding by evaluating the increase in blockiness of the image due to hiding. Numerical results demonstrating the efficacy of these techniques are provided. Kaushal Solanki, Kenneth Sullivan, Upamanyu Madhow, B. S. Manjunath, Shivkumar Chandrasekaran |
ICIP (2) | 4 |
| 2005 | The multiRANSAC algorithm and its application to detect planar homographiesabstractA RANSAC based procedure is described for detecting inliers corresponding to multiple models in a given set of data points. The algorithm we present in this paper (called multiRANSAC) on average performs better than traditional approaches based on the sequential application of a standard RANSAC algorithm followed by the removal of the detected set of inliers. We illustrate the effectiveness of our approach on a synthetic example and apply it to the problem of identifying multiple world planes in pairs of images containing dominant planar structures. Marco Zuliani, Charles S. Kenney, B. S. Manjunath |
ICIP (3) | 3 |
| 2005 | Region of interest extraction and virtual camera control based on panoramic video capturingabstractWe present a system for automatically extracting the region of interest (ROI) and controlling virtual cameras' control based on panoramic video. It targets applications such as classroom lectures and video conferencing. For capturing panoramic video, we use the FlyCam system that produces high resolution, wide-angle video by stitching video images from multiple stationary cameras. To generate conventional video, a region of interest can be cropped from the panoramic video. We propose methods for ROI detection, tracking, and virtual camera control that work in both the uncompressed and compressed domains. The ROI is located from motion and color information in the uncompressed domain and macroblock information in the compressed domain, and tracked using a Kalman filter. This results in virtual camera control that simulates human controlled video recording. The system has no physical camera motion and the virtual camera parameters are readily available for video indexing. Xinding Sun, Jonathan Foote, Don Kimber, B. S. Manjunath |
IEEE Trans. Multim. | 4 |
| 2004 | Drums and Curve DescriptorsabstractIn this paper we present a new physically motivated curve descriptor based on the solution of Helmholtz’s equation. The descriptor satisfies the six principles set by MPEG-7: it has a good retrieval accuracy, it is compact, it can be applied in general contexts, it has a reasonable computational complexity, it is robust and provides an hierarchical representation of the curve from coarse to fine. Moreover this descriptor generalizes straightforwardly to three dimensional surfaces. We tested the performance of the descriptor in the context of affine invariant curve matching using the multiview curve dataset (MCD), which consists of 40 curves (extracted from the MPEG-7 shape dataset) imaged under 14 different points of view. The results we obtained show that the proposed descriptor satisfies the MPEG-7 requirements and presents some advantages over some of the commonly used curve descriptors. 1 Marco Zuliani, Charles S. Kenney, Sitaram Bhagavathy, B. S. Manjunath |
BMVC | 4 |
| 2004 | Interactive segmentation using curve evolution and relevance feedback
Motaz A. El Saban, B. S. Manjunath |
ICIP | 2 |
| 2004 | Estimating and undoing rotation for print-scan resilient data hidingabstractThis paper proposes a method to hide information into images that achieves robustness against printing and scanning with blind decoding. A significant contribution of this paper is a technique to estimate and undo rotation. The method is based on the fact that laser printers use an ordered digital halftoning algorithm for printing. Using the proposed hiding method, several hundred information bits can be embedded into 512/spl times/512 images with perfect recovery against the print-scan operation. Moreover, the hidden images also survive other attacks such as Gaussian or median filtering, scaling or aspect ratio change, heavy JPEG compression, and rows and/or columns removal. Kaushal Solanki, Upamanyu Madhow, B. S. Manjunath, Shivkumar Chandrasekaran |
ICIP | 3 |
| 2004 | Steganalysis of quantization index modulation data hidingabstractQuantization index modulation (QIM) techniques have been gaining popularity in the data hiding community because of their robustness and information-theoretic optimally against a large class of attacks. In this paper, we consider detecting the presence of QIM hidden data, which is an important consideration when data hiding is used for covert communication, or steganography. For a given host distribution, we are able to quantify detectability compactly in terms of a parameter related to the robustness of the hiding scheme to attacks. Using detection theory we show that QIM quickly transitions from easily detectable to virtually undetectable as this parameter varies. We also obtain performance benchmarks for QIM hiding in images, indicating that a scheme designed to be robust to say a moderate degree of JPEG compression, should be easily detectable. While practical application of detection theory to images is difficult because of statistical variations across images, we employ supervised learning to show that standard QIM schemes for images are indeed quite easily detectable. However, it remains an open issue as to whether it is possible to devise QIM variants that are less vulnerable to steganalysis. Kenneth Sullivan, Zhiqiang Bi, Upamanyu Madhow, Shivkumar Chandrasekaran, B. S. Manjunath |
ICIP | 5 |
| 2004 | Affine-invariant curve matchingabstractIn this paper, we propose, an affine-invariant-method for describing and matching curves. This is important since affine transformations are often used to model perspective distortions. More specifically, we propose a new definition of the shape of a curve that characterizes a curve independently of the effects introduced by affine distortions. By combining this definition with a rotation-invariant shape descriptor, we show how it is possible to describe a curve in an intrinsically affine-invariant manner. To validate our procedure we built a database of shapes subject to perspective distortions and plotted the precision-recall curve for this dataset. Finally an application of our method is shown in the context of wide baseline matching. Marco Zuliani, Sitaram Bhagavathy, B. S. Manjunath, Charles S. Kenney |
ICIP | 3 |
| 2004 | Cortina: a system for large-scale, content-based web image retrievalabstractRecent advances in processing and networking capabilities of computers have led to an accumulation of immense amounts of multimedia data such as images. One of the largest repositories for such data is the World Wide Web (WWW). We present Cortina, a large-scale image retrieval system for the WWW. It handles over 3 million images to date. The system retrieves images based on visual features and collateral text. We show that a search process which consists of an initial query-by-keyword or query-by-image and followed by relevance feedback on the visual appearance of the results is possible for large-scale data sets. We also show that it is superior to the pure text retrieval commonly used in large-scale systems. Semantic relationships in the data are explored and exploited by data mining, and multiple feature spaces are included in the search process. Till Quack, Ullrich J. Mönich, Lars Thiele, B. S. Manjunath |
ACM Multimedia | 4 |
| 2004 | Robust image-adaptive data hiding using erasure and error correctionabstractInformation-theoretic analyses for data hiding prescribe embedding the hidden data in the choice of quantizer for the host data. In this paper, we propose practical realizations of this prescription for data hiding in images, with a view to hiding large volumes of data with low perceptual degradation. The hidden data can be recovered reliably under attacks, such as compression and limited amounts of image tampering and image resizing. The three main findings are as follows. 1) In order to limit perceivable distortion while hiding large amounts of data, hiding schemes must use image-adaptive criteria in addition to statistical criteria based on information theory. 2) The use of local criteria to choose where to hide data can potentially cause desynchronization of the encoder and decoder. This synchronization problem is solved by the use of powerful, but simple-to-implement, erasures and errors correcting codes, which also provide robustness against a variety of attacks. 3) For simplicity, scalar quantization-based hiding is employed, even though information-theoretic guidelines prescribe vector quantization-based methods. However, an information-theoretic analysis for an idealized model is provided to show that scalar quantization-based hiding incurs approximately only a 2-dB penalty in terms of resilience to attack. Kaushal Solanki, Noah Jacobsen, Upamanyu Madhow, B. S. Manjunath, Shivkumar Chandrasekaran |
IEEE Trans. Image Process. | 4 |
| 2003 | Nearest Neighbor Search for Relevance FeedbackabstractWe introduce the problem of repetitive nearest neighbor search in relevance feedback and propose an efficient search scheme for high dimensional feature spaces. Relevance feedback learning is a popular scheme used in content based image and video retrieval to support high-level concept queries. The paper addresses those scenarios in which a similarity or distance matrix is updated during each iteration of the relevance feedback search and a new set of nearest neighbors is computed. This repetitive nearest neighbor computation in high dimensional feature spaces is expensive, particularly when the number of items in the data set is large. In this context, we suggest a search algorithm that supports relevance feedback for the general quadratic distance metric. The scheme exploits correlations between two consecutive nearest neighbor sets thus significantly reducing the overall search complexity. Detailed experimental results are provided using 60 dimensional texture feature dataset. Jelena Tesic, B. S. Manjunath |
CVPR (2) | 2 |
| 2003 | On the Rayleigh nature of Gabor filter outputsabstractTexture has been recognized as an important visual primitive in image analysis. A widely used texture descriptor, which is part, of the MPEG-7 standard, is that computed using multiscale Gabor filters. The high dimensionality and computational complexity of this descriptor adversely affect the efficiency of content-based retrieval systems. We propose a modified texture descriptor that has comparable performance, but with nearly half the dimensionality and less computational expense. This gain is based on a claim that the distribution of (absolute values of) filter outputs have a strong tendency to be Rayleigh. Experimental results show that the dimensionality can be reduced by almost 50%, with a tradeoff of less than 3% on the error rate. Furthermore, it is easy to compute the new feature using the old, one, without having to repeat the computationally expensive filtering step. We also propose a new normalization method that improves similarity retrieval and indexing efficiency. Sitaram Bhagavathy, Jelena Tesic, B. S. Manjunath |
ICIP (3) | 3 |
| 2003 | Object localization using texture motifs and Markov random fieldsabstractThis work presents a novel approach to object localization in complex imagery. In particular, the spatial extents of objects characterized by distinct spatial signatures at multiple scales are estimated by using statistical models to control a simple region growing process. Texture motifs are used to model the spatial signatures at the smallest, or pixel, scale. Markov random fields are used to model the spatial signatures at the larger, or motif, scale. These models are used to iteratively expand a bounding box to approximate the spatial extent of an object. The approach is applied to localizing geo-spatial objects in high-resolution panchromatic aerial imagery. Shawn D. Newsam, Sitaram Bhagavathy, B. S. Manjunath |
ICIP (2) | 3 |
| 2003 | Video region segmentation by spatio-temporal watershedsabstractIn this paper, we propose a video region segmentation scheme combining spatio-temporal edges and watershed techniques. We consider the video sequence as a 3-D volume and compute color edges within this volume. These color edges form a vector field that is in turn used to obtain an edge function. This edge function is used as a topological surface for a watershed grouping stage. Considering the video as a 3-D volume results in a batch segmentation instead of the traditional frame-by-frame segmentation. The main advantages of this approach are: 1) exploiting the time continuity in the frame sequence , 2) avoiding problems in tracking regions from frame to frame, 3) using a fast watershed-based method informing the final video regions. Preliminary experimental results are very promising. Motaz A. El Saban, B. S. Manjunath |
ICIP (1) | 2 |
| 2003 | Joint source-channel coding scheme for image-in-image data hidingabstractWe consider the problem of hiding images in images. In addition to the usual design constraints such as imperceptible host degradation and robustness in presence of variety of attacks, we impose the condition that the quality of the recovered signature image should be better if the attack is milder. We present a simple hybrid analog-digital hiding technique for this purpose. The signature image is compressed efficiently (using JPEG) into a sequence of bits, which is hidden using a previously proposed digital hiding scheme. The residual error between the original and compressed signature image is then hidden using an analog hiding scheme. The results show (perceptual as well as mean-square error) improvement as the attack becomes milder. Kaushal Solanki, Onkar Dabeer, B. S. Manjunath, Upamanyu Madhow, Shivkumar Chandrasekaran |
ICIP (2) | 3 |
| 2003 | LLRT based detection of LSB hidingabstractIn this paper we consider a hypothesis testing approach for detection of hiding in the least significant bit (LSB). This steganalysis problem is a composite hypothesis testing problem. We state a regularity condition on the image histogram, which reduces this problem to a simple hypothesis testing problem. We then develop a number of simple practical tests based on the estimation of the optimal log likelihood ratio statistic. We show that our tests significantly outperform Stegdetect, a popular hypothesis test available in the literature. Our approach also leads to good estimates of the hiding rate. Kenneth Sullivan, Onkar Dabeer, Upamanyu Madhow, B. S. Manjunath, Shivkumar Chandrasekaran |
ICIP (1) | 4 |
| 2003 | Image segmentation using multi-region stability and edge strengthabstractA novel scheme for image segmentation is presented. An image segmentation criterion is proposed that groups similar pixels together to form regions. This criterion is formulated as a cost function. Using gradient-descent methods, which lead to a curve evolution equation that segments the image into multiple homogenous regions, minimizes this cost function. Homogeneity is specified through a pixel-to-pixel similarity measure, which is defined by the user and can be adaptive based on the current application. To improve the performance of the system, an edge function is also used to adjust the speed of the competing curves. The proposed method can be easily applied to vector valued images such as texture and color images without a significant addition to computational complexity. Baris Sumengen, B. S. Manjunath, Charles S. Kenney |
ICIP (3) | 2 |
| 2003 | A semantic representation for image retrievalabstractRobust semantic labeling of image regions is a basic problem in representing and retrieving image/video content. We propose an SVM-MRF framework to model features and their spatial distributions, leading towards a "semantic" representation. Eigenfeatures of Gabor wavelet features and Gaussian mixture model are used for feature clustering. Since similar feature vectors in one cluster can come from several different semantic classes, SVM is applied to represent conditioned feature vector distributions within each cluster, and a Markov random field is used to model the spatial distributions of the semantic labels. A semantic layout representation is proposed to describe the semantics of the images. Experiments show that this method can improve semantic labeling and is useful in similarity search. B. S. Manjunath |
ICIP (2) | 2 |
| 2003 | A Condition Number for Point Matching with Application to Registration and Postregistration Error EstimationabstractSelecting salient points from two or more images for computing correspondence is a well-studied problem in image analysis. This paper describes a new and effective technique for selecting these tiepoints using condition numbers, with application to image registration and mosaicking. Condition numbers are derived for point-matching methods based on minimizing windowed objective functions for 1) translation, 2) rotation-scaling-translation (RST), and 3) affine transformations. Our principal result is that the condition numbers satisfy K/sub Trans/ /spl les/ K/sub RST/ /spl les/ K/sub Affine/. That is, if a point is ill-conditioned with respect to point-matching via translation, then it is also unsuited for matching with respect to RST and affine transforms. This is fortunate since K/sub Trans/ is easily computed whereas K/sub RST/ and K/sub Affine/ are not. The second half of the paper applies the condition estimation results to the problem of identifying tiepoints in pairs of images for the purpose of registration. Once these points have been matched (after culling outliers using a RANSAC-like procedure), the registration parameters are computed. The postregistration error between the reference image and the stabilized image is then estimated by evaluating the translation between these images at points exhibiting good conditioning with respect to translation. The proposed method of tiepoint selection and matching using condition number provides a reliable basis for registration. The method has been tested on a large number of diverse collection of images - multidate Landsat images, aerial images, aerial videos, and infrared images. A Web site where the users can try our registration software is available and is being actively used by researchers around the world. Charles S. Kenney, B. S. Manjunath, Marco Zuliani, Gary A. Hewer, Alan Van Nevel |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2002 | High-volume data hiding in images: Introducing perceptual criteria into quantization based embeddingabstractInformation-theoretic analyses for data hiding prescribe embedding the hidden data in the choice of quantizer for the host data. In this paper, we consider a suboptimal implementation of this prescription, with a view to hiding high volumes of data in images with low perceptual degradation. Our two main findings are as follows: (a) In order to limit perceptual distortion while hiding large amounts of data, the hiding scheme must use perceptual criteria in addition to information-theoretic guidelines. (b) By focusing on “benign” JPEG compression attacks, we are able to attain very high volumes of embedded data, comparable to information-theoretic capacity estimates for the more malicious Additive White Gaussian Noise (AWGN) attack channel, using relatively simple embedding techniques. Kaushal Solanki, Noah Jacobsen, Shivkumar Chandrasekaran, Upamanyu Madhow, B. S. Manjunath |
ICASSP | 5 |
| 2002 | Modeling object classes in aerial images using hidden Markov modelsabstractA canonical model is proposed for object classes in aerial images. This model is motivated by the observation that geographic regions of interest are characterized by collections of texture motifs corresponding to geographic processes. Furthermore, the spatial arrangement of the motifs is an important discriminating characteristic. In our approach, the states of a hidden Markov model (HMM) correspond to the texture motifs and the state transitions correspond to the spatial arrangement of the motifs. A one-dimensional approach reduces the computational complexity. The model is shown to be effective in characterizing objects of interest in spatial datasets in terms of their underlying texture motifs. The potential of the model for identifying the classes of unlabeled objects is demonstrated. Shawn D. Newsam, Sitaram Bhagavathy, B. S. Manjunath |
ICIP (1) | 3 |
| 2002 | Image segmentation using curve evolution and flow fieldsabstractAn image segmentation scheme that utilizes image-based flow fields in a curve evolution framework is presented. Geometric curve evolution methods require an edge function and a vector field with certain characteristics that are obtained from the image itself. A vector field borrowed from the edgeflow segmentation (Ma and Manjunath 2000) method is utilized both to obtain an edge function and to guide the curve evolution towards the object boundaries. This vector field is computed from the image using intensity, texture and color features. The proposed method integrates well-tested image features with the well-studied curve evolution methods thus achieving better segmentation results. Baris Sumengen, B. S. Manjunath, Charles S. Kenney |
ICIP (1) | 2 |
| 2002 | Representation of motion activity in hierarchical levels for video indexing and filteringabstractA method for video indexing and filtering based on motion activity characteristics in hierarchical levels is proposed. To extract motion activity information, an MPEG (MPEG-1/2) video is first adaptively segmented into hierarchical levels with fixed percentage of original video length based on P-frame macroblock motion information. Three motion activity characteristics - motion intensity which represents the degree of change in motion, motion intensity histogram which represents the temporal statistics of motion intensity, and spatial descriptor which represents the spatial attribute of motion, are then computed to represent different levels of video. The descriptors from different levels are used selectively in different steps of video indexing and filtering. Experimental results show the proposed method is fast and effective, and provides a powerful video indexing and filtering tool. Xinding Sun, Ajay Divakaran, B. S. Manjunath |
ICIP (1) | 3 |
| 2002 | Panoramic capturing and recognition of human activityabstractThis paper presents a unified approach to human activity capturing and recognition. It targets applications such as a speaker walking, turning around, sitting and getting up from a chair in a classroom setting. A panoramic camera capturing system is designed for video capture. Virtual camera control outputs the region of interest (ROI) video that covers the speaker. Given an ROI sequence, the virtual camera control parameters are used for the recognition of activities like walking, and the motion parameters of each frame are used for the recognition of other activities like turning around, sitting down and getting up etc. For motion parameter based recognition, the likelihood of the motion parameters is represented using a multivariate Gaussian model. The temporal change of the likelihood is characterized using a continuous density hidden Markov model (HMM). Experimental results show that the method works well in recognizing the above mentioned human body activities. Xinding Sun, B. S. Manjunath |
ICIP (2) | 2 |
| 2002 | Detecting path intersections in panoramic videoabstractGiven panoramic video taken along a self-intersecting path, we present a method for detecting the intersection points. This allows "virtual tours" to be synthesized by splicing the panoramic video at the intersection points. Spatial intersections are detected by finding the best-matching panoramic images from a number of nearby candidates. Each panoramic image is segmented into horizontal strips. Each strip is averaged in the vertical direction. The Fourier coefficients of the resulting 1-D data capture the rotation-invariant horizontal texture of each panoramic image. The distance between two panoramic images is calculated as the sum of the distances between their strip texture pairs at the same row positions. The intersection is chosen as the two candidate panoramic images that have the minimum distance. Xinding Sun, Don Kimber, Jonathan Foote, B. S. Manjunath |
ICME (2) | 4 |
| 2002 | Scalable spatial event representationabstractThis work introduces a conceptual representation for complex spatial arrangements of image features in large multimedia datasets. A novel data structure, termed the spatial event cube (SEC), is formed from the co-occurrence matrices of perceptually classified features with respect to specific spatial relationships. A visual thesaurus constructed using supervised and unsupervised learning techniques is used to label the image features. SECs can be used to not only visualize the dominant spatial arrangements of feature classes but also discover non-obvious configurations. SECs also provide the framework for high-level data mining techniques such as using the generalized association rule approach. Experimental results are provided for a large dataset of aerial images. Jelena Tesic, Shawn D. Newsam, B. S. Manjunath |
ICME (2) | 3 |
| 2002 | System for automatic registration of remote sensing imagesabstractDescribes a system for automatic and semi-automatic registration/mosaic of remote sensing images. Information provided by the user can be used to speed up processing or to avoid mismatched control points. A statistical procedure is used to characterize good and bad registrations. Based on this "good fit-bad fit" statistical test the user can stop, modify the parameters, or continue the processing. Several tests have been performed by registering optical, radar, multi-sensor, high-resolution images and video sequences. We have included very difficult image registration examples in order to show the strengths and limits of our system. An online registration system demo containing several examples can be executed using a Web browser. Dmitry V. Fedorov, Leila M. G. Fonseca, Charles S. Kenney, B. S. Manjunath |
IGARSS | 4 |
| 2001 | Category-based image retrievalabstractThis work presents a novel approach to content-based image retrieval in categorical multimedia databases. The images are indexed using a combination of text and content descriptors. The categories are viewed as semantic clusters of images and are used to confine the search space. Keywords are used to identify candidate categories. Content-based retrieval is performed in these categories using multiple image features. Relevance feedback is used to learn the user's intent-query specification and feature-weighting-with minimal user-interface abstraction. The method is applied to a large number of images collected from a popular categorical structure on the World Wide Web. Results show that efficient and accurate performance is achievable by exploiting the semantic classification represented by the categories. The relevance feedback loop allows the content descriptor weightings to be determined without exposing the calculations to the user. Shawn D. Newsam, Baris Sumengen, B. S. Manjunath |
ICIP (3) | 3 |
| 2001 | Recording the region of interest from FlyCam panoramic videoabstractA novel method for region of interest tracking and recording video is presented. The proposed method is based on the FlyCam system (Foote et al., 2000), which produces high-resolution and wide-angle video sequences by stitching the video frames from multiple stationary cameras. The method integrates tracking and recording processes, and targets applications such as classroom lectures and videoconferencing. First, the region of interest (which typically covers the speaker) is tracked using a Kalman filter. Then, the Kalman filter estimation results are used for virtual camera control and to record the video. The system has no physical camera motion and the virtual camera parameters are readily available for video indexing. The proposed system has been implemented for real-time recording of lectures and presentations. Xinding Sun, Jonathan Foote, Don Kimber, B. S. Manjunath |
ICIP (1) | 4 |
| 2001 | Panoramic video capturing and compressed domain virtual camera controlabstractA system for capturing panoramic video and a novel method for corresponding compressed domain virtual camera control is presented. It targets applications such as classroom lectures and video conferencing. The proposed method is based on the FlyCam panoramic video system that is designed to produce high resolution and wide-angle video sequences by stitching the video pictures from multiple stationary cameras. The panoramic video sequence is compressed into an MPEG-2 stream for delivery. The proposed method integrates region of Interest (ROI) detection, tracking, and virtual camera control, and works on compressed domain information only. It first detects the ROI in the P (predictive coded) picture using only the macroblock type information, It then up-samples this detection result to obtain the ROI of the whole video stream. The ROI is tracked using a Kalman filter. The Kalman filter estimation results are used for virtual camera control that simulates human controlled video recording. The system has no physical camera motion and the virtual camera parameters are readily available for video indexing. The proposed system has been implemented for real time processing. Xinding Sun, Jonathan Foote, Don Kimber, B. S. Manjunath |
ACM Multimedia | 4 |
| 2001 | Adaptive nearest neighbor search for relevance feedback in large image databasesabstractRelevance feedback is often used in refining similarity retrievals in image and video databases. Typically this involves modification to the similarity metrics based on the user feedback and recomputing a set of nearest neighbors using the modified similarity values. Such nearest neighbor computations are expensive given that typical image features, such as color and texture, are represented in high dimensional spaces. Search complexity is a ciritcal issue while dealing with large databases and this issue has not received much attention in relevance feedback research. Most of the current methods report results on very small data sets, of the order of few thousand items, where a sequential (and hence exhaustive search) is practical. The main contribution of this paper is a novel algorithm for adaptive nearest neigbor computations for high dimensional feature vectors and when the number of items in the databse is large. The proposed method exploits the correlations between two consecutive nearest neighbor searches when the underlying similarity metric is changing, and filters out a significant number of candidates ina two stage search and retrieval process, thus reducing the number of I/O accesses to the database. Detailed experimental results are provided using a set of about 700,000 images. Comparision to the existing method shows an order of magnitude overall imporovement. Peng Wu 0005, B. S. Manjunath |
ACM Multimedia | 2 |
| 2001 | Unsupervised Segmentation of Color-Texture Regions in Images and VideoabstractA method for unsupervised segmentation of color-texture regions in images and video is presented. This method, which we refer to as JSEG, consists of two independent steps: color quantization and spatial segmentation. In the first step, colors in the image are quantized to several representative classes that can be used to differentiate regions in the image. The image pixels are then replaced by their corresponding color class labels, thus forming a class-map of the image. The focus of this work is on spatial segmentation, where a criterion for "good" segmentation using the class-map is proposed. Applying the criterion to local windows in the class-map results in the "J-image," in which high and low values correspond to possible boundaries and interiors of color-texture regions. A region growing method is then used to segment the image based on the multiscale J-images. A similar approach is applied to video sequences. An additional region tracking scheme is embedded into the region growing process to achieve consistent segmentation and tracking results, even for scenes with nonrigid object motion. Experiments show the robustness of the JSEG algorithm on real images and video. Yining Deng, B. S. Manjunath |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2001 | Color and texture descriptorsabstractThis paper presents an overview of color and texture descriptors that have been approved for the Final Committee Draft of the MPEG-7 standard. The color and texture descriptors that are described in this paper have undergone extensive evaluation and development during the past two years. Evaluation criteria include effectiveness of the descriptors in similarity retrieval, as well as extraction, storage, and representation complexities. The color descriptors in the standard include a histogram descriptor that is coded using the Haar transform, a color structure histogram, a dominant color descriptor, and a color layout descriptor. The three texture descriptors include one that characterizes homogeneous texture regions and another that represents the local edge distribution. A compact descriptor that facilitates texture browsing is also defined. Each of the descriptors is explained in detail by their semantics, extraction and usage. The effectiveness is documented by experimental results. B. S. Manjunath, Jens-Rainer Ohm, Vinod V. Vasudevan, Akio Yamada |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2001 | An efficient color representation for image retrievalabstractA compact color descriptor and an efficient indexing method for this descriptor are presented. The target application is similarity retrieval in large image databases using color. Colors in a given region are clustered into a small number of representative colors. The feature descriptor consists of the representative colors and their percentages in the region. A similarity measure similar to the quadratic color histogram distance measure is defined for this descriptor. The representative colors can be indexed in the three-dimensional (3-D) color space thus avoiding the high-dimensional indexing problems associated with the traditional color histogram. For similarity retrieval, each representative color in the query image or region is used independently to find regions containing that color. The matches from all of the query colors are then combined to obtain the final retrievals. An efficient indexing scheme for fast retrieval is presented. Experimental results show that this compact descriptor is effective and compares favorably with the traditional color histogram in terms of overall computational complexity. Yining Deng, B. S. Manjunath, Charles S. Kenney, Michael S. Moore, Hyundoo Shin |
IEEE Trans. Image Process. | 2 |
| 2001 | Peer group image enhancementabstractPeer group image processing identifies a "peer group" for each pixel and then replaces the pixel intensity with the average over the peer group. Two parameters provide direct control over which image features are selectively enhanced: area (number of pixels in the feature) and window diameter (window size needed to enclose the feature). A discussion is given of how these parameters determine which features in the image are smoothed or preserved. We show that the Fisher discriminant can be used to automatically adjust the peer group averaging (PGA) parameters at each point in the image. This local parameter selection allows smoothing over uniform regions while preserving features like corners and edges. This adaptive procedure extends to multilevel and color forms of PGA. Comparisons are made with a variety of standard filtering techniques and an analysis is given of computational complexity and convergence issues. Charles S. Kenney, Yining Deng, B. S. Manjunath, Gary A. Hewer |
IEEE Trans. Image Process. | 3 |
| 2000 | Dimensionality Reduction for Image RetrievalabstractDimensionality reduction methods are of interest in applications such as content based image and video retrieval. In large multimedia databases, it may not be practical to search through the entire database in order to retrieve the nearest neighbors of a query. Good data structures for similarity search and indexing are needed, and the existing data structures do not scale well for the high dimensional multimedia descriptors. We investigate the use of weighted multi-dimensional scaling (WMDS) for dimensionality reduction. The main objective of the WMDS is to preserve the local topology of the high dimensional space, i.e., to map the nearest neighbors in the high dimensional space to nearest neighbors in the lower dimensional space. In addition to the well known retrieval accuracy as a measure of performance, we propose two additional measures that take into account the ordinal relationships among the nearest neighbors. Experimental results are given. Peng Wu 0005, B. S. Manjunath, H. D. Shin |
ICIP | 2 |
| 2000 | A texture descriptor for browsing and similarity retrieval
Peng Wu 0005, B. S. Manjunath, Shawn D. Newsam, H. D. Shin |
Signal Process. Image Commun. | 2 |
| 2000 | EdgeFlow: a technique for boundary detection and image segmentationabstractA novel boundary detection scheme based on "edge flow" is proposed in this paper. This scheme utilizes a predictive coding model to identify the direction of change in color and texture at each image location at a given scale, and constructs an edge flow vector. By propagating the edge flow vectors, the boundaries can be detected at image locations which encounter two opposite directions of flow in the stable state. A user defined image scale is the only significant control parameter that is needed by the algorithm. The scheme facilitates integration of color and texture into a single framework for boundary detection. Segmentation results on a large and diverse collections of natural images are provided, demonstrating the usefulness of this method to content based image retrieval. Wei-Ying Ma, B. S. Manjunath |
IEEE Trans. Image Process. | 2 |
| 2000 | Guest editorial introduction to the special issue on image and video processing for digital libraries
B. S. Manjunath, Thomas S. Huang, A. Murat Tekalp, HongJiang Zhang |
IEEE Trans. Image Process. | 1 |
| 1999 | Color Image SegmentationabstractIn this work, a new approach to fully automatic color image segmentation, called JSEG, is presented. First, colors in the image are quantized to several representing classes that can be used to differentiate regions in the image. Then, image pixel colors are replaced by their corresponding color class labels, thus forming a class-map of the image. A criterion for "good" segmentation using this class-map is proposed. Applying the criterion to local windows in the class-map results in the "J-image", in which high and low values correspond to possible region boundaries and region centers, respectively. A region growing method is then used to segment the image based on the multi-scale J-images. Experiments show that JSEG provides good segmentation results on a variety of images. Yining Deng, B. S. Manjunath, Hyundoo Shin |
CVPR | 2 |
| 1999 | Subset Selection for Active Object RecognitionabstractThis paper presents an algorithm for constructing object representations suitable for recognition. The system automatically selects a representative subset of the views of the object while constructing the eigenspace basis. These views are actively located for object identification and pose determination. All processing is performed on-line. The camera is actively positioned during both representation and recognition. When tested with 240 views for each of seven objects, the system achieves 100% accurate object recognition and pose determination. These results are shown to degrade gracefully as conditions deteriorate. Jay Winkeler, B. S. Manjunath, Shivkumar Chandrasekaran |
CVPR | 2 |
| 1999 | An efficient low-dimensional color indexing scheme for region-based image retrievalabstractIn this work, an efficient low-dimensional color indexing scheme for region-based image retrieval is presented. The colors in each image region are first quantized so that only a small number of cluster centroids are needed to represent the region color information. The proposed color feature descriptor consists of these quantized colors and their percentages in the region. A similarity distance measure is defined and shown to be equivalent to the quadratic color histogram distance measure. The quantized colors are indexed in the 3-D color space so that high-dimensional indexing can be avoided. During the search process, each quantized color in the query is used as a separate cue to find matches containing that color. The matches from all the query colors are then joined to obtain the final retrievals. Experimental results show that the proposed scheme is fast and accurate compared to the color histogram approach. Yining Deng, B. S. Manjunath |
ICASSP | 2 |
| 1999 | Data Hiding in VideoabstractWe propose a video data embedding scheme in which the embedded signature data is reconstructed without knowing the original host video. The proposed method enables a high rate of data embedding and is robust to motion compensated coding, such as MPEG-2. Embedding is based on texture masking and utilizes a multi-dimensional lattice structure for encoding signature information. Signature data is embedded in individual video frames using the block DCT. The embedded frames are then MPEG-2 coded. At the receiver both the host and signature images are recovered from the embedded bit stream. We present examples of embedding image and video in video. Jong Jin Chae, B. S. Manjunath |
ICIP (1) | 2 |
| 1999 | NeTra: A Toolbox for Navigating Large Image Databases
Wei-Ying Ma, B. S. Manjunath |
Multim. Syst. | 2 |
| 1999 | Rotation-invariant texture classification using a complete space-frequency modelabstractA method of rotation-invariant texture classification based on a complete space-frequency model is introduced. A polar, analytic form of a two-dimensional (2-D) Gabor wavelet is developed, and a multiresolution family of these wavelets is used to compute information-conserving microfeatures. From these microfeatures a micromodel, which characterizes spatially localized amplitude, frequency, and directional behavior of the texture, is formed. The essential characteristics of a texture sample, its macrofeatures, are derived from the estimated selected parameters of the micromodel. Classification of texture samples is based on the macromodel derived from a rotation invariant subset of macrofeatures. In experiments, comparatively high correct classification rates were obtained using large sample sets. George M. Haley, B. S. Manjunath |
IEEE Trans. Image Process. | 2 |
| 1998 | Color Image Embedding using Multidimensional Lattice StructuresabstractThis paper describes a robust data embedding scheme which uses noise resilient channel codes based on a multidimensional lattice structure. Compared to prior work in digital watermarking, the proposed scheme can handle a significantly larger quantity of signature data such as gray-scale or color images. A trade-off between the quantity of hidden data and the quality of the watermarked image is achieved by varying the number of quantization levels for the signature, and a scale factor for data embedding. Experimental results on signature recovery from JPEG compressed watermarked images show that good quality reconstruction is possible even when the images are lossy compressed by as much as 85%. Potential applications of this method include, in addition to watermarking, digital data hiding for security and for bit stream control and manipulation. Jong Jin Chae, Debargha Mukherjee, B. S. Manjunath |
ICIP (1) | 3 |
| 1998 | Reversible Wavelet and Spectral Transforms for Lossless Compression of Color Images
Norbert Strobel, Sanjit K. Mitra, B. S. Manjunath |
ICIP (3) | 3 |
| 1998 | A Texture Thesaurus for Browsing Large Aerial PhotographsabstractA texture-based image retrieval system for browsing large-scale aerial photographs is presented. The salient components of this system include texture feature extraction, image segmentation and grouping, learning similarity measure, and a texture thesaurus model for fast search and indexing. The texture features are computed by filtering the image with a bank of Gabor filters. This is followed by a texture gradient computation to segment each large airphoto into homogeneous regions. A hybrid neural network algorithm is used to learn the visual similarity by clustering patterns in the feature space. With learning similarity, the retrieval performance improves significantly. Finally, a texture image thesaurus is created by combining the learning similarity algorithm with a hierarchical vector quantization scheme. This thesaurus facilitates the indexing process while maintaining a good retrieval performance. Experimental results demonstrate the robustness of the overall system in searching over a large collection of airphotos and in selecting a diverse collection of geographic features such as housing developments, parking lots, highways, and airports. © 1998 John Wiley & Sons, Inc. Wei-Ying Ma, B. S. Manjunath |
J. Am. Soc. Inf. Sci. | 2 |
| 1998 | NeTra-V: toward an object-based video representationabstractWe present a prototype video analysis and retrieval system, called NeTra-V, that is being developed to build an object-based video representation for functionalities such as search and retrieval of video objects. A region-based content description scheme using low-level visual descriptors is proposed. In order to obtain regions for local feature extraction, a new spatio-temporal segmentation and region-tracking scheme is employed. The segmentation algorithm uses all three visual features: color, texture, and motion in the video data. A group processing scheme similar to the one in the MPEG-2 standard is used to ensure the robustness of the segmentation. The proposed approach can handle complex scenes with large motion. After segmentation, regions are tracked through the video sequence using extracted local features. The results of tracking are sequences of coherent regions, called "subobjects". Subobjects are the fundamental elements in our low-level content description scheme, which can be used to obtain meaningful physical objects in a high-level content description scheme. Experimental results illustrating segmentation and retrieval are provided. Yining Deng, B. S. Manjunath |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 1998 | Variational image segmentation using boundary functionsabstractA general variational framework for image approximation and segmentation is introduced. By using a continuous "line-process" to represent edge boundaries, it is possible to formulate a variational theory of image segmentation and approximation in which the boundary function has a simple explicit form in terms of the approximation function. At the same time, this variational framework is general enough to include the most commonly used objective functions. Application is made to Mumford-Shah type functionals as well as those considered by Geman and others. Employing arbitrary Lp norms to measure smoothness and approximation allows the user to alternate between a least squares approach and one based on total variation, depending on the needs of a particular image. Since the optimal boundary function that minimizes the associated objective functional for a given approximation function can be found explicitly, the objective functional can be expressed in a reduced form that depends only on the approximating function. From this a partial differential equation (PDE) descent method, aimed at minimizing the objective functional, is derived. The method is fast and produces excellent results as illustrated by a number of real and synthetic image problems. Gary A. Hewer, Charles S. Kenney, B. S. Manjunath |
IEEE Trans. Image Process. | 3 |
| 1997 | Edge Flow: A Framework of Boundary Detection and Image SegmentationabstractA novel boundary detection scheme based on "edge flow" is proposed in this paper. This scheme utilizes a predictive coding model to identify the direction of change in color and texture at each image location at a given scale, and constructs an edge flow vector. By iteratively propagating the edge flow, the boundaries can be detected at image locations which encounter two opposite directions of flow in the stable state. A user defined image scale is the only significant control parameter that is needed by the algorithm. The scheme facilitates integration of color and texture into a single framework for boundary detection. Wei-Ying Ma, B. S. Manjunath |
CVPR | 2 |
| 1997 | Dimensionality Reduction Using Multi-Dimensional Scaling for Content-Based RetrievalabstractThere has been much interest recently in image content based retrieval, with applications to digital libraries and image database accessing. One approach to this problem is to base retrieval from the database upon feature vectors which characterize the image texture. Since feature vectors are often high dimensional, multi-dimensional scaling, or non-linear principal components analysis (PCA) may be useful in reducing feature vector size, and therefore computation time. We have investigated a variant of the non-linear PCA algorithm described by Webb (see Pattern Recognition, vol.28, no.6, p.753-9, 1995) and its usefulness in the database retrieval problem. The results are quite impressive; in an experiment using an aerial photo database, the feature vector length was reduced by a factor of 10 without significantly reducing the retrieval performance. Morris Beatty, B. S. Manjunath |
ICIP (2) | 2 |
| 1997 | Content-Based Search of Video Using Color, Texture, and MotionabstractWe present an implementation of a system for content-based search and retrieval of video based on low-level visual features. The system consists of three parts, automatic video partition, feature extraction, video search and retrieval. Three primary features, color texture and motion are used for indexing. They are represented by color histograms, Gabor texture features, and motion histograms. Most of the processing is done directly in the MPEG compressed domain. Testing on sports and movie databases have shown good retrieval performance. Yining Deng, B. S. Manjunath |
ICIP (2) | 2 |
| 1997 | NeTra: A Toolbox for Navigating Large Image DatabasesabstractWe present an implementation of NeTra, a prototype image retrieval system that uses color texture, shape and spatial location information in segmented image database. A distinguishing aspect of this system is its incorporation of a robust automated image segmentation algorithm that allows object or region based search. Image segmentation significantly improves the quality of image retrieval when images contain multiple complex objects. Other important components of the system include an efficient color representation, and indexing of color, texture, and shape features for fast search and retrieval. This representation allows the user to compose interesting queries such as "retrieve all images that contain regions that have the color of object A, texture of object B, shape of object C, and lie in the upper one-third of the image" where the individual objects could be regions belonging to different images. Wei-Ying Ma, B. S. Manjunath |
ICIP (1) | 2 |
| 1997 | Model-Based Detection and Correction of Corrupted Wavelet CoefficientsabstractImage decomposition based on the discrete wavelet transform (DWT) has been proposed for efficient storage and progressive transmission of images for visual browsing in digital image libraries. Although the compression aspects of the DWT have been carefully researched, reconstruction errors due to corrupted wavelet coefficients have received less attention. In this paper we consider the problem of bit errors affecting uniformly quantized wavelet coefficients. The proposed method, which is based on a local image model, simultaneously detects and masks corrupted wavelet coefficients. Norbert Strobel, Sanjit K. Mitra, B. S. Manjunath |
ICIP (1) | 3 |
| 1997 | An Eigenspace Update Algorithm for Image Analysis
Shivkumar Chandrasekaran, B. S. Manjunath, Yuan-Fang Wang, Jay Winkeler, Henry Zhang |
CVGIP Graph. Model. Image Process. | 2 |
| 1996 | Texture Features and Learning SimilarityabstractThis paper addresses two important issues related to texture pattern retrieval: feature extraction and similarity search. A Gabor feature representation for textured images is proposed, and its performance in pattern retrieval is evaluated on a large texture image database. These features compare favorably with other existing texture representations. A simple hybrid neural network algorithm is used to learn the similarity by simple clustering in the texture feature space. With learning similarity the performance of similar pattern retrieval improves significantly. An important aspect of this work is its application to real image data. Texture feature extraction with similarity learning is used to search through large aerial photographs. Feature clustering enables efficient search of the database as our experimental results indicate. Wei-Ying Ma, B. S. Manjunath |
CVPR | 2 |
| 1996 | Image segmentation via functionals based on boundary functionsabstractA general variational framework for image approximation and segmentation is introduced in which the boundary function has a simple explicit form in terms of the approximation function. At the same time, this variational framework is general enough to include the most commonly used objective functions. Since the optimal boundary function, that minimizes the associated objective functional for a given approximation function, can be found explicitly, the objective functional can be expressed in a reduced form that depends only on the approximating function. From this a partial differential equation descent method, aimed at minimizing the objective functional, is derived. The method is fast and produces excellent results as illustrated by a number of real and synthetic image problems. Gary A. Hewer, Charles S. Kenney, B. S. Manjunath |
ICIP (1) | 3 |
| 1996 | Browsing large satellite and aerial photographsabstractImage content based retrieval in the Alexandria digital library project has focussed on texture and color features for querying the database. A robust texture feature extraction algorithm and a fast segmentation scheme have been developed. The texture features are computed by filtering the image with a bank of Gabor filters. This is followed by a clustering scheme to create a texture based feature dictionary, which is then used to search and retrieve similar looking patterns from other images. Experimental results demonstrate the robustness of the overall system in searching over a large collection of airphotos and in selecting a surprisingly diverse collection of geographic features such as housing developments, parking lots, highways, and airports. B. S. Manjunath, Wei-Ying Ma |
ICIP (2) | 1 |
| 1996 | Texture-Based Pattern Retrieval from Image Databases
Wei-Ying Ma, B. S. Manjunath |
Multim. Tools Appl. | 2 |
| 1996 | Texture Features for Browsing and Retrieval of Image DataabstractImage content based retrieval is emerging as an important research area with application to digital libraries and multimedia databases. The focus of this paper is on the image processing aspects and in particular using texture information for browsing and retrieval of large image data. We propose the use of Gabor wavelet features for texture analysis and provide a comprehensive experimental evaluation. Comparisons with other multiresolution texture features using the Brodatz texture database indicate that the Gabor features provide the best pattern retrieval accuracy. An application to browsing large air photos is illustrated. B. S. Manjunath, Wei-Ying Ma |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1996 | A new approach to image feature detection with applications
B. S. Manjunath, Chandra Shekhar 0002, Rama Chellappa |
Pattern Recognit. | 1 |
| 1995 | Rotation-invariant texture classification using modified Gabor filtersabstractA method of rotation invariant texture classification based on a joint space-frequency model is introduced. Multiresolution filters, based on a truly analytic form of a polar 2-D Gabor (1946) wavelet, are used to compute spatial frequency-specific but spatially localized microfeatures. These microfeatures constitute an approximate basis set for the representation of the texture sample. The essential characteristics of a texture sample, its macrofeatures, are derived from the statistics of its microfeatures. A texture is modeled as a multivariate Gaussian distribution of macrofeatures. Classification is based on a rotation invariant subset of macrofeatures. George M. Haley, B. S. Manjunath |
ICIP | 2 |
| 1995 | A comparison of wavelet transform features for texture image annotationabstractA comparison of different wavelet transform based texture features for content based search and retrieval is made. These include the conventional orthogonal and bi-orthogonal wavelet transforms, tree-structured decompositions, and the Gabor wavelet transforms. Issues discussed include image processing complexity, texture classification and discrimination, and suitability for developing indexing techniques. Wei-Ying Ma, B. S. Manjunath |
ICIP | 2 |
| 1995 | Multisensor Image Fusion Using the Wavelet Transform
B. S. Manjunath, Sanjit K. Mitra |
CVGIP Graph. Model. Image Process. | 2 |
| 1995 | Automated segmentation of brain MR images
C. Tsai, B. S. Manjunath, R. Jagadeesan |
Pattern Recognit. | 2 |
| 1995 | A contour-based approach to multisensor image registrationabstractImage registration is concerned with the establishment of correspondence between images of the same scene. One challenging problem in this area is the registration of multispectral/multisensor images. In general, such images have different gray level characteristics, and simple techniques such as those based on area correlations cannot be applied directly. On the other hand, contours representing region boundaries are preserved in most cases. The authors present two contour-based methods which use region boundaries and other strong edges as matching primitives. The first contour matching algorithm is based on the chain-code correlation and other shape similarity criteria such as invariant moments. Closed contours and the salient segments along the open contours are matched separately. This method works well for image pairs in which the contour information is well preserved, such as the optical images from Landsat and Spot satellites. For the registration of the optical images with synthetic aperture radar (SAR) images, the authors propose an elastic contour matching scheme based on the active contour model. Using the contours from the optical image as the initial condition, accurate contour locations in the SAR image are obtained by applying the active contour model. Both contour matching methods are automatic and computationally quite efficient. Experimental results with various kinds of image data have verified the robustness of the algorithms, which have outperformed manual registration in terms of root mean square error at the control points. B. S. Manjunath, Sanjit K. Mitra |
IEEE Trans. Image Process. | 2 |
| 1994 | Multi-Sensor Image Fusion using the Wavelet TransformabstractIn the image fusion scheme presented in this paper, the wavelet transforms of the input images are appropriately combined, and the new image is obtained by taking the inverse wavelet transform of the fused wavelet coefficients. An area-based maximum selection rule and a consistency verification step are used for feature selection. A performance measure using specially generated test images is also suggested.> B. S. Manjunath, Sanjit K. Mitra |
ICIP (1) | 2 |
| 1993 | A unified approach to boundary perception: edges, textures, and illusory contoursabstractA model consisting of a multistage system which extracts and groups salient features in the image at different spatial scales (or frequencies) is used. In the first stage, a Gabor wavelet decomposition provides a representation of the image which is orientation selective and has optimal localization properties in space and frequency. This decomposition is useful in detecting significant features such as step and line edges at different scales and orientations in the image. Following the wavelet transformation, local competitive interactions are introduced to reduce the effects of noise and changes in illumination. Interscale interactions help in localizing the line ends and corners, and play a crucial role in boundary perception. The final stage groups similar features, aiding in boundary completion. The different stages can be identified with processing by simple, complex, and hypercomplex cells in the visual cortex of mammals. Experimental results demonstrate the performance of this model in detecting boundaries (both real and illusory) in real and synthetic images. B. S. Manjunath, Rama Chellappa |
IEEE Trans. Neural Networks | 1 |
| 1992 | A feature based approach to face recognitionabstractA feature-based approach to face recognition in which the features are derived from the intensity data without assuming any knowledge of the face structure is presented. The feature extraction model is biologically motivated, and the locations of the features often correspond to salient facial features such as the eyes, nose, etc. Topological graphs are used to represent relations between features, and a simple deterministic graph-matching scheme that exploits the basic structure is used to recognize familiar faces from a database. Each of the stages in the system can be fully implemented in parallel to achieve real-time recognition. Experimental results for a 128*128 image with very little noise are evaluated.> B. S. Manjunath, Rama Chellappa, Christoph von der Malsburg |
CVPR | 1 |
| 1992 | A robust method for detecting image features with application to face recognition and motion correspondenceabstractThe authors present an approach to feature detection, which is a fundamental issue in many intermediate-level vision problems such as stereo, motion correspondence, image registration, etc. The approach is based on a scale-interaction model of the end-inhibition property exhibited by certain cells in the visual- cortex of mammals. These feature detector cells are responsive to short lines, line endings, corners and other such sharp changes in curvature. In addition, this method also provides a compact representation of feature information which is useful in shape recognition problems. Application to face recognition and motion correspondence are illustrated.> B. S. Manjunath, Chandra Shekhar 0002, Rama Chellappa, Christoph von der Malsburg |
ICPR (2) | 1 |
| 1991 | A computational approach to boundary detectionabstractA unified approach to boundary perception is presented. The model consists of a hierarchical system which extracts and groups salient features in the image at different spatial scales. In the first stage a Gabor wavelet decomposition provides a representation of the image which is orientation selective, has optimal localization properties, and provides a good model for early feature detection. Following this, local competitive interactions are introduced which help in reducing the effects of noise and illumination variations. Scale interactions help in localizing line ends and corners, and play an important role in boundary perception. The final stage groups similar features aiding in boundary completion. Experimental results on detecting edges, texture boundaries, and illusory contours are provided.> B. S. Manjunath, Rama Chellappa |
CVPR | 1 |
| 1991 | Unsupervised Texture Segmentation Using Markov Random Field ModelsabstractThe problem of unsupervised segmentation of textured images is considered. The only explicit assumption made is that the intensity data can be modeled by a Gauss Markov random field (GMRF). The image is divided into a number of nonoverlapping regions and the GMRF parameters are computed from each of these regions. A simple clustering method is used to merge these regions. The parameters of the model estimated from the clustered segments are then used in two different schemes, one being all approximation to the maximum a posterior estimate of the labels and the other minimizing the percentage misclassification error. The proposed approach is contrasted with the algorithm of S. Lakshamanan and H. Derin (1989), which uses a simultaneous parameter estimation and segmentation scheme. The results of the adaptive segmentation algorithm of Lakshamanan and Derin are compared with a simple nearest-neighbor classification scheme to show that if enough information is available, simple techniques could be used as alternatives to computationally expensive schemes.> B. S. Manjunath, Rama Chellappa |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |