EDBT 2026 Demo / reviewers in the wild / expert
Vijay Mahadevan
dblp:43/2321 · also Vijay S. Mahadevan
· DBLP profile ↗
30ranked-venue papers
8as first author
9since 2021 · last 2024
0000-0002-3337-2607ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 7 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSystems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | No Head Left Behind - Multi-Head Alignment Distillation for TransformersabstractKnowledge distillation aims at reducing model size without compromising much performance. Recent work has applied it to large vision-language (VL) Transformers, and has shown that attention maps in the multi-head attention modules of vision-language Transformers contain extensive intra-modal and cross-modal co-reference relations to be distilled. The standard approach is to apply a one-to-one attention map distillation loss, i.e. the Teacher's first attention head instructs the Student's first head, the second teaches the second, and so forth, but this only works when the numbers of attention heads in the Teacher and Student are the same. To remove this constraint, we propose a new Attention Map Alignment Distillation (AMAD) method for Transformers with multi-head attention, which works for a Teacher and a Student with different numbers of attention heads. Specifically, we soft-align different heads in Teacher and Student attention maps using a cosine similarity weighting. The Teacher head contributes more to the Student heads for which it has a higher similarity weight. Each Teacher head contributes to all the Student heads by minimizing the divergence between the attention activation distributions for the soft-aligned heads. No head is left behind. This distillation approach operates like cross-attention. We experiment on distilling VL-T5 and BLIP, and apply AMAD loss on their T5, BERT, and ViT sub-modules. We show, under vision-language setting, that AMAD outperforms conventional distillation methods on VQA-2.0, COCO captioning, and Multi30K translation datasets. We further show that even without VL pre-training, the distilled VL-T5 models outperform corresponding VL pre-trained VL-T5 models that are further fine-tuned by ground-truth signals, and that fine-tuning distillation can also compensate to some degree for the absence of VL pre-training for BLIP models. Tianyang Zhao 0004, Kunwar Yashraj Singh, Srikar Appalaraju, Peng Tang 0005, Vijay Mahadevan, R. Manmatha, Ying Nian Wu |
AAAI | 5 |
| 2024 | Enhancing Vision-Language Pre-Training with Rich SupervisionsabstractWe propose Strongly Supervised pre-training with ScreenShots (S4) - a novel pre-training paradigm for Vision-Language Models using data from large-scale web screenshot rendering. Using web screenshots unlocks a treasure trove of visual and textual cues that are not present in using image-text pairs. In S4, we leverage the inherent tree-structured hierarchy of HTML elements and the spatial localization to carefully design 10 pre-training tasks with large scale annotated data. These tasks resemble down- stream tasks across different domains and the annotations are cheap to obtain. We demonstrate that, compared to current screenshot pre-training objectives, our innovative pre-training method significantly enhances performance of image-to-text model in nine varied and popular downstream tasks - up to 76.1% improvements on Table Detection, and at least 1 % on Widget Captioning. Kunyu Shi, Pengkai Zhu, Edouard Belval, Oren Nuriel, Srikar Appalaraju, Shabnam Ghadar, Zhuowen Tu, Vijay Mahadevan, Stefano Soatto |
CVPR | 9 |
| 2024 | DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding ModelsabstractSungnyun Kim, Haofu Liao, Srikar Appalaraju, Peng Tang, Zhuowen Tu, Ravi Kumar Satzoda, R. Manmatha, Vijay Mahadevan, Stefano Soatto. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Sungnyun Kim, Haofu Liao, Srikar Appalaraju, Peng Tang 0005, Zhuowen Tu, Ravi Kumar Satzoda, R. Manmatha, Vijay Mahadevan, Stefano Soatto |
EMNLP | 8 |
| 2024 | ICDAR 2024 Competition on Recognition and VQA on Handwritten Documents
Ajoy Mondal, Vijay Mahadevan, R. Manmatha, C. V. Jawahar |
ICDAR (6) | 2 |
| 2023 | PolyFormer: Referring Image Segmentation as Sequential Polygon GenerationabstractIn this work, instead of directly predicting the pixel-level segmentation masks, the problem of referring image seg-mentation is formulated as sequential polygon generation, and the predicted polygons can be later converted into segmentation masks. This is enabled by a new sequence-to-sequence framework, Polygon Transformer (PolyFormer), which takes a sequence of image patches and text query to-kens as input, and outputs a sequence of polygon vertices autoregressively. For more accurate geometric localization, we propose a regression-based decoder, which predicts the precise floating-point coordinates directly, without any co-ordinate quantization error. In the experiments, PolyFormer outperforms the prior art by a clear margin, e.g., 5.40% and 4.52% absolute improvements on the challenging Re-fCOCO+ and RefCOCOg datasets. It also shows strong generalization ability when evaluated on the referring video segmentation task without fine-tuning, e.g., achieving competitive 61.5% J&F on the Ref-DAVIS17 dataset. Zhaowei Cai, Ravi Kumar Satzoda, Vijay Mahadevan, R. Manmatha |
CVPR | 6 |
| 2023 | DocTr: Document Transformer for Structured Information Extraction in DocumentsabstractWe present a new formulation for structured information extraction (SIE) from visually rich documents. We address the limitations of existing IOB tagging and graph-based formulations, which are either overly reliant on the correct ordering of input text or struggle with decoding a complex graph. Instead, motivated by anchor-based object detectors in computer vision, we represent an entity as an anchor word and a bounding box, and represent entity linking as the association between anchor words. This is more robust to text ordering, and maintains a compact graph for entity linking. The formulation motivates us to introduce 1) a Document Transformer (DocTr) that aims at detecting and associating entity bounding boxes in visually rich documents, and 2) a simple pre-training strategy that helps learn entity detection in the context of language. Evaluations on three SIE benchmarks show the effectiveness of the proposed formulation, and the overall approach outperforms existing solutions. Haofu Liao, Aruni Roy Chowdhury, Ankan Bansal, Zhuowen Tu, Ravi Kumar Satzoda, R. Manmatha, Vijay Mahadevan |
ICCV | 9 |
| 2021 | End-to-end Piece-wise Unwarping of Document ImagesabstractDocument unwarping attempts to undo physical deformations of the paper and recover a ’flatbed’ scanned document-image for downstream tasks such as OCR. Current state-of-the-art relies on global unwarping of the document which is not robust to local deformation changes. Moreover, a global unwarping often produces spurious warping artifacts in less warped regions to compensate for severe warps present in other parts of the document. In this paper, we propose the first end-to-end trainable piece-wise unwarping1method that predicts local deformation fields and stitches them together with global information to obtain an improved unwarping. The proposed piece-wise formulation results in 4% improvement in terms of multi-scale structural similarity (MS-SSIM) and shows better performance in terms of OCR metrics, character error rate (CER) and word error rate (WER) compared to the state-of-the-art. Sagnik Das, Kunwar Yashraj Singh, Jon Wu, Erhan Bas, Vijay Mahadevan, Rahul Bhotika, Dimitris Samaras |
ICCV | 5 |
| 2021 | Visual Relationship Detection Using Part-and-Sum Transformers with Composite QueriesabstractComputer vision applications such as visual relationship detection and human object interaction can be formulated as a composite (structured) set detection problem in which both the parts (subject, object, and predicate) and the sum (triplet as a whole) are to be detected in a hierarchical fashion. In this paper, we present a new approach, denoted Part-and-Sum detection Transformer (PST), to perform end-to-end visual composite set detection. Different from existing Transformers in which queries are at a single level, we simultaneously model the joint part and sum hypotheses/interactions with composite queries and attention modules. We explicitly incorporate sum queries to enable better modeling of the part-and-sum relations that are absent in the standard Transformers. Our approach also uses novel tensor-based part queries and vector-based sum queries, and models their joint interaction. We report experiments on two vision tasks, visual relationship detection and human object interaction and demonstrate that PST achieves state of the art results among single-stage models, while nearly matching the results of custom designed two-stage models. Zhuowen Tu, Haofu Liao, Vijay Mahadevan, Stefano Soatto |
ICCV | 5 |
| 2021 | LayoutTransformer: Layout Generation and Completion with Self-attentionabstractWe address the problem of scene layout generation for diverse domains such as images, mobile applications, documents, and 3D objects. Most complex scenes, natural or human-designed, can be expressed as a meaningful arrangement of simpler compositional graphical primitives. Generating a new layout or extending an existing layout re- quires understanding the relationships between these primitives. To do this, we propose LayoutTransformer, a novel framework that leverages self-attention to learn contextual relationships between layout elements and generate novel layouts in a given domain. Our framework allows us to generate a new layout either from an empty set or from an initial seed set of primitives, and can easily scale to support an arbitrary of primitives per layout. Furthermore, our analyses show that the model is able to automatically capture the semantic properties of the primitives. We propose simple improvements in both representation of layout primitives, as well as training methods to demonstrate competitive performance in very diverse data domains such as object bounding boxes in natural images (COCO bounding box), documents (PubLayNet), mobile applications (RICO dataset) as well as 3D shapes (Part-Net). Code and other materials will be made available at https://kampta.github.io/layout. Kamal Gupta 0002, Justin Lazarow, Alessandro Achille, Larry Davis 0001, Vijay Mahadevan, Abhinav Shrivastava |
ICCV | 5 |
| 2019 | Scalable, High-Order Continuity Across Block Boundaries of Functional Approximations Computed in ParallelabstractWe investigate the representation of discrete scientific data with a Ckfunctional model, where Ckdenotes k-th order continuity, in a distributed-memory parallel setting. The Multivariate Functional Approximation (MFA) model is a piecewise-continuous functional approximation based on multi-variate high-dimensional B-splines. When computing an MFA approximation in parallel over multiple blocks in a spatial domain decomposition, the interior of each block will be Ck, k being the B-spline polynomial degree, but discontinuities exist across neighboring block boundaries. We present an efficient and scalable solution that involves blending neighboring approximations to ensure Ckcontinuity across block boundaries. We show that after decomposing the domain in structured, overlapping blocks and approximating blocks independently to high degrees of accuracy, we can extend the local solution, in a postprocessing step, to the global domain by using compact, multidimensional smoothstep functions. We prove that this approach, which can be viewed as an extended partition of unity approximation method, is scalable on high-performance computing architectures. Iulian R. Grindeanu, Tom Peterka, Vijay Mahadevan, Youssef S. G. Nashed |
CLUSTER | 3 |
| 2017 | Array-based, parallel hierarchical mesh refinement algorithms for unstructured meshes
Navamita Ray, Iulian R. Grindeanu, Xinglin Zhao, Vijay Mahadevan, Xiangmin Jiao |
Comput. Aided Des. | 4 |
| 2016 | VLAD3: Encoding Dynamics of Deep Features for Action RecognitionabstractPrevious approaches to action recognition with deep features tend to process video frames only within a small temporal region, and do not model long-range dynamic information explicitly. However, such information is important for the accurate recognition of actions, especially for the discrimination of complex activities that share sub-actions, and when dealing with untrimmed videos. Here, we propose a representation, VLAD for Deep Dynamics (VLAD3), that accounts for different levels of video dynamics. It captures short-term dynamics with deep convolutional neural network features, relying on linear dynamic systems (LDS) to model medium-range dynamics. To account for long-range inhomogeneous dynamics, a VLAD descriptor is derived for the LDS and pooled over the whole video, to arrive at the final VLAD3representation. An extensive evaluation was performed on Olympic Sports, UCF101 and THUMOS15, where the use of the VLAD3representation leads to state-of-the-art results. Yingwei Li 0001, Weixin Li 0002, Vijay Mahadevan, Nuno Vasconcelos |
CVPR | 3 |
| 2015 | Construction and evaluation of ontological tag trees
Chetan Kumar Verma, Vijay Mahadevan, Nikhil Rasiwasia, Gaurav Aggarwal, Alejandro Jaimes, Sujit Dey |
Expert Syst. Appl. | 2 |
| 2014 | Cluster Canonical Correlation AnalysisabstractIn this paper we present cluster canonical correlation analysis (cluster-CCA) for joint dimensionality reduction of two sets of data points. Unlike the standard pairwise correspondence between the data points, in our problem each set is partitioned into multiple clusters or classes, where the class labels define correspondences between the sets. Cluster-CCA is able to learn discriminant low dimensional representations that maximize the correlation between the two sets while segregating the different classes on the learned space. Furthermore, we present a kernel extension, kernel cluster canonical correlation analysis (cluster-KCCA) that extends cluster-CCA to account for non-linear relationships. Cluster-(K)CCA is shown to be computationally efficient, the complexity being similar to standard (K)CCA. By means of experimental evaluation on benchmark datasets, cluster-(K)CCA is shown to achieve state of the art performance for cross-modal retrieval tasks. Nikhil Rasiwasia, Dhruv Mahajan 0001, Vijay Mahadevan, Gaurav Aggarwal |
AISTATS | 3 |
| 2014 | Learning Optimal Seeds for Diffusion-Based Salient Object DetectionabstractIn diffusion-based saliency detection, an image is partitioned into superpixels and mapped to a graph, with superpixels as nodes and edge strengths proportional to superpixel similarity. Saliency information is then propagated over the graph using a diffusion process, whose equilibrium state yields the object saliency map. The optimal solution is the product of a propagation matrix and a saliency seed vector that contains a prior saliency assessment. This is obtained from either a bottom-up saliency detector or some heuristics. In this work, we propose a method to learn optimal seeds for object saliency. Two types of features are computed per superpixel: the bottom-up saliency of the superpixel region and a set of mid-level vision features informative of how likely the superpixel is to belong to an object. The combination of features that best discriminates between object and background saliency is then learned, using a large-margin formulation of the discriminant saliency principle. The propagation of the resulting saliency seeds, using a diffusion process, is finally shown to outperform the state of the art on a number of salient object detection datasets. Vijay Mahadevan, Nuno Vasconcelos |
CVPR | 2 |
| 2014 | Anomaly Detection and Localization in Crowded ScenesabstractThe detection and localization of anomalous behaviors in crowded scenes is considered, and a joint detector of temporal and spatial anomalies is proposed. The proposed detector is based on a video representation that accounts for both appearance and dynamics, using a set of mixture of dynamic textures models. These models are used to implement 1) a center-surround discriminant saliency detector that produces spatial saliency scores, and 2) a model of normal behavior that is learned from training data and produces temporal saliency scores. Spatial and temporal anomaly maps are then defined at multiple spatial scales, by considering the scores of these operators at progressively larger regions of support. The multiscale scores act as potentials of a conditional random field that guarantees global consistency of the anomaly judgments. A data set of densely crowded pedestrian walkways is introduced and used to evaluate the proposed anomaly detector. Experiments on this and other data sets show that the latter achieves state-of-the-art anomaly detection results. Weixin Li 0002, Vijay Mahadevan, Nuno Vasconcelos |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2013 | Biologically Inspired Object Tracking Using Center-Surround Saliency MechanismsabstractA biologically inspired discriminant object tracker is proposed. It is argued that discriminant tracking is a consequence of top-down tuning of the saliency mechanisms that guide the deployment of visual attention. The principle of discriminant saliency is then used to derive a tracker that implements a combination of center-surround saliency, a spatial spotlight of attention, and feature-based attention. In this framework, the tracking problem is formulated as one of continuous target-background classification, implemented in two stages. The first, or learning stage, combines a focus of attention (FoA) mechanism, and bottom-up saliency to identify a maximally discriminant set of features for target detection. The second, or detection stage, uses a feature-based attention mechanism and a target-tuned top-down discriminant saliency detector to detect the target. Overall, the tracker iterates between learning discriminant features from the target location in a video frame and detecting the location of the target in the next. The statistics of natural images are exploited to derive an implementation which is conceptually simple and computationally efficient. The saliency formulation is also shown to establish a unified framework for classifier design, target detection, automatic tracker initialization, and scale adaptation. Experimental results show that the proposed discriminant saliency tracker outperforms a number of state-of-the-art trackers in the literature. Vijay Mahadevan, Nuno Vasconcelos |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2012 | On the connections between saliency and trackingabstractA model connecting visual tracking and saliency has recently been proposed. This model is based on the saliency hypothesis for tracking which postulates that tracking is achieved by the top-down tuning, based on target features, of discriminant center-surround saliency mechanisms over time. In this work, we identify three main predictions that must hold if the hypothesis were true: 1) tracking reliability should be larger for salient than for non-salient targets, 2) tracking reliability should have a dependence on the defining variables of saliency, namely feature contrast and distractor heterogeneity, and must replicate the dependence of saliency on these variables, and 3) saliency and tracking can be implemented with common low level neural mechanisms. We confirm that the first two predictions hold by reporting results from a set of human behavior studies on the connection between saliency and tracking. We also show that the third prediction holds by constructing a common neurophysiologically plausible architecture that can computationally solve both saliency and tracking. This architecture is fully compliant with the standard physiological models of V1 and MT, and with what is known about attentional control in area LIP, while explaining the results of the human behavior experiments. Vijay Mahadevan, Nuno Vasconcelos |
NIPS | 1 |
| 2011 | Maximum Covariance Unfolding : Manifold Learning for Bimodal DataabstractWe propose maximum covariance unfolding (MCU), a manifold learning algorithm for simultaneous dimensionality reduction of data from different input modalities. Given high dimensional inputs from two different but naturally aligned sources, MCU computes a common low dimensional embedding that maximizes the cross-modal (inter-source) correlations while preserving the local (intra-source) distances. In this paper, we explore two applications of MCU. First we use MCU to analyze EEG-fMRI data, where an important goal is to visualize the fMRI voxels that are most strongly correlated with changes in EEG traces. To perform this visualization, we augment MCU with an additional step for metric learning in the high dimensional voxel space. Second, we use MCU to perform cross-modal retrieval of matched image and text samples from Wikipedia. To manage large applications of MCU, we develop a fast implementation based on ideas from spectral graph theory. These ideas transform the original problem for MCU, one of semidefinite programming, into a simpler problem in semidefinite quadratic linear programming. Vijay Mahadevan, Chi Wah Wong, José Costa Pereira, Tom Liu, Nuno Vasconcelos, Lawrence K. Saul |
NIPS | 1 |
| 2011 | Generalized Stauffer-Grimson background subtraction for dynamic scenes
Antoni B. Chan, Vijay Mahadevan, Nuno Vasconcelos |
Mach. Vis. Appl. | 2 |
| 2010 | Anomaly detection in crowded scenesabstractA novel framework for anomaly detection in crowded scenes is presented. Three properties are identified as important for the design of a localized video representation suitable for anomaly detection in such scenes: (1) joint modeling of appearance and dynamics of the scene, and the abilities to detect (2) temporal, and (3) spatial abnormalities. The model for normal crowd behavior is based on mixtures of dynamic textures and outliers under this model are labeled as anomalies. Temporal anomalies are equated to events of low-probability, while spatial anomalies are handled using discriminant saliency. An experimental evaluation is conducted with a new dataset of crowded scenes, composed of 100 video sequences and five well defined abnormality categories. The proposed representation is shown to outperform various state of the art anomaly detection techniques. Vijay Mahadevan, Weixin Li 0002, Viral Bhalodia, Nuno Vasconcelos |
CVPR | 1 |
| 2010 | On the design of robust classifiers for computer visionabstractThe design of robust classifiers, which can contend with the noisy and outlier ridden datasets typical of computer vision, is studied. It is argued that such robustness requires loss functions that penalize both large positive and negative margins. The probability elicitation view of classifier design is adopted, and a set of necessary conditions for the design of such losses is identified. These conditions are used to derive a novel robust Bayes-consistent loss, denoted Tangent loss, and an associated boosting algorithm, denoted TangentBoost. Experiments with data from the computer vision problems of scene classification, object tracking, and multiple instance learning show that TangentBoost consistently outperforms previous boosting algorithms. Hamed Masnadi-Shirazi, Vijay Mahadevan, Nuno Vasconcelos |
CVPR | 2 |
| 2010 | Motion vector refinement for FRUC using saliency and segmentationabstractMotion-Compensated Frame Interpolation (MCFI) is a technique used extensively for increasing the temporal frequency of a video sequence. In order to obtain a high quality interpolation, the motion field between frames must be well-estimated. However, many current techniques for determining the motion are prone to errors in occlusion regions, as well as regions with repetitive structure. An algorithm is proposed for improving both the objective and subjective quality of MCFI by refining the motion vector field. A Discriminant Saliency classifier is employed to determine regions of the motion field which are most important to a human observer. These regions are refined using a multi-stage motion vector refinement which promotes candidates based on their likelihood given a local neighborhood. For regions which fall below the saliency threshold, frame segmentation is used to locate regions of homogeneous color and texture via Normalized Cuts. Motion vectors are promoted such that each homogeneous region has a consistent motion. Experimental results demonstrate an improvement over previous methods in both objective and subjective picture quality. Natan Jacobson, Yen-Lin Lee, Vijay Mahadevan, Nuno Vasconcelos, Truong Q. Nguyen |
ICME | 3 |
| 2010 | Spatiotemporal Saliency in Dynamic ScenesabstractA spatiotemporal saliency algorithm based on a center-surround framework is proposed. The algorithm is inspired by biological mechanisms of motion-based perceptual grouping and extends a discriminant formulation of center-surround saliency previously proposed for static imagery. Under this formulation, the saliency of a location is equated to the power of a predefined set of features to discriminate between the visual stimuli in a center and a surround window, centered at that location. The features are spatiotemporal video patches and are modeled as dynamic textures, to achieve a principled joint characterization of the spatial and temporal components of saliency. The combination of discriminant center-surround saliency with the modeling power of dynamic textures yields a robust, versatile, and fully unsupervised spatiotemporal saliency algorithm, applicable to scenes with highly dynamic backgrounds and moving cameras. The related problem of background subtraction is treated as the complement of saliency detection, by classifying nonsalient (with respect to appearance and motion dynamics) points in the visual field as background. The algorithm is tested for background subtraction on challenging sequences, and shown to substantially outperform various state-of-the-art techniques. Quantitatively, its average error rate is almost half that of the closest competitor. Vijay Mahadevan, Nuno Vasconcelos |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2010 | A Novel Approach to FRUC Using Discriminant Saliency and Frame SegmentationabstractMotion-compensated frame interpolation (MCFI) is a technique used extensively for increasing the temporal frequency of a video sequence. In order to obtain a high quality interpolation, the motion field between frames must be well-estimated. However, many current techniques for determining the motion are prone to errors in occlusion regions, as well as regions with repetitive structure. We propose an algorithm for improving both the objective and subjective quality of MCFI by refining the motion vector field. We first utilize a discriminant saliency classifier to determine which regions of the motion field are most important to a human observer. These regions are refined using a multistage motion vector refinement (MVR), which promotes motion vector candidates based upon their likelihood given a local neighborhood. For regions which fall below the saliency-threshold, a frame segmentation is used to locate regions of homogeneous color and texture via normalized cuts. Motion vectors are promoted such that each homogeneous region has a consistent motion. Experimental results demonstrate an improvement over previous frame rate up-conversion (FRUC) methods in both objective and subjective picture quality. Natan Jacobson, Yen-Lin Lee, Vijay Mahadevan, Nuno Vasconcelos, Truong Q. Nguyen |
IEEE Trans. Image Process. | 3 |
| 2009 | Saliency-based discriminant trackingabstractWe propose a biologically inspired framework for visual tracking based on discriminant center surround saliency. At each frame, discrimination of the target from the background is posed as a binary classification problem. From a pool of feature descriptors for the target and background, a subset that is most informative for classification between the two is selected using the principle of maximum marginal diversity. Using these features, the location of the target in the next frame is identified using top-down saliency, completing one iteration of the tracking algorithm. We also show that a simple extension of the framework to include motion features in a bottom-up saliency mode can robustly identify salient moving objects and automatically initialize the tracker. The connections of the proposed method to existing works on discriminant tracking are discussed. Experimental results comparing the proposed method to the state of the art in tracking are presented, showing improved performance. Vijay Mahadevan, Nuno Vasconcelos |
CVPR | 1 |
| 2008 | Background subtraction in highly dynamic scenesabstractA new algorithm is proposed for background subtraction in highly dynamic scenes. Background subtraction is equated to the dual problem of saliency detection: background points are those considered not salient by suitable comparison of object and background appearance and dynamics. Drawing inspiration from biological vision, saliency is defined locally, using center-surround computations that measure local feature contrast. A discriminant formulation is adopted, where the saliency of a location is the discriminant power of a set of features with respect to the binary classification problem which opposes center to surround. To account for both motion and appearance, and achieve robustness to highly dynamic backgrounds, these features are spatiotemporal patches, which are modeled as dynamic textures. The resulting background subtraction algorithm is fully unsupervised, requires no training stage to learn background parameters, and depends only on the relative disparity of motion between the center and surround regions. This makes it insensitive to camera motion. The algorithm is tested on challenging video sequences, and shown to outperform various state-of-the-art techniques for background subtraction. Vijay Mahadevan, Nuno Vasconcelos |
CVPR | 1 |
| 2008 | Improved Detection of the Central Reflex in Retinal Vessels Using a Generalized Dual-Gaussian Model and Robust Hypothesis TestingabstractThis updates an earlier publication by the authors describing a robust framework for detecting vasculature in noisy retinal fundus images. We improved the handling of the "central reflex" phenomenon in which a vessel has a "hollow" appearance. This is particularly pronounced in dual-wavelength images acquired at 570 and 600 nm for retinal oximetry. It is prominent in the 600 nm images that are sensitive to the blood oxygen content. Improved segmentation of these vessels is needed to improve oximetry. We show that the use of a generalized dual-Gaussian model for the vessel intensity profile instead of the Gaussian yields a significant improvement. Our method can account for variations in the strength of the central reflex, the relative contrast, width, orientation, scale, and imaging noise. It also enables the classification of regular and central reflex vessels. The proposed method yielded a sensitivity of 72% compared to 38% by the algorithm of Can et al., and 60% by the robust detection based on a single-Gaussian model. The specificity for the methods were 95%, 97%, and 98%, respectively. Harihar Narasimha-Iyer, Vijay Mahadevan, James M. Beach, Badrinath Roysam |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2007 | The discriminant center-surround hypothesis for bottom-up saliencyabstractThe classical hypothesis, that bottom-up saliency is a center-surround process, is combined with a more recent hypothesis that all saliency decisions are optimal in a decision-theoretic sense. The combined hypothesis is denoted as discriminant center-surround saliency, and the corresponding optimal saliency architecture is derived. This architecture equates the saliency of each image location to the discriminant power of a set of features with respect to the classification problem that opposes stimuli at center and surround, at that location. It is shown that the resulting saliency detector makes accurate quantitative predictions for various aspects of the psychophysics of human saliency, including non-linear properties beyond the reach of previous saliency models. Furthermore, it is shown that discriminant center-surround saliency can be easily generalized to various stimulus modalities (such as color, orientation and motion), and provides optimal solutions for many other saliency problems of interest for computer vision. Optimal solutions, under this hypothesis, are derived for a number of the former (including static natural images, dense motion fields, and even dynamic textures), and applied to a number of the latter (the prediction of human eye fixations, motion-based saliency in the presence of ego-motion, and motion-based saliency in the presence of highly dynamic backgrounds). In result, discriminant saliency is shown to predict eye fixations better than previous models, and produce background subtraction algorithms that outperform the state-of-the-art in computer vision. Dashan Gao 0001, Vijay Mahadevan, Nuno Vasconcelos |
NIPS | 2 |
| 2004 | Robust model-based vasculature detection in noisy biomedical imagesabstractThis paper presents a set of algorithms for robust detection of vasculature in noisy retinal video images. Three methods are studied for effective handling of outliers. The first method is based on Huber's censored likelihood ratio test. The second is based on the use of a alpha-trimmed test statistic. The third is based on robust model selection algorithms. All of these algorithms rely on a mathematical model for the vasculature that accounts for the expected variations in intensity/texture profile, width, orientation, scale, and imaging noise. These unknown parameters are estimated implicitly within a robust detection and estimation framework. The proposed algorithms are also useful as nonlinear vessel enhancement filters. The proposed algorithms were evaluated over carefully constructed phantom images, where the ground truth is known a priori, as well as clinically recorded images for which the ground truth was manually compiled. A comparative evaluation of the proposed approaches is presented. Collectively, these methods outperformed prior approaches based on Chaudhuri et al. (1989) matched filtering, as well as the verification methods used by prior exploratory tracing algorithms, such as the work of Can et aL (1999). The Huber censored likelihood test yielded the best overall improvement, with a 145.7% improvement over the exploratory tracing algorithm, and a 43.7% improvement in detection rates over the matched filter. Vijay Mahadevan, Harihar Narasimha-Iyer, Badrinath Roysam, Howard L. Tanenbaum |
IEEE Trans. Inf. Technol. Biomed. | 1 |