Tanveer F. Syeda-Mahmood

dblp:11/5386 · also Tanveer Fathima Syeda-Mahmood · DBLP profile ↗
← Back
84ranked-venue papers
36as first author
13since 2021 · last 2026
0000-0003-0059-3208ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 48 · 26 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 35 · 7 first-author · 10 since 2021Artificial intelligence and machine learning · 33 · 21 first-author · 4 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author
YearPublicationVenuePosition
2026 HNOSeg-XS: Extremely Small Hartley Neural Operator for Efficient and Resolution-Robust 3D Image Segmentation
abstract
In medical image segmentation, convolutional neural networks (CNNs) and transformers are dominant. For CNNs, given the local receptive fields of convolutional layers, long-range spatial correlations are captured through consecutive convolutions and pooling. However, as the computational cost and memory footprint can be prohibitively large, 3D models can only afford fewer layers than 2D models with reduced receptive fields and abstract levels. For transformers, although long-range correlations can be captured by multi-head attention, its quadratic complexity with respect to input size is computationally demanding. Therefore, either model may require input size reduction to allow more filters and layers for better segmentation. Nevertheless, given their discrete nature, models trained with patch-wise training or image downsampling may produce suboptimal results when applied on higher resolutions. To address this issue, here we propose the resolution-robust HNOSeg-XS architecture. We model image segmentation by learnable partial differential equations through the Fourier neural operator which has the zero-shot super-resolution property. By replacing the Fourier transform by the Hartley transform and reformulating the problem in the frequency domain, we created the HNOSeg-XS model, which is resolution robust, fast, memory efficient, and extremely parameter efficient. When tested on the BraTS'23, KiTS'23, and MVSeg'23 datasets with a Tesla V100 GPU, HNOSeg-XS showed its superior resolution robustness with fewer than 34.7k model parameters. It also achieved the overall best inference time (<0.24 s) and memory efficiency (<1.8 GiB) compared to the tested CNN and transformer models. The code repository is available at https://github.com/IBM/multimodal-3d-image-segmentation.
Ken C. L. Wong, Hongzhi Wang 0002, Tanveer F. Syeda-Mahmood
IEEE Trans. Medical Imaging3
2025 Phrase-Grounded Fact-Checking for Automatically Generated Chest X-Ray Reports
Razi Mahmood, Diego Machado Reyes, Joy T. Wu, Parisa Kaviani, Ken C. L. Wong, Niharika D'Souza, Mannudeep K. Kalra, Ge Wang 0001, Pingkun Yan, Tanveer F. Syeda-Mahmood
MICCAI (7)10
2024 Vector Quantization with Sorting Transformation
abstract
Vector quantization is a nearest neighbor representation based compression technique for vector data. It creates a collection of codewords to represent the entire vector space. Each vector data is then represented by its nearest neighbor codeword, where the distance between them is the compression error. To improve nearest neighbor representation for vector quantization, we propose to apply sorting transformation to vector data such that members within each vector are sorted. We show that among all permutation transformations, the sorting transformation minimizes L2 distance and maximizes similarity measures such as cosine similarity and Pearson correlation for vector data. Applying sorting transformation with vector quantization can substantially reduce compression errors. Meanwhile, it incurs storage overhead for saving the sorting permutation for each compressed vector. Through experimental validation on compression and nearest neighbor retrieval, we show that this is a beneficial trade-off for vector quantization on low dimensional vectors, a common scenario for vector quantization applications.
Hongzhi Wang 0002, Tanveer F. Syeda-Mahmood
IEEE Big Data2
2024 Modelling-based joint embedding of histology and genomics using canonical correlation analysis for breast cancer survival prediction
abstract
Traditional approaches to predicting breast cancer patients' survival outcomes were based on clinical subgroups, the PAM50 genes, or the histological tissue's evaluation. With the growth of multi-modality datasets capturing diverse information (such as genomics, histology, radiology and clinical data) about the same cancer, information can be integrated using advanced tools and have improved survival prediction. These methods implicitly exploit the key observation that different modalities originate from the same cancer source and jointly provide a complete picture of the cancer. In this work, we investigate the benefits of explicitly modelling multi-modality data as originating from the same cancer under a probabilistic framework. Specifically, we consider histology and genomics as two modalities originating from the same breast cancer under a probabilistic graphical model (PGM). We construct maximum likelihood estimates of the PGM parameters based on canonical correlation analysis (CCA) and then infer the underlying properties of the cancer patient, such as survival. Equivalently, we construct CCA-based joint embeddings of the two modalities and input them to a learnable predictor. Real-world properties of sparsity and graph-structures are captured in the penalized variants of CCA (pCCA) and are better suited for cancer applications. For generating richer multi-dimensional embeddings with pCCA, we introduce two novel embedding schemes that encourage orthogonality to generate more informative embeddings. The efficacy of our proposed prediction pipeline is first demonstrated via low prediction errors of the hidden variable and the generation of informative embeddings on simulated data. When applied to breast cancer histology and RNA-sequencing expression data from The Cancer Genome Atlas (TCGA), our model can provide survival predictions with average concordance-indices of up to 68.32% along with interpretability. We also illustrate how the pCCA embeddings can be used for survival analysis through Kaplan-Meier curves.
Vaishnavi Subramanian, Tanveer F. Syeda-Mahmood, Minh N. Do
Artif. Intell. Medicine2
2024 Multimodal Machine Learning in Image-Based and Clinical Biomedicine: Survey and Prospects
abstract
Machine learning (ML) applications in medical artificial intelligence (AI) systems have shifted from traditional and statistical methods to increasing application of deep learning models. This survey navigates the current landscape of multimodal ML, focusing on its profound impact on medical image analysis and clinical decision support systems. Emphasizing challenges and innovations in addressing multimodal representation, fusion, translation, alignment, and co-learning, the paper explores the transformative potential of multimodal models for clinical predictions. It also highlights the need for principled assessments and practical implementation of such models, bringing attention to the dynamics between decision support systems and healthcare providers and personnel. Despite advancements, challenges such as data biases and the scarcity of "big data" in many biomedical domains persist. We conclude with a discussion on principled innovation and collaborative efforts to further the mission of seamless integration of multimodal ML models into biomedical practice.
Elisa Warner, Joonsang Lee, William Hsu, Tanveer F. Syeda-Mahmood, Charles E. Kahn Jr., Olivier Gevaert, Arvind Rao
Int. J. Comput. Vis.4
2024 Fusing modalities by multiplexed graph neural networks for outcome prediction from medical data and beyond
Niharika S. D'Souza, Hongzhi Wang 0002, Andrea Giovannini, Antonio Foncubierta-Rodríguez, Kristen L. Beck, Orest B. Boyko, Tanveer F. Syeda-Mahmood
Medical Image Anal.7
2023 Image-Based Soil Organic Carbon Remote Sensing from Satellite Images with Fourier Neural Operator and Structural Similarity
abstract
Soil organic carbon (SOC) sequestration is the transfer and storage of atmospheric carbon dioxide in soils, which plays an important role in climate change mitigation. SOC concentration can be improved by proper land use, thus it is beneficial if SOC can be estimated at a regional or global scale. As multispectral satellite data can provide SOC-related information such as vegetation and soil properties at a global scale, estimation of SOC through satellite data has been explored as an alternative to manual soil sampling. Although existing studies show promising results, they are mainly based on pixel-based approaches with traditional machine learning methods, and convolutional neural networks (CNNs) are uncommon. To study the use of CNNs on SOC remote sensing, here we propose the FNO-DenseNet based on the Fourier neural operator (FNO). By combining the advantages of the FNO and DenseNet, the FNO-DenseNet outperformed the FNO in our experiments with hundreds of times fewer parameters. The FNO-DenseNet also outperformed a pixel-based random forest by 18% in the mean absolute percentage error.
Ken C. L. Wong, Levente J. Klein, Ademir Ferreira da Silva, Hongzhi Wang 0002, Tanveer F. Syeda-Mahmood
IGARSS6
2023 HartleyMHA: Self-attention in Frequency Domain for Resolution-Robust and Parameter-Efficient 3D Image Segmentation
Ken C. L. Wong, Hongzhi Wang 0002, Tanveer F. Syeda-Mahmood
MICCAI (4)3
2022 Towards Automatic Prediction of Outcome in Treatment of Cerebral Aneurysms
Ashutosh Jadhav, Satyananda Kashyap, Hakan Bulu, Ronak Dholakia, Tanveer F. Syeda-Mahmood, Hussain Rangwala, Mehdi Moradi
AMIA5
2022 Fusing Modalities by Multiplexed Graph Neural Networks for Outcome Prediction in Tuberculosis
Niharika S. D'Souza, Hongzhi Wang 0002, Andrea Giovannini, Antonio Foncubierta-Rodríguez, Kristen L. Beck, Orest B. Boyko, Tanveer F. Syeda-Mahmood
MICCAI (8)7
2022 Improving Neural Models for Radiology Report Retrieval with Lexicon-based Automated Annotation
abstract
Many clinical informatics tasks that are based on electronic health records (EHR) need relevant patient cohorts to be selected based on findings, symptoms and diseases.Frequently, these conditions are described in radiology reports which can be retrieved using information retrieval (IR) methods.The latest of these techniques utilize neural IR models such as BERT trained on clinical text.However, these methods still lack semantic understanding of the underlying clinical conditions as well as ruled out findings, resulting in poor precision during retrieval.In this paper we combine clinical finding detection with supervised query match learning.Specifically, we use lexicon-driven concept detection to detect relevant findings in sentences.These findings are used as queries to train a Sentence-BERT (SBERT) model using triplet loss on matched and unmatched query-sentence pairs.We show that the proposed supervised training task remarkably improves the retrieval performance of SBERT.The trained model generalizes well to unseen queries and reports from different collections.
Luyao Shi, Tanveer F. Syeda-Mahmood, Tyler Baldwin
NAACL-HLT2
2021 Semantic Expansion of Clinician Generated Data Preferences for Automatic Patient Data Summarization
Ashutosh Jadhav, Tyler Baldwin, Joy T. Wu, Vandana V. Mukherjee, Tanveer F. Syeda-Mahmood
AMIA5
2021 Image-Derived Phenotype Extraction for Genetic Discovery via Unsupervised Deep Learning in CMR Images
Rodrigo Bonazzola, Nishant Ravikumar, Rahman Attar, Enzo Ferrante, Tanveer F. Syeda-Mahmood, Alejandro F. Frangi
MICCAI (5)5
2020 Combining Deep Learning and Knowledge-driven Reasoning for Chest X-Ray Findings Detection
Ashutosh Jadhav, Ken C. L. Wong, Joy T. Wu, Mehdi Moradi, Tanveer F. Syeda-Mahmood
AMIA5
2020 Extracting and Learning Fine-grained Labels from Chest Radiographs
Tanveer F. Syeda-Mahmood, Ken C. L. Wong, Joy T. Wu, Ashutosh Jadhav, Orest B. Boyko
AMIA1
2020 AI Accelerated Human-in-the-loop Structuring of Radiology Reports
Joy T. Wu, Ali Bin Syed, Hassan M. Ahmad, Anup Pillai, Yaniv Gur, Ashutosh Jadhav, Daniel Gruhl, Linda Kato, Mehdi Moradi, Tanveer F. Syeda-Mahmood
AMIA10
2020 Chest X-Ray Report Generation Through Fine-Grained Label Learning
Tanveer F. Syeda-Mahmood, Ken C. L. Wong, Yaniv Gur, Joy T. Wu, Ashutosh Jadhav, Satyananda Kashyap, Alexandros Karargyris, Anup Pillai, Arjun Sharma, Ali Bin Syed, Orest B. Boyko, Mehdi Moradi
MICCAI (2)1
2019 Automated Detection and Type Classification of Central Venous Catheters in Chest X-Rays
Vaishnavi Subramanian, Hongzhi Wang 0002, Joy T. Wu, Ken C. L. Wong, Arjun Sharma, Tanveer F. Syeda-Mahmood
MICCAI (6)6
2018 Generalized Extraction and Classification of Span-Level Clinical Phrases
Tyler Baldwin, Vandana V. Mukherjee, Tanveer F. Syeda-Mahmood
AMIA4
2018 Improving the Path from Diagnoses to Documentation: A Cognitive Review Tool for Clinical Notes and Administrative Records
Joy T. Wu, Tyler Baldwin, David Beymer, Vandana V. Mukherjee, Tanveer F. Syeda-Mahmood
AMIA6
2018 Classification of radiology reports by modality and anatomy: A comparative study
Marina Bendersky, Joy T. Wu, Tanveer F. Syeda-Mahmood
BIBM3
2018 Bimodal Network Architectures for Automatic Generation of Image Annotation from Text
Mehdi Moradi, Ali Madani, Yaniv Gur, Tanveer F. Syeda-Mahmood
MICCAI (1)5
2018 3D Segmentation with Exponential Logarithmic Loss for Highly Unbalanced Object Sizes
Ken C. L. Wong, Mehdi Moradi, Tanveer F. Syeda-Mahmood
MICCAI (3)4
2018 Building medical image classifiers with very limited data using segmentation networks
Ken C. L. Wong, Tanveer F. Syeda-Mahmood, Mehdi Moradi
Medical Image Anal.2
2017 Efficient Clinical Concept Extraction in Electronic Medical Records
abstract
Automatic identification of clinical concepts in electronic medical records (EMR) is useful not only in forming a complete longitudinal health record of patients, but also in recovering missing codes for billing, reducing costs, finding more accurate clinical cohorts for clinical trials, and enabling better clinical decision support. Existing systems for clinical concept extraction are mostly knowledge-driven, relying on exact match retrieval from original or lemmatized reports, and very few of them are scaled up to handle large volumes of complex, diverse data. In this demonstration we will showcase a new system for real-time detection of clinical concepts in EMR. The system features a large vocabulary of over 5.6 million concepts. It achieves high precision and recall, with good tolerance to typos through the use of a novel prefix indexing and subsequence matching algorithm, along with a recursive negation detector based on efficient, deep parsing. Our system has been tested on over 12.9 million reports of more than 200 different types, collected from 800,000+ patients. A comparison with the state of the art shows that it outperforms previous systems in addition to being the first system to scale to such large collections.
Deepika Kakrania, Tyler Baldwin, Tanveer F. Syeda-Mahmood
AAAI4
2017 A Multi-atlas Approach to Region of Interest Detection for Medical Image Classification
Hongzhi Wang 0002, Mehdi Moradi, Yaniv Gur, Prasanth Prasanna, Tanveer F. Syeda-Mahmood
MICCAI (3)5
2017 Building Disease Detection Algorithms with Very Small Numbers of Positive Samples
Ken C. L. Wong, Alexandros Karargyris, Tanveer F. Syeda-Mahmood, Mehdi Moradi
MICCAI (3)3
2016 Automatic Generation of Conditional Diagnostic Guidelines
Tyler Baldwin, Tanveer F. Syeda-Mahmood
AMIA3
2016 A Cross-Modality Neural Network Transform for Semi-automatic Medical Image Annotation
Mehdi Moradi, Yaniv Gur, Mohammadreza Negahdar, Tanveer F. Syeda-Mahmood
MICCAI (2)5
2016 Identifying Patients at Risk for Aortic Stenosis Through Learning from Multimodal Data
abstract
In this paper we present a new method of uncovering patients with aortic valve diseases in large electronic health record systems through learning with multimodal data. The method automatically extracts clinically-relevant valvular disease features from five multimodal sources of information including structured diagnosis, echocardiogram reports, and echocardiogram imaging studies. It combines these partial evidence features in a random forests learning framework to predict patients likely to have the disease. Results of a retrospective clinical study from a 1000 patient dataset are presented that indicate that over 25 % new patients with moderate to severe aortic stenosis can be automatically discovered by our method that were previously missed from the records. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves.
Tanveer F. Syeda-Mahmood, Yanrong Guo, Mehdi Moradi, David Beymer, Deepta Rajan, Yaniv Gur, Mohammadreza Negahdar
MICCAI (3)1
2016 Metric hashing forests
Sailesh Conjeti, Amin Katouzian, Anees Kazi, Sepideh Mesbah, David Beymer, Tanveer F. Syeda-Mahmood, Nassir Navab
Medical Image Anal.6
2015 Learning the Correlation Between Images and Disease Labels Using Ambiguous Learning
Tanveer F. Syeda-Mahmood, Ritwik Kumar, Colin B. Compas
MICCAI (2)1
2014 Automatic Detection of Dilated Cardiomyopathy in Cardiac Ultrasound Videos
Razi Mahmood, Tanveer F. Syeda-Mahmood
AMIA2
2014 Segmentation of Anatomical Structures in Four-Chamber View Echocardiogram Images
abstract
Automatic generation of a cardiac atlas from echocardiogram images requires that a good segmentation of all anatomical structures be available. In this paper, we present an algorithm for automatic segmentation of all relevant anatomical regions visible in four-chamber view echocardiogram images. Specifically, we propose a two-pass segmentation algorithm in which we first process the image to identify the heart muscle (bright) and chamber regions (dark) by adapting an edge-weighted centroidal Voronoi tessellation (AEWCVT) algorithm. We then partition the resulting bright and dark regions into approximately convex regions using a new convexity pursuit segmentation algorithm. Experimental comparison with several region segmentation algorithms for general imaging shows that our method outperforms these algorithms as it is able to better adapt to known cardiac anatomical structures.
Patrick McNeillie, Tanveer F. Syeda-Mahmood
ICPR3
2013 Mining Echocardiography Workflows for Disease Discriminative Patterns
Ritwik Kumar, Tanveer F. Syeda-Mahmood, David Beymer, Colin B. Compas, Karen Brannon
AMIA2
2013 Complex local phase based subjective surfaces (CLAPSS) and its application to DIC red blood cell image segmentation
Taoyi Chen, Yong Zhang 0050, Changhong Wang 0003, Zhenshen Qu, Fei Wang 0002, Tanveer F. Syeda-Mahmood
Neurocomputing6
2012 AngioViewer: A Tool for Assessing the State of Coronary Artery Disease
David Beymer, Tanveer F. Syeda-Mahmood, Fei Wang 0002, Ritwik Kumar, Yong Zhang 0050, Robert J. Lundstrom, Taylor Holve, Navid Shafaee, Edward McNulty
AMIA2
2012 Multimodal Informatics Platform for Clinical Decision Support¬
Tanveer F. Syeda-Mahmood, David Beymer, Fei Wang 0002, Ritwik Kumar, Yong Zhang 0050
AMIA1
2012 Finding Similar 2D X-Ray Coronary Angiograms
Tanveer F. Syeda-Mahmood, Fei Wang 0002, Ritwik Kumar, David Beymer, Yong Zhang 0050, Robert J. Lundstrom, Edward McNulty
MICCAI (3)1
2010 Shape-based similarity retrieval of Doppler images for clinical decision support
abstract
Flow Doppler imaging has become an integral part of an echocardiographic exam. Automated interpretation of flow doppler imaging has so far been restricted to obtaining hemodynamic information from velocity-time profiles depicted in these images. In this paper we exploit the shape patterns in Doppler images to infer the similarity in valvular disease labels for purposes of automated clinical decision support. Specifically, we model the similarity in appearance of Doppler images from the same disease class as a constrained non-rigid translation transform of the velocity envelopes embedded in these images. The shape similarity between two Doppler images is then judged by recovering the alignment transform using a variant of dynamic shape warping. Results of similarity retrieval of doppler images for cardiac decision support on a large database of images are presented.
Tanveer F. Syeda-Mahmood, Pavan Turaga, David Beymer, Fei Wang 0002, Arnon Amir, Hayit Greenspan, Kilian M. Pohl
CVPR1
2010 Automatic Selection of Keyframes from Angiogram Videos
abstract
In this paper we address the problem of automatic selection of important vessel-depicting key frames within 2D angiography videos. Two different methods of frame selection are described, one based on Frangi filter, and the other based on detecting parallel curves formed from edges in angiography images. Results are shown by comparison to physician annotation of such key frames on 2D coronary angiograms.
Tanveer F. Syeda-Mahmood, David Beymer, Fei Wang 0002, A. Mahmood, Robert J. Lundstrom, N. Shafee, Taylor Holve
ICPR1
2009 Echocardiogram view classification using edge filtered scale-invariant motion features
abstract
In an 2D echocardiogram exam, an ultrasound probe samples the heart with 2D slices. Changing the orientation and position on the probe changes the slice viewpoint, altering the cardiac anatomy being imaged. The determination of the probe viewpoint forms an essential step in automatic cardiac echo image analysis. In this paper we present a system for automatic view classification that exploits cues from both cardiac structure and motion in echocardiogram videos. In our framework, each image from the echocardiogram video is represented by a set of novel salient features. We locate these features at scale invariant points in the edge-filtered motion magnitude images and encode them using local spatial, textural and kinetic information. Training in our system involves learning a hierarchical feature dictionary and parameters of a pyramid matching kernel based support vector machine. While testing, each image, classified independently, casts a votes towards parent video classification and the viewpoint with maximum votes wins. Through experiments on a large database of echocardiograms obtained from both diseased and control subjects, we show that our technique consistently outperforms state-of-the-art methods in the popular four-view classification test. We also present results for eight-view classification to demonstrate the scalability of our framework.
Ritwik Kumar, Fei Wang 0002, David Beymer, Tanveer F. Syeda-Mahmood
CVPR4
2009 Disease-Specific Extraction of Text from Cardiac Echo Videos for Decision Support
abstract
Echo videos are an important modality for cardiac decision support. In addition to describing the shape and motion of the heart, they capture important diagnostic measurements as textual feature-value pairs that are good indicators of the underlying disease. In this paper, we describe reliable extraction of such textual information through selective image processing and region extraction prior to using an OCR engine. We then use tabular layout analysis rules to recover measurement attribute value pairs from the recognized text in videos. The measurement feature-value pairs are used to retrieve matching videos from a database. A ranked list of matching diseases is then obtained through collaborative filtering. Results are demonstrated on a large echo video database of patients with various diseases.
Tanveer F. Syeda-Mahmood, David Beymer, Arnon Amir
ICDAR1
2009 Information Extraction from Multimodal ECG Documents
abstract
With the rise of tools for clinical decision support, there is an increased need for automatic processing of electrocardiograms (ECG) documents. In fact, many systems have already been developed to perform signal processing tasks such as 12-lead off-line ECG analysis and real-time patient monitoring. All these applications require an accurate detection of the heart rate of the ECG. In this paper, we present the idea that the image form of ECG is actually a better medium to detect periodicity in ECG. When the ECG trace is scanned or rendered in videos, the peaks of the waveform (R-wave) is often traced thicker due to pixel dithering. We exploit the pixel thickness information, for the first time, as a reliable feature for determining periodicity. Results are presented on a database of 16,613 12-channel ECG waveforms, which demonstrate robustness and accuracy of our image-based period detection method on these ECGs of various cardiovascular diseases. 94.5% of bradycardia and tachycardia patient records are correctly identified using our estimated heart period as the disease criteria.
Fei Wang 0002, Tanveer F. Syeda-Mahmood, David Beymer
ICDAR2
2009 Closed-Form Jensen-Renyi Divergence for Mixture of Gaussians and Applications to Group-Wise Shape Registration
Fei Wang 0002, Tanveer F. Syeda-Mahmood, Baba C. Vemuri, David Beymer, Anand Rangarajan 0001
MICCAI (1)2
2009 Guest Editors' Introduction to the Special Section on Award Winning Papers from the IEEE CS Conference on Computer Vision and Pattern Recognition (CVPR)
abstract
The three articles in this special section are selected papers from the IEEE CS Conference on Computer Vision and Pattern Recognition that was held in Anchorage, AL, in June 2008.
Kim Boyer, Mubarak Shah, Tanveer F. Syeda-Mahmood
IEEE Trans. Pattern Anal. Mach. Intell.3
2008 Shape-Based Retrieval of Heart Sounds for Disease Similarity Detection
Tanveer F. Syeda-Mahmood, Fei Wang 0002
ECCV (2)1
2008 Shape-based matching of heart sounds
abstract
In this paper, we present an approach to matching heart sounds based on modeling the morphological variations of audio envelopes through a constrained nonrigid translation transform. Similar heart sounds are then retrieved by recovering the corresponding alignment transform using a variant of shape-based dynamic time warping. Results of comparison with other audio retrieval methods are reported on a large database of heart sounds.
Tanveer F. Syeda-Mahmood, Fei Wang 0002
ICPR1
2007 Unsupervised Clustering using Multi-Resolution Perceptual Grouping
abstract
Clustering is a common operation for data partitioning in many practical applications. Often, such data distributions exhibit higher level structures which are important for problem characterization, but are not explicitly discovered by existing clustering algorithms. In this paper, we introduce multi-resolution perceptual grouping as an approach to unsupervised clustering. Specifically, we use the perceptual grouping constraints of proximity, density, contiguity and orientation similarity. We apply these constraints in a multi-resolution fashion, to group sample points in high dimensional spaces into salient clusters. We present an extensive evaluation of the clustering algorithm against state-of-the-art supervised and unsupervised clustering methods on large datasets.
Tanveer F. Syeda-Mahmood, Fei Wang 0002
CVPR1
2007 Characterizing Spatio-temporal Patterns for Disease Discrimination in Cardiac Echo Videos
Tanveer F. Syeda-Mahmood, Fei Wang 0002, David Beymer, M. London, R. Reddy
MICCAI (1)1
2006 SEMAPLAN: Combining Planning with Semantic Matching to Achieve Web Service Composition
Rama Akkiraju, Biplav Srivastava, Anca Ivan, Richard Goodwin, Tanveer F. Syeda-Mahmood
AAAI5
2006 FormPad: A Camera-Assisted Digital Notepad
Tanveer F. Syeda-Mahmood, Thomas G. Zimmerman
ACCV (2)1
2006 SEMAPLAN: Combining Planning with Semantic Matching to Achieve Web Service Composition
abstract
In this paper, we present a novel algorithm to compose Web services in the presence of semantic ambiguity by combining semantic matching and AI planning algorithms. Specifically, we use cues from domain-independent and domain-specific ontologies to compute an overall semantic similarity score between ambiguous terms. This semantic similarity score is used by AI planning algorithms to guide the searching process when composing services. Experimental results indicate that planning with semantic matching produces better results than planning or semantic matching alone. The solution is suitable for semi-automated composition tools or directory browsers
Rama Akkiraju, Biplav Srivastava, Anca Ivan, Richard Goodwin, Tanveer F. Syeda-Mahmood
ICWS5
2006 Concept-based electronic health records: opportunities and challenges
abstract
Healthcare is a data-rich but information-poor domain. Terabytes of multimedia medical data are being generated on a monthly basis in a typical healthcare organization in order to document patients' health status and care process. Government and health-related organizations are pushing for fully electronic, cross-institution, integrated Electronic Health Records to provide a better, cost effective and more complete access to this data. However, provision of efficient access to the content of such records for timely and decision-enabling information extraction will not be available. Such a capability is essential for providing efficient decision support and objective evidence to clinicians. In addition researchers, medical students, patients, and payers could also benefit from it. We present the idea of concept-based multimedia health records, which aims at organizing the health records at the information level. We will explore the opportunities and possibilities that such an organization will provide, what role the field of multimedia content management could play to materialize this type of health record organization, and what the challenges will be in the quest for realizing the idea.We believe that the field of multimedia can play a very active role in taking healthcare information systems to the next level by facilitating the access to decision-enabling information for different types of users in healthcare. Our goal is to share with the community our thoughts on where the field of multimedia content management research should be focusing its attention to have a fundamental impact on the practice of medicine.
Shahram Ebadollahi, Anni Coden, Michael A. Tanenblatt, Shih-Fu Chang, Tanveer F. Syeda-Mahmood, Arnon Amir
ACM Multimedia5
2005 Searching Service Repositories by Combining Semantic and Ontological Matching
abstract
In this paper, we explore the use of domain-independent and domain-specific ontologies to find matching service descriptions. The domain-independent relationships are derived using an English thesaurus after tokenization and part-of-speech tagging. The domain-specific ontological similarity is derived by an inference on the semantic annotations associated with Web service descriptions. Matches due to the two cues are combined to determine an overall semantic similarity score. By combining multiple cues, we show that better relevancy results can be obtained for service matches from a large repository, than could be obtained using any one cue alone.
Tanveer F. Syeda-Mahmood, Gauri Shah, Rama Akkiraju, Anca Ivan, Richard Goodwin
ICWS1
2005 Validating cardiac echo diagnosis through video similarity
abstract
Video data is increasingly being used in medical diagnosis. Due to the quality of the video and the complexities of underlying motion captured, it is difficult for an in-experienced physician/radiologist to describe motion abnormalities in a crisp way, leading to possible errors in diagnosis. In this paper, we present a method of capturing video similarity and its use for diagnosis verification during decision support. Specifically, we describe the motion information in videos using average velocity curves. Second-order motion statistics are extracted from average velocity curves and serve as features for computing video similarity. Given a new video sample already labeled with a diagnosis, a neighborhood of similar videos is assembled from the training set and their diagnosis labels are used to verify the diagnosis.
Tanveer F. Syeda-Mahmood, Dulce B. Ponceleon, Jing Yang 0005
ACM Multimedia1
2004 Content-based retrieval in gene expression databases
abstract
Research in the field of content-based retrieval has primarily focused on image, video an audio information. In this paper, we demonstrate content-based retrieval in a new data domain called gene expression data derive from gene chip images. In particular, we consider the problem of retrieving functionally similar genes from a database based on the pattern of variation of the expression of genes over time. Specifically, we model the time-varying gene expression patterns as curves, an analyze similarity between gene profiles by the relative amounts of twists an turns produce in a higher-imensional curve forme from the projection of the individual gene profiles. Scale-space analysis is use to detect the sharp twists an turns an their relative strength with respect to the component curves is estimated to form a shape similarity measure between gene profiles. The higher-dimensional curves also form prototypical escriptions of the individual gene profiles, serving as a way to index the database using clustering. Functionally similar genes are then identified using scale-space distance metric on the cluster prototypes.
Tanveer F. Syeda-Mahmood
ACM Multimedia1
2004 Searching databases for sematically-related schemas
abstract
In this paper, we address the problem of searching schema databases for semantically-related schemas. We first give a method of finding semantic similarity between pair-wise schemas based on tokenization, part-of-speech tagging, word expansion, and ontology matching. We then address the problem of indexing the schema database through a semantic hash table. Matching schemas in the database are found by hashing the query attributes and recording peaks in the histogram of schema hits. Results indicated a 90% improvement in search performance while maintaining high precision and recall.
Gauri Shah, Tanveer F. Syeda-Mahmood
SIGIR2
2004 CVIU special issue on event detection in video
Tanveer F. Syeda-Mahmood, Ismail Haritaoglu, Thomas S. Huang
Comput. Vis. Image Underst.1
2003 Detecting salient changes in genomic signals
abstract
The functional state of an organism is determined largely by the pattern of the expression of its genes. Salient changes in variation in expression of genes can give clues about important events, such as the onset of a disease. We address the problem of detecting salient inflection points in genomic signals. These are detected using a scale-space decomposition, to be the negative-going zero-crossings of the second derivative of smoothed signals that are preserved over an automatically selected scale. The utility of salient change detection is demonstrated in the automatic identification of the regulatory phase for genes active in the mitotic cell cycle of budding yeast.
Tanveer F. Syeda-Mahmood
ICASSP (2)1
2003 View-invariant Alignment and Matching of Video Sequences
abstract
In this paper, we propose a novel method to establish temporal correspondence between the frames of two videos. 3D epipolar geometry is used to eliminate the distortion generated by the projection from 3D to 2D. Although the fundamental matrix contains the extrinsic property of the projective geometry between views, it is sensitive to noise. Therefore, we propose the use of a rank constraint of corresponding points in two views to measure the similarity between trajectories. This rank constraint shows more robustness and avoids computation of the fundamental matrix. A dynamic programming approach using the similarity measurement is proposed to find the non-linear time-warping function for videos containing human activities. In this way, videos of different individuals taken at different times and from distinct viewpoints can be synchronized. A temporal pyramid of trajectories is applied to improve the accuracy of the view-invariant dynamic time-warping approach. We show various applications of this approach such as video synthesis, human action recognition, and computer aider training. Compared to state-of-the-art techniques, our method shows a great improvement. 1.
Cen Rao, Alexei Gritai, Mubarak Shah, Tanveer F. Syeda-Mahmood
ICCV4
2003 Invariance in motion analysis of videos
abstract
In this paper, we propose an approach that retrieves motion of objects from the videos based on the dynamic time warping of view invariant characteristics. The motion is represented as a sequence of dynamic instants and intervals, which are automatically computed using the spatiotemporal curvature of the trajectory of moving object in the videos. Dynamic Time Warping (DTW) method matches trajectories using a view invariant similarity measure. Our system is able to incrementally learn different actions without any initialization mode, therefore it can work in an unsupervised manner. The retrieval of relevant videos can be easily performed by computing a simple distance metric. This paper makes two fundamental contribution to view invariant video retrieval: (1) Dynamic Instant detection in trajectories of moving objects acquired from video. (2) View-invariant Dynamic Time Warping to measure similarity between two trajectories of actions performed by different persons and from different viewpoints. Although the learning algorithm is relatively simple in our approach, we can achieve high recognition rate because of the view-invariant representation and the similarity measure using DTW.
Cen Rao, Mubarak Shah, Tanveer F. Syeda-Mahmood
ACM Multimedia3
2002 Retrieving actions embedded in video
abstract
In a number of applications including surveillance, there is a need to reliably retrieve an action-depicting segment in a video. This is an enormously difficult problem due to the variability in an action's appearance when seen at different times. It requires reliable object and action segmentation, and robust methods for indexing the action content in a video. In this paper, we present a novel approach to action retrieval that extracts salient action events in query and database videos. These events serve as anchor points to initiate action recognition. Actions are recognized by forming a spatio-temporal shape for an action called the action cylinder. Robust recognition is achieved by recovering the viewpoint transformation and time correspondence between a query action and a given action segment in the video. We demonstrate the versatility of our method for the retrieving of complex actions within videos.
Tanveer F. Syeda-Mahmood
ACM Multimedia1
2001 Learning video browsing behavior and its application in the generation of video previews
abstract
With more and more streaming media servers becoming commonplace, streaming video has now become a popular medium of instruction, advertisement, and entertainment. With such prevalence comes a new challenge to the servers: Can they track browsing behavior of users to determine what interest users? Learning this information is potentially valuable not only for improved customer tracking and context-sensitive e-commerce, but also in the generation of fast previews of videos for easy pre-downloads. In this paper, we present a formal learning mechanism to track video browsing behavior of users. This information is then used to generate fast video previews. Specifically, we model the states a user transitions while browsing through videos to be the hidden states of a Hidden Markov Model. We estimate the parameters of the HMM using maximum likelihood estimation for each sample observation sequence of user interaction with videos. Video previews are then formed from interesting segments of the video automatically inferred from an analysis of the browsing states of viewers. Audio coherence in the previews is maintained by selecting clips spanning complete clauses containing topically significant spoken phrases. The utility of learning video browsing behavior is demonstrated through user studies and experiments.
Tanveer F. Syeda-Mahmood, Dulce B. Ponceleon
ACM Multimedia1
2000 Indexing for Topics in Videos Using Foils
abstract
A long-standing goal of distance learning has been to provide a quality of learning comparable to the face-to-face environment of a traditional classroom for teaching or training. One of the fundamental problems in achieving this goal is providing effective ways of high-level semantic querying such as for the retrieval of relevant learning material relating to a topic of discussion. In this paper we present a method of identifying video segments relating to a topic of discussion by indexing videos using the image and text content of foils. Specifically, we present a novel method of locating and recognizing foil images in video using the color and spatial layout geometry of their regions. We then search the audio associated with video based on the text content of the foil to identify related video segments in which concepts represented on a foil are heard. Finally, we combine the results of foil image and text search of video exploiting their time co-occurrence. The resulting identification of topics is evaluated in the domain of classroom lectures and talks.
Tanveer F. Syeda-Mahmood
CVPR1
2000 CueVideo: A System for Cross-Modal Search and Browse of Video Databases
abstract
The detection and recognition of events is a challenging problem in video databases. It involves cross-linking and combining information available in multiple modalities such as audio, video and associated text metadata. CueVideo is a system designed for the discovery and recognition of specific events called topics of discussion through advanced video summarization and cross-modal indexing. It supports search for relevant video content through several modes of video summarization including storyboards, moving storyboards and time-scale modified audio summarization. It also enables the recognition and indexing of topical events through cross-model search of audio and video content based on text and image queries respectively.
Tanveer F. Syeda-Mahmood, Savitha Srinivasan, Arnon Amir, Dulce B. Ponceleon, Brian Blanchard, Dragutin Petkovic
CVPR1
2000 Detecting topical events in digital video
abstract
The detection of events is essential to high-level semantic querying of video databases. It is also a very challenging problem requiring the detection and integration of evidence for an event available in multiple information modalities, such as audio, video and language. This paper focuses on the detection of specific types of events, namely, topic of discussion events that occur in classroom/lecture environments. Specifically, we present a query-driven approach to the detection of topic of discussion events with foils used in a lecture as a way to convey a topic. In particular, we use the image content of foils to detect visual events in which the foil is displayed and captured in the video stream. The recognition of a f0il in video f~ames exploits the color and spatial layout of regions on foils using a technique called region hashing. Next, we use the textual phrases listed on a foil as an indication of a topic, and detect topical audio events as places in the audio track where the best evidence for the topical phrases was heard. Finally, we use a probabilistic model of event likelihood to combine the results of visual and audio event detection that exploits their time cooccurrence. The resulting identification of topical events is evaluated in the domain of classroom lectures and talks.
Tanveer F. Syeda-Mahmood, Savitha Srinivasan
ACM Multimedia1
2000 On describing color and shape information in images
Tanveer F. Syeda-Mahmood, Dragutin Petkovic
Signal Process. Image Commun.1
1999 Locating Indexing Structures in Engineering Drawing Databases Using Location Hashing
abstract
Fast and robust localization of image regions that are likely to come from a query object is a critical operation needed in several image database applications. In this paper, we address this problem in the context of locating indexing structures in engineering drawing databases. Specifically, we present a general 2d pattern localization technique called location hashing and apply it to the problem of localizing title block structures in engineering drawings. Location hashing is a variation of geometric hashing that directly identifies locations in images of a database that are likely to contain a query object, possibly depicted under different imaging conditions. We also describe an engineering drawing indexing system that combines title block localization with test recognition to enable automatic index creation of engineering drawing databases.
Tanveer F. Syeda-Mahmood
CVPR1
1999 Extracting Indexing Keywords from Image Structures in Engineering Drawings
abstract
A critical operation in the creation of databases of electronic versions of scanned engineering drawings is the automatic extraction of indexing text information from image structures, called title blocks. This paper addresses the problem of locating title block regions and the subsequent extraction of indexing keywords from such regions. A general technique of 2D pattern localization in unsegmented images, called location hashing, is used to locate title blocks. An engineering drawing indexing system then combines title block localization with text recognition to enable indexing text keyword extraction.
Tanveer F. Syeda-Mahmood
ICDAR1
1999 Multimedia access and retrieval: the state of the art and future directions (panel session)
abstract
Several years have passed since the research topic of content based multimedia retrieval emerged.We have witnessed the burgeoning research activities into a plenitude of new indexing, retrieval, and filtering tools for images, video, audio, music, graphics, and their combinations with text-based information.Exciting research opportunities arise when integrating knowledge from multiple disciplines, such as media content processing, database, information retrieval, and machine user interface.In the commercial domain, we have also witnessed several impressive efforts moving technologies into practical arenas.This panel includes experts from industry, research labs, and academia.The panel will assess the state of the art and articulate the important future directions in the general field of multimedia access and retrieval.
Shih-Fu Chang, Gwendal Auffret, Jonathan Foote, Chung-Sheng Li, Behzad Shahraray, Tanveer F. Syeda-Mahmood, HongJiang Zhang
ACM Multimedia (1)6
1999 CueVideo: automated multimedia indexing and retrieval
abstract
No abstract available.
Dulce B. Ponceleon, Arnon Amir, Savitha Srinivasan, Tanveer F. Syeda-Mahmood, Dragutin Petkovic
ACM Multimedia (2)4
1999 Detecting Perceptually Salient Texture Regions in Images
Tanveer F. Syeda-Mahmood
Comput. Vis. Image Underst.1
1999 Indexing of Technical Line Drawing Databases
abstract
Image indexing, namely, the problem of retrieving content information from images in response to queries, is a key problem underlying the operations in image databases. We present a method of indexing for 3D object queries in a database of a class of images called technical line drawings. Indexing is achieved as a combination of query-specific region selection and object recognition. The selection phase isolates relevant images and the regions in these images that are likely to contain the queried object. This is done using text information in the query and a grouping mechanism that is guaranteed to isolate single-object containing regions for the class of technical line drawing images. The grouping mechanism is an adaptation of Waltz relaxation to an extended junction set derived by analyzing the physically plausible ways in which interpretation lines interact with object contours. Model-based object recognition then confirms the presence of the part at the selected location using geometrical description of the queried 3D object. Results are shown that indicate that query-specific selection is very effective for reducing the search during indexing while lowering the chance of false positives and negatives.
Tanveer F. Syeda-Mahmood
IEEE Trans. Pattern Anal. Mach. Intell.1
1998 Global Integration of Visual Databases
abstract
Different visual databases have been designed in various locations. The global integration of such databases can enable users to access data across the world in a transparent manner. In this paper, we investigate an approach to the design and creation of an integrated information system which supports global visual query access to various visual databases over the Internet. Specifically, a metaserver, including a hierarchical metadatabase, a metasearch agent and a query manager, is designed to support such an integration. The metadatabase houses abstracted data about individual remote visual databases. To support visual content-based queries, the abstracted data in the metadatabase reflect the semantics of each visual database. The query manager extracts the feature contents from the queries. The metasearch agent processes the queries by matching their feature contents with the metadata. A list of relevant database sites is derived for efficient retrieval of the query in the selected databases. The performance of the system is refined based on the user's feedback. The proposed system is implemented using Java in a Web-based environment.
Wendy Chang, Deepak Murthy, Aidong Zhang 0001, Tanveer F. Syeda-Mahmood
ICDE4
1997 Efficient Resource Selection in Distributed Visual Information Systems
abstract
made available at other locations for purposes ranting from diagnostics, distance learning, to electronic c;rnI merce.Accessing these repositories in a distributed setting over either a proprietary or a public network poses a number of challenges.With the advances in multimedia databases and the popularization of the Internet, it is now possible to access large image and video repositories distributed throughout the world.One of the challenging problems in such an access is how the information in the respective databases can be summarized to enable an intelligent selection of relevant database sites based on visual queries.This paper presents an approach to this problem that is based on image content-based indexing of a metadatabase at a query distribution server.The metadatabase records a summary of the visual content of the images per database through image templates and statistical features characterizing the similarity value distributions of images.The selection of databases is done by searching the metadatabase using a histogram-based ranking algorithm that uses query similarity to a template and the features of databases associated with the template.The database selection mechanism has been implemented as a metaserver and extensive experiments have been done to demonstrate the effectiveness of the database ranking algorithm.
Wendy Chang, Gholamhosein Sheikholeslami, Aidong Zhang 0001, Tanveer F. Syeda-Mahmood
ACM Multimedia4
1997 Data- and Model-Driven Selection Using Parallel Line Groups
Tanveer F. Syeda-Mahmood
Comput. Vis. Image Underst.1
1997 Data and Model-Driven Selection Using Color Regions
Tanveer F. Syeda-Mahmood
Int. J. Comput. Vis.1
1997 Supporting Content-Based Retrieval in Large Image Database Systems
Edward Remias, Gholamhosein Sheikholeslami, Aidong Zhang 0001, Tanveer F. Syeda-Mahmood
Multim. Tools Appl.4
1996 Recognizing similarity through a constrained non-rigid transform
abstract
The recognition of the class or category of an object based on shape similarity, is an important problem in image databases. Categorizing objects not only helps in efficient image database organization for faster indexing but also allows shape similarity-based querying. The recognition of category is, however, a difficult problem since member objects of a class can show considerable variation in the size and position of individual features even when the overall shape similarity is maintained. In this paper we present an approach to recognizing the class or category of an object in the case where the similarity between member objects is specified by a constrained non-rigid transform. The class is characterized by a single model or prototype consisting of a set of non-overlapping regions and a set of motion (direction and extent) constraints that capture the relation between members of the class. The recognition of category is done by using region correspondence between model and image and recovering the constrained non-rigid transform corresponding to a member of the class that is nearest in shape to the image.
Tanveer F. Syeda-Mahmood
ICPR1
1996 Indexing colored surfaces in images
abstract
We present a method of indexing color images based on color surface reflectance classes. Specifically, we analyse the distributions from members of a surface color class in a chrominance-based color space and show that they form collinear clusters under identical illumination. We also design an optimal color space based on the Fisher discriminant criterion to maximally separate members of different color surface classes. Finally, color indexing is achieved by finding regions in unsegmented images that belong to a queried color surface class using the location and orientation information of the color distributions in the optimal color space. Results are shown that indicate that using both location and orientation information of surface color in the optimal color space lowers the number of false positives and negatives in indexing.
Tanveer F. Syeda-Mahmood, Yong-Qing Cheng
ICPR1
1994 Data and model-driven selection using closely-spaced parallel-line groups
abstract
Selecting regions in an image likely to come from a single object is important for reducing the amount of searching involved in object recognition. Such selections can be purely based on image data (data-driven), or based on the knowledge of the model object (model-driven). In this paper, we present methods for data- and model-driven selection by grouping closely-spaced parallel lines in images. Data-driven selection is achieved by selecting salient line groups that emphasize the likelihood of the groups coming from single objects. Model-driven selection is achieved by selectively generating image line groups that are likely to be the projections of the model groups, taking into account the effect of occlusions, illumination changes and imaging errors. We also present results that indicate a vast improvement in the search performance of a recognition system that is integrated with parallel fine group-based selection.>
Tanveer F. Syeda-Mahmood
CVPR1
1993 Model-driven Selection using Texture
abstract
In this paper we explore the use of texture or pattern information on a 3D object as a cue to isolate regions in an image that are likely to come from the object. We develop a representation of texture based on the linear prediction (LP) spectrum that allows the recognition of the model texture under changes in orientation and occlusions. The candidate matching image regions are obtained without detailed segmentation by a technique called overlapping window analysis. This analysis, under some conditions, guarantees the existence of a window spanning only the model texture regardless of its position and orientation which is sufficient for the recognition of the model texture using the LP spectrum representation. Finally, we evaluate the utility of texture-based selection in combination with other cues such as color in the context of reducing the search involved in recognition. 1.
Tanveer F. Syeda-Mahmood
BMVC1
1992 Data and Model-Driven Selection using Color Regions
Tanveer F. Syeda-Mahmood
ECCV1