VLDB 2026 Research / reviewers in the wild / expert
Milind R. Naphade
dblp:n/MilindRNaphade
· DBLP profile ↗
62ranked-venue papers
27as first author
1since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 47 · 25 first-authorArtificial intelligence and machine learning · 13 · 4 first-author · 1 since 2021Databases, data management, data science and information retrieval · 7Applied, interdisciplinary, general and emerging computing · 4Human-computer interaction and ubiquitous computing · 3
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Language models and text generation · 38% Planning, search and constraint satisfaction · 38% Question answering and dialogue systems · 12% | |
| Human-computer interaction and pervasive computing
3 papers |
Ubiquitous computing and smart environments · 41% Health and well-being technologies · 29% User interface design and tools · 16% | |
| Computer graphics and multimedia
9 papers |
Multimedia analysis and retrieval · 92% Multimedia systems and quality of experience · 8% | |
| Databases, data mining, and information retrieval
4 papers |
Data mining · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Smart cities and intelligent transportation · 81% Energy systems and smart grids · 19% |
Topics — the 27 heaviest of 35, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › agentic language model
tool-augmented language models |
0.9 | 1 | 2025 | T1: A Tool-Oriented Conversational Dataset for Multi-Turn Agentic Planning · NeurIPS 2025 |
Health and well-being technologies
behavior change |
0.3 | 2 | 2013 | The dubuque electricity portal: evaluation of a city-scale residential electricity consumption feedback system · CHI 2013 The dubuque water portal: evaluation of the uptake, use and impact of residential water consumption feedback · CHI 2012 |
Natural language and speech › Question answering and dialogue systems
conversational agents |
0.3 | 1 | 2025 | T1: A Tool-Oriented Conversational Dataset for Multi-Turn Agentic Planning · NeurIPS 2025 |
Ubiquitous computing and smart environments › smart buildings
energy consumption feedback |
0.2 | 1 | 2013 | The dubuque electricity portal: evaluation of a city-scale residential electricity consumption feedback system · CHI 2013 |
User interface design and tools
feedback systems |
0.2 | 1 | 2013 | The dubuque electricity portal: evaluation of a city-scale residential electricity consumption feedback system · CHI 2013 |
Multimedia analysis and retrieval › multimedia analysis › multimedia content description › multimedia semantics
semantic concept detection |
0.2 | 4 | 2005 | Learning the semantics of multimedia queries and concepts from a small number of examples · ACM Multimedia 2005 On the detection of semantic concepts at TRECVID · ACM Multimedia 2004 Semantic representation: search and mining of multimedia content · KDD 2004 |
Collaborative and social computing › social computing
social comparison |
0.1 | 1 | 2012 | The dubuque water portal: evaluation of the uptake, use and impact of residential water consumption feedback · CHI 2012 |
Multimedia analysis and retrieval
video annotation |
0.1 | 3 | 2005 | Visual Concepts for News Story Tracking: Analyzing and Exploiting the NIST TRECVID Video Annotation Experiment · CVPR (1) 2005 User-trainable video annotation using multimodal cues · SIGIR 2003 MPEG-7 video automatic labeling system · ACM Multimedia 2003 |
Smart cities and intelligent transportation
demand prediction |
0.1 | 1 | 2011 | Multi-granular demand forecasting in SmarterWater · UbiComp 2011 |
Data mining
pattern mining |
0.1 | 1 | 2011 | Activity analysis based on low sample rate smart meters · KDD 2011 |
Multimedia analysis and retrieval
video retrieval |
0.1 | 2 | 2001 | A probabilistic framework for semantic video indexing, filtering, and retrieval · IEEE Trans. Multim. 2001 Supporting audiovisual query using dynamic programming · ACM Multimedia 2001 |
Multimedia analysis and retrieval › video indexing
semantic video indexing |
0.1 | 2 | 2001 | A probabilistic framework for semantic video indexing, filtering, and retrieval · IEEE Trans. Multim. 2001 Probabilistic Semantic Video Indexing · NIPS 2000 |
Machine learning › Representation and self-supervised learning
multi-view learning |
0.1 | 1 | 2005 | Semi-Supervised Cross Feature Learning for Semantic Concept Detection in Videos · CVPR (1) 2005 |
Machine learning › Learning paradigms
semi-supervised learning |
0.1 | 1 | 2005 | Semi-Supervised Cross Feature Learning for Semantic Concept Detection in Videos · CVPR (1) 2005 |
Computer vision › Video understanding and tracking › video classification
video concept detection |
0.1 | 1 | 2005 | Semi-Supervised Cross Feature Learning for Semantic Concept Detection in Videos · CVPR (1) 2005 |
Multimedia analysis and retrieval
news story tracking |
0.1 | 1 | 2005 | Visual Concepts for News Story Tracking: Analyzing and Exploiting the NIST TRECVID Video Annotation Experiment · CVPR (1) 2005 |
Multimedia analysis and retrieval › multimedia retrieval › content-based retrieval
query-by-example |
0.1 | 1 | 2005 | Learning the semantics of multimedia queries and concepts from a small number of examples · ACM Multimedia 2005 |
Multimedia systems and quality of experience › multimedia data modeling
multimedia representation |
0.0 | 1 | 2004 | Semantic representation: search and mining of multimedia content · KDD 2004 |
Energy systems and smart grids › building energy management
home energy management |
0.0 | 1 | 2011 | Activity analysis based on low sample rate smart meters · KDD 2011 |
Multimedia analysis and retrieval › interactive retrieval
relevance feedback |
0.0 | 1 | 2001 | Supporting audiovisual query using dynamic programming · ACM Multimedia 2001 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
factor graphs |
0.0 | 1 | 2000 | Probabilistic Semantic Video Indexing · NIPS 2000 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
0.0 | 1 | 2000 | Probabilistic Semantic Video Indexing · NIPS 2000 |
Data mining › text mining › information extraction
concept mining |
0.0 | 1 | 2005 | Visual Concepts for News Story Tracking: Analyzing and Exploiting the NIST TRECVID Video Annotation Experiment · CVPR (1) 2005 |
Data mining › multimodal data mining
multimedia data mining |
0.0 | 1 | 2004 | Semantic representation: search and mining of multimedia content · KDD 2004 |
Computer vision › Image recognition and object detection › object detection
anchor-based detection |
0.0 | 1 | 2003 | MPEG-7 video automatic labeling system · ACM Multimedia 2003 |
Computer vision › Image recognition and object detection
object detection |
0.0 | 1 | 2003 | MPEG-7 video automatic labeling system · ACM Multimedia 2003 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
bayesian network |
0.0 | 1 | 2001 | A probabilistic framework for semantic video indexing, filtering, and retrieval · IEEE Trans. Multim. 2001 |
Methods — techniques the papers use, named apart from their topics
large language model · 0.9caching · 0.9statistical framework · 0.4household behavior modeling · 0.4fixture characteristic modeling · 0.4survey · 0.3interviews · 0.3field study · 0.3positive and unlabeled learning · 0.2biased support vector machine · 0.2multi-resolution prediction · 0.1normalized cut · 0.1laplacian eigenmaps · 0.1information gain feature selection · 0.1semantic concept detection · 0.1feature extraction · 0.1classification · 0.1support vector machine · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | T1: A Tool-Oriented Conversational Dataset for Multi-Turn Agentic PlanningabstractLarge Language Models (LLMs) have demonstrated impressive capabilities as intelligent agents capable of solving complex problems. However, effective planning in scenarios involving dependencies between API or tool calls-particularly in multi-turn conversations-remains a significant challenge. To address this, we introduce T1, a tool-augmented, multi-domain, multi-turn conversational dataset specifically designed to capture and manage inter-tool dependencies across diverse domains. T1 enables rigorous evaluation of agents' ability to coordinate tool use across nine distinct domains (4 single domain and 5 multi-domain) with the help of an integrated caching mechanism for both short- and long-term memory, while supporting dynamic replanning-such as deciding whether to recompute or reuse cached results. Beyond facilitating research on tool use and planning, T1 also serves as a benchmark for evaluating the performance of open-weight and proprietary large language models. We present results powered by T1-Agent highlighting their ability to plan and reason in complex, tool-dependent scenarios. Amartya Chakraborty, Paresh Dashore, Nadia Bathaee, Anmol Jain, Sambit Sahu, Milind R. Naphade, Genta Indra Winata |
NeurIPS | 8 |
| 2013 | The dubuque electricity portal: evaluation of a city-scale residential electricity consumption feedback systemabstractThis paper describes the Dubuque Electricity Portal, a city-scale system aimed at supporting voluntary reductions of electricity consumption. The Portal provided each household with fine-grained feedback on its electricity use, as well as using incentives, comparisons, and goal setting to encourage conservation. Logs, a survey and interviews were used to evaluate the user experience of the Portal during a 20-week pilot with 765 volunteer households. Although the volunteers had already made a wide range of changes to conserve electricity prior to the pilot, those who used the Portal decreased their electricity use by about 3.7%. They also reported increased understanding of their usage, and reported taking an array of actions - both changing their behavior and their electricity infrastructure. The paper discusses the experience of the system's users, and describes challenges for the design of ECF systems, including balancing accessibility and security, a preference for time-based visualizations, and the advisability of multiple modes of feedback, incentives and information presentation. Thomas Erickson, Ming Li 0009, Younghun Kim, Ajay Deshpande, Sambit Sahu, Tian Chao, Noi Sukaviriya, Milind R. Naphade |
CHI | 8 |
| 2013 | Heat pump detection from coarse grained smart meter data with positive and unlabeled learningabstractRecent advances in smart metering technology enable utility companies to have access to tremendous amount of smart meter data, from which the utility companies are eager to gain more insight about their customers. In this paper, we aim to detect electric heat pumps from coarse grained smart meter data for a heat pump marketing campaign. However, appliance detection is a challenging task, especially given a very low granularity and partial labeled even unlabeled data. Traditional methods install either a high granularity smart meter or sensors at every appliance, which is either too expensive or requires technical expertise. We propose a novel approach to detect heat pumps that utilizes low granularity smart meter data, prior sales data and weather data. In particular, motivated by the characteristics of heat pump consumption pattern, we extract novel features that are highly relevant to heat pump usage from smart meter data and weather data. Under the constraint that only a subset of heat pump users are available, we formalize the problem into a positive and unlabeled data classification and apply biased Support Vector Machine (BSVM) to our extracted features. Our empirical study on a real-world data set demonstrates the effectiveness of our method. Furthermore, our method has been deployed in a real-life setting where the partner electric company runs a targeted campaign for 292,496 customers. Based on the initial feedback, our detection algorithm can successfully detect substantial number of non-heat pump users who were identified heat pump users with the prior algorithm the company had used. Hongliang Fei, Younghun Kim, Sambit Sahu, Milind R. Naphade, Sanjay K. Mamidipalli, John Hutchinson |
KDD | 4 |
| 2012 | The dubuque water portal: evaluation of the uptake, use and impact of residential water consumption feedbackabstractThe Dubuque Water Portal is a system aimed at supporting voluntary reductions of water consumption that is intended to be deployed city-wide. It provides each household with fine-grained, near real time feedback on their water consumption, as well as using techniques like social comparison, weekly games, and news and chat to encourage water conservation. This study used logs, a survey and interviews to evaluate a 15-week pilot with 303 households. It describes the Portal's design, and discusses its adoption, use and impacts. The system resulted in a 6.6% decrease in water consumption, and the paper employs qualitative methods to look at the ways in which the Portal was (or wasn't) effective in supporting its users and enabling them to reduce their consumption. The paper concludes with a discussion of design implications for residential feedback systems, and possible engagement models. Thomas Erickson, Mark Podlaseck, Sambit Sahu, Tian Chao, Milind R. Naphade |
CHI | 6 |
| 2011 | Trip analyzer through smartphone appsabstractBroad usage of Smartphones and mobile apps enables both individual trip summary and regional travel demand analysis. In this paper, we describe a trip analysis system, as part of smarter transit service. This trip analysis system consists of mobile apps and a centralized analyzer. It identifies the travel mode and purpose of the trips sensed by mobile devices, provides trip summaries and insights to mobile subscribers, and generates meaningful patterns to support traffic operation planning and transit system design. It is developed and deployed to the Smartphones of the volunteers in Dubuque, IA, to serve both the volunteers and the transit agencies. Preliminary evaluation has demonstrated the applicability of the design. Ming Li 0009, Sambit Sahu, Milind R. Naphade |
GIS | 4 |
| 2011 | Multi-granular demand forecasting in SmarterWaterabstractIn this paper we describe the multi-resolution water consumption prediction based on the SmarterWater system. This prediction service provides household consumption projection and regional demand forecasting for both short-term and med-term. Water consumption prediction, together with the other functions in the SmarterWater service, has been deployed to Dubuque, IA. Consumption behavior change after accessing the service has been observed. Ming Li 0009, Sambit Sahu, Milind R. Naphade, Feng Chen 0001 |
UbiComp | 4 |
| 2011 | Activity analysis based on low sample rate smart metersabstractActivity analysis disaggregates utility consumption from smart meters into specific usage that associates with human activities. It can not only help residents better manage their consumption for sustainable lifestyle, but also allow utility managers to devise conservation programs. Existing research efforts on disaggregating consumption focus on analyzing consumption features with high sample rates (mainly between 1 Hz ~ 1MHz). However, many smart meter deployments support sample rates at most 1/900 Hz, which challenges activity analysis with occurrences of parallel activities, difficulty of aligning events, and lack of consumption features. We propose a novel statistical framework for disaggregation on coarse granular smart meter readings by modeling fixture characteristics, household behavior, and activity correlations. This framework has been implemented into two approaches for different application scenarios, and has been deployed to serve over 300 pilot households in Dubuque, IA. Interesting activity-level consumption patterns have been identified, and the evaluation on both real and synthetic datasets has shown high accuracy on discovering washer and shower. Feng Chen 0001, Bingsheng Wang, Sambit Sahu, Milind R. Naphade, Chang-Tien Lu |
KDD | 5 |
| 2011 | Smarter Water Management: A Challenge for Spatio-Temporal Network Databases
KwangSoo Yang, Shashi Shekhar 0001, Sambit Sahu, Milind R. Naphade |
SSTD | 5 |
| 2010 | Regional behavior change detection via local spatial scanabstractRegional human behavior change refers to the scenarios that people in a certain area exhibit significant behavior deviation from their neighbors and their own past. This regional pattern usually reveals underlying changes of living environment, such as regional development, immigration, disease breakout; or uncovers demographic information from special events, for instance, start/end of school holidays, or religious holidays. Statistically significant behavior changes contain both temporal and spatial characteristics. In this paper, we propose local spatial scan statistic to identify regional behavior changes. To accelerate local search, spatial index is modified to provide data-driven clusters and scalable data access. Base on the restricted spatial index, we provide both exact and approximated approaches to compute local spatial scan. Simulation analysis and case studies on water bills of 15K households validated the efficiency and effectiveness of these approaches on identifying regional behavior changes. Feng Chen 0001, Sambit Sahu, Milind R. Naphade |
GIS | 4 |
| 2009 | Ranking Mortgage Origination Applications Using Customer, Product, Environment and Workflow AttributesabstractIn this paper, we analyze the performance of an end-to-end Mortgage Origination (MO) process. The process begins with the submission of a mortgage application by an applicant to a lender and ends with one of the following outcomes: closing, i.e., loan approved by the lender and accepted by the applicant or non-closing, i.e., loan either rejected by the lender, or approved by the lender and not accepted by the applicant. Ranking mortgage applications by their predicted likelihood of closing at various steps in the process is useful for process efficiency and identification of actionable insights to convert applications likely to non-close into those that are likely to close. To build models for ranking applications at any step of the MO process, we take into account customer and product specific attributes of the applications as well as environment attributes and the history of the applications or workflow.The large state-space of the workflow makes the ranking problem challenging. We propose two workflow attributes, each with a state-space of dimension one, based on the number of visits to any step and a particular step (re-work) respectively. We find that incorporating these workflow attributes into the density modeling technique that we develop results in improvement of 4:8 percent in Average Precision over models that only incorporate customer, product and environment attributes. The simple and scalable density modeling technique allows for easy identification of applications that are likely to non-close and consequent corrective action such as change in the attributes of the mortgage product being offered. Further, our results indicate that the model is comparable to Support Vector Machines and superior to Logistic Regression for ranking. Qihong Shao, Anshul Sheopuri, Milind R. Naphade, Chitra Dorai, Jane Hoffman |
IEEE CLOUD | 3 |
| 2007 | A Greedy Performance Driven Algorithm for Decision Fusion LearningabstractWe propose a greedy performance driven algorithm for learning how to fuse across multiple classification and search systems. We assume a scenario when many such systems need to be fused to generate the final ranking. The algorithm is inspired from Ensemble Learning but takes that idea further for improving generalization capability. Fusion learning is applied to leverage text, visual and model based modalities for 2005 TRECVID query retrieval task. Experiments using the well established retrieval effectiveness measure of mean average precision reveal that our proposed algorithm improves over naive baseline (fusion with equal weights) as well as over Caruana's original algorithm (NACHOS) by 36% and 46% respectively. Dhiraj Joshi, Milind R. Naphade, Apostol Natsev |
ICIP (6) | 2 |
| 2007 | A Generalized Multiple Instance Learning Algorithm for Iterative Distillation and Cross-Granular Propagation of Video AnnotationsabstractVideo annotation is an expensive but necessary task for most vision and learning problems that require building models of visual semantics. This annotation gets prohibitively expensive especially when annotation has to happen at finer grained levels of regions in the videos. One way around the finer grained annotation dilemma is to support annotation at coarser granularity and then propagate this annotation to the finer granularity in a concept-dependent way. In this paper we propose a new generalized multiple instance learning algorithm that can work with any underlying density modeling techniques, and help propagate semantic concepts provided at the coarse granularity of video key-frames to finer grained regions. Our experiments on the NIST TRECVID common annotation corpus reveal improvement in annotation propagation accuracy between 3% to a dramatic 161%. Feng Kang, Milind R. Naphade |
ICIP (2) | 2 |
| 2006 | A Generalized Multiple Instance Learning Algorithm with Multiple Selection Strategies for Cross Granular LearningabstractStatistical learning techniques provide a robust framework for learning representations of semantic concepts from multimedia features. The bottleneck is the number of training samples needed to construct robust models. This is particularly expensive when the annotation needs to happen at finer granularity. We present a novel approach where the annotations may be entered at coarser spatial granularity while the concept may still be learnt at finer granularity. This can speed up annotation significantly. Using the multiple instance learning paradigm, we show that it is possible to learn representations of concepts occurring at the regional level by using annotations for several images. We present a generalized multiple instance learning algorithm with three variations in the strategy to select the most likely positive instance from a positively annotated bag. Furthermore, we show how the three strategies can be combined to improve upon any single strategy and demonstrate 15% performance improvement over any single strategy using a few regional semantic concepts from the TRECVID 2003 benchmark corpus. Feng Kang, Milind R. Naphade |
ICIP | 2 |
| 2006 | Semantic Multimedia Retrieval using Lexical Query Expansion and Model-Based RerankingabstractWe present methods for improving text search retrieval of visual multimedia content by applying a set of visual models of semantic concepts from a lexicon of concepts deemed relevant for the collection. Text search is performed via queries of words or fully qualified sentences, and results are returned in the form of ranked video clips. Our approach involves a query expansion stage, in which query terms are compared to the visual concepts for which we independently build classifier models. We leverage a synonym dictionary and WordNet similarities during expansion. Results over each query are aggregated across the expanded terms and ranked. We validate our approach on the TRECVID 2005 broadcast news data with 39 concepts specifically designed for this genre of video. We observe that concept models improve search results by nearly 50% after model-based re-ranking of text-only search. We also observe that purely model-based retrieval significantly outperforms text-based retrieval on non-named entity queries Alexander Haubold, Apostol Natsev, Milind R. Naphade |
ICME | 3 |
| 2006 | Video News Shot Labeling Refinement via Shot Rhythm ModelsabstractWe present a three-step post-processing method for increasing the precision of video shot labels in the domain of television news. First, we demonstrate that news shot sequences can be characterized by rhythms of alternation (due to dialogue), repetition (due to persistent background settings), or both. Thus a temporal model is necessarily third-order Markov. Second, we demonstrate that the output of feature detectors derived from machine learning methods (in particular, from SVMs) can be converted into probabilities in a more effective way than two suggested existing methods. This is particularly true when detectors are errorful due to sparse training sets, as is common in this domain. Third, we demonstrate that a straightforward application of the Viterbi algorithm on a third-order FSM, constructed from observed transition probabilities and converted feature detector outputs, can refine feature label precision at little cost. We show that on a test corpus of TRECVID 2005 news videos annotated with 39 LSCOM-lite features, the mean increase in the measure of average precision (AP) was 4%, with some of the rarer and more difficult features having relative increases in AP of as much as 67% John R. Kender, Milind R. Naphade |
ICME | 2 |
| 2005 | Visual Concepts for News Story Tracking: Analyzing and Exploiting the NIST TRECVID Video Annotation ExperimentabstractIn the summer of 2003, using an interactive intelligent tool, over 100 researchers in video understanding annotated from the NIST TRECVID database over 62 hours of news video spanning six months of 1998. These 47K shots with 43 3 K labels from over 1000 visual concept categories comprise the largest publicly available ground truth for this domain. Our analysis of this data, combining the tools of statistical natural language processing, machine learning, and computer vision, finds significant novel statistical patterns that can be exploited for the accurate tracking of the episodes of a given news story over time, by using semantic labels that are solely visual. We find that the ground "truth" is very muddy, but by using the feature selection tool of information gain, we extract 14 reliable visual concepts with mid-frequency use; all but one are visual concepts that refer to settings, rather than actors, objects, or events. We discover that the probability of another episode of a named story to recur after a gap of d days is proportional to 1/(d + 1). We define a novel similarity measure incorporating both semantic and temporal properties between episodes i and j as: Dice(i, j)/(1 + gap(i, j)). We exploit a low-level computer vision technique, normalized cut (Laplacian eigenmaps), for clustering these episodes into stories, and in the process document a weakness of this popular technique. We use these empirical results to make specific recommendations on how better visual semantic ontologies for news stories, and how better video annotation tools, should be designed. John R. Kender, Milind R. Naphade |
CVPR (1) | 2 |
| 2005 | Semi-Supervised Cross Feature Learning for Semantic Concept Detection in VideosabstractFor large scale automatic semantic video characterization, it is necessary to learn and model a large number of semantic concepts. But a major obstacle to this is the insufficiency of labeled training samples. Multi-view semi-supervised learning algorithms such as co-training may help by incorporating a large amount of unlabeled data. However, one of their assumptions requiring that each view be sufficient for learning is usually violated in semantic concept detection. In this paper, we propose a novel multi-view semi-supervised learning algorithm called semi-supervised cross feature learning (SCFL). The proposed algorithm has two advantages over co-training. First, SCFL can theoretically guarantee its performance not being significantly degraded even when the assumption of view sufficiency fails. Also, SCFL can also handle additional views of unlabeled data even when these views are absent from the training data. As demonstrated in the TRECVID '03 semantic concept extraction task, the proposed SCFL algorithm not only significantly outperforms the conventional co-training algorithms, but also comes close to achieving the performance when the unlabeled set were to be manually annotated and used for training along with the labeled data set. Milind R. Naphade |
CVPR (1) | 2 |
| 2005 | A generalized multiple instance learning algorithm for large scale modeling of multimedia semanticsabstractStatistical learning techniques provide a robust framework for learning representations of semantic concepts from multimedia features. The bottleneck is the number of training samples needed to construct robust models. This is particularly expensive when the annotation needs to happen at finer granularity. We present a novel approach where the annotations may be entered at coarser spatial granularity while the concept may still be learnt at finer granularity. This can speed up annotation significantly. Using the multiple instance learning paradigm, we show that it is possible to learn representations of concepts occurring at the regional level by using annotations for several images. We present a generalized multiple instance learning algorithm that can scale to a large number of training samples as well as a large number of instances per bag. The algorithm also provides the ability to plug in different density modeling or regression techniques. Using the TREC 2001 Corpus we demonstrate the superior performance of the proposed algorithm over the existing diverse density algorithm. Milind R. Naphade, John R. Smith |
ICASSP (5) | 1 |
| 2005 | Co-training non-robust classifiers for video semantic concept detectionabstractSemantic video characterization by automatic metadata tagging is increasingly popular. While some of these concepts are unimodal manifest in image or audio modalities, a large number of such concepts are multimodal manifest in both the image and the audio modalities. Further while some concepts like outdoors and face occur sufficiently in terms of frequency of occurrence in training sets, a large number are rarer to find thus making them difficult to detect during automatic annotation. Semi-supervised learning algorithms such as co-training may help by incorporating a large amount of unlabeled data, which holds the promise of allowing the redundant information across views to improve the learning performance. Unfortunately, this promise has not been realized in multimedia content analysis partly because the models built using the labeled data alone are not too robust and their noisy classification of the unlabeled data set compounds problems faced by the co-training algorithm. In this paper we analyze whether a judicious application of co-training for automatically labeling some of the unlabeled samples and reinducting them into the training set along with manual quality control can help improve the detection performance. We report our findings in the context of the TRECVID 2003 common annotation corpus. Milind R. Naphade |
ICIP (1) | 2 |
| 2005 | Ontology Design for Video Semantic ThreadsabstractWe propose that, at the highest level of video understanding, the human needs for meaning and the methodologies to extract it are both universal and generic. One must develop an ontology, then develop analyzers that learn the statistical correlates of that ontology, and finally use the analyzers to tie together common occurrences across individual videos. The first step towards adapting the ontology to the genre is the design of automated tools to assist in the annotation of the ground truth; these tools in turn provide feedback on the appropriateness of the filters and the ontology. We support this hypothesis by presenting and discussing some experiments conducted on the NIST TRECVID 2003 video corpus. We also validate this hypothesis by showing the connection between story tracking in our multimedia news and topic detection and tracking in the NIST TDT natural language effort. At the highest level, we find that our annotation tool shows that semantic concepts tend to cluster reliably into a few significant semantic dimensions. For news videos specifically, two of these clusters measure "presidentiality" and "outdoor-ness" John R. Kender, Milind R. Naphade |
ICME | 2 |
| 2005 | Multi-Modal Video Concept Extraction Using Co-TrainingabstractFor large scale automatic semantic video characterization, it is necessary to learn and model a large number of semantic concepts. A major obstacle to this is the insufficiency of labeled training samples. Semi-supervised learning algorithms such as co-training may help by incorporating a large amount of unlabeled data, which allows the redundant information across views to improve the learning performance. Although co-training has been successfully applied in several domains, it has not been used to detect video concepts before. In this paper, we extend co-training to the domain of video concept detection and investigate different strategies of co-training as well as their effects to the detection accuracy. We demonstrate performance based on the guideline of the TRECVID' 03 semantic concept extraction task. Milind R. Naphade |
ICME | 2 |
| 2005 | Learning the semantics of multimedia queries and concepts from a small number of examplesabstractIn this paper we unify two supposedly distinct tasks in multimedia retrieval. One task involves answering queries with a few examples. The other involves learning models for semantic concepts, also with a few examples. In our view these two tasks are identical with the only differentiation being the number of examples that are available for training. Once we adopt this unified view, we then apply identical techniques for solving both problems and evaluate the performance using the NIST TRECVID benchmark evaluation data [15]. We propose a combination hypothesis of two complementary classes of techniques, a nearest neighbor model using only positive examples and a discriminative support vector machine model using both positive and negative examples. In case of queries, where negative examples are rarely provided to seed the search, we create pseudo-negative samples. We then combine the ranked lists generated by evaluating the test database using both methods, to create a final ranked list of retrieved multimedia items. We evaluate this approach for rare concept and query topic modeling using the NIST TRECVID video corpus.In both tasks we find that applying the combination hypothesis across both modeling techniques and a variety of features results in enhanced performance over any of the baseline models, as well as in improved robustness with respect to training examples and visual features. In particular, we observe an improvement of 6% for rare concept detection and 17% for the search task. Apostol Natsev, Milind R. Naphade, Jelena Tesic |
ACM Multimedia | 2 |
| 2004 | Multimodal video search techniques: late fusion of speech-based retrieval and visual content-based retrievalabstractThis paper describes multimodal systems for ad-hoc search constructed by IBM for the TRECVID 2003 benchmark of search systems for broadcast video. These systems all use a late fusion of independently developed speech-based and visual content-based retrieval systems and outperform our individual retrieval systems on both manual and interactive search tasks. For the manual task, our best system used a query-dependent linear weighting between speech-based and image-based retrieval systems. This system has mean average precision (MAP) performance 20% above our best unimodal system for manual search. For the interactive task, where the user has full knowledge of the query topic and the performance of the individual search systems, our best system used an interlacing approach. The user determines the (subjectively) optimal weights A and B for the speech-based and image-based systems, where the multimodal result set is aggregated by combining the top A documents from system A followed by top B documents of system B and then repeating this process until the desired result set size is achieved. This multimodal interactive search has MAP 40% above our best unimodal interactive search system. Arnon Amir, Giridharan Iyengar, Ching-Yung Lin, Milind R. Naphade, Apostol Natsev, Chalapathy Neti, Harriet J. Nock, John R. Smith, Belle L. Tseng |
ICASSP (3) | 4 |
| 2004 | Over-complete representation and fusion for semantic concept detectionabstractAutomatic semantic concept detection in images is a promising tool for alleviating the user effort in annotating and cataloging digital media collections. It enables automatic identification of people, places and objects, for enhanced indexing and searching of home photographs, for example. While constructing robust semantic detectors has been shown feasible for global generic concepts with a sufficient number of good training examples (e.g., indoors, outdoors), many interesting concepts, such as face, people, occur at subpicture granularity, occupy only a portion of the image and therefore frequently have training examples with a reduced signal-to-noise ratio. Such regional concepts are harder to detect due to imperfections in automatic image segmentation algorithms leading to inaccurate object boundaries and low-level feature ambiguities. In this paper we focus on the problem of boosting detection performance of existing regional concept detectors by exploiting detection redundancy. Specifically, we propose to use the same detector multiple times to evaluate and combine multiple detection hypotheses for the same content-but at different content granularities-in order to reduce detection sensitivity to segmentation errors. We validate the approach using support vector machine classifiers for 14 regional semantic concepts from the NISTTRFCVID 2003 common annotation lexicon and show performance improvements of multigranular detection and fusion. Apostol Natsev, Milind R. Naphade, John R. Smith |
ICIP | 2 |
| 2004 | Multi-granular detection of regional semantic conceptsabstractA large number of interesting visual semantic concepts occur at a sub-frame granularity in images and occupy one or more regions at the sub-frame level. Detecting these concepts is a challenge due to segmentation imperfections. We propose multi-granular detection of visual concepts that have regional support. We build a single set of support vector machine based binary concept models from the training set with manually marked up regions. In this paper, we show that detection can be significantly improved by scoring these models over multiple granularities in the test set images, where the regions are automatically detected as a preprocessing step in detection. Using 27 regional semantic concepts from the NIST TRECVID 2003 common annotation lexicon and the corpus, we demonstrate that multi-granular detection leads to improvement in detection. Milind R. Naphade, Apostol Natsev, Ching-Yung Lin, John R. Smith |
ICME | 1 |
| 2004 | Active learning for simultaneous annotation of multiple binary semantic conceptsabstractA model-based approach to video analysis requires annotated corpora. Video annotation, however is a very expensive process. Tools that allow users to annotate video shots with scenes, events, and objects should minimize user interaction. These tools should particularly leverage redundancy in content and advances in machine learning and human computer intelligence to reduce the amount of human interaction needed to annotate large corpora. As corpora sizes and the lexicon grows, this is increasingly relevant. Active learning can play a critical role in reducing the amount of supervision. We apply active learning to the simultaneous annotation of multiple binary concepts. The challenge is to minimize the total number of samples to be annotated across all concepts. Preliminary experiments with the simultaneous annotation of two concepts outdoors and indoors using the TRECVID corpus are promising and reduce annotation workload significantly. Milind R. Naphade, John R. Smith |
ICME | 1 |
| 2004 | Semantic representation: search and mining of multimedia contentabstractSemantic understanding of multimedia content is critical in enabling effective access to all forms of digital media data. By making large media repositories searchable, semantic content descriptions greatly enhance the value of such data. Automatic semantic understanding is a very challenging problem and most media databases resort to describing content in terms of low-level features or using manually ascribed annotations. Recent techniques focus on detecting semantic concepts in video, such as indoor, outdoor, face, people, nature, etc. This approach works for a fixed lexicon for which annotated training examples exist. In this paper we consider the problem of using such semantic concept detection to map the video clips into semantic spaces. This is done by constructing a model vector that acts as a compact semantic representation of the underlying content. We then present experiments in the semantic spaces leveraging such information for enhanced semantic retrieval, classification, visualization, and data mining purposes. We evaluate these ideas using a large video corpus and demonstrate significant performance gains in retrieval effectiveness. Apostol Natsev, Milind R. Naphade, John R. Smith |
KDD | 2 |
| 2004 | On the detection of semantic concepts at TRECVIDabstractSemantic multimedia management is necessary for the effective and widespread utilization of multimedia repositories and realizing the potential that lies untapped in the rich multimodal information content. This challenge has driven researchers to devise new algorithms and systems that enable automatic or semi-automatic tagging of large scale multimedia content with rich semantics. An emerging research area is the detection of a predetermined set of semantic concepts that can act as semantic filters and aid in search, and manipulation. The NIST TRECVID benchmark has responded by creating a task that has evaluated the performance of concept detection. Within the scope of this benchmark task, this paper studies trends in the emerging concept detection systems, architectures and algorithms. It also analyzes strategies that have yielded reasonable success, and challenges and gaps that lie ahead. Milind R. Naphade, John R. Smith |
ACM Multimedia | 1 |
| 2004 | A multi-modal system for the retrieval of semantic video events
Arnon Amir, Sankar Basu, Giridharan Iyengar, Ching-Yung Lin, Milind R. Naphade, John R. Smith, Savitha Srinivasan, Belle L. Tseng |
Comput. Vis. Image Underst. | 5 |
| 2004 | On supervision and statistical learning for semantic multimedia analysis
Milind R. Naphade |
J. Vis. Commun. Image Represent. | 1 |
| 2003 | VideoAL: a novel end-to-end MPEG-7 video automatic labeling systemabstractIn this paper, we describe a novel end-to-end video automatic labeling system, which accepts MPEG-I sequence inputs and generates MPEG-7 XML metadata files based on the prior established anchor models. Seven modules were developed for the system: shot segmentation, region segmentation, annotation, feature extraction, model learning, classification, and XML rendering. The performance of this system has been tested in the NIST TREC-2002 video concept detection benchmark. The proposed system performs best in the mean average precision out of 18 worldwide participants. Ching-Yung Lin, Belle L. Tseng, Milind R. Naphade, Apostol Natsev, John R. Smith |
ICIP (3) | 3 |
| 2003 | Exploring semantic dependencies for scalable concept detectionabstractSemantic concept detection from multimedia features enables high-level access to multimedia content. While constructing robust detectors is feasible for concepts with sufficient training samples, concepts with fewer training samples are hard to train efficiently. Comparable performance may be possible if the dependence of these concepts on the ones that can be robustly modeled is exploited. In this paper we show this phenomenon using the TREC Video 2002 Corpus as a test bed. Using a basic set of 12 semantic concepts modeled with support vector machines, we predict presence of 4 other concepts. We then compare the performance of these predictors with direct SVM models for these 4 concepts and observe improvements of up to 150% in average precision. Milind R. Naphade, Apostol Natsev, John R. Smith |
ICIP (3) | 1 |
| 2003 | Learning visual models of semantic conceptsabstractStatistical machine learning provides a computational framework for mapping low level media features to high level semantics concepts. In this paper we expose the challenges that these techniques face. Using support vector machine (SVM) classification we build models for 34 semantic concepts for the TREC 2002 benchmark corpus. We study the effect of number of examples available for training with respect to their impact on detection. We also examine low level feature fusion as well as parameter sensitivity with SVM classifiers. Milind R. Naphade, John R. Smith |
ICIP (2) | 1 |
| 2003 | Learning regional semantic concepts from incomplete annotationabstractFor multimedia retrieval to be effective, the semantic gap needs to be bridged. Statistical learning techniques provide a robust framework for learning representations of semantic concepts from visual features. The bottleneck is the need to annotate a large number of training samples to construct robust models. We present a novel approach where the annotations may be entered at coarser spatial granularity while the concept may still be learnt at finer granularity. This can speed up annotation significantly and provide bootstrapping. We show that it is possible to learn representations of concepts occurring at the regional level by using annotations for several images, where the annotations are provided only at the global level. The disambiguation can be handled by the multiple instance learning paradigm. We demonstrate this using the TREC 2001 corpus for the concept sky. Milind R. Naphade, John R. Smith |
ICIP (2) | 1 |
| 2003 | Interactive search fusion methods for video database retrievalabstractIn this paper, we investigate a new method for video database retrieval using interactive search fusion. Recent video analysis techniques have enabled the extraction of a variety of descriptors of features, concepts, clusters, classification results, speech and textual terms, MPEG-7 metadata, and so on. However, given an information need users are faced with a daunting task of trying to formulate queries over these multiple disparate data sources in order to retrieve the desired video content. In this paper, we explore a novel approach based on search fusion in which the user interactively builds a query by sequentially choosing among the descriptors and data sources and by selecting from various combining and score aggregation functions to fuse results of individual searches. For example, the system allows building of queries such as "retrieve video clips that have color of beach scenes, the detection of sky, and detection of water". In this paper we present the search fusion method and evaluate the performance on a large video database. John R. Smith, Alejandro Jaimes, Ching-Yung Lin, Milind R. Naphade, Apostol Natsev, Belle L. Tseng |
ICIP (1) | 4 |
| 2003 | Normalized classifier fusion for semantic visual concept detectionabstractIn this paper, we describe our classifier fusion framework for the visual concept detections of NIST TREC-2002 video retrieval benchmark. A normalized ensemble fusion is explored to improve overall performance by incorporating normalization of confidence scores, aggregation via combiner function, and an optimize selection. The normalized classifier fusion shows significant detection improvements for our visual concepts. Belle L. Tseng, Ching-Yung Lin, Milind R. Naphade, Apostol Natsev, John R. Smith |
ICIP (2) | 3 |
| 2003 | Epi-SPIRE: a system for environmental and public health activity monitoringabstractHealth activity monitoring (HAM) has received increasing attention due to the rapid advances of both hardware and software technologies and strong environmental and public health needs. In this paper, we describe the architecture and implementation of the Epi-SPIRE prototype, which is a novel health activity monitoring system that generates alerts from environmental, behavioral, and public health data sources. A model-based approach is used to develop disease and behavior models from multi-modal heterogeneous data sources. Furthermore, a model-based indexing technique has been developed to speed up the data access and retrieval. This system has been successfully applied to various genuine and simulated diseases outbreaks scenarios'. Chung-Sheng Li, Charu C. Aggarwal, Murray Campbell, Yuan-Chi Chang, Gregory Glass, Vijay S. Iyengar, Mahesh Joshi, Ching-Yung Lin, Milind R. Naphade, John R. Smith, Belle L. Tseng, Min Wang 0001, Kun-Lung Wu, Philip S. Yu |
ICME | 9 |
| 2003 | A framework for moderate vocabulary semantic visual concept detectionabstractExtraction of semantic features from visual concepts is essential for meaningful content management in terms of filtering, searching and retrieval. Recently, machine learning techniques have been shown to provide a computational framework to map low level features to high level semantics. In this paper we expose these techniques to the challenge of supporting a moderately large lexicon of semantic concepts. Using the TREC 2002 benchmark corpus for training and validation we investigate a support vector machine based learning system for modeling 34 visual concepts. The detection results show excellent performance for a set of concepts with moderately large training samples. Promising performance is also observed for concepts with few training concepts. Milind R. Naphade, Ching-Yung Lin, Apostol Natsev, Belle L. Tseng, John R. Smith |
ICME | 1 |
| 2003 | Multimedia semantic indexing using model vectorsabstractIn this paper we propose a novel method for multimedia semantic indexing using model vectors. Model vectors provide a semantic signature for multimedia documents by capturing the detection of concepts broadly across a lexicon using a set of independent binary classifiers. While recent techniques have been developed for detecting simple generic concepts such as indoors, outdoors, nature, manmade, faces, people, speech, music, and so forth [W.H. Adams et al., November 2002], these labels directly support only a small number of queries. Model vectors address the problem of answering queries for which relationships to specific concepts is either unknown or indirect by developing a basis across across the lexicon. In the simplest case, each model vector dimension corresponds to the confidence score by which a corresponding concept from the lexicon is detected. However, we show how other information such as relevance, reliability and concept correlation can also be incorporated. Overall, the model vectors can be used in a variety of methods for multimedia indexing, including model-based retrieval, relevance feedback searching and concept querying. In this paper, we present the model vector method and study different strategies for computing and comparing model vectors. We empirically evaluate the retrieval effectiveness of the model vector approach compared to other search methods in a large video retrieval testbed. John R. Smith, Milind R. Naphade, Apostol Natsev |
ICME | 2 |
| 2003 | MPEG-7 video automatic labeling systemabstractIn this demo, we show a novel end-to-end video automatic labeling system, which accepts MPEG-1 sequence inputs and generates MPEG-7 XML metadata files. Detections are based on the prior established anchor models. This system has two parts: model training process and labeling process. They are comprised of seven modules: Shot Segmentation, Region Segmentation, Annotation, Feature Extraction, Model Learning, Classification, and XML Rendering. Ching-Yung Lin, Belle L. Tseng, Milind R. Naphade, Apostol Natsev, John R. Smith |
ACM Multimedia | 3 |
| 2003 | User-trainable video annotation using multimodal cuesabstractThis paper describes progress towards a general framework for incorporating multimodal cues into a trainable system for automatically annotating user-defined semantic concepts in broadcast video. Models of arbitrary concepts are constructed by building classifiers in a score space defined by a pre-deployed set of multimodal models. Results show annotation for user-defined concepts both in and outside the pre-deployed set is competitive with our best video-only models on the TREC Video 2002 corpus. An interesting side result shows speech-only models give performance comparable to our best video-only models for detecting visual concepts such as "outdoors", "face" and "cityscape". Ching-Yung Lin, Milind R. Naphade, Apostol Natsev, Chalapathy Neti, John R. Smith, Belle L. Tseng, Harriet J. Nock, W. H. Adams |
SIGIR | 2 |
| 2002 | A statistical modeling approach to content based retrievalabstractStatistical modeling for content based retrieval is examined in the context of recent TREC Video benchmark exercise. The TREC Video exercise can be viewed as a test bed for evaluation and comparison of a variety of different algorithms on a set of high-level queries for multimedia retrieval. We report on the use of techniques adopted from statistical learning theory. Our method, as in most statistical methods, depend on training of models based on large data sets. A plethora of statistical models such the Gaussian mixture models, support vector machines etc. can be thought of, only a few of which are exploited in this preliminary report. Training requires a large amount of annotated (labeled) data. Thus, we explore use of active learning for the annotation engine that minimizes the number of training samples to be labeled for satisfactory performance. Sankar Basu, Milind R. Naphade, John R. Smith |
ICASSP | 2 |
| 2002 | Modeling semantic concepts to support query by keywords in videoabstractSupporting semantic queries is a challenging problem in video retrieval. We propose the use of a lexicon of semantic concepts for handling the queries. We also propose automatic modeling of lexicon items using probabilistic techniques. We use Gaussian mixture models to build computational representations for a variety of semantic concepts including rocket-launch, outdoor greenery, sky etc. Training requires a large amount of annotated (labeled) data. Using the TREC Video test bed we compare the performance of this system supporting query by keywords with the conventional approach of query by example. Results demonstrate significant gains in performance using the automatically learnt models of semantic concepts. Milind R. Naphade, Sankar Basu, John R. Smith, Ching-Yung Lin, Belle L. Tseng |
ICIP (1) | 1 |
| 2002 | Discovering recurrent events in video using unsupervised methodsabstractProduction videos such as news, sports and movies have a definitive structure that involves short term interaction as well as long term correlation. This structure in video can be captured by models that take into consideration the short term statistics as well as long term recurrence. We investigate the application of probabilistic models that capture this structure. The novel approach is to characterize the short term events in video by models that can account for temporal support in terms of piecewise stationary signals with transitions, These short term events can then be embedded within another temporal model that accounts for transitions between these event and thus characterizes long term history. This also leads to the detection of recurring events in video using a monolithic model. The proposed approach is an unsupervised algorithm for event detection and it can be used for summarization, similarity based matching and enhanced browsing. Milind R. Naphade, Thomas S. Huang |
ICIP (2) | 1 |
| 2002 | Video texture indexing using spatio-temporal waveletsabstractWe present a new compact spatio-temporal texture descriptor designed for indexing dynamic video content. Video texture provides a way to characterize spatio-temporal features such as those corresponding to splashing water, flying birds, blowing trees, and so forth, which are not easily characterized by static feature descriptors. The video texture descriptor measures the 3D wavelet energy corresponding to spatio-temporal-frequency subbands of video segments. The wavelet energy captures the texture patterns as they unfold simultaneously along the spatial and temporal dimensions. We describe the process for extracting video texture features and investigate different methods of constructing video texture descriptors from 3D wavelet energy. We evaluate the video texture descriptors in retrieval experiments and show performance improvements compared to methods based on traditional texture descriptors. Milind R. Naphade, Ching-Yung Lin, John R. Smith |
ICIP (2) | 1 |
| 2002 | Interactive content-based retrieval of videoabstractWe describe a system for content-based retrieval of video that involves a series of query interactions with the user. The proposed approach allows the user, iteratively and selectively, to integrate different feature- and model-based methods of querying in the search process. This allows the user to choose among different retrieved content, features and matching dimensions, and classifiers, as appropriate, given the query objective and interim retrieval results. We investigate several approaches for integrating featureand model-based queries and results in successive query rounds including iterative filtering, score aggregation, and relevance feedback searching. We describe experimental results of applying the interactive content-based retrieval method to an automatically indexed corpus of 11 hours of video. John R. Smith, Sankar Basu, Ching-Yung Lin, Milind R. Naphade, Belle L. Tseng |
ICIP (1) | 4 |
| 2002 | Learning semantic multimedia representations from a small set of examplesabstractWe approach the problem of semantic multimedia retrieval as a supervised learning problem. Defining a lexicon of a small number of interesting semantic concepts we can handle a number of semantic queries. Since the number of interesting concepts available for training is usually small we explore discriminant learning techniques. In particular, we examine the use of kernel based methods and demonstrate impressive retrieval performance using semantic concepts like rocket, outdoor, greenery, sky and face. We also show that loosely coupled multimodal events can be detected based on the late fusion of detection of related auditory and visual concepts. Using a Bayesian network for inference we show how a rocket-launch event can be detected based on the detection of a related visual concept (rocket object) and a related auditory concept (explosion/blast-off). Milind R. Naphade, Ching-Yung Lin, John R. Smith |
ICME (2) | 1 |
| 2002 | Factor graph framework for semantic video indexingabstractVideo query by semantic keywords is one of the most challenging research issues in video data management. To go beyond low-level similarity and access video data content by semantics, we need to bridge the gap between the low-level representation and high-level semantics. This is a difficult multimedia understanding problem. We formulate this problem as a probabilistic pattern-recognition problem for modeling semantics in terms of concepts and context. To map low-level features to high-level semantics, we propose probabilistic multimedia objects (multijects). Examples of multijects in movies include explosion, mountain, beach, outdoor, music, etc. Semantic concepts in videos interact and appear in context. To model this interaction explicitly, we propose a network of multijects (multinet). To model the multinet computationally, we propose a factor graph framework which can enforce spatio-temporal constraints. Using probabilistic models for multijects, rocks, sky, snow, water-body, and forestry/greenery, and using a factor graph as the multinet, we demonstrate the application of this framework to semantic video indexing. We demonstrate how detection performance can be significantly improved using the multinet to take inter-conceptual relationships into account. Our experiments using a large video database consisting of clips from several movies and based on a set of five semantic concepts reveal a significant improvement in detection performance by over 22%. We also show how the multinet is extended to take temporal correlation into account. By constructing a dynamic multinet, we show that the detection performance is further enhanced by as much as 12%. With this framework, we show how keyword-based query and semantic filtering is possible for a predetermined set of concepts. Milind R. Naphade, Igor Kozintsev, Thomas S. Huang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2002 | Extracting semantics from audio-visual content: the final frontier in multimedia retrievalabstractMultimedia understanding is a fast emerging interdisciplinary research area. There is tremendous potential for effective use of multimedia content through intelligent analysis. Diverse application areas are increasingly relying on multimedia understanding systems. Advances in multimedia understanding are related directly to advances in signal processing, computer vision, pattern recognition, multimedia databases, and smart sensors. We review the state-of-the-art techniques in multimedia retrieval. In particular, we discuss how multimedia retrieval can be viewed as a pattern recognition problem. We discuss how reliance on powerful pattern recognition and machine learning techniques is increasing in the field of multimedia retrieval. We review the state-of-the-art multimedia understanding systems with particular emphasis on a system for semantic video indexing centered around multijects and multinets. We discuss how semantic retrieval is centered around concepts and context and the various mechanisms for modeling concepts and context. Milind R. Naphade, Thomas S. Huang |
IEEE Trans. Neural Networks | 1 |
| 2001 | Recognizing high-level audio-visual concepts using contextabstractThe recognition of high-level semantics from audio-visual data is a challenging multimedia understanding problem. The difficulty mainly lies in the gap that exists between low level media features and high level semantic concepts. In an attempt to bridge this gap we proposed a probabilistic framework for semantic understanding. The components of this framework are probabilistic multimedia objects and a graphical network of such objects. We show how the framework supports detection of multiple high-level concepts, which enjoy spatial and temporal support. More importantly, we show why context matters and how it can be modeled. Using a factor graph framework, we model context and use it to improve the detection of sites, objects and events. Using the concepts outdoor and a flying-helicopter we demonstrate how the factor graph multinet models context. Using ROC curves and probability of error curves we support the intuition that context should help. Milind R. Naphade, Thomas S. Huang |
ICIP (3) | 1 |
| 2001 | Duration Dependent Input Output Markov Models For Audio-Visual Event DetectionabstractDetecting semantic events from audio-visual data with Spatiotemporal support is a challenging multimedia Understanding problem. The difficulty lies in the gap that exists between low level media features and high level semantic concept. We present a duration dependent input output Markov model (DDIOMM) to detect events based on multiple modalities. The DDIOMM combines the ability to model nonexponential duration densities with the mapping of input sequences to output sequences. In spirit it resembles the IOHMMs [1] as well as inhomogeneousHMMs [2]. We use the DDIOMM to model the audio-visual event explosion. We compare the detection performance of the DDIOMM with the IOMM as well as the HMM. Experiments reveal that modeling of duration improves detection performance. Milind R. Naphade, Ashutosh Garg 0001, Thomas S. Huang |
ICME | 1 |
| 2001 | Classifying Motion Picture Soundtrack For Video IndexingabstractWe investigate a method for classification of patterns with temporal support. This method combines the ability of a nonlinear-kernel based classifier (in the form of a support vector machine) to discriminate and the ability of a first order Markov chain to model temporal transitions. We apply this to the task of classifying motion picture soundtrack. Experiments with classification of the soundtrack into speech and non-speech audio patterns reveal improvement in classification performance using this proposed method over HMM-based classification as well as SVM-based classification. Using a normalized margin obtained from the SVM and mapping it to a non-negative confidence measure bounded by 1, we attempt to alter the classification of patterns close to the separating boundary, by using the constraints on the transition between the two classes. Sound track classification with semantic classes can help browse and index a video efficiently. Milind R. Naphade, Roy Wang, Thomas S. Huang |
ICME | 1 |
| 2001 | Supporting audiovisual query using dynamic programmingabstractA necessary capability for content-based retrieval is to support the paradigm of query by example. Most systems for video retrieval support queries using image sequences only. We present an algorithm for matching multimodal (audio-visual) patterns for the purpose of content-based video retrieval. The novel ability of our approach to use the information content in multiple media coupled with a strong emphasis on temporal similarity differentiates it from the state-of-the-art in content-based retrieval. At the core of the pattern matching scheme is a dynamic programming algorithm, which leads to a significant improvement in performance. Coupling the use of audio with video this algorithm can be applied to grouping of shots based on audio-visual similarity. We also support relevance feedback. The user can provide feedback to the system, by choosing clips, which are closer to the user's desired target. The system then automatically adjusts the relative weights or relevance of the media and fetches different sets of target clips accordingly. It is our observation that a few iterations of such feedback are generally sufficient, for retrieving the desired video clips. Milind R. Naphade, Roy Wang, Thomas S. Huang |
ACM Multimedia | 1 |
| 2001 | Video retrieval and relevance feedback in the context of a post-integration modelabstractVideo can be viewed as the integration of several heterogeneous media interwoven in a temporally close-coupled fashion. To support the retrieval of video segments under the query-by-example scenario, we purport in this paper a post-integration model that integrates low-level media types to identify visually/auditorily similar video segments. The model also allows flexible relevance feedback on the user end to further improve the speed and accuracy of searching in a video database. As its name implies, the post-integration model first treats models of the underlying media as independent processes and then combines distance scores from each of the underlying media at the later stage. This decoupling allows the usage of efficient algorithms to fast and accurately compare thousands of video clips in a moderate-to-large-sized video database. It also naturally facilitates improved user interaction by fast dynamic weight adjustment on the media distance scores in a multiple query succession. We explicit our models and dynamic weight adjustment schemes in this paper and offer demonstration with a working system on a large video database set. Roy Wang, Milind R. Naphade, Thomas S. Huang |
MMSP | 2 |
| 2001 | A probabilistic framework for semantic video indexing, filtering, and retrievalabstractSemantic filtering and retrieval of multimedia content is crucial for efficient use of the multimedia data repositories. Video query by semantic keywords is one of the most difficult problems in multimedia data retrieval. The difficulty lies in the mapping between low-level video representation and high-level semantics. We therefore formulate the multimedia content access problem as a multimedia pattern recognition problem. We propose a probabilistic framework for semantic video indexing, which call support filtering and retrieval and facilitate efficient content-based access. To map low-level features to high-level semantics we propose probabilistic multimedia objects (multijects). Examples of multijects in movies include explosion, mountain, beach, outdoor, music etc. Semantic concepts in videos interact and to model this interaction explicitly, we propose a network of multijects (multinet). Using probabilistic models for six site multijects, rocks, sky, snow, water-body forestry/greenery and outdoor and using a Bayesian belief network as the multinet we demonstrate the application of this framework to semantic indexing. We demonstrate how detection performance can be significantly improved using the multinet to take interconceptual relationships into account. We also show how the multinet can fuse heterogeneous features to support detection based on inference and reasoning. Milind R. Naphade, Thomas S. Huang |
IEEE Trans. Multim. | 1 |
| 2000 | Multimedia Understanding: Challenges in the New MillenniumabstractMultimedia understanding is a fast emerging interdisciplinary research area with tremendous potential to increase the effective use of multimedia content. Diverse application areas are increasingly relying on multimedia understanding systems. Advances in multimedia understanding are related directly to advances in various disciplines including signal processing, computer vision, pattern recognition, multimedia databases and smart sensors. In this paper we present a perspective on the state-of-the-art in multimedia understanding systems and also discuss emerging trends in such systems in the new millennium. A generic framework is discussed for multimedia understanding systems and a semantic video indexing framework is given. Milind R. Naphade, Thomas S. Huang |
ICIP | 1 |
| 2000 | Inferring Semantic Concepts for Video Indexing and RetrievalabstractThis paper proposes a novel probabilistic framework for semantic indexing and retrieval in digital video. The components of the framework are probabilistic multimedia objects (multijects) and a network of such objects (multinet). The main contribution is a Bayesian multinet which enhances the detection performance of individual multijects and supports inference of concepts that are not observed directly in the multiple media. This inference is based on their relation with observable concepts. We develop multijects for detecting sites (locations) in video and integrate the multijects using a multinet in the form of a Bayesian network. We also use the site multijects and the multinet to infer the presence of the multiject Outdoor which has no direct support in media features. Milind R. Naphade, Thomas S. Huang |
ICIP | 1 |
| 2000 | Learning Sparse Multiple Cause ModelsabstractMultiple cause models (MCM) are a way to describe patterns as a superposition of a selection of cause patterns. In contrast to clustering methods and dimensionality reduction, multiple cause models are capable of turning local features on and off and this makes them a more realistic model for many types of data. However, inference and learning in general multiple cause models takes an amount of time that is exponential in the number of causes. We present an approximate inference algorithm that examines only sparse cause patterns, i.e., those configurations of causes where only a small number of causes are active at a time. This leads to an approximate EM algorithm that maximizes a lower bound on the likelihood of a data set. We show that this sparse multiple cause model can model different types of human facial expression patterns. Performance comparison of the MCM classifier with the SNoW (sparse network of winnows) architecture and the nearest neighbor classifier reveals significant improvement in classification accuracy using the MCM classifier. Milind R. Naphade, Lawrence S. Chen, Thomas S. Huang, Brendan J. Frey |
ICPR | 1 |
| 2000 | Semantic Video Indexing Using a Probabilistic FrameworkabstractProposes a probabilistic framework for semantic video indexing. The components of the framework are multijects and multinets. Multijects are probabilistic multimedia objects representing semantic features or concepts. A multinet is a probabilistic network of multijects which accounts for the interaction between concepts. The main contribution of the paper is the application of a graphical probabilistic framework to build the multinet. The multinet enhances the detection performance of individual multijects, provides a unified framework for integrating multiple modalities and supports inference of unobservable concepts based on their relation with observable concepts. We develop multijects for detecting sites (locations) in video and integrate the multijects using multinet in the form of a Bayesian network. Detection performance is significantly improved using the multinet. Milind R. Naphade, Thomas S. Huang |
ICPR | 1 |
| 2000 | Probabilistic Semantic Video IndexingabstractWe propose a novel probabilistic framework for semantic video in(cid:173) dexing. We define probabilistic multimedia objects (multijects) to map low-level media features to high-level semantic labels. A graphical network of such multijects (multinet) captures scene con(cid:173) text by discovering intra-frame as well as inter-frame dependency relations between the concepts. The main contribution is a novel application of a factor graph framework to model this network. We model relations between semantic concepts in terms of their co-occurrence as well as the temporal dependencies between these concepts within video shots. Using the sum-product algorithm [1] for approximate or exact inference in these factor graph multinets, we attempt to correct errors made during isolated concept detec(cid:173) tion by forcing high-level constraints. This results in a significant improvement in the overall detection performance. Milind R. Naphade, Igor Kozintsev, Thomas S. Huang |
NIPS | 1 |
| 1998 | Probabalistic Multimedia Objects (Multijects): A Novel Approach to Video Indexing and Retrieval in Multimedia Systems
Milind R. Naphade, Trausti T. Kristjansson, Brendan J. Frey, Thomas S. Huang |
ICIP (3) | 1 |
| 1998 | A High-Performance Shot Boundary Detection Algorithm using Multiple CuesabstractA central step in content-based video retrieval is the temporal segmentation of video. An application independent approach to video segmentation is to detect temporally contiguous segments without significant content change between successive frames. Each such segment is termed a shot. A high-performance shot boundary detection-based video segmentation algorithm is proposed. The technique uses unsupervised clustering on a multiple feature input space, followed by a heuristic elimination process to detect, with almost perfect accuracy, shot boundaries in the video. With an extremely high accuracy coupled with a very small number of false positives, this algorithm outperforms most of the existing techniques. Milind R. Naphade, Rajiv Mehrotra, A. Müfit Ferman, Jim Warnick, Thomas S. Huang, A. Murat Tekalp |
ICIP (1) | 1 |