Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

John R. Smith

dblp:s/JohnRSmith · also John Smith 0001 · DBLP profile ↗
← Back
126ranked-venue papers
29as first author
4since 2021 · last 2023
0000-0001-5877-4859ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 100 · 21 first-author · 4 since 2021Artificial intelligence and machine learning · 19 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 15 · 6 first-authorSystems, architecture and hardware · 6 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2Computer networks · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
27 papers
Multimedia analysis and retrieval · 96% Multimedia systems and quality of experience · 2% Audio and music processing · 2%
Artificial intelligence
14 papers
Image recognition and object detection · 20% Deep learning architectures and training · 17% Video understanding and tracking · 16%
Databases, data mining, and information retrieval
21 papers
Information retrieval · 37% Data mining · 32% Query processing and optimization · 16%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computational science and engineering · 100%

Topics — the 30 heaviest of 108, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Multimedia analysis and retrieval › multimedia analysis
human-centric multimedia analysis
2.242023
HCMA '23: 4th International Workshop on Human-Centric Multimedia Analysis · ACM Multimedia 2023
HCMA'22: 3rd International Workshop on Human-Centric Multimedia Analysis · ACM Multimedia 2022
HUMA'21: 2nd International Workshop on Human-centric Multimedia Analysis · ACM Multimedia 2021
Machine learning › Deep learning architectures and training
foundation model
0.712023
Foundation Model for Material Science · AAAI 2023
Computational science and engineering › materials science
materials discovery
0.712023
Foundation Model for Material Science · AAAI 2023
Computational science and engineering
materials science
0.712023
Foundation Model for Material Science · AAAI 2023
Multimedia analysis and retrieval › multimedia analysis
multimodal behavior analysis
0.712023
HCMA '23: 4th International Workshop on Human-Centric Multimedia Analysis · ACM Multimedia 2023
Multimedia analysis and retrieval › action recognition
action detection
0.612022
HCMA'22: 3rd International Workshop on Human-Centric Multimedia Analysis · ACM Multimedia 2022
Machine learning › Transfer learning and domain adaptation › few-shot learning
cross-domain few-shot learning
0.412020
A Broader Study of Cross-Domain Few-Shot Learning · ECCV (27) 2020
Computer vision › Video understanding and tracking
action recognition
0.412019
Automatic Curation of Sports Highlights Using Multimodal Excitement Features · IEEE Trans. Multim. 2019
Multimedia analysis and retrieval › video summarization
video highlight detection
0.412019
Automatic Curation of Sports Highlights Using Multimodal Excitement Features · IEEE Trans. Multim. 2019
Computer vision › Image recognition and object detection
attribute-based recognition
0.422014
Modeling Attributes from Category-Attribute Proportions · ACM Multimedia 2014
Designing Category-Level Attributes for Discriminative Visual Recognition · CVPR 2013
Multimedia analysis and retrieval
multimodal fusion
0.322017
IBM High-Five: Highlights From Intelligent Video Engine · ACM Multimedia 2017
Optimal multimodal fusion for multimedia data analysis · ACM Multimedia 2004
Information retrieval
ranking
0.322015
Top Rank Supervised Binary Coding for Visual Search · ICCV 2015
Imbalanced RankBoost for efficiently ranking large-scale image/video collections · CVPR 2009
Multimedia analysis and retrieval › affective computing
affective video content analysis
0.312017
Harnessing A.I. for Augmenting Creativity: Application to Movie Trailer Creation · ACM Multimedia 2017
Multimedia analysis and retrieval › video summarization
highlight video generation
0.312017
IBM High-Five: Highlights From Intelligent Video Engine · ACM Multimedia 2017
Multimedia analysis and retrieval › multimedia analysis
multimodal video analysis
0.312017
Harnessing A.I. for Augmenting Creativity: Application to Movie Trailer Creation · ACM Multimedia 2017
Multimedia analysis and retrieval › video summarization
trailer generation
0.312017
Harnessing A.I. for Augmenting Creativity: Application to Movie Trailer Creation · ACM Multimedia 2017
Multimedia analysis and retrieval
video summarization
0.312017
IBM High-Five: Highlights From Intelligent Video Engine · ACM Multimedia 2017
Computer vision › Image recognition and object detection › image classification
fine-grained image classification
0.212016
Learning to Make Better Mistakes: Semantics-aware Visual Food Recognition · ACM Multimedia 2016
Computer vision › Image recognition and object detection
food recognition
0.212016
Learning to Make Better Mistakes: Semantics-aware Visual Food Recognition · ACM Multimedia 2016
Machine learning › Optimization for machine learning
convex optimization
0.212015
Low-Rank Similarity Metric Learning in High Dimensions · AAAI 2015
Machine learning › Deep learning architectures and training › regularization › spectral regularization
nuclear norm regularization
0.212015
Low-Rank Similarity Metric Learning in High Dimensions · AAAI 2015
Information retrieval
image retrieval
0.212015
Top Rank Supervised Binary Coding for Visual Search · ICCV 2015
Machine learning and data management
metric learning
0.212015
Low-Rank Similarity Metric Learning in High Dimensions · AAAI 2015
Information retrieval
similarity learning
0.212015
Low-Rank Similarity Metric Learning in High Dimensions · AAAI 2015
Machine learning › Representation and self-supervised learning › representation learning › semantic representation learning
attribute learning
0.212014
Modeling Attributes from Category-Attribute Proportions · ACM Multimedia 2014
Machine learning › Time series and sequential data › anomaly detection
one-class classification
0.212014
Unsupervised One-Class Learning for Automatic Outlier Removal · CVPR 2014
Data mining
anomaly detection
0.212014
Unsupervised One-Class Learning for Automatic Outlier Removal · CVPR 2014
Data mining › semi-supervised learning
learning from label proportions
0.212014
Modeling Attributes from Category-Attribute Proportions · ACM Multimedia 2014
Data mining › anomaly detection
outlier removal
0.212014
Unsupervised One-Class Learning for Automatic Outlier Removal · CVPR 2014
Multimedia analysis and retrieval
video retrieval
0.232011
Visual memes in social media: tracking real-world news in YouTube videos · ACM Multimedia 2011
Imbalanced RankBoost for efficiently ranking large-scale image/video collections · CVPR 2009
VideoZoom Spatio-Temporal Video Browser · IEEE Trans. Multim. 1999

Methods — techniques the papers use, named apart from their topics

multimodal learning · 1.9structure generation · 1.3property prediction · 1.3multimodal analysis · 1.3human behavior understanding · 1.3foundation model · 1.3multimodal fusion · 1.0classifier learning with reduced annotation · 0.8alternating direction method of multipliers · 0.7statistical modeling · 0.6multimodal sentiment analysis · 0.6multimodal sensing · 0.6large-scale computing · 0.6trace norm regularization · 0.4SVD · 0.4crowd reaction analysis · 0.3random walk smoothing · 0.2multi-task loss · 0.2
YearPublicationVenuePosition
2023 Foundation Model for Material Science
abstract
Foundation models (FMs) are achieving remarkable successes to realize complex downstream tasks in domains including natural language and visions. In this paper, we propose building an FM for material science, which is trained with massive data across a wide variety of material domains and data modalities. Nowadays machine learning models play key roles in material discovery, particularly for property prediction and structure generation. However, those models have been independently developed to address only specific tasks without sharing more global knowledge. Development of an FM for material science will enable overarching modeling across material domains and data modalities by sharing their feature representations. We discuss fundamental challenges and required technologies to build an FM from the aspects of data preparation, model development, and downstream tasks.
Seiji Takeda, Akihiro Kishimoto, Lisa Hamada, Daiju Nakano, John R. Smith
AAAI5
2023 HCMA '23: 4th International Workshop on Human-Centric Multimedia Analysis
abstract
Understanding human interactions within diverse media contexts has emerged as a fundamental challenge. The explosive growth of multimedia data not only provides opportunities for human-centirc analysis but also increases the complexity of processing multimodal data. To address this pivotal challenge and explore its multifaceted dimensions, the Fourth International Workshop on Human-Centric Multimedia Analysis is concentrated on the tasks of human-centric analysis with multimedia and multimodal information. By delving into the nuances of human behavior within multimedia, this workshop aims to uncover novel insights, showcase innovative methodologies, and discuss future directions. With a spotlight on cutting-edge research and a focus on real-world applications, the workshop seeks to equip researchers and practitioners with the tools and knowledge to navigate the intricacies of human-centric multimedia analysis.
Jingkuan Song, Wu Liu 0005, Xinchen Liu, Dingwen Zhang, Chaowei Fang, Hongyuan Zhu 0002, Wenbing Huang 0001, John R. Smith, Xin Wang 0019
ACM Multimedia8
2022 HCMA'22: 3rd International Workshop on Human-Centric Multimedia Analysis
abstract
The Third International Workshop on Human-Centric Multimedia Analysis concentrates on the tasks of human-centric analysis with multimedia and multimodal information. It involves multiple tasks such as face detection and recognition, human body pattern analysis, person re-identification, human action detection, etc. Today, multiple multimedia sensing technologies and large-scale computing infrastructures are emerging at a rapid velocity a wide variety of big multi-modality data for human-centric analysis, which provides rich knowledge to help tackle these challenges. Researchers have strived to push the limits of human-centric multimedia analysis in a wide variety of applications, such as intelligent surveillance, retailing, fashion design, and services. Therefore, this workshop aims to provide a platform to bridge the gap between the communities of human analysis and multimedia.
Dingwen Zhang, Chaowei Fang, Wu Liu 0005, Xinchen Liu, Jingkuan Song, Hongyuan Zhu 0002, Wenbing Huang 0001, John R. Smith
ACM Multimedia8
2021 HUMA'21: 2nd International Workshop on Human-centric Multimedia Analysis
abstract
The Second International Workshop on Human-centric Multimedia Analysis is focused on human-centric analysis using multimedia information. The human-centric multimedia analysis is one of the fundamental and challenging problems of multimedia understanding. It involves various human-centric analysis tasks like face recognition, human pose estimation, person re-identification, human action recognition, person tracking, human-computer interaction, etc. Nowadays, various multimedia sensing devices and large-scale computing infrastructures are generating a wide variety of multi-modality data at a rapid velocity, which supplies rich knowledge to tackle these challenges for human-centric analysis. Researchers and engineers have strived to push the limits of human-centric multimedia analysis in a wide variety of applications, such as smart city, retailing, intelligent manufacturing, and public services. To this end, our workshop aims to provide a platform to promote exchanges and integration for the fields of human analysis and multimedia.
Wu Liu 0005, Xinchen Liu, Jingkuan Song, Dingwen Zhang, Wenbing Huang 0001, Junbo Guo, John R. Smith
ACM Multimedia7
2020 A Broader Study of Cross-Domain Few-Shot Learning
Yunhui Guo, Noel Codella, Leonid Karlinsky, James V. Codella, John R. Smith, Kate Saenko, Tajana Rosing, Rogério Feris
ECCV (27)5
2020 HUMA'20: 1st International Workshop on Human-Centric Multimedia Analysis
abstract
The First International Workshop on Human-Centric MultimediaAnalysis is concentrated on the tasks of human-centric analysis with multimedia and multimodal information. It is one of the fundamental and challenging problems of multimedia understanding. The human-centric multimedia analysis involves multiple tasks such as face detection and recognition, human body pattern analysis, person re-identification, human action detection, person tracking,human-object interaction, and so on. Today, multiple multimedia sensing technologies and large-scale computing infrastructures are producing at a rapid velocity a wide variety of big multi-modality data for human-centric analysis, which provides rich knowledge to help tackle these challenges. Researchers have strived to push the limits of human-centric multimedia analysis in a wide variety of applications, such as intelligent surveillance, retailing, fashion design, and services. Therefore, this workshop aims to provide a platform to bridge the gap between the communities of human analysis and multimedia.
Wu Liu 0005, Chuang Gan 0001, Jingkuan Song, Dingwen Zhang, Wenbing Huang 0001, John R. Smith
ACM Multimedia6
2019 Automatic Curation of Sports Highlights Using Multimodal Excitement Features
abstract
The production of sports highlight packages summarizing a game's most exciting moments is an essential task for broadcast media. Yet, it requires labor-intensive video editing. We propose a novel approach for auto-curating sports highlights, and demonstrate it to create a first of a kind, real-world system for the editorial aid of golf and tennis highlight reels. Our method fuses information from the players’ reactions (action recognition such as high-fives and fist pumps), players’ expressions (aggressive, tense, smiling, and neutral), spectators (crowd cheering), commentator (tone of the voice and word analysis), and game analytics to determine the most interesting moments of a game. We accurately identify the start and end frames of key shot highlights with additional metadata, such as the player's name and the whole number, or analysts input allowing personalized content summarization and retrieval. In addition, we introduce new techniques for learning our classifiers with reduced manual training data annotation by exploiting the correlation of different modalities. Our work has been demonstrated at a major golf tournament (2017 Masters) and two major international tennis tournaments (2017 Wimbledon and U.S. Open), successfully extracting highlights through the course of the sporting events. For the 2017 Masters, 54% of the clips selected by our system overlapped with the official highlights reels. Furthermore, user studies showed that 90% of the non-overlapping ones were of the same quality of the official clips for the 2017 Masters, while the automatic selection of clips for highlights of 2017 Wimbledon and 2017 US Open agreed with human preferences 80% and 84.2% of the time, respectively.
Michele Merler, Khoi-Nguyen C. Mac, Dhiraj Joshi, Quoc-Bao Nguyen, Stephen Hammer, John Kent, Jinjun Xiong, Minh N. Do, John R. Smith, Rogério Feris
IEEE Trans. Multim.9
2017 IBM High-Five: Highlights From Intelligent Video Engine
abstract
We introduce a novel multi-modal system for auto-curating golf highlights that fuses information from players' reactions (celebration actions), spectators (crowd cheering), and commentator (tone of the voice and word analysis) to determine the most interesting moments of a game. The start of a highlight is determined with additional metadata (player's name and the hole number), allowing personalized content summarization and retrieval. Our system was demonstrated at Masters 2017, a major golf tournament, generating real-time highlights from four live video streams over four days.
Dhiraj Joshi, Michele Merler, Quoc-Bao Nguyen, Stephen Hammer, John Kent, John R. Smith, Rogério Feris
ACM Multimedia6
2017 Harnessing A.I. for Augmenting Creativity: Application to Movie Trailer Creation
abstract
In this paper, we describe the first-ever machine human collaboration at creating a real movie trailer (officially released by 20th Century Fox). We introduce an intelligent system designed to understand and encode patterns and types of emotions in horror movies that are useful in trailers. We perform multi-modal semantics extraction including audio visual sentiments and scene analysis and employ a statistical approach to model the key defining components that characterize horror movie trailers. The system was applied on a full-length feature film, "Morgan'' released in 2016 where the system identified 10 moments as best candidates for a trailer. We partnered with a professional filmmaker who arranged and edited each of the moments together to construct a comprehensive trailer completing the entire processing as well as the final trailer assembly within 24 hours. We discuss disruptive opportunities for the film industry and the tremendous media impact of the AI trailer. We confirm the effectiveness of our trailer with a very supportive user study. Finally based on our close interaction with the film industry, we also introduce and investigate the novel paradigm of tropes within the context of movies for advancing content creation.
John R. Smith, Dhiraj Joshi, Benoit Huet, Winston H. Hsu, Jozef Cota
ACM Multimedia1
2017 Leveraging multiple cues for recognizing family photos
Xiaolong Wang 0006, Guodong Guo, Michele Merler, Noel Codella, M. V. Rohith, John R. Smith, Chandra Kambhamettu
Image Vis. Comput.6
2016 Oracle Performance for Visual Captioning
Nicolas Ballas, Kyunghyun Cho, John R. Smith, Yoshua Bengio
BMVC4
2016 A closer look at Faster R-CNN for vehicle detection
abstract
Faster R-CNN achieves state-of-the-art performance on generic object detection. However, a simple application of this method to a large vehicle dataset performs unimpressively. In this paper, we take a closer look at this approach as it applies to vehicle detection. We conduct a wide range of experiments and provide a comprehensive analysis of the underlying structure of this model. We show that through suitable parameter tuning and algorithmic modification, we can significantly improve the performance of Faster R-CNN on vehicle detection and achieve competitive results on the KITTI vehicle dataset. We believe our studies are instructive for other researchers investigating the application of Faster R-CNN to their problems and datasets.
Quanfu Fan, Lisa M. Brown, John R. Smith
Intelligent Vehicles Symposium3
2016 Learning to Make Better Mistakes: Semantics-aware Visual Food Recognition
abstract
We propose a visual food recognition framework that integrates the inherent semantic relationships among fine-grained classes. Our method learns semantics-aware features by formulating a multi-task loss function on top of a convolutional neural network (CNN) architecture. It then refines the CNN predictions using a random walk based smoothing procedure, which further exploits the rich semantic information. We evaluate our algorithm on a large "food-in-the-wild" benchmark, as well as a challenging dataset of restaurant food dishes with very few training images. The proposed method achieves higher classification accuracy than a baseline which directly fine-tunes a deep learning network on the target dataset. Furthermore, we analyze the consistency of the learned model with the inherent semantic relationships among food categories. Results show that the proposed approach provides more semantically meaningful results than the baseline method, even in cases of mispredictions.
Hui Wu 0009, Michele Merler, Rosario Uceda-Sosa, John R. Smith
ACM Multimedia4
2015 Low-Rank Similarity Metric Learning in High Dimensions
abstract
Metric learning has become a widespreadly used tool in machine learning. To reduce expensive costs brought in by increasing dimensionality, low-rank metric learning arises as it can be more economical in storage and computation. However, existing low-rank metric learning algorithms usually adopt nonconvex objectives, and are hence sensitive to the choice of a heuristic low-rank basis. In this paper, we propose a novel low-rank metric learning algorithm to yield bilinear similarity functions. This algorithm scales linearly with input dimensionality in both space and time, therefore applicable to high-dimensional data domains. A convex objective free of heuristics is formulated by leveraging trace norm regularization to promote low-rankness. Crucially, we prove that all globally optimal metric solutions must retain a certain low-rank structure, which enables our algorithm to decompose the high-dimensional learning task into two steps: an SVD-based projection and a metric learning problem with reduced dimensionality. The latter step can be tackled efficiently through employing a linearized Alternating Direction Method of Multipliers. The efficacy of the proposed algorithm is demonstrated through experiments performed on four benchmark datasets with tens of thousands of dimensions.
Wei Liu 0005, Cun Mu, Rongrong Ji, Shiqian Ma, John R. Smith, Shih-Fu Chang
AAAI5
2015 Top Rank Supervised Binary Coding for Visual Search
abstract
In recent years, binary coding techniques are becoming increasingly popular because of their high efficiency in handling large-scale computer vision applications. It has been demonstrated that supervised binary coding techniques that leverage supervised information can significantly enhance the coding quality, and hence greatly benefit visual search tasks. Typically, a modern binary coding method seeks to learn a group of coding functions which compress data samples into binary codes. However, few methods pursued the coding functions such that the precision at the top of a ranking list according to Hamming distances of the generated binary codes is optimized. In this paper, we propose a novel supervised binary coding approach, namely Top Rank Supervised Binary Coding (Top-RSBC), which explicitly focuses on optimizing the precision of top positions in a Hamming-distance ranking list towards preserving the supervision information. The core idea is to train the disciplined coding functions, by which the mistakes at the top of a Hamming-distance ranking list are penalized more than those at the bottom. To solve such coding functions, we relax the original discrete optimization objective with a continuous surrogate, and derive a stochastic gradient descent to optimize the surrogate objective. To further reduce the training time cost, we also design an online learning algorithm to optimize the surrogate objective more efficiently. Empirical studies based upon three benchmark image datasets demonstrate that the proposed binary coding approach achieves superior image search accuracy over the state-of-the-arts.
Dongjin Song, Wei Liu 0005, Rongrong Ji, David A. Meyer 0001, John R. Smith
ICCV5
2015 You are what you tweet...pic! gender prediction based on semantic analysis of social media images
abstract
We propose a method to extract user attributes from the pictures posted in social media feeds, specifically gender information. While traditional approaches rely on text analysis or exploit visual information only from the user profile picture or colors, we propose to look at the distribution of semantics in the pictures coming from the whole feed of a person to estimate gender. In order to compute such semantic distribution, we trained models from existing visual taxonomies to recognize objects, scenes and activities, and applied them to the images in each user's feed. Experiments conducted on a set of ten thousand twitter users and their collection of half a million images revealed that the gender signal can indeed be extracted from the users image feed (75.6% accuracy). Furthermore, the combination of visual cues resulted almost as strong as textual analysis in predicting gender, while providing complementary information that can be employed to further boost gender prediction accuracy to 88% when combined with textual data. As a byproduct of our investigation, we were also able to extrapolate the semantic categories of posted pictures mostly correlated to males and females.
Michele Merler, Liangliang Cao, John R. Smith
ICME3
2015 Multi-facet Learning using Deep Convolutional Neural Network for Person-Related Categories in Photos
abstract
This paper proposes to leverage multiple facets of person photos to improve the training of deep neural networks. Existing studies usually require a lot of labeled images to train deep convolutional networks. Our study suggests exploring multiple datasets and learning effective representation to learn related visual concepts. The practice of learning from multiple facets implicitly enforces to share features for image recognition. We show deep neural network benefits from the learning of multiple person-related categories in photos. Faceted classification systems learn from multiple resources, and alleviate the overfitting problems in deep learning. Moreover, by exploring multiple taxonomies of an object, it provides a finer annotation for the query images.
Liangliang Cao, Zhicheng Yan 0001, John R. Smith
ICMR3
2014 Unsupervised One-Class Learning for Automatic Outlier Removal
abstract
Outliers are pervasive in many computer vision and pattern recognition problems. Automatically eliminating outliers scattering among practical data collections becomes increasingly important, especially for Internet inspired vision applications. In this paper, we propose a novel one-class learning approach which is robust to contamination of input training data and able to discover the outliers that corrupt one class of data source. Our approach works under a fully unsupervised manner, differing from traditional one-class learning supervised by known positive labels. By design, our approach optimizes a kernel-based max-margin objective which jointly learns a large margin one-class classifier and a soft label assignment for inliers and outliers. An alternating optimization algorithm is then designed to iteratively refine the classifier and the labeling, achieving a provably convergent solution in only a few iterations. Extensive experiments conducted on four image datasets in the presence of artificial and real-world outliers demonstrate that the proposed approach is considerably superior to the state-of-the-arts in obliterating outliers from contaminated one class of images, exhibiting strong robustness at a high outlier proportion up to 60%.
Wei Liu 0005, Gang Hua 0001, John R. Smith
CVPR3
2014 Automated Medical Image Modality Recognition by Fusion of Visual and Text Information
Noel Codella, Jonathan H. Connell, Sharath Pankanti, Michele Merler, John R. Smith
MICCAI (2)5
2014 Modeling Attributes from Category-Attribute Proportions
abstract
Attribute-based representation has been widely used in visual recognition and retrieval due to its interpretability and cross-category generalization properties. However, classic attribute learning requires manually labeling attributes on the images, which is very expensive, and not scalable. In this paper, we propose to model attributes from category-attribute proportions. The proposed framework can model attributes without attribute labels on the images. Specifically, given a multi-class image datasets with N categories, we model an attribute, based on an N-dimensional category-attribute proportion vector, where each element of the vector characterizes the proportion of images in the corresponding category having the attribute. The attribute learning can be formulated as a learning from label proportion (LLP) problem. Our method is based on a newly proposed machine learning algorithm called $\propto$SVM. Finding the category-attribute proportions is much easier than manually labeling images, but it is still not a trivial task. We further propose to estimate the proportions from multiple modalities such as human commonsense knowledge, NLP tools, and other domain knowledge. The value of the proposed approach is demonstrated by various applications including modeling animal attributes, visual sentiment attributes, and scene attributes.
Felix X. Yu, Liangliang Cao, Michele Merler, Noel Codella, Tao Chen 0015, John R. Smith, Shih-Fu Chang
ACM Multimedia6
2013 Learning Locally-Adaptive Decision Functions for Person Verification
abstract
This paper considers the person verification problem in modern surveillance and video retrieval systems. The problem is to identify whether a pair of face or human body images is about the same person, even if the person is not seen before. Traditional methods usually look for a distance (or similarity) measure between images (e.g., by metric learning algorithms), and make decisions based on a fixed threshold. We show that this is nevertheless insufficient and sub-optimal for the verification problem. This paper proposes to learn a decision function for verification that can be viewed as a joint model of a distance metric and a locally adaptive thresholding rule. We further formulate the inference on our decision function as a second-order large-margin regularization problem, and provide an efficient algorithm in its dual from. We evaluate our algorithm on both human body verification and face verification problems. Our method outperforms not only the classical metric learning algorithm including LMNN and ITML, but also the state-of-the-art in the computer vision community.
Zhen Li 0028, Shiyu Chang, Feng Liang 0002, Thomas S. Huang, Liangliang Cao, John R. Smith
CVPR6
2013 Designing Category-Level Attributes for Discriminative Visual Recognition
abstract
Attribute-based representation has shown great promises for visual recognition due to its intuitive interpretation and cross-category generalization property. However, human efforts are usually involved in the attribute designing process, making the representation costly to obtain. In this paper, we propose a novel formulation to automatically design discriminative "category-level attributes", which can be efficiently encoded by a compact category-attribute matrix. The formulation allows us to achieve intuitive and critical design criteria (category-separability, learn ability) in a principled way. The designed attributes can be used for tasks of cross-category knowledge transfer, achieving superior performance over well-known attribute dataset Animals with Attributes (AwA) and a large-scale ILSVRC2010 dataset (1.2M images). This approach also leads to state-of-the-art performance on the zero-shot learning task on AwA.
Felix X. Yu, Liangliang Cao, Rogério Feris, John R. Smith, Shih-Fu Chang
CVPR4
2013 Large-scale video event classification using dynamic temporal pyramid matching of visual semantics
abstract
Video event classification and retrieval has recently emerged as a challenging research topic. In addition to the variation in appearance of visual content and the large scale of the collections to be analyzed, this domain presents new and unique challenges in the modeling of the explicit temporal structure and implicit temporal trends of content within the video events. In this study, we present a technique for video event classification that captures temporal information over semantics using a scalable and efficient modeling scheme. An architecture for partitioning videos into a linear temporal pyramid, using segments of equal length and segments determined by the patterns of the underlying data, is applied over a rich underlying semantic description at the frame level using a taxonomy of nearly 1000 concepts containing 500,000 training images. Forward model selection with data bagging is used to prune the space of temporal features and data for efficiency. The system is implemented in the Hadoop Map-Reduce environment for arbitrary scalability. Our method is applied to the TRECVID Multimedia Event Detection 2012 task. Results demonstrate a significant boost in performance of over 50%, in terms of mean average precision, compared to common max or average pooling, and 17.7% compared to more complex pooling strategies that ignore temporal content.
Noel Codella, Gang Hua 0001, Liangliang Cao, Michele Merler, Leiguang Gong, Matthew L. Hill, John R. Smith
ICIP7
2013 Learning by focusing: A new framework for concept recognition and feature selection
abstract
In this paper, we develop a new method for feature selection and category learning. We first introduce two observations from our experiments: (1) It is easier to distinguish two concepts than to learn an isolated concept. (2) To distinguish different concept pairs we can find different selections of optimal features. These two observations may partly explain the success of human vision learning, especially why an infant can simultaneously capture distinguished visual features when learning new concepts. Based on these two observations, we developed a new learning-by-focusing method which first constructs focalized concept discriminators for pairs of concepts, and then builds nonlinear classifiers using the discrimination scores. We build datasets for four concept structure: vehicle, human affliction, sports, and animals, and experiments on all the four datasets verify the success of our new approach.
Liangliang Cao, Leiguang Gong, John R. Kender, Noel Codella, John R. Smith
ICME5
2013 Massive-scale multimedia semantic modeling
abstract
Visual data is exploding! 500 billion consumer photos are taken each year world-wide, 633 million photos taken per year in NYC alone. 120 new video-hours are uploaded on YouTube per minute. The explosion of digital multimedia data is creating a valuable open source for insights. However, the unconstrained nature of 'image/video in the wild' makes it very challenging for automated computer-based analysis. Furthermore, the most interesting content in the multimedia files is often complex in nature reflecting a diversity of human behaviors, scenes, activities and events. To address these challenges, this tutorial will provide a unified overview of the two emerging techniques: Semantic modeling and Massive scale visual recognition, with a goal of both introducing people from different backgrounds to this exciting field and reviewing state of the art research in the new computational era.
John R. Smith, Liangliang Cao
ACM Multimedia1
2013 Riding the multimedia big data wave
abstract
In this talk we present a perspective across multiple industry problems, including safety and security, medical, Web, social and mobile media, and motivate the need for large-scale analysis and retrieval of multimedia data. We describe a multi-layer architecture that incorporates capabilities for audio-visual feature extraction, machine learning and semantic modeling and provides a powerful framework for learning and classifying contents of multimedia data. We discuss the role semantic ontologies for representing audio-visual concepts and relationships, which are essential for training semantic classifiers. We discuss the importance of using faceted classification schemes in particular for organizing multimedia semantic concepts in order to achieve effective learning and retrieval. We also show how training and scoring of multimedia semantics can be implemented on big data distributed computing platforms to address both massive-scale analysis and low-latency processing. We describe multiple efforts at IBM on image and video analysis and retrieval, including IBM Multimedia Analysis and Retrieval System (IMARS), and show recent results for semantic-based classification and retrieval. We conclude with future directions for improving analysis of multimedia through interactive and curriculum-based techniques for multimedia semantics-based learning and retrieval.
John R. Smith
SIGIR1
2013 Tracking Large-Scale Video Remix in Real-World Events
abstract
Content sharing networks, such as YouTube, contain traces of both explicit online interactions (such as likes, comments, or subscriptions), as well as latent interactions (such as quoting, or remixing, parts of a video). We propose visual memes, or frequently re-posted short video segments, for detecting and monitoring such latent video interactions at scale. Visual memes are extracted by scalable detection algorithms that we develop, with high accuracy. We further augment visual memes with text, via a statistical model of latent topics. We model content interactions on YouTube with visual memes, defining several measures of influence and building predictive models for meme popularity. Experiments are carried out with over 2 million video shots from more than 40,000 videos on two prominent news events in 2009: the election in Iran and the swine flu epidemic. In these two events, a high percentage of videos contain remixed content, and it is apparent that traditional news media and citizen journalists have different roles in disseminating remixed content. We perform two quantitative evaluations for annotating visual memes and predicting their popularity. The proposed joint statistical model of visual memes and words outperforms an alternative concurrence model, with an average error of 2% for predicting meme volume and 17% for predicting meme lifespan.
Lexing Xie, Apostol Natsev, Xuming He 0001, John R. Kender, Matthew L. Hill, John R. Smith
IEEE Trans. Multim.6
2012 Scene Aligned Pooling for Complex Video Recognition
Liangliang Cao, Yadong Mu, Apostol Natsev, Shih-Fu Chang, Gang Hua 0001, John R. Smith
ECCV (2)6
2012 Video Event Detection Using Temporal Pyramids of Visual Semantics with Kernel Optimization and Model Subspace Boosting
abstract
In this study, we present a system for video event classification that generates a temporal pyramid of static visual semantics using minimum-value, maximum-value, and average-value aggregation techniques. Kernel optimization and model subspace boosting are then applied to customize the pyramid for each event. SVM models are independently trained for each level in the pyramid using kernel selection according to 3-fold cross-validation. Kernels that both enforce static temporal order and permit temporal alignment are evaluated. Model subspace boosting is used to select the best combination of pyramid levels and aggregation techniques for each event. The NIST TRECVID Multimedia Event Detection (MED) 2011 dataset was used for evaluation. Results demonstrate that kernel optimizations using both temporally static and dynamic kernels together achieves better performance than any one particular method alone. In addition, model sub-space boosting reduces the size of the model by 80%, while maintaining 96% of the performance gain.
Noel Codella, Apostol Natsev, Gang Hua 0001, Matthew L. Hill, Liangliang Cao, Leiguang Gong, John R. Smith
ICME7
2012 Mining Multimedia Data for Meaning - (Extended Abstract)
John R. Smith
MMM1
2011 Tracking Visual Memes in Rich-Media Social Communities
Lexing Xie, Apostol Natsev, John R. Kender, Matthew L. Hill, John R. Smith
ICWSM5
2011 Visual memes in social media: tracking real-world news in YouTube videos
abstract
We propose visual memes, or frequently reposted short video segments, for tracking large-scale video remix in social media. Visual memes are extracted by novel and highly scalable detection algorithms that we develop, with over 96% precision and 80% recall. We monitor real-world events on YouTube, and we model interactions using a graph model over memes, with people and content as nodes, and meme postings as links. This allows us to define several measures of influence. These abstractions, using more than two million video shots from several large-scale event datasets, enable us to quantify and efficiently extract several important observations: over half of the videos contain re-mixed content, which appears rapidly; video view counts, particularly high ones, are poorly correlated with the virality of content; the influence of traditional news media versus citizen journalists varies from event to event; iconic single images of an event are easily extracted; and content that will have long lifespan can be predicted within a day after it first appears. Visual memes can be applied to a number of social media scenarios: brand monitoring, social buzz tracking, ranking content and users, among others.
Lexing Xie, Apostol Natsev, John R. Kender, Matthew L. Hill, John R. Smith
ACM Multimedia5
2010 Design and evaluation of an effective and efficient video copy detection system
abstract
We consider the end-to-end system design and evaluation of an efficient and effective system for video copy detection that bridges the gap between computationally expensive methods and practical applications. We use a compact SIFT-based bag-of-words fingerprint (which we call a SIFTogram), requiring only 1000 bytes per second of video, and show that beyond the descriptor choice, many variables can affect performance. We also consider a complementary color-based descriptor, which contrary to popular recent belief, performs better than SIFTogram on some transforms. We emphasize robustness with respect to the most common transformations on content sharing sites, and report a 99.3% detection rate with 0 false alarms on one such transform category from a standardized evaluation. We perform an evaluation of the system using two TRECVID benchmark datasets, and examine the trade-off between speed and accuracy relative to other TRECVID submissions.
Apostol Natsev, Matthew L. Hill, John R. Smith
ICME3
2010 Video genetics: a case study from YouTube
abstract
We explore in a single but large case study how videos within YouTube, competing for view counts, are like organisms within an ecology, competing for survival. We develop this analogy, whose core idea shows that short video clips, best detected across videos as near-duplicate keyframes, behave similarly to genes. We report work in progress, on a dataset of 5.4K videos with 210K keyframes on a single topic, which traces sequences, not bags, of "near-dups" over time, both within videos and across them. We demonstrate their utility to: cleanse responses to queries contaminated by over-eager YouTube query expansion; separate videos temporally according to their responses to external events; track the evolution and lifespan of continuing video "stories"; automatically locate video summaries already present within a video ecology; quickly verify video copying via a direct application of the Smith-Waterman algorithm used in genetics - which also provides useful feedback for tuning the near-dup detection and clustering process; and quickly classify videos via a kind of Lempel-Ziv encoding into the categories of news, monologue, dialogue, and slideshow. We demonstrate a number of novel visualizations of this large dataset, including a direct use of the Matlab black-body "hot" false-color map, together with the GraphViz package, to display the gene-like inheritance of viral properties of keyframes. We further speculate that, as with genes, there are "functional roles" for semantic categories of clips, and, as with species, there are differing rates of "genetic drift" for each video genre.
John R. Kender, Matthew L. Hill, Apostol Natsev, John R. Smith, Lexing Xie
ACM Multimedia4
2010 Probabilistic visual concept trees
abstract
This paper presents probabilistic visual concept trees, a model for large visual semantic taxonomy structures and its use in visual concept detection. Organizing visual semantic knowledge systematically is one of the key challenges towards large-scale concept detection, and one that is complementary to optimizing visual classification for individual concepts. Semantic concepts have traditionally been treated as isolated nodes, a densely-connected web, or a tree. Our analysis shows that none of these models are sufficient in modeling the typical relationships on a real-world visual taxonomy, and these relationships belong to three broad categories -- semantic, appearance and statistics. We propose probabilistic visual concept trees for modeling a taxonomy forest with observation uncertainty. As a Bayesian network with parameter constraints, this model is flexible enough to account for the key assumptions in all three types of taxonomy relations, yet it is robust enough to accommodate expansion or deletion in a taxonomy. Our evaluation results on a large web image dataset show that the classification accuracy has considerably improved upon baselines without, or with only a subset of concept relationships
Lexing Xie, Jelena Tesic, Apostol Natsev, John R. Smith
ACM Multimedia5
2009 Imbalanced RankBoost for efficiently ranking large-scale image/video collections
abstract
Ranking large scale image and video collections usually expects higher accuracy on top ranked data, while tolerates lower accuracy on bottom ranked ones. In view of this, we propose a rank learning algorithm, called Imbalanced RankBoost, which merges RankBoost and iterative thresholding into a unified loss optimization framework. The proposed approach provides a more efficient ranking process by iteratively identifying a cutoff threshold in each boosting iteration, and automatically truncating ranking feature computation for the data ranked below. Experiments on the TRECVID 2007 high-level feature benchmark show that the proposed approach outperforms RankBoost in terms of both ranking effectiveness and efficiency. It achieves an up to 21% improvement in terms of mean average precision, or equivalently, a 6-fold speedup in the ranking process.
Michele Merler, John R. Smith
CVPR3
2009 The 1st workshop on large-scale multimedia retrieval and mining (LS-MMRM'09)
abstract
This workshop, as the first of its kind, aims to bring together researchers and industrial practitioners interested in large-scale multimedia data retrieval and mining. The workshop will provide a venue for the participants to explore a variety of aspects and applications on how advanced multimedia analysis techniques can be leveraged to address the challenges in large-scale data collections.
John R. Smith, Qi Tian 0001, Rahul Sukthankar
ACM Multimedia2
2009 Evaluating application mapping scenarios on the Cell/B.E
abstract
Abstract Applications running on multicore platforms are difficult to program, and even more difficult to optimize, mainly due to (1) the several layers where the optimizations occur and (2) the multitude of available resources to be exploited in parallel. Although low‐level optimizations only target code running on individual cores, high‐level optimizations (e.g. data‐ and task‐parallelism) target the overall application performance. In this paper, we focus on the latter, by evaluating possible mapping scenarios of a real application on a heterogeneous multicore processor. Specifically, we focus on analyzing the impact of combining data‐ and task‐parallelism for a multimedia analysis application running on the Cell Broadband Engine (Cell/B.E.). We find that both low‐level and high‐level optimizations are important for the overall application speed‐up. However, we show that a speed‐up factor of over 20 for the application running on Cell/B.E. can only be obtained if core utilization is increased by combining data‐ and task‐parallelism. Thus, we consider this case study essential for building expertise in both application optimization and performance analysis for multicore platforms. Copyright © 2008 John Wiley & Sons, Ltd.
Ana Lucia Varbanescu, Henk J. Sips, Kenneth A. Ross, Apostol Natsev, John R. Smith, Lurng-Kuo Liu
Concurr. Comput. Pract. Exp.6
2007 Digital Media Indexing on the Cell Processor
abstract
We present a case study of developing a digital media indexing application, code-named MARVEL, on the STI cell broadband engine (CBE) processor. There are two aspects of the target application that require significant computing power: image analysis for feature extraction, and support vector machine (SVM) based pattern classification for concept detection. We discuss the mapping of a large application like MARVEL onto a multicore processor, and show how feature extraction and concept detection can be implemented on the CBE. We discuss how the synergistic processing units of a CBE can be used to gain dramatic performance improvements. The empirical results of our experiments, conducted on a Cell blade running at 3.2 GHz, show that the CBE provides a significant performance speed-up in our digital media indexing application.
Lurng-Kuo Liu, Apostol Natsev, Kenneth A. Ross, John R. Smith, Ana Lucia Varbanescu
ICME5
2007 Data Modeling Strategies for Imbalanced Learning in Visual Search
abstract
In this paper we examine a novel approach to the difficult problem of querying video databases using visual topics with few examples. Typically with visual topics, the examples are not sufficiently diverse to create a robust model of the user's need. As a result, direct modeling using the provided topic examples as training data is inadequate. Otherwise, systems resort to multiple content-based searches using each example in turn, which typically provides poor results. We propose a new technique of leveraging unlabeled data to expand the diversity of the topic examples as well as provide a robust set of negative examples that allow direct modeling. The approach intelligently models a pseudo-negative space using unbiased and biased methods for data sampling and data selection. We apply the proposed method in a fusion framework to improve discriminative support vector machine modeling, and improve the overall system performance. The result is an enhanced performance over any of the baseline models, as well as improved robustness with respect to training examples, visual features, and visual support of video topics in TRECVID. The proposed method outperforms a baseline retrieval approach by more than 18% on the TRECVID 2006 video collection and query topics.
Jelena Tesic, Apostol Natsev, Lexing Xie, John R. Smith
ICME4
2007 An Effective Strategy for Porting C++ Applications on Cell
abstract
In this paper we present a solution for efficient porting of sequential C++ applications on the Cell B.E. processor. We present our step-by-step approach, focusing on its generality, we provide a set of code templates and optimization guidelines to support the porting, and we include a set of equations to estimate the performance gain of the new application. As a case-study, we show the use of our solution on a multimedia content analysis application, named MARVEL. The results of our experiments with MARVEL prove the significant performance increase in favor of the application running on Cell when compared with the reference implementation.
Ana Lucia Varbanescu, Henk J. Sips, Kenneth A. Ross, Lurng-Kuo Liu, Apostol Natsev, John R. Smith
ICPP7
2007 Model-shared subspace boosting for multi-label classification
abstract
Typical approaches to the multi-label classification problem require learning an independent classifier for every label from all the examples and features. This can become a computational bottleneck for sizeable datasets with a large label space. In this paper, we propose an efficient and effective multi-label learning algorithm called model-shared subspace boosting (MSSBoost) as an attempt to reduce the information redundancy in the learning process. This algorithm automatically finds, shares and combines a number of base models across multiple labels, where each model is learned from random feature subspace and boots trap data samples. The decision functions for each label are jointly estimated and thus a small number of shared subspace models can support the entire label space. Our experimental results on both synthetic data and real multimedia collections have demonstrated that the proposed algorithm can achieve better classification performance than the non-ensemble baselineclassifiers with a significant speedup in the learning and prediction processes. It can also use a smaller number of base models to achieve the same classification performance as its non-model-shared counterpart.
Jelena Tesic, John R. Smith
KDD3
2006 ResGrid: A Grid-aware Toolkit for Reservoir Uncertainty Analysis
abstract
Many efforts in Grid communities have focused on middleware research and development. However, Grid application-level tools are needed which can build higherlevel functionality on top of core middleware services. We work with specific classes of scientific applications and present a Grid-aware toolkit ResGrid for reservoir uncertainty analysis. With the help of the ResGrid, a reservoir engineer can transparently take advantage of Grid resources and services for compute-intensive and dataintensive uncertainty analysis as well as enforce the understanding of reservoir modeling. In this paper, the ResGrid is introduced in terms of overview, architecture, and implementation status.
Zhou Lei 0001, Dayong Huang, Archit Kulshrestha, Santiago Peña, Gabrielle Allen, Christopher D. White, Richard Duff, John R. Smith, Subhash Kalla
CCGRID9
2006 Visual Event Detection using Multi-Dimensional Concept Dynamics
abstract
A novel framework is introduced for visual event detection. Visual events are viewed as stochastic temporal processes in the semantic concept space. In this concept-centered approach to visual event modeling, the dynamic pattern of an event is modeled through the collective evolution patterns of the individual semantic concepts in the course of the visual event. Video clips containing different events are classified by employing information about how well their dynamics in the direction of each semantic concept matches those of a given event. Results indicate that such a data-driven statistical approach is in fact effective in detecting different visual events such as exiting car, riot, and airplane flying
Shahram Ebadollahi, Lexing Xie, Shih-Fu Chang, John R. Smith
ICME4
2006 Semantic Labeling of Multimedia Content Clusters
abstract
In this paper we present a novel approach for labeling clusters of multimedia content that leverages supervised classification techniques in conjunction with unsupervised clustering. Recent research has produced significant results for automatic tagging of video content such as broadcast news. For example, powerful techniques have been demonstrated in the context of the NIST TRECVID video retrieval benchmark [1]. However, the information needs of users typically span a range of semantic concepts. One of the challenges of these multimedia retrieval systems is to organize the video data in such a way that allows the user to most efficiently navigate the semantic space for the video data set. One important tool for video data organization is clustering. However, clustering results cannot be leveraged effectively when they are not labeled. We propose to build on clustering by aggregating the automatically tagged semantics. We propose and compare four techniques for labeling the clusters and evaluate the performance compared to human labeled ground-truth. We present examples of the cluster labeling results obtained on the BBC stock shots from the TRECVID-2005 video data set.
Jelena Tesic, John R. Smith
ICME2
2005 A generalized multiple instance learning algorithm for large scale modeling of multimedia semantics
abstract
Statistical learning techniques provide a robust framework for learning representations of semantic concepts from multimedia features. The bottleneck is the number of training samples needed to construct robust models. This is particularly expensive when the annotation needs to happen at finer granularity. We present a novel approach where the annotations may be entered at coarser spatial granularity while the concept may still be learnt at finer granularity. This can speed up annotation significantly. Using the multiple instance learning paradigm, we show that it is possible to learn representations of concepts occurring at the regional level by using annotations for several images. We present a generalized multiple instance learning algorithm that can scale to a large number of training samples as well as a large number of instances per bag. The algorithm also provides the ability to plug in different density modeling or regression techniques. Using the TREC 2001 Corpus we demonstrate the superior performance of the proposed algorithm over the existing diverse density algorithm.
Milind R. Naphade, John R. Smith
ICASSP (5)2
2005 What is the state of our community?
abstract
10.1145/1101149.1101297
Yong Rui, Ramesh Jain 0001, Nicolas D. Georganas, HongJiang Zhang, Klara Nahrstedt, John R. Smith, Mohan Kankanhalli
ACM Multimedia6
2005 A web-based system for collaborative annotation of large image and video collections: an evaluation and user study
abstract
Annotated collections of images and videos are a necessary basis for the successful development of multimedia retrieval systems. The underlying models of such systems rely heavily on quality and availability of large training collections. The annotation of large collections, however, is a time-consuming and error prone task as it has to be performed by human annotators. In this paper we present the IBM Efficient Video Annotation (EVA) system, a server-based tool for semantic concept annotation of large video and image collections. It is optimised for collaborative annotation and includes features such as workload sharing and support in conducting inter-annotator analysis. We discuss initial results of an ongoing user-evaluation of this system. The results are based on data collected during the 2005 TRECVID Annotation Forum, where more than 100 annotators have been using the system.
Timo Volkmer, John R. Smith, Apostol Natsev
ACM Multimedia2
2005 Special Issue on Image Understanding for Digital Photographs
Jiebo Luo 0001, Thomas S. Huang, John R. Smith, HongJiang Zhang
Pattern Recognit.3
2005 Selected papers from the ACM multimedia conference 2003
abstract
No abstract available.
Thomas Plagemann, Prashant J. Shenoy, John R. Smith
ACM Trans. Multim. Comput. Commun. Appl.3
2004 Multimodal video search techniques: late fusion of speech-based retrieval and visual content-based retrieval
abstract
This paper describes multimodal systems for ad-hoc search constructed by IBM for the TRECVID 2003 benchmark of search systems for broadcast video. These systems all use a late fusion of independently developed speech-based and visual content-based retrieval systems and outperform our individual retrieval systems on both manual and interactive search tasks. For the manual task, our best system used a query-dependent linear weighting between speech-based and image-based retrieval systems. This system has mean average precision (MAP) performance 20% above our best unimodal system for manual search. For the interactive task, where the user has full knowledge of the query topic and the performance of the individual search systems, our best system used an interlacing approach. The user determines the (subjectively) optimal weights A and B for the speech-based and image-based systems, where the multimodal result set is aggregated by combining the top A documents from system A followed by top B documents of system B and then repeating this process until the desired result set size is achieved. This multimodal interactive search has MAP 40% above our best unimodal interactive search system.
Arnon Amir, Giridharan Iyengar, Ching-Yung Lin, Milind R. Naphade, Apostol Natsev, Chalapathy Neti, Harriet J. Nock, John R. Smith, Belle L. Tseng
ICASSP (3)8
2004 Over-complete representation and fusion for semantic concept detection
abstract
Automatic semantic concept detection in images is a promising tool for alleviating the user effort in annotating and cataloging digital media collections. It enables automatic identification of people, places and objects, for enhanced indexing and searching of home photographs, for example. While constructing robust semantic detectors has been shown feasible for global generic concepts with a sufficient number of good training examples (e.g., indoors, outdoors), many interesting concepts, such as face, people, occur at subpicture granularity, occupy only a portion of the image and therefore frequently have training examples with a reduced signal-to-noise ratio. Such regional concepts are harder to detect due to imperfections in automatic image segmentation algorithms leading to inaccurate object boundaries and low-level feature ambiguities. In this paper we focus on the problem of boosting detection performance of existing regional concept detectors by exploiting detection redundancy. Specifically, we propose to use the same detector multiple times to evaluate and combine multiple detection hypotheses for the same content-but at different content granularities-in order to reduce detection sensitivity to segmentation errors. We validate the approach using support vector machine classifiers for 14 regional semantic concepts from the NISTTRFCVID 2003 common annotation lexicon and show performance improvements of multigranular detection and fusion.
Apostol Natsev, Milind R. Naphade, John R. Smith
ICIP3
2004 Multimodal information fusion for video concept detection
abstract
Video media carries multimodal information including visual, audio, textual data. Considerable research has been focused on utilizing multimodal features for better understanding of video content. However, many problems remain such as how to combine multimodal features and what are the effects of different combinations. In this paper, we propose to find the optimal combination of multimodal information in order to improve the performance of video concept detection using two methods, one is gradient-descent-optimization linear fusion and the other is super-kernel nonlinear fusion. Gradient-descent-optimization linear fusion learns an optimal weighted linear combination of single modalities based on fusing individual kernel matrices with gradient descent techniques. Super-kernel nonlinear fusion trains separate classifiers for single modalities as the first step. Once individual models have been designed, super-kernel nonlinear fusion learns an optimal nonlinear combination of individual models by fusing single-modality classifiers. Our experiments show that both methods improve performance significantly on TREC-Video 2003 benchmarks.
Yi Wu 0005, Ching-Yung Lin, Edward Y. Chang, John R. Smith
ICIP4
2004 Content transcoding middleware for pervasive geospatial intelligence access
abstract
We describe a novel content transcoding middleware for accessing military geospatial intelligence information in real-time. Intelligence information, including maps and location, category and properties of object targets, is adapted for various pervasive devices such as laptop, personal digital assistant (PDA), cellular phone, etc. The middleware is deployed as proxies on the Web using the IBM Websphere Transcoding Publisher (WTP) platform, which facilitates the middleware management. We developed several Java-based plug-ins and Extensible Stylesheet Language (XSL) stylesheets for content transcoding. A prototype has been established and real experiments have demonstrated the effectiveness of this novel middleware.
Ching-Yung Lin, Apostol Natsev, Belle L. Tseng, Matthew L. Hill, John R. Smith, Chung-Sheng Li
ICME5
2004 Multi-granular detection of regional semantic concepts
abstract
A large number of interesting visual semantic concepts occur at a sub-frame granularity in images and occupy one or more regions at the sub-frame level. Detecting these concepts is a challenge due to segmentation imperfections. We propose multi-granular detection of visual concepts that have regional support. We build a single set of support vector machine based binary concept models from the training set with manually marked up regions. In this paper, we show that detection can be significantly improved by scoring these models over multiple granularities in the test set images, where the regions are automatically detected as a preprocessing step in detection. Using 27 regional semantic concepts from the NIST TRECVID 2003 common annotation lexicon and the corpus, we demonstrate that multi-granular detection leads to improvement in detection.
Milind R. Naphade, Apostol Natsev, Ching-Yung Lin, John R. Smith
ICME4
2004 Active learning for simultaneous annotation of multiple binary semantic concepts
abstract
A model-based approach to video analysis requires annotated corpora. Video annotation, however is a very expensive process. Tools that allow users to annotate video shots with scenes, events, and objects should minimize user interaction. These tools should particularly leverage redundancy in content and advances in machine learning and human computer intelligence to reduce the amount of human interaction needed to annotate large corpora. As corpora sizes and the lexicon grows, this is increasingly relevant. Active learning can play a critical role in reducing the amount of supervision. We apply active learning to the simultaneous annotation of multiple binary concepts. The challenge is to minimize the total number of samples to be annotated across all concepts. Preliminary experiments with the simultaneous annotation of two concepts outdoors and indoors using the TRECVID corpus are promising and reduce annotation workload significantly.
Milind R. Naphade, John R. Smith
ICME2
2004 Ontology-based multi-classification learning for video concept detection
abstract
In this paper, an ontology-based multi-classification learning algorithm is adopted to detect concepts in the NIST TREC-2003 video retrieval benchmark which defines 133 video concepts, organized hierarchically and each video data can belong to one or more concepts. The algorithm consists of two steps. In the first step, each single concept model is constructed independently. In the second step, ontology-based concept learning improves the accuracy of the individual concept by considering the possible influence relations between concepts based on a predefined ontology hierarchy. The advantage of ontology learning is that its influence path is based on an ontology hierarchy, which has real semantic meanings. Besides semantics, it also considers the data correlation to decide the exact influence assigned to each path, which makes the influence more flexible according to data distribution. This learning algorithm can be used for multiple topic document classification such as Internet documents and video documents. We demonstrate that precision-recall can be significantly improved by taking ontology into account
Yi Wu 0005, Belle L. Tseng, John R. Smith
ICME3
2004 Semantic video clustering across sources using bipartite spectral clustering
abstract
Data clustering is an important technique for visual data management. Most previous work focuses on clustering video data within single sources. We address the problem of clustering across sources, and propose novel spectral clustering algorithms for multisource clustering problems. Spectral clustering is a new discriminative method realizing clustering by partitioning data graphs. We represent multi-source data as bipartite or K-partite graphs, and investigate the spectral clustering algorithm under these representations. The algorithms are evaluated using the TRECVID-2003 corpus with semantic features extracted from speech transcripts and visual concept recognition results from videos. The experiments show that the proposed bipartite clustering algorithm significantly outperforms the regular spectral clustering algorithm in capturing cross-source associations.
DongQing Zhang, Ching-Yung Lin, Shih-Fu Chang, John R. Smith
ICME4
2004 Semantic representation: search and mining of multimedia content
abstract
Semantic understanding of multimedia content is critical in enabling effective access to all forms of digital media data. By making large media repositories searchable, semantic content descriptions greatly enhance the value of such data. Automatic semantic understanding is a very challenging problem and most media databases resort to describing content in terms of low-level features or using manually ascribed annotations. Recent techniques focus on detecting semantic concepts in video, such as indoor, outdoor, face, people, nature, etc. This approach works for a fixed lexicon for which annotated training examples exist. In this paper we consider the problem of using such semantic concept detection to map the video clips into semantic spaces. This is done by constructing a model vector that acts as a compact semantic representation of the underlying content. We then present experiments in the semantic spaces leveraging such information for enhanced semantic retrieval, classification, visualization, and data mining purposes. We evaluate these ideas using a large video corpus and demonstrate significant performance gains in retrieval effectiveness.
Apostol Natsev, Milind R. Naphade, John R. Smith
KDD3
2004 On the detection of semantic concepts at TRECVID
abstract
Semantic multimedia management is necessary for the effective and widespread utilization of multimedia repositories and realizing the potential that lies untapped in the rich multimodal information content. This challenge has driven researchers to devise new algorithms and systems that enable automatic or semi-automatic tagging of large scale multimedia content with rich semantics. An emerging research area is the detection of a predetermined set of semantic concepts that can act as semantic filters and aid in search, and manipulation. The NIST TRECVID benchmark has responded by creating a task that has evaluated the performance of concept detection. Within the scope of this benchmark task, this paper studies trends in the emerging concept detection systems, architectures and algorithms. It also analyzes strategies that have yielded reasonable success, and challenges and gaps that lie ahead.
Milind R. Naphade, John R. Smith
ACM Multimedia2
2004 Optimal multimodal fusion for multimedia data analysis
abstract
Considerable research has been devoted to utilizing multimodal features for better understanding multimedia data. However, two core research issues have not yet been adequately addressed. First, given a set of features extracted from multiple media sources (e.g., extracted from the visual, audio, and caption track of videos), how do we determine the best modalities? Second, once a set of modal-ities has been identified, how do we best fuse them to map to se-mantics? In this paper, we propose a two-step approach. The first step finds statistically independent modalities from raw features. In the second step, we use super-kernel fusion to determine the optimal combination of individual modalities. We carefully ana-lyze the tradeoffs between three design factors that affect fusion performance: modality independence, curse of dimensionality, and fusion-model complexity. Through analytical and empirical studies, we demonstrate that our two-step approach, which achieves a care-ful balance of the three design factors, can improve class-prediction accuracy over traditional techniques.
Yi Wu 0005, Edward Y. Chang, Kevin Chen-Chuan Chang, John R. Smith
ACM Multimedia4
2004 A multi-modal system for the retrieval of semantic video events
Arnon Amir, Sankar Basu, Giridharan Iyengar, Ching-Yung Lin, Milind R. Naphade, John R. Smith, Savitha Srinivasan, Belle L. Tseng
Comput. Vis. Image Underst.6
2004 Multimedia database management systems
John R. Smith, Tong Zhang 0007, Shih-Fu Chang
J. Vis. Commun. Image Represent.2
2004 Video personalization and summarization system for usage environment
Belle L. Tseng, Ching-Yung Lin, John R. Smith
J. Vis. Commun. Image Represent.3
2004 Making the Threshold Algorithm Access Cost Aware
abstract
Assume a database storing N objects with d numerical attributes or feature values. All objects in the database can be assigned an overall score that is derived from their single feature values (and the feature values of a user-defined query). The problem considered here is then to efficiently retrieve the k objects with minimum (or maximum) overall score. The well-known threshold algorithm (TA) was proposed as a solution to this problem. TA views the database as a set of d sorted lists storing the feature values. Even though TA is optimal with regard to the number of accesses, its overall access cost can be high since, in practice, some list accesses may be more expensive than others. We therefore propose to make TA access cost aware by choosing the next list to access such that the overall cost is minimized. Our experimental results show that this overall cost is close to the optimal cost and significantly lower than the cost of prior approaches.
Christian A. Lang, Yuan-Chi Chang, John R. Smith
IEEE Trans. Knowl. Data Eng.3
2004 A Wavelet Framework for Adapting Data Cube Views for OLAP
abstract
This article presents a method for adaptively representing multidimensional data cubes using wavelet view elements in order to more efficiently support data analysis and querying involving aggregations. The proposed method decomposes the data cubes into an indexed hierarchy of wavelet view elements. The view elements differ from traditional data cube cells in that they correspond to partial and residual aggregations of the data cube. The view elements provide highly granular building blocks for synthesizing the aggregated and range-aggregated views of the data cubes. We propose a strategy for selectively materializing alternative sets of view elements based on the patterns of access of views. We present a fast and optimal algorithm for selecting a non-expansive set of wavelet view elements that minimizes the average processing cost for supporting a population of queries of data cube views. We also present a greedy algorithm for allowing the selective materialization of a redundant set of view element sets which, for measured increases in storage capacity, further reduces processing costs. Experiments and analytic results show that the wavelet view element framework performs better in terms of lower processing and storage cost than previous methods that materialize and store redundant views for online analytical processing (OLAP).
John R. Smith, Chung-Sheng Li, Anant Jhingran
IEEE Trans. Knowl. Data Eng.1
2003 VideoAL: a novel end-to-end MPEG-7 video automatic labeling system
abstract
In this paper, we describe a novel end-to-end video automatic labeling system, which accepts MPEG-I sequence inputs and generates MPEG-7 XML metadata files based on the prior established anchor models. Seven modules were developed for the system: shot segmentation, region segmentation, annotation, feature extraction, model learning, classification, and XML rendering. The performance of this system has been tested in the NIST TREC-2002 video concept detection benchmark. The proposed system performs best in the mean average precision out of 18 worldwide participants.
Ching-Yung Lin, Belle L. Tseng, Milind R. Naphade, Apostol Natsev, John R. Smith
ICIP (3)5
2003 Exploring semantic dependencies for scalable concept detection
abstract
Semantic concept detection from multimedia features enables high-level access to multimedia content. While constructing robust detectors is feasible for concepts with sufficient training samples, concepts with fewer training samples are hard to train efficiently. Comparable performance may be possible if the dependence of these concepts on the ones that can be robustly modeled is exploited. In this paper we show this phenomenon using the TREC Video 2002 Corpus as a test bed. Using a basic set of 12 semantic concepts modeled with support vector machines, we predict presence of 4 other concepts. We then compare the performance of these predictors with direct SVM models for these 4 concepts and observe improvements of up to 150% in average precision.
Milind R. Naphade, Apostol Natsev, John R. Smith
ICIP (3)3
2003 Learning visual models of semantic concepts
abstract
Statistical machine learning provides a computational framework for mapping low level media features to high level semantics concepts. In this paper we expose the challenges that these techniques face. Using support vector machine (SVM) classification we build models for 34 semantic concepts for the TREC 2002 benchmark corpus. We study the effect of number of examples available for training with respect to their impact on detection. We also examine low level feature fusion as well as parameter sensitivity with SVM classifiers.
Milind R. Naphade, John R. Smith
ICIP (2)2
2003 Learning regional semantic concepts from incomplete annotation
abstract
For multimedia retrieval to be effective, the semantic gap needs to be bridged. Statistical learning techniques provide a robust framework for learning representations of semantic concepts from visual features. The bottleneck is the need to annotate a large number of training samples to construct robust models. We present a novel approach where the annotations may be entered at coarser spatial granularity while the concept may still be learnt at finer granularity. This can speed up annotation significantly and provide bootstrapping. We show that it is possible to learn representations of concepts occurring at the regional level by using annotations for several images, where the annotations are provided only at the global level. The disambiguation can be handled by the multiple instance learning paradigm. We demonstrate this using the TREC 2001 corpus for the concept sky.
Milind R. Naphade, John R. Smith
ICIP (2)2
2003 Interactive search fusion methods for video database retrieval
abstract
In this paper, we investigate a new method for video database retrieval using interactive search fusion. Recent video analysis techniques have enabled the extraction of a variety of descriptors of features, concepts, clusters, classification results, speech and textual terms, MPEG-7 metadata, and so on. However, given an information need users are faced with a daunting task of trying to formulate queries over these multiple disparate data sources in order to retrieve the desired video content. In this paper, we explore a novel approach based on search fusion in which the user interactively builds a query by sequentially choosing among the descriptors and data sources and by selecting from various combining and score aggregation functions to fuse results of individual searches. For example, the system allows building of queries such as "retrieve video clips that have color of beach scenes, the detection of sky, and detection of water". In this paper we present the search fusion method and evaluate the performance on a large video database.
John R. Smith, Alejandro Jaimes, Ching-Yung Lin, Milind R. Naphade, Apostol Natsev, Belle L. Tseng
ICIP (1)1
2003 Normalized classifier fusion for semantic visual concept detection
abstract
In this paper, we describe our classifier fusion framework for the visual concept detections of NIST TREC-2002 video retrieval benchmark. A normalized ensemble fusion is explored to improve overall performance by incorporating normalization of confidence scores, aggregation via combiner function, and an optimize selection. The normalized classifier fusion shows significant detection improvements for our visual concepts.
Belle L. Tseng, Ching-Yung Lin, Milind R. Naphade, Apostol Natsev, John R. Smith
ICIP (2)5
2003 Semi-automatic, data-driven construction of multimedia ontologies
abstract
In this paper we investigate semi-automatic construction of multimedia ontologies using a data-driven approach. We start with a collection of videos for which we wish to build an ontology (an explicit specification of a domain). Each video is pre-processed: scene cut detection, automatic speech recognition (ASR), and metadata extraction are performed. In addition we automatically index the videos based on visual content by extracting syntactic (e.g., color, texture, etc.) and semantic features (e.g., face, landscape, etc.). We then combine standard tools for ontology engineering and tools in content-based retrieval to semi-automatically build ontologies. In the first stage we process the text information available with the videos (ASR, metadata, and annotations, if any). Stop words (e.g., a, on, the) are eliminated and statistics (e.g., frequency, TFIDF, and entropy) are computed for all terms. Based on this data we manually select concepts and relationships to include in the ontology. Then we use content-based retrieval tools to assign multimedia entities (e.g., shots, videos, collections of videos) to concepts, properties, or relationships in the ontology, and to select multimedia entities as concepts, relationships, or properties in the ontology. We explore this methodology to construct multimedia ontologies from 24 hours of educational films from the 1940s-1960s used in the TREC video retrieval benchmark and discuss the problems encountered and future directions.
Alejandro Jaimes, John R. Smith
ICME2
2003 Epi-SPIRE: a system for environmental and public health activity monitoring
abstract
Health activity monitoring (HAM) has received increasing attention due to the rapid advances of both hardware and software technologies and strong environmental and public health needs. In this paper, we describe the architecture and implementation of the Epi-SPIRE prototype, which is a novel health activity monitoring system that generates alerts from environmental, behavioral, and public health data sources. A model-based approach is used to develop disease and behavior models from multi-modal heterogeneous data sources. Furthermore, a model-based indexing technique has been developed to speed up the data access and retrieval. This system has been successfully applied to various genuine and simulated diseases outbreaks scenarios'.
Chung-Sheng Li, Charu C. Aggarwal, Murray Campbell, Yuan-Chi Chang, Gregory Glass, Vijay S. Iyengar, Mahesh Joshi, Ching-Yung Lin, Milind R. Naphade, John R. Smith, Belle L. Tseng, Min Wang 0001, Kun-Lung Wu, Philip S. Yu
ICME10
2003 A framework for moderate vocabulary semantic visual concept detection
abstract
Extraction of semantic features from visual concepts is essential for meaningful content management in terms of filtering, searching and retrieval. Recently, machine learning techniques have been shown to provide a computational framework to map low level features to high level semantics. In this paper we expose these techniques to the challenge of supporting a moderately large lexicon of semantic concepts. Using the TREC 2002 benchmark corpus for training and validation we investigate a support vector machine based learning system for modeling 34 visual concepts. The detection results show excellent performance for a set of concepts with moderately large training samples. Promising performance is also observed for concepts with few training concepts.
Milind R. Naphade, Ching-Yung Lin, Apostol Natsev, Belle L. Tseng, John R. Smith
ICME5
2003 Active selection for multi-example querying by content
abstract
Multi-example content-based retrieval (MECBR) is the process of querying content by specifying multiple query examples with single query iteration. MECBR attempts to mitigate some of the semantic limitations of traditional relevance feedback or CBR techniques by allowing multiple query examples and thus a more accurate modeling of the user's query need. It also attempts to minimize the burden on the user, as compared to relevance feedback methods, by eliminating the need for user feedback and limiting all interaction into a single query specification step. Multi-example content-based retrieval is therefore a simple alternative for modeling low- and mid-level semantics without the need for heavy user interaction or extensive training, as in interactive feedback systems or complex statistical modeling approaches. In this paper, we describe the MECBR technique in some detail and study methods for active selection of query examples and query features. In particular, we propose and investigate techniques for automatic query example selection, feature selection, and feature fusion. We compare different approaches and evaluate performance of different parameter settings through an extensive empirical study. We also compare MECBR performance to that of explicitly built semantic models using state-of-the-art support vector machines (SVM). We find that lightweight MECBR performs up to 60% better for rare concepts and only 12% to 25% worse for frequent concepts, as compared to heavy-weight SVM modeling! This shows that MECBR is not only a viable lightweight alternative to statistical semantic modeling but is also preferred for very diverse or rare-class semantic modeling situations.
Apostol Natsev, John R. Smith
ICME2
2003 Multimedia semantic indexing using model vectors
abstract
In this paper we propose a novel method for multimedia semantic indexing using model vectors. Model vectors provide a semantic signature for multimedia documents by capturing the detection of concepts broadly across a lexicon using a set of independent binary classifiers. While recent techniques have been developed for detecting simple generic concepts such as indoors, outdoors, nature, manmade, faces, people, speech, music, and so forth [W.H. Adams et al., November 2002], these labels directly support only a small number of queries. Model vectors address the problem of answering queries for which relationships to specific concepts is either unknown or indirect by developing a basis across across the lexicon. In the simplest case, each model vector dimension corresponds to the confidence score by which a corresponding concept from the lexicon is detected. However, we show how other information such as relevance, reliability and concept correlation can also be incorporated. Overall, the model vectors can be used in a variety of methods for multimedia indexing, including model-based retrieval, relevance feedback searching and concept querying. In this paper, we present the model vector method and study different strategies for computing and comparing model vectors. We empirically evaluate the retrieval effectiveness of the model vector approach compared to other search methods in a large video retrieval testbed.
John R. Smith, Milind R. Naphade, Apostol Natsev
ICME1
2003 Improved text overlay detection in videos using a fusion-based classifier
abstract
In this paper, classifier fusion is adopted to demonstrate improved performance for our text overlay detections in the NIST TREC-2002 video retrieval benchmark. A normalized ensemble fusion is explored to combine two text overlay detection models. The fusion incorporates normalization of confidence scores, aggregation via combiner function, and an optimize selection. The proposed fusion classifier resulted best out of 11 detectors submitted to the NIST text overlay detection benchmarking and its average precision performance is 227% of the second best detector in the benchmark.
Belle L. Tseng, Ching-Yung Lin, DongQing Zhang, John R. Smith
ICME4
2003 MPEG-7 video automatic labeling system
abstract
In this demo, we show a novel end-to-end video automatic labeling system, which accepts MPEG-1 sequence inputs and generates MPEG-7 XML metadata files. Detections are based on the prior established anchor models. This system has two parts: model training process and labeling process. They are comprised of seven modules: Shot Segmentation, Region Segmentation, Annotation, Feature Extraction, Model Learning, Classification, and XML Rendering.
Ching-Yung Lin, Belle L. Tseng, Milind R. Naphade, Apostol Natsev, John R. Smith
ACM Multimedia5
2003 Searching dynamically bundled goods with pairwise relations
abstract
Economics research has long recognized that bundling enables savings in production and transaction costs, promotes complementary among the bundle components and sorts consumers according to their valuations. Sellers employ market analysis and intelligence to extract the most surplus. In the age of electronic commerce with low product information access cost, buyers can take advantage of the benefits of bundling by performing dynamic composition of goods from multiple companies offering heterogeneous products and services. These goods, with the proper mix of sources and quantity, may offer additional discounts and benefits, which would not have risen should purchase decisions were made independently. A prominent example is packaged travel, which often involves air, hotel and car rentals. An optimal travel package search not only takes advantage of the lowest available prices of air, hotel and car rental individually but also exploits various discounts through business partnerships between service providers.Today's database infrastructure to support the search of dynamically bundled goods, however, is insufficient. The complex search operations involving cross join of many product categories with hundreds or thousands of offerings can be formulated as SQL queries. But executing these queries in a traditional database is inefficient. This paper proposes an I/O conscious, dynamic programming based algorithm for bundle search. The proposed algorithm finds the top-K combinations of goods abstracted by a linear relationship graph. Experimental results indicate that the proposed algorithm achieves more than two orders of magnitude speedup over cross join, and it is more than an order of-magnitude faster than the simple dynamic programming solution. The performance gap further widens as the number of product categories and the number of offerings within each category increase. This paper characterizes the computational and I/O complexity of the proposed algorithm and suggests extensions to search bundles with more complex relationships.
Yuan-Chi Chang, Chung-Sheng Li, John R. Smith
EC3
2003 User-trainable video annotation using multimodal cues
abstract
This paper describes progress towards a general framework for incorporating multimodal cues into a trainable system for automatically annotating user-defined semantic concepts in broadcast video. Models of arbitrary concepts are constructed by building classifiers in a score space defined by a pre-deployed set of multimodal models. Results show annotation for user-defined concepts both in and outside the pre-deployed set is competitive with our best video-only models on the TREC Video 2002 corpus. An interesting side result shows speech-only models give performance comparable to our best video-only models for detecting visual concepts such as "outdoors", "face" and "cityscape".
Ching-Yung Lin, Milind R. Naphade, Apostol Natsev, Chalapathy Neti, John R. Smith, Belle L. Tseng, Harriet J. Nock, W. H. Adams
SIGIR5
2003 Introduction to the special issue on conceptual and dynamical aspects of multimedia content description
Ali J. Tabatabai, Sethuraman Panchanathan, John R. Smith, Hiroshi Yasuda, Siegfried Handschuh
IEEE Trans. Circuits Syst. Video Technol.3
2002 A statistical modeling approach to content based retrieval
abstract
Statistical modeling for content based retrieval is examined in the context of recent TREC Video benchmark exercise. The TREC Video exercise can be viewed as a test bed for evaluation and comparison of a variety of different algorithms on a set of high-level queries for multimedia retrieval. We report on the use of techniques adopted from statistical learning theory. Our method, as in most statistical methods, depend on training of models based on large data sets. A plethora of statistical models such the Gaussian mixture models, support vector machines etc. can be thought of, only a few of which are exploited in this preliminary report. Training requires a large amount of annotated (labeled) data. Thus, we explore use of active learning for the annotation engine that minimizes the number of training samples to be labeled for satisfactory performance.
Sankar Basu, Milind R. Naphade, John R. Smith
ICASSP3
2002 Modeling semantic concepts to support query by keywords in video
abstract
Supporting semantic queries is a challenging problem in video retrieval. We propose the use of a lexicon of semantic concepts for handling the queries. We also propose automatic modeling of lexicon items using probabilistic techniques. We use Gaussian mixture models to build computational representations for a variety of semantic concepts including rocket-launch, outdoor greenery, sky etc. Training requires a large amount of annotated (labeled) data. Using the TREC Video test bed we compare the performance of this system supporting query by keywords with the conventional approach of query by example. Results demonstrate significant gains in performance using the automatically learnt models of semantic concepts.
Milind R. Naphade, Sankar Basu, John R. Smith, Ching-Yung Lin, Belle L. Tseng
ICIP (1)3
2002 Video texture indexing using spatio-temporal wavelets
abstract
We present a new compact spatio-temporal texture descriptor designed for indexing dynamic video content. Video texture provides a way to characterize spatio-temporal features such as those corresponding to splashing water, flying birds, blowing trees, and so forth, which are not easily characterized by static feature descriptors. The video texture descriptor measures the 3D wavelet energy corresponding to spatio-temporal-frequency subbands of video segments. The wavelet energy captures the texture patterns as they unfold simultaneously along the spatial and temporal dimensions. We describe the process for extracting video texture features and investigate different methods of constructing video texture descriptors from 3D wavelet energy. We evaluate the video texture descriptors in retrieval experiments and show performance improvements compared to methods based on traditional texture descriptors.
Milind R. Naphade, Ching-Yung Lin, John R. Smith
ICIP (2)3
2002 Interactive content-based retrieval of video
abstract
We describe a system for content-based retrieval of video that involves a series of query interactions with the user. The proposed approach allows the user, iteratively and selectively, to integrate different feature- and model-based methods of querying in the search process. This allows the user to choose among different retrieved content, features and matching dimensions, and classifiers, as appropriate, given the query objective and interim retrieval results. We investigate several approaches for integrating featureand model-based queries and results in successive query rounds including iterative filtering, score aggregation, and relevance feedback searching. We describe experimental results of applying the interactive content-based retrieval method to an automatically indexed corpus of 11 hours of video.
John R. Smith, Sankar Basu, Ching-Yung Lin, Milind R. Naphade, Belle L. Tseng
ICIP (1)1
2002 Universal MPEG content access using compressed-domain system stream editing techniques
abstract
An MPEG system layer compressed-domain editing technique is proposed to facilitate the delivery and integration of multiple segments of MPEG files, residing on remote databases. Various multimedia applications, including retrieval and summarization, split MPEG files into small segments along shot boundaries and store them separately. This traditional method requires extra management and storage payload, provides only fixed segmentations, and may not be play smoothly. In order to solve this problem, our MPEG system-domain editing tool directly extracts video-audio information from the original MPEG sources and combines them to generate a single MPEG file. Manipulated wholly in the system bitstream domain, this method does not require decoding, re-encoding, and re-synchronization of audio and video data. Thus, it operates in real-time and provides great flexibility. This composite MPEG file can be transmitted and displayed through general Web interfaces. The proposed method is applied to our video retrieval, video summarization, and video editing systems, and has shown its great advantages.
Ching-Yung Lin, Belle L. Tseng, John R. Smith
ICME (2)3
2002 Learning semantic multimedia representations from a small set of examples
abstract
We approach the problem of semantic multimedia retrieval as a supervised learning problem. Defining a lexicon of a small number of interesting semantic concepts we can handle a number of semantic queries. Since the number of interesting concepts available for training is usually small we explore discriminant learning techniques. In particular, we examine the use of kernel based methods and demonstrate impressive retrieval performance using semantic concepts like rocket, outdoor, greenery, sky and face. We also show that loosely coupled multimodal events can be detected based on the late fusion of detection of related auditory and visual concepts. Using a Bayesian network for inference we show how a rocket-launch event can be detected based on the detection of a related visual concept (rocket object) and a related auditory concept (explosion/blast-off).
Milind R. Naphade, Ching-Yung Lin, John R. Smith
ICME (2)3
2002 A study of image retrieval by anchoring
abstract
Anchoring is a technique for representing objects by their distances to a few well chosen anchors, or vantage points. It can be used in content-based image retrieval for computing image similarity as a function of distances to a fixed set of representative images. Since the number of anchors is usually small, this leads to a reduced dimensionality for similarity searching, enables efficient indexing, and avoids potentially expensive similarity computations in the original feature domain, while guaranteeing lack of false dismissals. Anchoring is therefore surprisingly simple, yet effective, and flavors of it have seen application in speech recognition, audio classification, protein homology detection, and shape matching. In this paper, we describe the anchoring technique in some detail and study its properties, both from an empirical and an analytical standpoint. In particular, we investigate issues in baseline distance selection, anchor selection, and number of anchors. We compare different approaches and evaluate performance of different parameter settings. We also propose two new anchor selection heuristics which may overcome some of the drawbacks of the currently used greedy selection methods.
Apostol Natsev, John R. Smith
ICME (2)2
2002 Spatial and feature normalization for content-based retrieval
abstract
We explore methods for spatial and feature normalization of visual descriptors for content-based retrieval (CBR). A great many descriptors have been developed for characterizing features such as color, texture, edges, and so forth. In addition, numerous methods have also been proposed for extracting descriptors from whole images or regions. Furthermore, different options are possible for normalizing descriptor values for matching. We study different spatial and feature normalization strategies that include extracting descriptors from different spatial partitionings and normalizing descriptor values based on metric-space considerations or statistics of image collections. We empirically evaluate the relative efficacy in an image retrieval testbed.
John R. Smith, Apostol Natsev
ICME (1)1
2002 Real-time video surveillance for traffic monitoring using virtual line analysis
abstract
A real-time video surveillance is presented for traffic monitoring of vehicle volume on major highways. Determining traffic volume automatically and in real-time assists drivers to dynamically plan their trips more efficiently. Our traffic monitoring system uses the virtual line graph to facilitate the detection of vehicles, classification of vehicle types, tracking of individual vehicles, and subsequently an accurate count of the number of vehicles. The virtual line analyzer detects vehicles as they cross a virtual boundary. The goal of this traffic monitoring system is to provide a real-time and accurate vehicle counter while taking advantage of stationary Web-cams, fixed highways and lanes, and deterministic vehicle characteristics.
Belle L. Tseng, Ching-Yung Lin, John R. Smith
ICME (2)3
2001 MPEG-7 MDS Content Description Tools and Applications
Ana B. Benitez, Di Zhong, Shih-Fu Chang, John R. Smith
CAIP4
2001 Solarspire: querying temporal solar imagery by content
abstract
In this paper, we describe a novel content-based retrieval application which permits astrophysicists to search large image sequence archives for solar phenomenon, such as solar flares, based on the spatio-temporal behavior of the solar phenomenon. Specifically, images are preprocessed to identify bright and dark spots based on their relative intensity with respect to their neighboring regions. Temporally persistent objects are then extracted from the collection of spots, and their spatio-temporal behavior represented as intensity and size time series. Users define a query in terms of a model of spatio-temporal behaviors through a Web-based interface. The stored intensity and size time series are searched, and series segments that match the specified specified spatio-temporal behavior are returned. The benchmark results based on 2500 satellite images show that the proposed methodology demonstrated better than 85% accuracy on a solar phenomenon previously identified by astrophysicists.
Matthew L. Hill, Vittorio Castelli, Chung-Sheng Li, Yuan-Chi Chang, Lawrence D. Bergman, John R. Smith, Barbara J. Thompson
ICIP (1)6
2001 Multi-object multi-feature content based search using MPEG-7
abstract
We describe methods for content-based searching of images using MPEG-7 descriptions. The search problems range from matching of images based on global features to matching based on multiple objects, multiple features, and structural or semantic constraints and relationships. We provide a taxonomy of the different searching and matching problems and present query methods for each type. Furthermore, we examine methods for computing approximate answers for some of the searching problems in order to allow a trade-off of query response time and precision.
John R. Smith, Yuan-Chi Chang, Chung-Sheng Li
ICIP (3)1
2001 An e-Marketplace Infrastructure for Model-Based Matchmaking between Consumers and Providers of Multimodal Earth Science Data
abstract
As the earth science data and information products begin to proliferate due to the increased number of earth observing instruments and platforms, it has become increasingly difficult for the end consumer to leverage the wide variety of available earth science data and information products. In this paper, we propose an innovative infrastructure to enable the consumers to locate and tradeoff possible alternative earth science data and information sources in an electronic marketplace setting. Specifically, this architecture provides mechanisms to annotate the requests and offerings of the data and information products, to decompose the concepts of the requests and offerings to facilitate the matchmaking and inferencing. Based on the knowledge models developed for each application domain and science discipline, the matchmaking mechanism will be able to fuse and combine multiple alternative data and information sources so that the quality of the results can be maximized while the cost for data acquisition is minimizing.
Chung-Sheng Li, Yuan-Chi Chang, John R. Smith
ICME3
2001 IMKA: a multimedia organization system combining perceptual and semantic knowledge
abstract
In the demo, we present the IMKA system, which implements the innovative approach of integrating perceptual information such as low-level features and images, and symbolic information such as words to represent the knowlege associated with a large multimedia collection for multimedia organization and retrieval. The IMKA system utilizes the unique MediaNet framework, which greatly extends existing knowlege representation tools in the text domain (e.g., semantic networks and WordNet) and the multimedia domain (e.g. Multimedia Thesaurus) by combining perceptual and semantic concepts in the same network and by supporting perceptual and semantic relationships among concepts exemplified by different media. It also brings the level of multimedia retrieval closer to users' needs by translating low-level feature queries to high-level semantic queries and vice versa. We will demonstrate the process of constructing the MediaNet knowledge base and new ways of searching multimedia in the IMKA system by presenting the current implementation of the IMKA system that uses image collections from online sources.
Ana B. Benitez, Shih-Fu Chang, John R. Smith
ACM Multimedia3
2001 CPU/power-constrained mobile devices
abstract
Due to the limited processing capability, memory constraints, and the power budget of mobile clients, multimedia coders and/or decoders are often difficult to implement on wireless handheld PDAs. In this Universal Tuner project, we designed and implemented a wireless video streaming system that transcodes MPEG-1/2 videos or live TV broadcasting videos to the BW or indexed color Palm OS devices. In our system, the complexity of multimedia compression and decompression algorithms is adaptively partitioned between the encoder and decoder. A mobile client would selectively disable or reenable stages of the algorithm to adapt to the device's effective processing capability. Our variable-complexity strategy of selective disabling of modules supports graceful degradation of the complexity of multimedia coding and decoding into a mobile client's low-power mode, i.e. the clock frequency of its next-generation low power CPU has been scaled down to conserve power. We modified the structure of the standard motion-compensated DCT video codecs to implement a simplified the encoder on a PC server and the decoder on a complexity-constrained PDA viewing client.
Richard Han 0001, Ching-Yung Lin, John R. Smith, Belle L. Tseng, Vida Ha
ACM Multimedia3
2001 MPEG-7 Standard for Multimedia Databases
abstract
The Moving Picture Experts Group (MPEG) is developing a new standard called the “Multimedia Content Description Interface,” also known as MPEG-7. The goal of MPEG-7 is to enable fast and effective searching and filtering of multimedia content. The effort is being driven by requirements taken from a large number of applications related to multimedia databases, interactive media services (music, TV programs), video libraries, and so forth. MPEG-7 is achieving this goal by developing an XML-Schema based standard for describing features of multimedia content. In this tutorial, we study the emerging MPEG-7 standard and describe the new challenges for MPEG-7 multimedia databases.
John R. Smith
SIGMOD Conference1
2001 Supporting Incremental Join Queries on Ranked Inputs
Apostol Natsev, Yuan-Chi Chang, John R. Smith, Chung-Sheng Li, Jeffrey Scott Vitter
VLDB3
2001 Quantitative assessment of image retrieval effectiveness
abstract
Abstract Content‐based retrieval (CBR) promises to greatly improve capabilities for searching for images based on semantic features and visual appearance. However, developing a framework for evaluating image retrieval effectiveness remains a significant challenge. Difficulties include determining how matching at different description levels affects relevance, designing meaningful benchmark queries of large image collections, and developing suitable quantitative metrics for measuring retrieval effectiveness. This article studies the problems of developing a framework and testbed for quantitative assessment of image retrieval effectiveness. In order to better harness the extensive research on CBR and improve capabilities of image retrieval systems, this article advocates the establishment of common image retrieval testbeds consisting of standardized image collections, benchmark queries, relevance assessments, and quantitative evaluation methods.
John R. Smith
J. Assoc. Inf. Sci. Technol.1
2001 MPEG-7 multimedia description schemes
abstract
MPEG-7 multimedia description schemes (MDSs) are metadata structures for describing and annotating audio-visual (AV) content. The description schemes (DSs) provide a standardized way of describing in XML the important concepts related to AV content description and content management in order to facilitate searching, indexing, filtering, and access. The DSs are defined using the MPEG-7 description definition language, which is based on the XML Schema language, and are instantiated as documents or streams. The resulting descriptions can be expressed in a textual form (i.e., human readable XML for editing, searching, filtering) or compressed binary form (i.e., for storage or transmission). In this paper, we provide an overview of the MPEG-7 MDSs and describe their targeted functionality and use in multimedia applications.
Philippe Salembier, John R. Smith
IEEE Trans. Circuits Syst. Video Technol.2
2000 Distributed application service for Internet information portal
abstract
As Internet information portals become prevalent for both Internet and Intranet, most existing Internet Application Server architectures are not scalable to support the large amount of personalization, customization and content adaptation required. We propose a framework to capture the information and content dissemination process. Furthermore, we propose a methodology to map this process to a distributed application server environment. By fully exploiting the intersections of user preference at multiple content processing stages, this new framework enables high hit ratio on processing, storage, and transmission of content and thus scales well to support a large number of clients.
Chung-Sheng Li, John R. Smith, Rakesh Mohan, Yuan-Chi Chang, Brad Topol, John Hind
ISCAS2
2000 The Onion Technique: Indexing for Linear Optimization Queries
abstract
This paper describes the Onion technique, a special indexing structure for linear optimization queries. Linear optimization queries ask for top-N records subject to the maximization or minimization of linearly weighted sum of record attribute values. Such query appears in many applications employing linear models and is an effective way to summarize representative cases, such as the top-50 ranked colleges. The Onion indexing is based on a geometric property of convex hull, which guarantees that the optimal value can always be found at one or more of its vertices. The Onion indexing makes use of this property to construct convex hulls in layers with outer layers enclosing inner layers geometrically. A data record is indexed by its layer number or equivalently its depth in the layered convex hull. Queries with linear weightings issued at run time are evaluated from the outmost layer inwards. We show experimentally that the Onion indexing achieves orders of magnitude speedup against sequential linear scan when N is small compared to the cardinality of the set. The Onion technique also enables progressive retrieval, which processes and returns ranked results in a progressive manner. Furthermore, the proposed indexing can be extended into a hierarchical organization of data to accommodate both global and local queries.
Yuan-Chi Chang, Lawrence D. Bergman, Vittorio Castelli, Chung-Sheng Li, Ming-Ling Lo, John R. Smith
SIGMOD Conference6
2000 SPIRE: A Progressive Content-Based Spatial Image Retrieval Engine
abstract
In this demo, we will show the implementation of a content-based SPatial Image Retrieval Engine (SPIRE) for multimodal unstructured data. This architecture provides a framework for retrieving multi-modal data including image, image sequence, time series and parametric data from large archives. Dramatic speedup (from a factor of 4 to 35) has been achieved for many search operations such as template matching, texture feature extraction. This framework has been applied and validated in solar flares and petroleum exploration in which spatial and spatial-temporal phenomena are located.
Chung-Sheng Li, Lawrence D. Bergman, Vittorio Castelli, John R. Smith
SIGMOD Conference4
2000 Object-based multimedia content description schemes and applications for MPEG-7
abstract
In this paper, we describe description schemes (DSs) for image, video, multimedia, home media, and archive content proposed to the MPEG-7 standard. MPEG-7 aims to create a multimedia content description standard in order to facilitate various multimedia searching and filtering applications. During the design process, special care was taken to provide simple but powerful structures that represent generic multimedia data. We use the extensible markup language (XML) to illustrate and exemplify the proposed DSs because of its interoperability and flexibility advantages. The main components of the image, video, and multimedia description schemes are object, feature classification, object hierarchy, entity-relation graph, code downloading, multi-abstraction levels, and modality transcoding. The home media description instantiates the former DSs proposing the 6-W semantic features for objects, and 1-P physical and 6-W semantic object hierarchies. The archive description scheme aims to describe collections of multimedia documents, whereas the former DSs only aim at individual multimedia documents. In the archive description scheme, the content of an archive is represented using multiple hierarchies of clusters, which may be related by entity-relation graphs. The hierarchy is a specific case of entity-relation graph using a containment relation. We explicitly include the hierarchy structure in our DSs because it is a natural way of defining composite objects, a more efficient structure for retrieval, and the representation structure used in MPEG-4. We demonstrate the feasibility and the efficiency of our description schemes by presenting applications that already use the proposed structures or will greatly benefit from their use. These applications are the visual apprentice, the AMOS-search system, a multimedia broadcast news browser, a storytelling system, and an image meta-search engine, MetaSEEk.
Ana B. Benitez, Seungyup Paek, Shih-Fu Chang, Atul Puri, John R. Smith, Chung-Sheng Li, Lawrence D. Bergman, Charles N. Judice
Signal Process. Image Commun.6
1999 An Adaptive View Element Framework for Multi-Dimensional Data Management
abstract
We present an adaptive wavelet view element framework for managing different types of multi-dimensional data in storage and retrieval applications. We consider the problems of multi-dimensional data compression, multi-resolution subregion access, selective materialization, progressive retrieval and similarity searching. The framework uses wavelets to partition the multi-dimensional data into view elements that form the building blocks for synthesizing views of the data. The view elements are organized and managed using different view element graphs. The graphs are used to guide cost-based view element selection algorithms for optimizing compression, access, retrieval and search performance.
John R. Smith, Chung-Sheng Li
CIKM1
1999 Scalable multimedia delivery for pervasive computing
abstract
Growing numbers of pervasive devices are gaining access to the Internet and other information sources. However, much of the rich multimedia content cannot be easily handled by the client devices with limited communication, processing, storage and display capabilities. In order to improve access, we are developing a system for scalable delivery of multimedia. The system uses an InfoPyramid for managing and manipulating multimedia content composed of video, images, audio and text. The InfoPyramid manages the different variations of media objects with different fidelities and modalities and generates and selects among the alternatives in order to adapt the delivery to different client devices. We describe a system for scalable multimedia delivery for a variety of client devices, including PDAs, HHCs, smart phones, TV browsers and color PCs.
John R. Smith, Rakesh Mohan, Chung-Sheng Li
ACM Multimedia (1)1
1999 Image Classification and Querying Using Composite Region Templates
John R. Smith, Chung-Sheng Li
Comput. Vis. Image Underst.1
1999 Integrated Spatial and Feature Image Query
John R. Smith, Shih-Fu Chang
Multim. Syst.1
1999 Adapting Multimedia Internet Content for Universal Access
abstract
Content delivery over the Internet needs to address both the multimedia nature of the content and the capabilities of the diverse client platforms the content is being delivered to. We present a system that adapts multimedia Web documents to optimally match the capabilities of the client device requesting it. This system has two key components. 1) A representation scheme called the InfoPyramid that provides a multimodal, multiresolution representation hierarchy for multimedia. 2) A customizer that selects the best content representation to meet the client capabilities while delivering the most value. We model the selection process as a resource allocation problem in a generalized rate distortion framework. In this framework, we address the issue of both multiple media types in a Web document and multiple resource types at the client. We extend this framework to allow prioritization on the content items in a Web document. We illustrate our content adaptation technique with a web server that adapts multimedia news stories to clients as diverse as workstations, PDA's and cellular phones.
Rakesh Mohan, John R. Smith, Chung-Sheng Li
IEEE Trans. Multim.2
1999 VideoZoom Spatio-Temporal Video Browser
abstract
We describe a system for browsing and interactively retrieving video over the Internet at multiple spatial and temporal resolutions. The VideoZoom system enables users to start with coarse, low-resolution views of the sequences and selectively zoom-in in space and time. VideoZoom decomposes the video sequences into a hierarchy of view elements, which are retrieved in a progressive fashion. The client browser incrementally builds the views by retrieving, caching, and assembling the view elements, as needed. By integrating browsing and retrieval into a single progressive retrieval paradigm, VideoZoom provides a new and useful system for accessing video over the Internet. VideoZoom is suitable for digital video libraries and a number of other applications in which streaming methods provide insufficient quality of video, video downloading introduces large latencies, and generating video summaries is difficult or not well integrated with video retrieval tasks.
John R. Smith
IEEE Trans. Multim.1
1998 Multimedia content description in the InfoPyramid
abstract
There is a growing need for developing a content description language for multimedia that improves searching, indexing and managing of the multimedia content. The MPEG group established the MPEG-7 effort to standardize the multimedia content interface. The proposed interface will bridge the gap between various types of content meta-data, such as content features, annotations, relationships, and the search engines. We develop a method of handling multimedia content description in a new multi-abstraction, multi-modal content representation framework called the InfoPyramid. The InfoPyramid facilitates the search, retrieval, manipulation, and transmission of multimedia data by providing a hierarchy for content descriptors. We illustrate the suitability of the InfoPyramid multimedia content description to MPEG-7 by examining four multimedia retrieval applications: a Web-image search engine, a satellite image retrieval system, an Internet content delivery system, and a TV news storage and retrieval system.
Chung-Sheng Li, Rakesh Mohan, John R. Smith
ICASSP3
1998 Content-based Transcoding of Images in the Internet
John R. Smith, Rakesh Mohan, Chung-Sheng Li
ICIP (3)1
1998 Dynamic Assembly of Views in Data Cubes
abstract
In this paper, we present a method for dynamically assembling views in multi-dimensional data cubes in order to more e#ciently support data analysis and querying involving aggregations. The proposed method decomposes the data cubes into an indexed hierarchy of view elements. The view elements di#er from traditional data cube cells in that they correspond to partial and residual aggregations of the data cube. The view elements provide highly granular building blocks for synthesizing the aggregated and rangeaggregated views of the data cubes. We propose a strategy for selecting and materializing the view elements based on the frequency of view access. This allows the dynamic adaptation of the view element sets to patterns of retrieval. We present a fast and optimal algorithm for selecting non-expansive view element sets that minimize the processing costs for generating a population of aggregated views. We also present a greedy algorithm for selecting redundant view element sets in order...
John R. Smith, Chung-Sheng Li, Vittorio Castelli, Anant Jhingran
PODS1
1997 Exploring Image Functionalities in WWW Applications- Development of Image/Video Search and Editing Engines
abstract
Image technologies provide critical components for achieving various Web-based multimedia applications. We identify key technical challenges in today's image/video applications on the WWW. We present two prototype systems, a video search engine (WebSEEk) and a compressed video editor (WebClip), to demonstrate technical challenges and viable solutions. We also discuss major issues in developing next-generation, scalable solutions for large-scale, distributed on-line environments.
Shih-Fu Chang, John R. Smith, Horace J. Meng
ICIP (3)2
1997 Joint adaptive space and frequency basis selection
abstract
We develop a new method for building a representation of an image from a library of basis elements that is facilitated by a joint adaptive space and frequency (JASF) graph. The JASF graph combines partitionable frequency expansion and spatial segmentation of the image, symmetrically. We demonstrate by using a rate-distortion framework for basis selection that the JASF graph improves compression performance over recent wavelet packet and double-tree methods by offering exponentially more bases in which to represent the images.
John R. Smith, Shih-Fu Chang
ICIP (3)1
1997 SaFe: a general framework for integrated spatial and feature image search
abstract
We present a system for querying for images by the spatial and feature attributes of regions. The system enables the user to find the images that contain an arrangement of regions similar to that diagrammed in a query image. We propose a general framework which allows for different types of features (e.g., color, texture, shape, motion) to be integrated with spatial information in the query process. We demonstrate that integrated spatial and feature querying improves image search capabilities over previous content-based image retrieval methods.
John R. Smith, Shih-Fu Chang
MMSP1
1997 Enhancing image search engines in visual information environments
abstract
The recent diverse environments and applications for image searching (e.g. Web image search engines) provide an enormous resource of information beyond the image pixels which can be used to improve the image search process. We explore several directions of enhancements that integrate the visual information with other information related to the images in the analysis and query processes. We demonstrate that these methods improve image search functionalities over non-integrated content-based methods.
John R. Smith, Shih-Fu Chang
MMSP1
1996 Automated binary texture feature sets for image retrieval
abstract
Digital image and video libraries require new algorithms for the automated extraction and indexing of salient image features. Texture features provide one important cue for the visual perception and discrimination of image content. We propose a new approach for automated content extraction that allows for efficient database searching using texture features. The algorithm automatically extracts texture regions from image spatial-frequency data which are represented by binary texture feature vectors. We demonstrate that the binary texture features provide excellent performance in image query response time while providing highly effective texture discriminability, accuracy in spatial localization and capability for extraction from compressed data representations. We present the binary texture feature extraction and indexing technique and examine searching by texture on a database of 500 images.
John R. Smith, Shih-Fu Chang
ICASSP1
1996 Local color and texture extraction and spatial query
abstract
In this paper we present a unified system for the extraction, representation and query of spatially localized color and texture regions. The system utilizes a back-projection of binary feature sets to identify and extract prominent regions. The binary feature sets provide an effective and easily indexable representation of color and texture. We also provide a mechanism for integrating features by combining the binary color and texture feature sets. This enables the extraction and representation of joint color and texture regions. Since all extracted regions are spatially localized, in image database queries the user can specify the locations and spatial boundaries of regions. We present the unified color and texture back-projection method and describe its implementation in the VisualSEEk content-based image retrieval system.
John R. Smith, Shih-Fu Chang
ICIP (3)1
1996 VisualSEEk: A Fully Automated Content-Based Image Query System
abstract
Article Free Access Share on VisualSEEk: a fully automated content-based image query system Authors: John R. Smith Department of Electrical Engineering and Center for Image Technology for New Media, Columbia University, New York, N.Y. Department of Electrical Engineering and Center for Image Technology for New Media, Columbia University, New York, N.Y.View Profile , Shih-Fu Chang Department of Electrical Engineering and Center for Image Technology for New Media, Columbia University, New York, N.Y. Department of Electrical Engineering and Center for Image Technology for New Media, Columbia University, New York, N.Y.View Profile Authors Info & Claims MULTIMEDIA '96: Proceedings of the fourth ACM international conference on MultimediaFebruary 1997 Pages 87–98https://doi.org/10.1145/244130.244151Online:01 February 1997Publication History 1,080citation3,873DownloadsMetricsTotal Citations1,080Total Downloads3,873Last 12 Months90Last 6 weeks11 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
John R. Smith, Shih-Fu Chang
ACM Multimedia1
1995 Frequency and spatially adaptive wavelet packets
abstract
We consider a method for image compression based on frequency and spatially adaptive wavelet packets. We present a new fast directed acyclic graph (DAG) structured decomposition, with both spatial segmentation and orthogonal frequency branching from each node. Whereas traditional wavelet packet decomposition adapts to a global frequency distribution, this technique finds the best joint spatial segmentation and local frequency basis. The algorithm is derived from the fast double tree algorithm proposed by Herley, et. al. (see IEE Transactions on Signal Processing, December 1993), for 1-D signals, with an extension to 2-D and modification to include spatial segmentation of frequency nodes. By collecting redundant nodes in this full adaptive tree, we have derived a directed acyclic graph (DAG) structure which contains the same number of nodes as the double tree, but includes new connections between nodes. We present the adaptive wavelet packet DAG algorithm and examine image compression performance on test images.
John R. Smith, Shih-Fu Chang
ICASSP1
1995 Single color extraction and image query
abstract
We propose a method for automatic color extraction and indexing to support color queries of image and video databases. This approach identifies the regions within images that contain colors from predetermined color sets. By searching over a large number of color sets, a color index for the database is created in a fashion similar to that for file inversion. This allows very fast indexing of the image collection by the color contents of the images. Furthermore, information about the identified regions, such as the color set, size, and location, enables a rich variety of queries that specify both color content and spatial relationships of regions. We present the single color extraction and indexing method and contrast it to other color approaches. We examine single and multiple color extraction and image query on a database of 3000 color images.
John R. Smith, Shih-Fu Chang
ICIP (3)1
1994 Transform Features for Texture Classification and Discrimination in Large Image Databases
abstract
Proposes a method for classification and discrimination of textures based on the energies of image subbands. The authors show that with this relatively simple feature set, effective texture discrimination can be achieved. In the paper, subband-energy feature sets extracted from the following typical image decompositions are compared: wavelet subband, uniform subband, discrete cosine transform (DCT), and spatial partitioning. The authors report that over 90% correct classification was attained using the feature set in classifying the full Brodatz [1965] collection of 112 textures. Furthermore, the subband energy-based feature set can be readily applied to a system for indexing images by texture content in image databases, since the features can be extracted directly from spatial-frequency decomposed image data. The authors also show that to construct a suitable space for discrimination, Fisher discrimination analysis (Dillon and Goldstein, 1984) can be used to compact the original features into a set of uncorrelated linear discriminant functions. This procedure makes it easier to perform texture-based searches in a database by reducing the dimensionality of the discriminant space. The authors also examine the effects of varying training class size, the number of training classes, the dimension of the discriminant space and number of energy measures used for classification. The authors hope that the performance for texture discrimination of these simple energy-based features will allow images in a database to be efficiently and effectively indexed by contents of their textured regions.>
John R. Smith, Shih-Fu Chang
ICIP (3)1
1989 The impact of supercomputing capabilities on U.S. materials science and technology
William D. Wilson, Robert J. Asaro, Robert W. Dutton, Juan M. Sanchez, David J. Srolovitz, Richard H. Boyd, William A. Goddard III, John R. Smith, Wilhelm G. Wolfer
Future Gener. Comput. Syst.8
1968 An Adaptive Threshold Logic Gate Using Capacitive Analog Weights
abstract
A physical realization of an adaptive threshold logic gate which can be used to realize linearly separable switching functions is presented. The system is completely electronic and is easily implemented in the sense that standard components may be used throughout. The memory circuit used for each weight of the device stabilizes the voltage across a capacitor by means of a sampling technique. Over the range that the capacitor voltage is stabilized, the net average capacitor current is zero at a discrete number of stable voltages which can be made numerous enough to approximate analog memory. The device exhibits both long-term stability and ease in weight adjustment. Training results are given for several linearly separable switching functions. This circuit shows that a practical adaptive threshold logic gate can be realized. The incorporation of integrated circuitry in the device presented would greatly enhance the feasibility of building more complex trainable machines.
John R. Smith, Cyrus O. Harbourt
IEEE Trans. Computers1